Multi-instance learning for peptide-MHC presentation prediction

The multi-instance learning approach addresses the limitations of existing peptide-MHC presentation methods by calibrating model confidence and efficiently using limited data, enhancing predictive accuracy for MHC class II molecules, which is essential for personalized vaccine design and immunotherapy.

JP7788046B2Active Publication Date: 2025-12-18NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023522514
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-13
Filing Date
2021-03-12
Publication Date
2025-12-18
Estimated Expiration
2041-03-12

AI Technical Summary

Technical Problem

Existing methods for predicting peptide-MHC presentation, particularly for MHC class II molecules, face challenges due to limited training data and inefficiencies in using mass spectrometry data, leading to suboptimal predictive performance.

Method used

A multi-instance learning (MIL) approach that uses a loss function to train a classifier, explicitly compensating for model confidence and calibrating probabilities, allowing for accurate prediction of peptide binding and presentation by MHC molecules, even with multiple alleles.

Benefits of technology

The proposed method enhances predictive accuracy by effectively utilizing limited data and transferring knowledge across MHC alleles, improving the identification of peptides likely to be presented on the cell surface, crucial for personalized T cell-based vaccine design and immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007788046000029
    Figure 0007788046000029
  • Figure 0007788046000030
    Figure 0007788046000030
  • Figure 0007788046000031
    Figure 0007788046000031
Patent Text Reader

Abstract

Disclosed embodiments of the present invention provide a computer-implemented method for predicting peptide binding and presentation by MHC molecules, the method comprising: collecting or generating training data comprising a set of MHC molecules present in a biological sample and a set of observed peptide sequences presented by at least one of the MHC molecules present in the biological sample, where it is unknown to which particular MHC molecules the peptide sequences bind, and the training data is organized as a plurality of bags, each bag having a set of training instances, where labels are known for the bags but unknown for the training instances; θ We train the MIL classifier f at the instance level and then apply the labels of multiple new instances to the MIL classifier f θ and / or predict the labels of multiple new bags by directly applying the MIL classifier f θ and making predictions by applying the criterion 1 to each instance in each bag and aggregating the results across all instances in each bag.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a computer-implemented method and system for predicting peptide binding and presentation by MHC molecules.

[0002] Additionally, the present invention relates to a computer-implemented method for performing multi-instance learning, or MIL. [Background technology]

[0003] The adaptive immune system plays a central role in the immune response against foreign molecules such as pathogens and cancer cells. It is divided into two main components: humoral immunity, which is involved in antibody production, and cellular immunity, which entails, among other things, the stimulation of cytotoxic CD8+ T cells.

[0004] Major histocompatibility complex (MHC) class II plays an important role in both humoral and cellular immunity (see for reference Murphy, K. and Weaver, C., 2016. Janeway's immunobiology. Garland science). The primary role of MHC class II is to bind to peptide sequences, which are short amino acid sequences derived from exogenous proteins, and then present them on the cell surface. This peptide-MHC complex stimulates CD4+ T cells or "helper T cells." The helper T cells can then stimulate either humoral or cellular immune response pathways.

[0005] MHC class II molecules are primarily found in "professional" antigen-presenting cells, such as dendritic cells. Each individual typically possesses two alleles of MHC class II molecules, each derived from the HLA-DQ and HLA-DP gene families, while up to 10 alleles derived from the HLA-DR gene family may be present (see for reference Choo, SY, 2007. The HLA system: genetics, immunology, clinical testing, and clinical implications. Yonsei medical journal, 48(1), pp. 11-23). ​​Individuals possess different MHC alleles, although importantly, some alleles are more common than others. Different versions of MHC alleles have different amino acid sequences and structures, and these differences affect which peptides an MHC allele binds to and presents on the cell surface.

[0006] Peptide presentation to T cells involves a series of processes. Key steps include binding between MHC molecules and peptides and presentation of peptide-MHC complexes on the cell surface. Mass spectrometry can be used to detect peptides eluted from the cell surface, thereby determining peptide presentation (see for reference Purcell, AW, Ramarathinam, SH, and Ternette, N., 2019. Mass spectrometry-based identification of MHC-bound peptides for immunopeptidomics. Nature protocols, 14(6), 1687). Such assays for hundreds of different MHC molecules have generated thousands of data points (see for reference Vita, R., Mahajan, S., Overton, JA, Dhanda, SK, Martini, S., Cantrell, JR, Wheeler, DK, Sette, A., and Peters, B., 2019. The immune epitope database (IEDB): 2018 update, Nucleic acids research, 47(D1), pp. D339-D343). As previously mentioned, each individual has multiple MHC class II molecules, and therefore, typical mass spectrometry experiments cannot accurately identify the MHC molecule that presented a particular peptide. Another limitation of mass spectrometry is that it can only show detected peptides, i.e., mass spectrometry cannot generate "negative" data points. Therefore, a significant challenge is to use this experimental data to train machine learning models to predict peptide-MHC presentation. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Murphy, K. and Weaver, C., 2016. Janeway's immunobiology. Garland science [Non-patent document 2] Choo, S.Y., 2007. The HLA system: genetics, immunology, clinical testing, and clinical implications. Yonsei medical journal, 48(1), pp. 11-23 [Non-patent document 3] Purcell, AW, Ramarathinam, SH, and Ternette, N., 2019. Mass spectrometry-based identification of MHC-bound peptides for immunopeptidomics. Nature protocols, 14(6), 1687 [Non-patent document 4] Vita, R., Mahajan, S., Overton, JA, Dhanda, SK, Martini, S., Cantrell, JR, Wheeler, DK, Sette, A., and Peters, B., 2019. The immune epitope database (IEDB): 2018 update, Nucleic acids research, 47(D1), pages D339-D343 [Non-patent document 5] Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. Proceedings of the 16th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1, pp. 2227-2237. [Non-patent document 6] Barlow, R.E., 1972. Statistical inference under order restrictions; the theory and application of isotonic regression (No. 04; QA278. 7, B3.) [Non-Patent Document 7] Reynisson, B., Alvarez, B., Paul, S., Peters, B., and Nielsen, M., 2020. NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research [Non-patent document 8] Jensen, K. K., Andreatta, M., Marcatili, P., Buus, S., Greenbaum, J. A., Yan, Z., Sette, A., Peters, B., and Nielsen, M. (2018). Improved methods for predicting peptide binding affnity to MHC class II molecules. Immunology, 154(3), pp. 394-406. [Non-Patent Document 9] G. L., Lin, H. H., Keskin, D. B., Reinherz, E. L., and Brusic, V. (2011). Dana-farber repository for machine learning in immunology. Journal of immunological methods, 374(1-2), pp. 18-25 [Non-Patent Document 10] Reynisson, B., Alvarez, B., Paul, S., Peters, B., and Nielsen, M. (2020). NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research, pp. 1-6 Summary of the Invention [Problem to be solved by the invention]

[0008] It is therefore an object of the present invention to improve and further develop methods and systems of the type described in the introduction so as to improve their predictive performance. [Means for solving the problem]

[0009] According to the present invention, the foregoing object is achieved by a computer-implemented method for predicting peptide binding and presentation by MHC molecules, comprising collecting or generating training data, the training data comprising a set of MHC molecules present in a biological sample and a set of observed peptide sequences presented by at least one of the MHC molecules present in the biological sample, where it is unknown to which particular one of the MHC molecules the peptide sequence binds, the training data being organized as a plurality of bags, each bag having a set of training instances, where labels are known for the bags but unknown for the training instances; and using a loss function to generate a MIL classifier f θ We train the MIL classifier f at the instance level and then apply the labels of multiple new instances to the MIL classifier f θ and / or predict the labels of multiple new bags by directly applying the MIL classifier f θ to each instance of each bag and aggregating the results across all instances of each bag.

[0010] The foregoing objects are further achieved by a computer-implemented method for performing multi-instance learning, i.e., MIL, including: collecting or generating training data, the training data including a plurality of bags, each bag having a set of training instances, where labels are known for the bags but unknown for the training instances; training a MIL classifier at the instance level by using a loss function that explicitly compensates for the model's confidence in model predictions during training, where each training instance from a positively labeled bag is weighted by a calibrated confidence function of the current model; and predicting labels of a plurality of new instances by directly applying the MIL classifier and / or predicting labels of a plurality of new bags by applying the MIL classifier to each instance in each bag and aggregating results across all instances in each bag.

[0011] In a further embodiment, a system for predicting peptide binding and presentation by MHC molecules comprises one or more processors configured to enable the execution of any of the methods according to the embodiments of the present invention, either alone or in combination.

[0012] In a further embodiment, a tangible, non-transitory computer-readable medium includes instructions that, when executed on one or more processors, cause the one or more processors to perform, alone or in combination, any of the methods according to embodiments of the present invention.

[0013] Embodiments of the present invention provide an MIL algorithm applied to peptide-MHC prediction in the case of multiple MHC alleles. Embodiments of the present invention allow for the efficient use of typical peptide-MHC mass spectrometry data with multiple labels of potential alleles. However, while this disclosure focuses on accurately predicting peptide binding and presentation by MHC alleles, a key step toward personalized T cell-based vaccine design and immunotherapy, embodiments of the present invention also relate to the application of the MIL algorithm in a variety of contexts.

[0014] In one embodiment, the present invention provides a computer-implemented method for performing multi-instance learning, the method including a first step of collecting or generating training data, where labels are known only for a bag of instances. The method may further include training a classifier at the instance level, where each training instance from the bag of positively labeled instances is weighted by a calibrated confidence of the current model in a loss function using the training data from the first step. The method may then include predicting labels for multiple new instances based on the trained MIL classifier by directly applying the instance-level classifier, or predicting labels for multiple new bags by applying the instance-level classifier to each instance and aggregating scores across all instances in the bags.

[0015] In one embodiment, the MIL classifier can be trained using a loss function that explicitly compensates for the model's confidence in the model predictions during training. In the same or other embodiments, it can be provided that the probabilities are calibrated to accurately reflect the current model's confidence by a probability calibration function. In this context, it can be allowed that the training instances in the bag that are positively labeled are weighted by the calibrated confidence level of the current model.

[0016] There are several ways in which the teaching of the present invention can be advantageously designed and further developed, and for this purpose reference is made on the one hand to the dependent claims and on the other hand to the following description of preferred embodiments of the invention, which are shown by way of example in the figures. In conjunction with the description of the preferred embodiments of the invention using the figures, generally preferred embodiments and further developments of the teaching are described. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 is a schematic diagram illustrating a prediction scheme based on experimentally obtained data, according to an embodiment of the present invention. [Figure 2] FIG. 1 is a schematic diagram illustrating bag label prediction by using a classifier to predict instance labels and by applying a pooling operation according to an embodiment of the present invention; [Figure 3] FIG. 2 is a schematic diagram illustrating a probability calibration function used to calibrate the belief of a model, according to one embodiment of the present invention. [Figure 4] FIG. 1 is a schematic diagram illustrating a loss function modified to approximate negative samples with negative sampling, according to an embodiment of the present invention; [Figure 5] FIG. 1 is a schematic diagram illustrating personalized cancer vaccine design according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] Predicting binding and presentation between MHC molecules and peptides is an important step toward T cell-based vaccine design and immunotherapy. Given the importance of the problem and the availability of data, many methods have been developed to predict MHC-peptide binding and peptide presentation. In some approaches, a single model is trained specifically for each MHC allele, while in other approaches, a single model covering all MHC alleles (pan model) is trained. The predictive performance of MHC class I models has reached a high level (auROC > 0.98; see Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. Proceedings of the 16th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1, pp. 2227–2237). On the other hand, models for class II still have limited performance. Despite recent progress, there is still a need for better performing models. One significant factor limiting MHC class II models is the limited amount of training data compared to class I. Therefore, models that can efficiently use the limited available data and transfer knowledge from other sources would be extremely beneficial.

[0019] As previously mentioned, predicting which peptides can or cannot be presented by which MHC molecules is crucial for neoantigen discovery and T cell-based vaccine design, among other health-related challenges. One important source of training data for such models is mass spectrometry, a technique that identifies short peptides presented on the cell surface due to available MHC molecules within the cell. As shown in the left portion of Figure 1, many mass spectrometry data sets 100 are generated using two or more MHC molecules within a cell, meaning that for a positive peptide 110 discovered using mass spectrometry, one or more MHC molecules 120a, 120b may be responsible for presenting this peptide 110.

[0020]

[0003] Embodiments of the present invention provide methods and systems for prioritizing peptides for inclusion in a vaccine for a particular individual based on the likelihood that they will be presented on the cell surface by MHC molecules. In one embodiment, prioritization is posed as a prediction problem, and a multi-instance learning (MIL) formulation is employed to solve it. While previous work has also formulated this as a MIL problem, embodiments of the present invention use a novel learning algorithm to explicitly compensate and calibrate model confidence during the learning process.

[0021] In standard supervised learning, a label is given for each input example. However, in some contexts, labels are instead assigned to sets or bags of inputs. In this setting, a bag of inputs is labeled positive if it contains at least one positive input; otherwise, the bag is labeled negative.

[0022] Therefore, according to one embodiment of the present invention, the training data for multi-instance learning, or MIL, is X={x1, x2,..., x N} and the associated bag labels are {y1,y2,...,y N Each bag can be defined as a set of instances, namely X i ={x i1 ,x i2 ,...,x im In MIL, each instance in the bag has the label y ij ∈{0,1} but remains unknown during training. i Only, that is, given by:

[0023]

number

[0024] MIL classifier f θ The new bag label θ We can learn to predict (X) (bag-level methods) or the label of an instance f θ (x ij ) (instance-level approach). In embodiments of the present invention, we focus on training a classifier to predict the label of an instance, i.e., the instance-level approach.

[0025] Using a classifier that predicts the label of an instance, the label of a bag can be predicted by applying a pooling operation h(·) to the predictions for all instances in the bag. f θ (X)=h(f θ (x i1 ),f θ (x i2 ),...,f θ (x im )) This is shown in the right part of Figure 1 and in Figure 2, where s i represents the peptide sequence, and A i ={a1,...,a m} is a set of m MHC molecules associated with a biological sample, and y i is, s i A i is a binary label indicating whether the antigen was found to be presented by any of the MHC molecules in

[0026] The definition of the problem requires that h(·) is a permutation-invariant function, i.e., the order of inputs to the function does not affect the result. θ is of the form

[0027]

number

[0028] where N is the number of bags and M is the number of instances in each bag. Note that in general, not all bags need to have the same number of instances.

[0029] According to some embodiments, the present invention provides methods and systems that include a multi-instance learning (MIL) approach based on the peptide-MHC presentation problem discussed above. The method can be implemented in two stages: an offline training stage and an online prediction stage. In the offline training stage, a predictive model is trained, which, in contrast to previous studies, explicitly compensates and calibrates the model's confidence. The trained model is then used during the online prediction stage.

[0030] According to some embodiments, the present invention provides methods and systems for predicting peptide binding and presentation by MHC molecules, configured to receive as input in an offline training phase a set of observed peptides presented by at least one MHC molecule present in a biological sample, as well as a set of MHC molecules present in the sample. However, as already explained above, it is unknown to which specific MHC molecule the peptide bound. This is exactly the kind of data generated by mass spectrometry experiments.

[0031] In one embodiment, standard techniques can be used to generate negative examples for training, but it should be noted that the applicability of the technique proposed according to the present invention does not depend on how the negative examples are created.

[0032] More specifically, the input is a set of triples {s i ,A i ,y i}, where s i is the peptide sequence, and A i ={a1,...,a m} is a set of m MHC molecules associated with a biological sample, and y i is, s i A i is a binary label indicating whether the antigen was found to be presented by any of the MHC molecules in

[0033] The goal of the offline training phase is to take X as input. i =(s i ,A i ) and i A machine learning model that correctly predicts f θ The key is to train the θOne example of a pre-trained model is the Bidirectional Encoder Representations from Transformers (BERT) model. However, as will be appreciated by those skilled in the art, other model types are possible as well. The only constraint is the probability p(y ij =1|s i ,a j ) that the model must provide.

[0034] According to one embodiment of the present invention, each peptide is associated with a bag of alleles. A bag is labeled positive if at least one of the alleles represents the peptide; otherwise, the bag is labeled negative. The training data can be modeled as a multi-instance learning (MIL) problem, where the i-th bag with m alleles is A i ={a i1 ,a i2 ,...,a im} and the corresponding peptide sequence is s i At each training step, we train every instance in the bag (a ij ,s i ) probability p(y ij =1|x ij ) into the neural network model f θ Using

[0035]

number

[0036] A symmetric pooling operator can be used to pool the bag predictions from the predictions of the instances in the bag. To incorporate the uncertainty of the deconvolution operation, at each training epoch, each positive data point i from the deconvolution is assigned a calibrated predicted probability of being positive.

[0037]

number

[0038] can be weighted by

[0039] According to one embodiment of the present invention, the following loss function is then used:

[0040]

number

[0041] The model parameters can be learned according to

[0042]

number

[0043] and

[0044]

number

[0045] is the number of previous training epochs of the model

[0046]

number

[0047] is the predicted probability of , C is the probability calibration function (Figure 3), w is the weighting value for the positive class to account for class imbalance, and x ij is a tuple(s i ,a i ), where s i is a peptide, and a i is A iAccording to the embodiment shown in Figure 3, the probability calibration function C takes as input the value of the current training epoch k of the model

[0048]

number

[0049] and the calibrated probability for the subsequent training epoch k+1 of the model.

[0050]

number

[0051] With regard to instance weighting, it should be noted that only instances in the positively labeled bag are weighted with the calibrated confidence of the model according to an embodiment of the present invention, but not negative samples (since there is no uncertainty associated with the negative class label). In this context, it can be realized that all negative samples are used, or negative sampling is performed if there are many negative bags and computation is limited.

[0052] In the given formulation, we incorporate all negative instances in all negative bags. However, this is computationally intractable when there are many negative bags. Therefore, according to an alternative embodiment, we can achieve approximation of negative instances with negative sampling, as shown in Figure 4. Therefore, we can rewrite the above loss function as follows:

[0053]

number

[0054] can be corrected as follows.

[0055] For computational reasons, rather than using all negative samples, we perform negative sampling using a probability distribution P i (X i According to one embodiment, in the MHC-peptide presentation problem, P i (X i )=P i (x i1 ,x i2 ,...,x im ) the following delta distribution

[0056]

number

[0057] You can choose to use:

[0058] That is, the method uses the most likely positive examples predicted from the negative bag by the current model.

[0059] In view of the above, a multi-instance learning (MIL) algorithm according to one embodiment of the present invention applied to peptide-MHC prediction in the case of multiple MHC alleles can be defined as follows. Algorithm: Probability Reweighted Multi-Instance Learning Input: Training data {X i ,y i} i∈1...N , but X i :={s i ,A i}, y i ∈{0,1}; Randomly initialize θ0 or move θ0 from the related task, θ k ←Select θ0 and w while non-convergent do: for

[0060]

number

[0061] k of: Bag label, current model

[0062]

number

[0063] Predict using Probability Calibration Model

[0064]

number

[0065] of

[0066]

number

[0067] Train using as input θ t ←θ k for

[0068]

number

[0069] Of which:

[0070]

number

[0071] end for θ k ←θ t end for return θ

[0072] It is important to note that, in contrast to the prior art, our loss function L(θ) explicitly compensates for the model's confidence in the model predictions during training. According to an embodiment of the present invention, this is

[0073]

number

[0074] is the predicted probability of

[0075]

number

[0076] Specifically, in existing methods, the loss is calculated by compensating for the incorrect instance x ij Therefore, the incorrect instance x ij f θ may be optimized.

[0077] Furthermore, embodiments of the present invention also extend the prior art by including a function C for calibrating the predicted probabilities.

[0078]

number

[0079] is the predicted logit (i.e., odds

[0080]

number

[0081] ) and the labels attached to the training set. By way of example, monotonic regression can be performed according to the method described in the following document: Barlow, RE, 1972. Statistical inference under order restrictions; the theory and application of isotonic regression (No. 04; QA278.7, B3.), the entire contents of which are hereby incorporated by reference. However, as will be understood by those skilled in the art, other methods, such as Platt scaling, can also be used. The applicability of the method proposed according to the present invention does not depend on the exact form of the calibration function.

[0082] The parameters of the model, θ, can then be trained to minimize this loss function using an appropriate optimization technique, e.g., f, such as in the case of the BERT model. θ If f is differentiable, a gradient descent algorithm or similar algorithm can be used. θ If f is not differentiable, Bayesian optimization or other black-box methods can be used. The applicability of the proposed method according to the present invention is θ is differentiable or not.

[0083] After the offline training phase is completed as described above, an online prediction phase can be performed. Specifically, after training, the model f θ takes X as input i Take the label y i That is, the model takes as input a peptide sequence and a set of MHC molecules, and the model predicts whether the peptide is presented by any of those MHC molecules. According to an embodiment, the MIL classifier f θUsing this method, predictions can be made for all combinations of peptide sequences and MHC molecules present in a biological sample, based on which peptides that are most likely to be presented can be determined as candidates for synthesis and inclusion in a personalized cancer vaccine.

[0084] In fact, peptide presentation by MHC molecules is only one (very important) step among many that ultimately leads to the creation of an effective cancer vaccine. Predictive models for many of these steps clearly do not require multi-instance learning problems. Therefore, the proposed approach according to embodiments of the present invention may be applicable to only a portion of the vaccine design process. Furthermore, it should be noted that the proposed approach requires a model that outputs some notion of probability, which is common in classification problems but less common in regression problems. Therefore, this approach may be of limited use to multi-instance regression learning problems. Furthermore, it should be noted that most probability calibration functions require the availability of all uncalibrated probabilities. Therefore, mini-batch optimization approaches that update a model after making predictions on only a few training samples may be incompatible with some embodiments of the present approach. Instead, in embodiments of the present invention, a calibration model is trained at the beginning of each epoch.

[0085] The current state of the art in multi-instance learning for peptide-MHC presentation is the work of Reynisson, B., Alvarez, B., Paul, S., Peters, B., and Nielsen, M., 2020. NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research, the entire contents of which are hereby incorporated by reference. However, their approach does not incorporate confidence weighting or calibration operations. Experiments will demonstrate that our approach outperforms the approach of Reynisson et al. on a variety of datasets.

[0086] MHC class II binding data According to an embodiment of the present invention, data from Jensen et al., 2018 (see Jensen, K.K., Andreatta, M., Marcatili, P., Buus, S., Greenbaum, J.A., Yan, Z., Sette, A., Peters, B., and Nielsen, M. (2018). Improved methods for predicting peptide binding affinity to MHC class II molecules. Immunology, 154(3), pp. 394-406, the entire contents of which are hereby incorporated by reference) was used to train the MHC class II binding model, as this data was designed to minimize overlap between the training and evaluation sets. The original data were collected from the Immune Epitope Database (IEDB, Vita, R., Mahajan, S., Overton, JA, Dhanda, SK, Martini, S., Cantrell, JR, Wheeler, DK, Sette, A., and Peters, B. (2019). The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Research, 47(D1), pp. D339–D343, accessed June 30, 2020) up to 2016. The data consist of 134,281 data points, covering HLA-DR, HLA-DQ, HLA-DP, and H-2 mouse MHC alleles. Affinity indexes were converted from IC50 to a value between 0 and 1 using the formula 1-log(IC50) / log(50000).

[0087] Data from Jensen et al. were collected from the IEDB up to 2016. Quantitative binding data were collected from the IEDB to benchmark against an independent dataset that had not been trained or validated using the model, and data already used in Jensen et al. was filtered out. Additionally, independent binding data were collected from the Dana-Farber repository (see for reference: GL, Lin, HH, Keskin, DB, Reinherz, EL, and Brusic, V. (2011). Dana-farber repository for machine learning in immunology. Journal of immunological methods, 374(1-2), pp. 18-25, the entire contents of which are incorporated herein by reference). Finally, 2413 additional MHC-peptide pairs covering 47 MHC class II alleles were collected.

[0088] MHC class II presentation data To train the MHC class II mass spectrometry presentation model, we used data curated from the following publication: Reynisson, B., Alvarez, B., Paul, S., Peters, B., and Nielsen, M. (2020). NetMHCpan-4.1 and NetMHCIIpan-4.0: Improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research, pp. 1–6, the entire contents of which are hereby incorporated by reference. Original data were curated from the IEDB and other public sources. The data cover 41 MHC class II alleles and peptide lengths ranging from 13 to 21. Each data point consists of a peptide ligand, a source protein, and a list of possible MHC class II alleles bound to the peptide. Data points that unambiguously represent only one MHC allele are called single-allelic data (SA), whereas data points that represent multiple potential alleles due to the nature of mass spectrometry experiments are called multi-allelic data (MA). Reynisson et al. selected negative peptides by randomly sampling from the UniProt database. The peptide lengths for the negatives were uniformly sampled from 13 to 21.

[0089] According to an embodiment of the present invention, the MIL problem is addressed using an instance-level approach. This approach maximizes model accuracy when predicting a single instance rather than the entire bag, compared to bag-level approaches. Implementation of the instance-level approach relies on correctly detecting key instances (positive instances in a positive bag). Therefore, a good instance-level model can be applied not only to MIL problems but also to single-instance learning problems. In fact, for peptide-MHC presentation problems, embodiments of the present invention use the same model jointly trained on single-instance and multi-instance data to maximize the use of existing data. Previous studies have shown that models that detect key instances also have good bag-level generalizability. However, while bag-level approaches may perform well at the bag level, they are not guaranteed to generalize well to the single-instance case. In biological applications, it is crucial that a model can correctly detect key instances.

[0090] In the following some further exemplary embodiments from several areas in which the present invention can be used are described.

[0091] Personalized Cancer Vaccine Design. This embodiment relates to a personalized cancer vaccine design system 500, shown generally in Figure 5, in which a model has been trained as described above. For prediction, a set of MHC molecules (generically labeled HLA, or human leukocyte antigen, typing 530 in Figure 5) from a biological sample 520 taken from a particular patient 510 is used, and a set of peptides 540 is based on mutations present in the cancer cells of the patient 510. The prediction is made using a trained MIL classifier f, shown in 550, for all combinations of peptide and MHC pairs for that patient, as previously described. θThe peptides that are most likely to be presented (i.e., have the highest scores as shown in FIG. 5) are then synthesized and included in a personalized cancer vaccine for that particular patient 510.

[0092] Immune Response Prediction. ELISpot is a widely used immune response assay that measures whether a specific peptide produces an immune response when combined with a biological sample, such as blood, from a patient infected with coronavirus. For example, interferon gamma is commonly measured using ELISpot. The immune response measurement from ELISpot is the result of the interaction between the peptide and at least one MHC molecule present in the sample. According to one embodiment, the MIL approach disclosed herein can also be used to train a model for predicting this immune response. The only difference compared to the above formulation is that the bag label is the result of an immune response assay. Such a model can also be used in a personalized cancer vaccine design system.

[0093] Histopathology-Based Cancer Diagnosis. Histopathology stains are created by taking tissue sections from biological samples and then staining them with chemicals such as hematoxylin and eosin. The stained images can then be used to identify features such as cell nuclei and extracellular support structures such as collagen. These stained images can also be used to train machine learning models to predict whether a particular tissue section contains cancer, i.e., to diagnose cancer. However, stained images are typically too large for current hardware to process in one go, so stained images are divided into "patches" for training. Typically, not all patches from a single stained image contain cancerous regions, even if other patches from that image contain cancerous regions.

[0094] This can also be thought of as a multi-instance learning problem, where a single stained image corresponds to each bag and the patches are the individual instances within the bag. The labels attached to the bags indicate whether cancer is present in that stained image. According to embodiments of the present invention, such predictive models can be used in cancer diagnosis systems.

[0095] Document Classification. The document classification task takes documents as input and classifies them into predefined categories. According to embodiments, MIL can be applied by considering paragraphs or sentences as instances and documents as bags. Exemplary labels can be document topics, such as "politics," "sports," or "science." Note that this example shows that the proposed approach according to embodiments of the present invention can be used to classify tasks with more than two classes with obvious modifications to the loss function. Furthermore, this example shows that the approach can be easily generalized to multi-label classification. For example, a document may relate to both "politics" and "sports." In this case, embodiments of the present invention can simply treat each label as a binary classification and replicate the loss function for each label.

[0096] In further embodiments, a system for predicting peptide binding and presentation by MHC molecules or a system for performing multi-instance learning comprises one or more processors configured to, alone or in combination, perform any of the methods according to embodiments of the present invention. In further embodiments, a tangible, non-transitory computer-readable medium comprises instructions that, when executed on one or more processors, cause the one or more processors to, alone or in combination, perform any of the methods according to embodiments of the present invention. The processors include one or more individual processors, each having one or more cores, and capable of accessing memory. Each of the separate processors can have the same or different architectures. The processors can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuits (e.g., application-specific integrated circuits (ASICs)), digital signal processors (DSPs), etc. The processors can be mounted on a common substrate or on multiple different substrates. A processor is configured to perform (e.g., configure to enable performance of) a particular function, method, or operation, at least when one of one or more separate processors is capable of performing the operations embodying that function, method, or operation. A processor may perform the operations embodying that function, method, or operation, for example, by executing code stored on memory (e.g., interpreting a script) and / or by passing data (traffic) through one or more ASICs. A processor may be configured to automatically perform any and all functions, methods, and operations disclosed herein. Thus, a processor may be configured to implement any (e.g., all) of the protocols, devices, mechanisms, systems, and methods described herein.For example, when this disclosure states that a method or device performs (or performs) task "X," such a statement should be understood to disclose that a processor is configured to perform task "X."

[0097] Each computer entity may include a memory. Memory may include volatile memory, nonvolatile memory, and any other medium capable of storing data. Volatile memory, nonvolatile memory, and any other type of memory may each include multiple different memory devices located in multiple separate locations, each with a different structure. Memory may include remotely hosted (e.g., cloud) storage. Examples of memory include non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any type of optical storage disk such as a DVD, magnetic storage, holographic storage, HDD, SSD, or any medium that can be used to store program code in the form of instructions or data structures. Any and all methods, functions, and operations described in this application may be embodied entirely in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable script) stored in memory.

[0098] Many modifications and other embodiments of the inventions described herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. It is to be understood, therefore, that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation. [Explanation of symbols]

[0099] 100 mass spectrometry data 110 Positive Peptides 120a MHC molecule 120b MHC molecule 500 Personalized Cancer Vaccine Design System 510 Specific patient 520 Biological samples 530 HLA, i.e. human leukocyte antigen, typing 540 Peptides

Claims

1. 1. A computer-implemented method for predicting peptide binding and presentation by MHC molecules, comprising: collecting or generating training data, the training data comprises a set of MHC molecules present in a biological sample and a set of observed peptide sequences presented by at least one of the MHC molecules present in the biological sample; It is unknown to which specific one of the MHC molecules the peptide sequence will bind; the training data is given in the form of a set of triples {si, Ai, yi}, where si is a peptide sequence, Ai = {a1,...,am} is a set of m MHC molecules associated with the biological sample, and yi is a binary label indicating whether si is found to be presented by any of the MHC molecules in Ai; the training data is organized as a plurality of bags, each bag having a set of training instances; labels are known for the plurality of bags but unknown for the training instances; Use the loss function to estimate the MIL classifier f θ training the sigma at an instance level, the MIL classifier f θ is trained to predict whether a particular peptide sequence is presented by any of the MHC molecules present in the biological sample; The parameters of the MIL classifier f θ are trained by a loss function L(θ) including a probability calibration function C to calibrate the model's confidence; The probability calibration function C is calculated at each training epoch k+1 by [Equation 1] Probability of [Equation 2] configured to predict [Equation 3] and x ij corresponds to a tuple (si, ai) where si is said peptide and ai is the j-th MHC molecule in A i ; Each training instance from the positively labeled bag is weighted by the probability calibration function C; The labels of multiple new instances are then passed to the MIL classifier f θ and / or predict the labels of multiple new bags by directly applying the MIL classifier f θ predicting by applying ∑ i = ∑ j ...

11. A computer-implemented method comprising:

2. The method of claim 1 , further comprising obtaining the training data from a mass spectrometry experiment.

3. In the post-training prediction stage, the MIL classifier f θ Peptide sequences i and one set of MHC molecules a i as an input; The MIL classifier f θ is applied to the input to obtain the peptide sequence s i predicting whether the protein is presented by any of those MHC molecules; 3. The method of claim 1 or 2, further comprising:

4. The MIL classifier f θ making predictions for all combinations of peptide sequences and MHC molecules present in said biological sample using determining the peptides most likely to be presented as candidates for synthesis and inclusion in a personalized cancer vaccine; 4. The method of claim 1, further comprising:

5. 10. A tangible, non-transitory computer-readable medium storing instructions that, when executed on one or more processors, enable the one or more processors to perform the method of any one of claims 1 to 4, either alone or in combination.

6. 1. A system for predicting peptide binding and presentation by MHC molecules, comprising one or more processors configured to enable the execution of methods, either alone or in combination, comprising: collecting or generating training data, the training data comprises a set of MHC molecules present in a biological sample and a set of observed peptide sequences presented by at least one of the MHC molecules present in the biological sample; It is unknown to which specific one of the MHC molecules the peptide sequence will bind; collecting or generating said training data in the form of a set of triples {si, Ai, yi}, where si is a peptide sequence, Ai = {a1,...,am} is a set of m MHC molecules associated with the biological sample, and yi is a binary label indicating whether si is found to be presented by any of the MHC molecules in Ai; organizing the training data into a plurality of bags, each bag having a set of training instances; an organization in which labels are known for the plurality of bags but unknown for the training instances; Use the loss function to estimate the MIL classifier f θ training the model at an instance level, the MIL classifier f θ is trained to predict whether a particular peptide sequence is presented by any of the MHC molecules present in the biological sample; The parameters of the MIL classifier f θ are trained by a loss function L(θ) including a probability calibration function C to calibrate the model's confidence; The probability calibration function C is calculated at each training epoch k+1 by [Equation 4] Probability of [Equation 5] configured to predict [Equation 6] and x ij corresponds to a tuple (si, ai) where si is said peptide and ai is the j-th MHC molecule in A i ; training, where each training instance from the positively labeled bag is weighted by the probability calibration function C; The labels of multiple new instances are then passed to the MIL classifier f θ and / or predict the labels of multiple new bags by directly applying the MIL classifier f θ predicting by applying ∑ i = ∑ j ... Including, the system.

Citation Information

Patent Citations

  • High-order semi-Restricted Boltzmann Machines and Deep Models for accurate peptide-MHC binding prediction

    US20150278441A1