Deep learning models for predicting the MHC class I or class II immunogenicity of tumor-specific neoantigens

By using deep learning models to predict the binding affinity and surface presentation probability of tumor-specific neoantigens to MHC molecules, the problem of insufficient prediction accuracy in existing methods is solved, enabling the design of highly efficient and personalized cancer vaccines.

JP7863114B2Active Publication Date: 2026-05-20AMAZON TECH INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
AMAZON TECH INC
Filing Date
2021-12-01
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing silico methods lack high positive predictive values ​​in predicting the MHC binding and surface presentation capabilities of tumor-specific neoantigens, resulting in poor performance in the design of personalized cancer vaccines.

Method used

A deep learning model was used to predict the binding affinity and surface presentation probability of tumor-specific neoantigens to MHC molecules by combining the amino acid sequence and HLA allele pseudo-sequence. The model performance was optimized by training the dataset, and efficient neoantigens were selected for personalized vaccine design.

Benefits of technology

This improves the accuracy of predicting tumor-specific neoantigens, ensuring that the selected neoantigens can effectively stimulate an immune response and enhance the efficacy of personalized cancer vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007863114000024
    Figure 0007863114000024
  • Figure 0007863114000025
    Figure 0007863114000025
  • Figure 0007863114000026
    Figure 0007863114000026
Patent Text Reader

Abstract

Disclosed herein is a method for predicting MHC class I or MHC class II immunogenicity of tumor-specific neoantigens by both predicting MHC class I or MHC class II binding affinity and predicting the likelihood that the tumor-specific neoantigens will be presented on the cell surface by MHC class I or class II proteins.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 139,074 filed on 19 January 2021, and all the contents of the said application are incorporated herein by reference. [Background technology]

[0002] Cancer is the leading cause of death globally, accounting for one in four deaths. (Siegel et al., CA: A Cancer Journal for Clinicians, 68:7-30(2018)). In 2018, 18.1 million people were newly diagnosed with cancer, and 9.6 million died from cancer-related causes. (Bray et al., CA: A Cancer Journal for Clinicians, 68(6):394-424). Numerous standard treatments exist in existing cancer therapy, including resection techniques (e.g., surgical procedures and radiation) and chemical techniques (e.g., chemotherapy). Unfortunately, these treatments often involve significant risks, toxic side effects, extremely high costs, and uncertainty regarding their effectiveness.

[0003] Cancer immunotherapy (e.g., cancer vaccines) is emerging as a promising cancer treatment. The goal of cancer immunotherapy is to utilize the immune system to selectively destroy cancer while maintaining healthy tissue intact. Conventional cancer vaccines typically target tumor-associated antigens. Tumor-associated antigens are normally present in normal tissues but are overexpressed in cancer. However, because these antigens are often present in normal tissues, immune tolerance may hinder the activation of the immune response. Several clinical trials targeting tumor-associated antigens have failed to demonstrate sustained and beneficial effects compared to standard treatment. Li et al., Ann Oncol., 28(Suppl 12): xii11-xii17(2017).

[0004] Neoantigens are attractive targets for cancer immunotherapy. Neoantigens are non-self proteins with individual specificity. Neoantigens originate from random somatic mutations in tumor cell genomes and are not expressed on the surface of normal cells. Because neoantigens are expressed only on tumor cells and do not induce central immune tolerance, cancer vaccines targeting cancer neoantigens have potential advantages such as reduced central immune tolerance and an improved safety profile.

[0005] The landscape of cancer mutations is complex, and tumor mutations are generally unique to each individual target. Most somatic mutations detected by sequencing do not lead to effective neoantigens. Only a small percentage of mutations in tumor DNA or tumor cells are transcribed and translated with sufficient precision to design a vaccine expected to be effective, and processed into tumor-specific neoantigens. Furthermore, not all neoantigens are immunogenic. In fact, the percentage of T cells that spontaneously recognize endogenous neoantigens is about 1% to 2%. See Paul et al., J. Immunol., 192, 5831-5839 (2013); Yewdell, Immunity, 25, 533-543 (2006). Of the approximately 1% of MHC-binding neoantigens mentioned above, only about 50% will be recognized by T cells, and only 30-40% will be spontaneously processed and capable of killing tumor cells. Ibid.

[0006] Current in silico methods primarily focus on modeling which neoantigen peptides bind to MHC-I or MHC-II molecules, or predicting which neoantigens are likely to be processed into short-chain peptides by tumor cells and presented by MHC class I / II molecules. Available tools lack the predictive accuracy to determine which of the presented peptides is immunogenic. As a result, existing methods have low positive predictive values. For example, in one study, three melanoma patients were each immunized with seven peptides that had an in vitro confirmed MHC binding affinity of less than 500 nM. Carreno et al., Science, 348, 803-808 (2015). Of the 21 peptides tested, only nine induced a T-cell response. Ibid. If a personalized vaccine containing neoantigen peptides is designed using methods with low positive predictive values, the therapeutic neoantigen administered to the patient is unlikely to have the ability to induce an immune response against the cancer in question.

[0007] Therefore, systematically identifying personalized neoantigens in cancer patients is a crucial requirement for the successful development of personalized cancer vaccines. Thus, efficiently and accurately predicting immunogenic neoantigen candidates with high positive predictive values ​​for personalized vaccines remains a challenge.

[0008] Summary of the Invention This disclosure relates to a novel method for predicting the MHC class I or MHC class II immunogenicity of tumor-specific neoantigens, comprising predicting both the binding affinity of the tumor-specific neoantigen and the likelihood that the tumor-specific neoantigen will be presented on the cell surface by an MHC class I or MHC class II protein. This method accurately identifies tumor-specific neoantigens with high accuracy that are likely to be processed by tumor cells into peptides that bind to the target MHC class I or MHC class II molecule, come into contact with T cell receptors, and ultimately become immunogenic. This method has high predictive accuracy for identifying neoantigens that will induce an immune response, which is important for developing effective personalized immunogenic compositions (e.g., cancer vaccines). This has been a limitation of existing methods.

[0009] Furthermore, the methods described herein outperform the gold standard predictors, the MHCflurry-1.4 binding affinity predictor and the MHCflurry-2.0 predictor. See the "Examples" section for further details. Each of these models is a distinct predictor. MHCflurry-1.4 is an allele-specific MHC class I binding predictor (O'Donnell et al., Cell System, 7:129-132 (2018)). The MHCflurry-2.0 predictor is a pan-allele predictor of peptides presented via MHC class I.

[0010] The inventors have further developed an immunogenicity assessment dataset to create a unique benchmark for directly evaluating the method's search capability, which directly assesses the method's ability to search for immunogenic peptides from a large pool of peptide candidates against a given MHC class I allele. This capability is essential for vaccine design based on personalized immunotherapy.

[0011] To more clearly explain the method for predicting the MHC class I immunogenicity of the above tumor-specific neoantigens, Figure 1 shows a schematic flowchart of this method.

[0012] The method for predicting the MHC class I or MHC class II immunogenicity of the above-mentioned tumor-specific neoantigen begins with obtaining the peptide sequence of the tumor-specific neoantigen and the corresponding facile region of the peptide sequence. The facile region may be the amino acid sequence immediately to the left of the tumor-specific neoantigen peptide or immediately to the right of the tumor-specific neoantigen peptide. For example, the facile region may be the amino acid sequence located at the C-terminus and / or N-terminus of the tumor-specific neoantigen peptide. Typically, the length of the facile region may be about 10 amino acids. For example, the length of the facile region immediately to the left of the tumor-specific neoantigen may be about 5 amino acids. For example, the length of the facile region immediately to the right of the tumor-specific neoantigen may be about 5 amino acids. Next, the peptide sequence and the facile region are encoded into a number vector. Each number vector contains the amino acid residues encoding the peptide of the tumor-specific neoantigen, the facile region, and the positions of the amino acid residues. An HLA allele pseudosequence representing the HLA allele is obtained. The length of the HLA pseudosequence may be at least about 20 to about 100 amino acids. Preferably, the length of the HLA pseudosequence is at least about 30 to 60 amino acids. The HLA allele pseudosequence is encoded into a corresponding vector. The HLA allele is type A, type B, or type C, DQ, DP, or DR.

[0013] Next, a neural network model is used to predict both the MHC class I or MHC class II binding affinity of tumor-specific neoantigens and the numerical probability that each target peptide will be presented on the cell surface by MHC class I or MHC. The neural network model may be a pan-allele model, an allele-specific model, a super-type-specific model, or a combination thereof.

[0014] First, the neural network model is trained on the training dataset to optimize its performance. The training dataset includes a peptide-MHC class I or MHC class II affinity measurement dataset and a cell surface peptide presentation dataset. Preferably, the neural network model is trained on positive training data and negative training data. The negative training data may include peptides that do not have MHC class I or class II binding affinity to tumor-specific neoantigens and / or are not presented on the cell surface by MHC class I or class II proteins.

[0015] The model input layer includes a number vector containing the peptide sequence and adjacent region of the tumor-specific neoantigen, as well as a number vector containing the HLA allele pseudosequence. Next, each of the number vectors is encoded in the amino acid embedding layer. Then, the neural network model flattens the amino acid embedding layer to create number vector representations of each peptide sequence of the tumor-specific neoantigen, the adjacent region of the peptide sequence, and the HLA allele pseudosequence.

[0016] To predict the MHC class I or MHC class II binding affinity of the above-mentioned peptide tumor-specific neoantigen, the peptide sequence of the tumor-specific neoantigen and the HLA allele pseudosequence are concatenated. The model further comprises applying one or more layers and / or one or more activation functions. For example, the model may include applying one or more binding layers. For example, the model may include applying a dropout layer. For example, the model may include applying an activation function. In some cases, the model may include applying one or more binding layers, one or more dropout layers, and / or an activation function. The output is a numerical score representing the peptide ligand-MHC class I or MHC class II binding affinity. To predict the probability that the above-mentioned tumor-specific neoantigen will be presented on the cell surface by an MHC class I or MHC class II protein, the target peptide sequence, the adjacent region of the peptide sequence, and the HLA allele pseudosequence are concatenated in a single number vector. The predicted peptide ligand-MHC class I or MHC class II binding affinity is also concatenated. The model further comprises applying one or more layers and / or one or more activation functions. For example, the above model may include the application of one or more binding layers. For example, the above model may include the application of a dropout layer. For example, the above model may include the application of an activation function. The above model may include the application of one or more binding layers, the application of a dropout layer, and / or the application of an activation function. The output is the probability, expressed as a number, that the peptide will be presented on the cell surface by an MHC class I or MHC class II protein. The MHC class I or MHC class II binding affinity of the above tumor-specific neoantigen, and the probability, expressed as a number, that the above tumor-specific neoantigen will be presented on the cell surface by an MHC class I or MHC class II protein are surrogates for the MHC class I immunogenicity of the tumor-specific neoantigen. Generally, MHC class I immunogenicity is CD8+ T cell immunogenicity.MHC class II immunogenicity is CD4+ T cell immunogenicity.

[0017] This method may further include validating the neural network by applying one or more ranking metrics to an immunogenicity validation dataset, ranking peptides for each allele in the immunogenicity validation dataset based on the predicted MHC class I binding affinity of the peptides and a probability represented by a numerical value indicating that the peptides will be presented on the cell surface by MHC class I proteins, and aggregating the ranking metrics for all alleles. The ranking metrics may be aggregated by using weighted allele frequencies. This method may further include calibrating the neural network. The calibration of the neural network model may be performed by probability calculation. By this calculation, the overall presentation probability of the target allele can be estimated.

[0018] Tumor-specific neoantigens predicted to be MHC class I or MHC class II immunogenic may be selected for use in an immunogenic composition. Usually, about 10 to about 20 tumor-specific neoantigens may be selected for use in an immunogenic composition.

Brief Description of the Drawings

[0019] [Figure 1] It is a model architecture diagram. Model input: a) peptide sequence (and adjacent regions); b) allele pseudo-sequence. Model output: a) predicted binding affinity; b) predicted presentation probability. Loss functions used for model training: a) MSE-with-inequalities loss for the prediction of binding affinity; b) binary focal loss for the prediction of presentation probability. [Figure 2]A is a graph showing the proportion of peptides that do not have partner peptides with similarity exceeding a threshold. B is a graph showing the distribution of peptide lengths across three datasets (affinity, presentation, and immunogenicity). [Figure 3] It is a graph showing the distribution of peptide-MHC binding affinity labels. The Y-axis represents the proportion of samples in the training data belonging to each bin. [Figure 4] It is a graph showing the distribution of HLA allele supertypes in dataset samples. [Figure 5A] It is a graph showing the sample distribution for each HLA allele, indicating the sample distribution of peptide MHC binding affinity. [Figure 5B] It is a graph showing the sample distribution for each HLA allele, indicating the sample distribution of peptide presentation on the cell surface. [Figure 5C] It is a graph showing the sample distribution for each HLA allele, indicating the sample distribution of T cell immunogenicity. [Figure 6] It is a graph showing the immunogenic allele frequency weights based on allele frequencies in the US population. [Figure 7] It is a graph showing the comparison of performance with different values of alpha (which determines the weight of each loss element). [Figure 8] A is a graph showing the correlation between the proportion with immunogenicity and the predicted probability of presentation. B is a graph showing the correlation between the proportion with immunogenicity and the predicted binding affinity. This experiment was conducted on an immunogenicity validation set where the ground truth immunogenicity labels were known and presentation and binding affinity were predicted using a pan-model. [Figure 9] It is a diagram showing an example of a provider network environment. [Figure 10] It is a block diagram of an exemplary provider network that provides storage services and hardware virtualization to customers according to some embodiments. [Figure 11] It is a block diagram showing an exemplary computer system. [Figure 12] This figure shows the model input. The model receives two sequences: a token sequence and a corresponding segment sequence. The above token sequence is: <cls>Tokens, allele pseudo-sequence tokens,<SE[> Tokens, n-neighbor tokens, peptides, c-neighbor tokens, and <eos>It is constructed by concatenating tokens. The above segment sequence provides an index indicating the segment to which the corresponding token belongs. [Figure 13] This is a schematic diagram of a Transformer layer consisting of a multi-head self-attention module followed by a feedforward module (composed of two linear layers with GELU activation functions sandwiched in between). Layer normalization is applied at the beginning of each module, and residual dropout is applied at the end of each component before the residual connection. [Modes for carrying out the invention]

[0020] This disclosure relates to a novel method for predicting the MHC class I or MHC class II immunogenicity of tumor-specific neoantigens, comprising predicting either the MHC class I or MHC class II binding affinity and predicting the likelihood that the tumor-specific neoantigens will be presented on the cell surface by MHC class I or MHC class II proteins. The novel method is preferably used to predict the MHC class I or MHC class II immunogenicity of tumor-specific neoantigens.

[0021] This method involves obtaining sequencing data for tumor-specific neoantigens and HLA allele pseudosequences representing HLA alleles. For example, sequencing data and peptide sequences for tumor-specific neoantigens can be obtained using exome, transcriptome, and / or whole-genome nucleotide sequencing. This method may further include encoding the peptide sequence and optionally adjacent regions of each tumor-specific neoantigen into corresponding numerical vectors. Each numerical vector contains information describing the amino acid residues that make up the peptide sequence and the positions of the amino acid residues.

[0022] The method may also include encoding the above HLA pseudosequences into numerical vectors. The method may also include inputting the above numerical vectors into a neural network model to predict both the MHC class I or MHC class II binding affinity of the tumor-specific neoantigens and the numerical probability that the corresponding peptide of each tumor-specific neoantigen will be presented on the cell surface by MHC class I or MHC class II proteins. Both of these predictions can be used as surrogate values ​​to predict MHC class I or MHC class II immunogenicity (e.g., CD8+ T cell immunogenicity or CD4+ T cell immunogenicity). After the above numerical vectors are input into the neural network model, the numerical vectors may be converted into an amino acid embedding layer, which may then be flattened to create numerical vector representations of the peptide sequences of each tumor-specific neoantigen, optionally the peptide-adjacent regions, and the HLA allele pseudosequences.

[0023] Next, the neural network model described above can be used to predict the MHC class I or MHC class II binding affinity of the tumor-specific neoantigen and the numerical probability that the tumor-specific neoantigen will be presented on the cell surface by an MHC class I or MHC class II protein. These predictions can be made by concatenating the peptide sequence of the tumor-specific neoantigen, the HLA allele pseudosequence, and optionally the peptide facies region. Then, one or more layers and / or functions may be applied. For example, one or more fully bound dense layers may be applied. For example, one or more dropout layers may be applied. For example, one or more activation functions may be applied. In embodiments, a combination of one or more fully bound dense layers, one or more dropout layers, and / or one or more activation functions may be applied. The output is a numerical score representing the MHC class I or MHC class II binding affinity of the tumor-specific neoantigen and / or the numerical probability that the peptide will be presented on the cell surface by an MHC class I or MHC class II protein. These predicted values ​​can be used as surrogates for immunogenicity. Next, immunogenic tumor-specific neoantigens can be selected for inclusion in personalized immunogenic compositions.

[0024] The predictions disclosed herein are identified based on a training dataset. The training dataset includes multiple samples. The training dataset may include a dataset of peptide MHC class I or MHC class II affinity measurements and a dataset of peptide presentation on the cell surface. The neural network model is preferably trained not only on positive training data but also on negative training data. The negative training data may include peptides that do not have MHC class I or MHC class II binding affinity to tumor-specific neoantigens and / or are not presented on the cell surface by MHC class I or MHC class II proteins.

[0025] Neoantigens predicted to be MHC class I or MHC class II immunogenic can be selected for inclusion in the immunogenic composition. Typically, about 10 to about 20 tumor-specific neoantigens can be selected for the immunogenic composition. For example, the immunogenic composition may contain about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 tumor-specific neoantigens.

[0026] I. Definition All publications and patents cited herein are incorporated in their entirety. To the extent that any incorporated material is inconsistent or contradictory to this specification, this specification takes precedence over any such material. No reference to any reference herein constitutes prior art to this disclosure. Various terms relating to aspects of this specification are used throughout this specification and the claims. Unless otherwise indicated, such terms should be given their common meaning in the art. Other terms specifically defined should be construed in a manner consistent with the definitions set forth herein.

[0027] In this specification, the singular forms "a," "an," and "the" include the plural form unless otherwise clearly indicated by the context. Terms such as "includes" and "etc." mean unrestricted inclusion unless otherwise indicated.

[0028] In this specification, the term “cancer” refers to a physiological condition in an object characterized by uncontrolled proliferation, immortality, metastatic ability, rapid growth and proliferation rates, and / or specific morphological features. Cancer often takes the form of a tumor or mass, but may also exist alone within the object's body or circulate in the bloodstream as independent cells, such as leukemia or lymphoma cells. The term cancer encompasses all types of cancer and metastases, including hematological malignancies, solid tumors, sarcomas, carcinomas, and other solid and non-solid tumors. Examples of cancer include, but are not limited to, carcinomas, lymphomas, blastomas, sarcomas, and leukemias. More specific examples of such cancers include squamous cell carcinoma, small cell lung cancer, non-small cell lung cancer, adenocarcinoma of the lung, squamous cell carcinoma of the lung, peritoneal cancer, hepatocellular carcinoma, gastrointestinal cancer, pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatocellular carcinoma, breast cancer (e.g., triple-negative breast cancer, hormone receptor-positive breast cancer), osteosarcoma, melanoma, colon cancer, colorectal cancer, endometrial (e.g., serous) or uterine cancer, salivary gland cancer, kidney cancer, liver cancer, prostate cancer, vulvar cancer, thyroid cancer, liver cancer, and various types of head and neck cancers. Triple-negative breast cancer refers to breast cancer in which the expression of the estrogen receptor (ER), progesterone receptor (PR), and Her2 / neu genes is negative. Hormone receptor-positive breast cancer refers to breast cancer in which at least one of ER and PR is positive and Her2 / neu (HER2) is negative.

[0029] In this specification, the term “neoantigen” refers to an antigen that has at least one change that distinguishes it from its corresponding parent antigen, for example, through mutation in tumor cells or tumor cell-specific post-translational modifications. Mutations can include frameshifts, indels, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic expression changes that give rise to neoantigens. Splice mutations can be considered mutations. Abnormal phosphorylation can be considered a tumor cell-specific post-translational modification. Splice antigens produced by the proteasome can also be considered tumor cell-specific post-translational modifications. See Lipe et al., Science, 354(6310):354:358 (2016). Generally, point mutations account for about 95% of mutations in tumors, with indel and frameshift mutations making up the remainder. See Snyder et al., N Engl J Med., 371:2189-2199 (2014).

[0030] In this specification, the term "tumor-specific neoantigen" refers to a neoantigen that is present in the target tumor cells or tissues but not in the target normal cells or tissues.

[0031] In this specification, the term “immunogenicity” means the ability to induce an immune response (e.g., a T-cell response, a B-cell response, or both).

[0032] In this specification, the term "HLA allele pseudosequence" refers to an amino acid sequence generated by an algorithm for representing the HLA allele amino acid sequence.

[0033] In this specification, the term "subject" means any animal, including but not limited to any mammal, such as humans, non-human primates, and rodents. In some embodiments, the mammal is a mouse. In some embodiments, the mammal is a human.

[0034] In this specification, the term “tumor cell” means any cell that is a cancer cell or derived from a cancer cell. The term “tumor cell” may also mean a cell that exhibits cancer-like characteristics, such as uncontrolled growth, resistance to antiproliferative signals, metastatic ability, and loss of the ability to undergo programmed cell death.

[0035] In this specification, the term “neural network” means a machine learning model for classification or regression, typically consisting of multiple layers of linear transformations followed by element-wise nonlinearities, and usually trained by stochastic gradient descent and backpropagation.

[0036] Any term not directly defined herein should be understood to have a general meaning associated with it as understood within the art of the present invention. Any term not directly defined herein should be understood to have multiple general meanings associated with it as understood within the art of the present invention. In this specification, certain terms are discussed to provide physicians with further guidance in descriptions of compositions, devices, methods, etc., of aspects of the present invention, and methods of manufacturing or using them. It will be understood that the same thing may be expressed in multiple ways. Therefore, for any one or more terms discussed herein, alternative words and synonyms may be used. Emphasis should not be placed on whether a term is detailed or discussed herein. Several synonyms or alternative methods, materials, etc., are provided. Where one or more synonyms or equivalents are listed, the use of other synonyms or equivalents is not excluded unless expressly stated otherwise. Examples, including examples of terms, are used for illustrative purposes only and do not limit the scope and meaning of aspects of the present invention herein.

[0037] Further explanation of the method and guidance for carrying out the method are provided herein. For the sake of clarity, further details and guidance are provided regarding a preferred embodiment of the prediction of MHC class I immunogenicity, which involves both predicting MHC class I binding affinity and predicting the likelihood that the tumor-specific neoantigen will be presented on the cell surface by MHC class I proteins. Further details and guidance are also intended to be relevant to the prediction of MHC class II immunogenicity.

[0038] II. Training The above-described neural network model may include training the neural network model on a training dataset to optimize its performance so that the neural network can predict the binding affinity of tumor-specific neoantigens to MHC class I or MHC class II proteins, and the probability that the tumor-specific neoantigens will be presented on the cell surface by MHC class I or MHC class II proteins.

[0039] The training dataset used in the method described herein includes multiple samples. The training data may include a variety of data. The training data may include MHC class I or MHC class II affinity measurement datasets for peptides, peptide presentation datasets for cell surfaces, and optionally immunogenicity datasets. The training data may include human rhinovirus data. Negative samples can be used for immunogenicity assessment. One or more neural network models can be trained using the training dataset. In embodiments, one or more neural network models can be trained. For example, at least two or more neural networks can be trained. At least about three, four, five, six, seven, eight, nine, or ten or more neural networks can be trained.

[0040] A dataset containing peptide MHC class I or MHC class II affinity measurement data may include experimentally measured binding affinity peptides to a specific MHC class I or MHC class II allele. The dataset can be obtained from one or more data sources, such as publicly available data sources. For example, the dataset can be obtained from the Immune Epitope Database ("IEDB," iedb.org). The training dataset can be further expanded based on one or more data sources. The training dataset may be further refined for the methods disclosed herein. For example, the training data may include predictions of binding affinity between the peptide and each associated MHC molecule. The experimentally measured binding affinity in the dataset may be quantitative (with an inequality of "="), qualitative (with an inequality of "<" or ">"), or a combination thereof. Quantitative data may include IC50 (mM) values. Qualitative datasets may be expressed as positive-high (e.g., binding affinity <100 nm), positive-medium (e.g., binding affinity <1,000 nm), positive-low (e.g., binding affinity <5,000 nm), or negative (e.g., binding affinity >5,000 nm). Training datasets, including MHC class I or MHC class II affinity measurement datasets, may contain peptides eluted from MHC and identified by mass spectrometry.

[0041] The MHC class I affinity measurement dataset may be further refined to retain a subset of predicted binding affinities for specific MHC class I peptide alleles. For example, entries for HLA-A, HLA-B, and / or HLA-C alleles may be retained. For example, peptides of a specific length may be retained. Peptides with a length of at least about 5 to about 20 amino acids may be retained. The length of the above peptides may be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids. The lengths of the peptides in the above training set may be the same or different, and the above lengths may vary depending on the type of MHC allele. The length of the above peptides is preferably about 5 to about 15 amino acids. Peptides containing post-translational modifications or non-standard amino acids may be removed.

[0042] The MHC class II affinity measurement dataset may be further refined to retain a subset of predicted binding affinities for specific MHC class II peptide alleles. For example, entries for HLA-DP, HLA-DQ, and / or HLA-DR alleles may be retained. For example, peptides of a specific length may be retained. Peptides with a length of at least about 5 amino acids to about 40 amino acids may be retained. The lengths of the above peptides may be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 amino acids. The lengths of the peptides in the above training set may be the same or different, and the above lengths may vary depending on the type of MHC allele. Typically, the lengths of the above peptides are about 13 to about 35 amino acids. Peptides containing post-translational modifications or non-standard amino acids may be removed.

[0043] The above MHC class I or MHC class II affinity measurement data may be provided as a regression model. In particular, a loss function may be used. Exemplary loss functions include the cross-entropy loss function, mean squared error, Huber loss, Kullback-Leibler, MAE(L1), MAE(L3), likelihood function, and hinge loss. In particular, a variation of the mean squared loss function may be used. The mean squared loss function is L BA-MSE It can be expressed as follows, where the measured value is accompanied by (>) or (<), and with respect to the processing of both quantitative and qualitative peptide-MHC binding affinity measurements in the dataset, the above measurement contributes to the loss only if it violates the inequality. The formula that can be used is,

number

[0044]

number

number

[0045] The above MHC class I or MHC class II affinity measurement datasets contain at least approximately 5,000 species, approximately 10,000 species, approximately 15,000 species, approximately 20,000 species, approximately 25,000 species, approximately 30,000 species, approximately 35,000 species, approximately 40,000 species, approximately 45,000 species, approximately 50,000 species, approximately 60,000 species, approximately 70,000 species, approximately 80,000 species, approximately 90,000 species, approximately 100,000 species, approximately 150,000 species, approximately 200,000 species, approximately 250,000 species, approximately 300,000 species, approximately 350,000 species, approximately 400,000 species, and approximately The dataset may include measurements of the binding affinity of 450,000, approximately 500,000, approximately 550,000, approximately 600,000, approximately 650,000, approximately 700,000, approximately 750,000, approximately 800,000, approximately 850,000, approximately 900,000, approximately 950,000, approximately 1,000,000, approximately 1,250,000, approximately 1,500,000, approximately 1,750,000, approximately 2,000,000, or more peptides to MHC class I or MHC class II peptide alleles. Generally, the above MHC class I or MHC class II affinity measurement datasets contain at least approximately 20,000 unique peptides.

[0046] The above-described peptide presentation dataset on the cell surface may include peptides known to be presented via HLA molecules. Cell surface peptides can be measured, for example, by peptide elution experiments or mass spectrometry data. The above-described peptide presentation dataset on the cell surface can be obtained from one or more data sources, such as publicly available data sources. For example, the Immune Epitope Database ("IEDB," iedb.org) or peptides produced by the SysteMHC project can be useful data sources. The above-described peptide presentation dataset on the cell surface may be further constructed experimentally. For example, peptides can be prepared by diluting peptides derived from cell lines expressing HLA peptides and analyzing the peptides by mass spectrometry. The above-described training dataset can be further expanded based on one or more data sources.

[0047] These training datasets are typically selected for the methods disclosed herein. Peptide sequences are generally represented as strings, where each letter represents an amino acid. A peptide sequence may be converted into a numerical vector containing information describing the amino acids of the peptide and their positions. The numerical vector may be a binary classification. For example, k i Peptide sequence p having individual amino acids i It is represented by a row vector (20-k) of 20 amino acids, where in the row vector, a single element corresponding to the alphabet of an amino acid at a specific position in the peptide sequence will have a value of 1. The remaining elements will have a value of 0. For example, if the amino acid alphabet is A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, and Y, the peptide sequence AFP of 3 amino acids will have a row vector of 60 elements and

number

[0048] A loss function may be used for the peptide presentation dataset on the cell surface described above. Examples of loss functions include cross-entropy loss function, mean squared error, Huber loss, Kullback-Leibler, MAE(L1), MAE(L3), likelihood function, and hinge loss. In particular, the cross-entropy loss function may be used for the peptide presentation dataset on the cell surface described above. In certain embodiments, Focal Loss binary classification may be used. Focal Loss binary classification can be used to reduce imbalances in the dataset described above. Focal Loss is L P-FL It can be expressed as L P-FL This is a weighted extension of the standard binary cross-entropy loss, which further emphasizes samples that are poorly classified. The formula for Focal Loss is as follows:

number

number

number

[0049] A ranking objective may be further used with the above-mentioned cell surface peptide presentation dataset for training. For example, N-way classification may be used for ranking-oriented training. N-way classification makes it possible to make positive and negative samples in the dataset compete. The samples can then be classified as samples where each of the N sets of samples is a positive sample (N = number of negative samples + 1). For N-way classification loss, a cross-entropy loss function or a Focal Loss function can be applied.

[0050] The above-mentioned peptide presentation datasets for the cell surface contain at least approximately 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 150,000, 200,000, 250,000, 300,000, and 350, The dataset may contain 000, approximately 400,000, approximately 450,000, approximately 500,000, approximately 550,000, approximately 600,000, approximately 650,000, approximately 700,000, approximately 750,000, approximately 800,000, approximately 850,000, approximately 900,000, approximately 950,000, approximately 1,000,000, approximately 1,250,000, approximately 1,500,000, approximately 1,750,000, approximately 2,000,000, or more peptides. A training dataset of more than 35,000 samples is preferred.

[0051] The above neural network model can be trained on all training data or on a portion of the training data. For example, the above neural network model can be trained on approximately 100% of the training data, approximately 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or less of the training dataset. The above neural network model can be trained on all of the above MHC class I or MHC class II affinity measurement sets and all of the above peptide presentation to cell surface training datasets. For example, the above neural network model can be trained on the above MHC class I or MHC class II affinity measurement datasets at approximately 100%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or lower, and / or the above peptide presentation training datasets to the cell surface at approximately 100%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or lower.

[0052] In one embodiment, cross-training may be performed on the training data of one or more training datasets. For example, cross-training may be performed on the MHC class I or MHC class II affinity measurement dataset and the peptide presentation dataset on the cell surface. Each dataset typically contains a single known target. For example, the MHC class I or MHC class II affinity measurement dataset contains peptide affinity, and the peptide presentation dataset on the cell surface contains peptides that may be presented on the cell surface by MHC class I or MHC class II proteins. To cross-train the training datasets, the target of each training set may be inferred. For example, peptides that present on the cell surface may be inferred to have high binding affinity values, and peptides that are not presented on the cell surface may be inferred to have low binding affinity. For example, peptides with high binding affinity may be inferred to present on the cell surface, and peptides with low binding affinity may be inferred not to present on the cell surface.

[0053] In one embodiment, autodistillation may be performed on the training data of one or more training datasets. Autodistillation can be performed by extracting estimates of binding affinity and presentation for multiple samples. These samples can be added to the weak label corresponding to the training dataset. Autodistillation can be performed using multi-alliative spectroscopy data. Autodistillation can be performed using positively presenting cells. For positively presenting cells in a training dataset that include unknown binding affinities, binding affinities can be estimated using an established model.

[0054] Preferably, the neural network model described above is trained on both positive and negative training data to limit bias. Training the network on an imbalanced dataset may introduce bias into the neural network model by learning more representations of the dominant class of data, potentially overlooking other classes. For example, a neural network model trained only on the positive training dataset may be biased to overestimate the MHC class I or MHC class II binding affinity of peptide tumor-specific neoantigens, or to overestimate the probability that tumor-specific neoantigens will be presented on the cell surface by MHC class I or MHC class II proteins. A neural network model trained only on the negative training set may be biased to underestimate the MHC class I or MHC class II binding affinity of peptide tumor-specific neoantigens, or to underestimate the probability that tumor-specific neoantigens will be presented on the cell surface by MHC class I or MHC class II proteins.

[0055] The above MHC class I or MHC class II affinity measurement datasets typically include both positive and negative training data. For example, positive training data might include binding affinity predictions classified as positive (e.g., binding affinity <5,000 nm). For example, negative training data might include binding affinity predictions that are negative (e.g., binding affinity >5,000 nm). Additional negative training data can be incorporated into the training set by extending the training dataset to include random peptides with low affinity, if necessary. For example, the random peptides may have a qualitative weak affinity target of approximately >20,000 nm.

[0056] The above-mentioned cell surface peptide presentation training dataset typically includes positive training data (e.g., peptides presented on the cell surface by MHC class I proteins) and does not include negative training data (e.g., peptides that cannot be presented on the cell surface by MHC class I proteins). If the above training dataset does not include negative training data, the positive training dataset may be used to generate a probabilistic negative training dataset (e.g., a negative training dataset derived from the positive training dataset). The negative training dataset can be generated by shuffling peptides that are "positive" for HLA alleles. The above peptides can be shuffled by changing their amino acid length (e.g., making the peptides longer or shorter). Alternatively, the amino acid sequence of a peptide can be modified, for example, by amino acid substitution, insertion, or deletion. Insertions include fusion of the amino terminus and / or carboxyl terminus, as well as intrasequence insertions or multiple amino acid residues. Deletions are characterized by the removal of one or more amino acid residues from the peptide sequence. Amino acid substitutions are usually single-residue substitutions, but may occur at multiple positions. Substitutions, deletions, insertions, or any combination thereof can be used to reach peptides that are not presented on the cell surface by MHC class I or MHC class II proteins. For example, a peptide sequence having the amino acid sequence AVGGGERRYIKL may be modified to CVGGGEHRYIMNNL.

[0057] In addition, HLA shuffling can be used, or combined with peptide shuffling, to generate negative training datasets. HLA alleles classified as "positive" (e.g., HLA alleles that present the corresponding peptide on the cell surface) may be replaced with different alleles that do not belong to the positive allele supertype.

[0058] In addition, HRV-negative sampling can be used, or combined with peptide shuffling and / or HLA shuffling, to generate negative training datasets.

[0059] The above training data may be further filtered to remove unnecessary peptides. For example, duplicate peptides (e.g., identical amino acid sequences) may be removed so that the training dataset contains unique peptides. Those skilled in the art will readily understand how to determine peptide identity (i.e., how to determine whether peptides are identical or different).

[0060] A trained neural network model may be validated using an immunogenicity dataset. Validating the neural network model may involve applying one or more ranking metrics to the immunogenicity dataset. Peptides in the immunogenicity validation dataset can be ranked based on their predicted MHC class I or MHC class II binding affinity and the numerical probability that the peptide will be presented on the cell surface by an MHC class I or MHC class II protein. The ranking metrics may aggregate all alleles. The ranking metrics may be aggregated using weighted allele frequencies.

[0061] In one embodiment, the neural network model may be trained using an unlabeled dataset. For example, the peptides in a peptide presentation dataset on the cell surface may be unlabeled. While not theoretically constrained, it is thought that an unlabeled dataset (e.g., peptide sequences) can provide a more accurate vector representation of the input peptide sequences.

[0062] III. Model Architecture This disclosure relates to the use of a neural network model to predict both the MHC class I or MHC class II binding affinity of tumor-specific neoantigens and the numerical probability, expressed for each target peptide, that the corresponding peptide will be presented on the cell surface (i.e., the surface of tumor cells) by an MHC class I or MHC class II protein. The neural network model is suitable for tumor-specific neoantigens that the neural network model has encountered or has not encountered in training.

[0063] The neural network model described above may be a single neural network comprising a set of nodes arranged in one or more layers. The nodes may be connected to other nodes via connections, each having associated parameters. The value at a particular node may be expressed as the sum of the values ​​of the nodes connected to that particular node, weighted by associated parameters mapped by the activation function associated with that particular node. The neural network model used in the method described herein may be a pan-allele model, an allele-specific model, a supertype-specific model, or a combination thereof.

[0064] In a particular embodiment, the method includes converting a peptide sequence of a tumor-specific neoantigen into a numerical vector. Typically, the peptide sequence is represented as a string (each letter representing an amino acid). The peptide sequence can be converted into a numerical vector containing information describing the amino acids of the peptide and the positions of those amino acids. The numerical vector may be a binary classification. For example, k i Peptide sequence p having individual amino acids i It is represented by a row vector (20-k) of 20 amino acids, where in the row vector, a single element corresponding to the alphabet of an amino acid at a specific position in the peptide sequence will have a value of 1. The remaining elements will have a value of 0. For example, if the amino acid alphabet is A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, and Y, then the peptide sequence AGQY of 4 amino acids will have a row vector of 80 elements and

number

[0065] The peptide sequence length of the tumor-specific neoantigen may be approximately 5 to 40 amino acids. For example, the peptide sequence length may be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 amino acids. MHC class I molecules bind to short-chain peptides. MHC class I molecules can generally accept peptides with a length of approximately 5 to 10 amino acids. In embodiments, the peptide sequence of the tumor-specific neoantigen is a short-chain peptide with a length of approximately 5 to 10 amino acids. MHC class II molecules bind to longer peptides. MHC class II molecules can generally accept peptides with a length of approximately 13 to 25 amino acids. In this embodiment, the peptide sequence of the tumor-specific neoantigen is a long peptide with a length of approximately 13 to 25 amino acids.

[0066] The peptide sequences of tumor-specific neoantigens may be identical or different in length. If the peptide sequences of tumor-specific neoantigens are of different lengths (for example, one peptide sequence is 7 amino acids long and the other is 15 amino acids long), padding characters may be added to the vector until each tumor-specific neoantigen peptide reaches the maximum peptide length including adjacent regions (for example, 15 amino acids). The padding characters can be added to the C-terminus or N-terminus of adjacent regions. For example, the padding characters may encode the amino acids A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, and Y.

[0067] The above-mentioned tumor-specific neoantigen peptide sequence may include sequences adjacent to the tumor-specific neoantigen peptide. These adjacent sequences may be immediately to the left of the tumor-specific neoantigen peptide sequence, immediately to the right of the tumor-specific neoantigen peptide sequence, or both.

[0068] In one embodiment, the peptide sequence of the tumor-specific neoantigen may include at least one C-terminal sequence adjacent to the tumor-specific neoantigen peptide within its source protein sequence, or at least one N-terminal sequence adjacent to the tumor-specific neoantigen peptide within its source protein sequence. Preferably, the peptide sequence of the tumor-specific neoantigen includes at least one C-terminal amino acid sequence adjacent to the tumor-specific neoantigen peptide and at least one N-terminal amino acid sequence adjacent to the tumor-specific neoantigen peptide.

[0069] The length of the adjacent region described above may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 amino acids or more. The length of the adjacent region immediately to the left of the tumor-specific neoantigen peptide described above may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 amino acids or more. The length of the adjacent region immediately to the right of the tumor-specific neoantigen peptide may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids or more. In one embodiment, the peptide sequence of the tumor-specific neoantigen includes an adjacent region immediately to the left of the tumor-specific neoantigen, which may have a length of up to approximately 10 amino acids, and / or an adjacent region immediately to the right of the tumor-specific neoantigen, which may have a length of up to approximately 10 amino acids. Preferably, the adjacent region includes an adjacent region immediately to the left of the tumor-specific neoantigen with a length of 5 amino acids, and an adjacent region immediately to the right of the tumor-specific neoantigen with a length of 5 amino acids. The adjacent region may also be encoded in the above number vector.

[0070] This method further includes converting HLA allele pseudosequences into numerical vectors. The above HLA allele pseudosequences represent HLA alleles. The length of the above HLA allele pseudosequences may be approximately 5 to 100 amino acids. For example, the lengths of the above HLA allele pseudosequences are 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 5 The length of the above HLA allele pseudosequence may be 2, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 amino acids. The length of the above HLA allele pseudosequence may be about 30 to about 60 amino acids. In certain embodiments of the methods disclosed herein, the length of the above HLA allele pseudosequence is about 40 to about 50 amino acids.

[0071] The input to the above neural network model may include (i) a vector containing the peptide sequence of the tumor-specific neoantigen and the adjacent region of the peptide sequence, and (ii) a vector containing the HLA allele pseudosequence. The input to the above neural network may optionally include a segment identifier sequence. The segment identifier sequence informs the model which segment each amino acid belongs to. Next, the vector containing the peptide sequence of the tumor-specific neoantigen and the adjacent region of the peptide sequence, and (ii) the vector containing the HLA allele pseudosequence, are encoded into one or more embedding layers. The embedding layers translate the high-dimensional vector containing the peptide sequence of the tumor-specific neoantigen and the adjacent region of the peptide sequence, and the vector containing the HLA allele pseudosequence, into a low-dimensional space. The embedding layers can be considered as the first layer of the above neural network model. The embedding layers may then be flattened to generate vector representations of each peptide sequence of the tumor-specific neoantigen, the peptide adjacent region, and the HLA allele pseudosequence.

[0072] To predict the MHC class I or MHC class II binding affinity of the above-mentioned peptide tumor-specific neoantigen, the tumor-specific peptide sequence and the HLA allele pseudosequence are ligated together. This means that the peptide sequence of the tumor-specific neoantigen and the HLA allele pseudosequence are linked to each other in a chain-like or continuous manner. To predict the MHC class I or MHC class II binding affinity, the adjacent regions do not need to be ligated. Although not necessary, ligating the adjacent regions may sometimes be desirable. Once the peptide sequence of the tumor-specific neoantigen and the HLA pseudosequence are ligated together, it becomes possible to apply one or more parameters (e.g., layers and / or functions).

[0073] Examples of applicable layers include, but are not limited to, fully connected dense layers, sequencing layers, activation layers, normalization layers, dropout layers, cropping layers, pooling layers and unpooling layers, combination layers, object detection layers, or generative adversarial network layers.

[0074] Examples of fully connected dense layers include 2D convolutional layers, 3D convolutional layers, 2D grouped convolutional layers, transposed 2D convolutional layers, transposed 3D convolutional layers, or fully connected dense layers. Examples of sequence layers include sequence input layers, LSTM layers, bidirectional LSTM layers, GRU layers, sequence folding layers, sequence unfolding layers, flattening layers, or word embedding layers. Examples of activation layers include ReLU layers, leaky ReLU layers, clipped ReLU layers, ELU activation layers, hyperbolic tangent activation layers, or PReLU layers. Examples of normalization layers, dropout layers, and cropping layers include batch normalization layers, group normalization layers, channel-specific local response normalization layers, dropout layers, 2D crop layers, 3D crop layers, 2D resize layers, and 3D resize layers. Examples of pooling and unpooling layers include average pooling layers, 3D layers, global average pooling layers, 3D global average pooling layers, maximum value pooling layers, 3D maximum value pooling layers, global maximum value pooling layers, or maximum value unpooling layers. Examples of combination layers include additive layers, multiplicative layers, depth-connected layers, and weighted average layers. Examples of object detection layers include ROI input layers, ROI maximum value pooling layers, ROI align layers, anchor box layers, region proposal layers, SSD merge layers, space-to-depth conversion layers, region proposal networks, focal loss layers, region proposal networks, and box regression.

[0075] In embodiments, one or more fully connected dense layers may be applied. The fully connected layer multiplies the input by a weight matrix and then adds a bias vector. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more fully connected dense layers can be applied. When predicting MHC class I or MHC class II binding affinity, it is preferable to apply at least three fully connected dense layers.

[0076] In embodiments, one or more activation layers (functions) may be applied. These activation functions can be assigned to an entire neuron or layer of such neurons. Exemplary activation functions that can be applied are the ELU activation function or the reLU layer. Other activation layers known to those skilled in the art may be applied. These activation functions can convert aggregated, weighted inputs from nodes into activations of nodes or outputs. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more activation layers (functions) may be applied. Typically, about 1, 2, 3, 4, or 5 activation functions may be applied.

[0077] One or more dropout layers may be applied. Dropout layers are beneficial in reducing overfitting, thereby leading to better results. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more dropout layers may be applied. Typically, about 1, 2, 3, 4, or 5 dropout layers may be applied.

[0078] In an exemplary neural network model, one or more fully connected dense layers, one or more activation functions, and one or more dropout layers may be applied. In a preferred neural network model, one or more fully connected dense layers, an activation function (e.g., an ELU activation function), and one or more dropout layers may be applied.

[0079] To enable better learning of the sequence representation, one or more LSTM layers or one or more bidirectional LTSM layers may be applied. Transformers may also be added. For example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more Transformer layers may be added. Transformers can add positional embeddings to the embedded amino acid sequence and may include one or more stacked encoder layers. For example, a Transformer may include one or more multi-head attention layers, one or more dropout layers, one or more normalization layers, one or more feedforward layers, or a combination thereof. An exemplary Transformer may include (1) a multi-head attention layer, (2) a dropout layer (ratio 0.1), (3) a normalization layer, (4) a feedforward layer (linear layer and ReLU layer), (5) a dropout layer (ratio 0.1), and (6) layer normalization.

[0080] By applying a regression model, the binding affinity of the above peptide tumor-specific neoantigen to MHC class I or class II can be predicted. In particular, a variation of the mean squares loss function can be used. The mean squares loss function is L BA-MSE This can be expressed as follows, where the measured value is accompanied by (>) or (<), and with respect to the processing of both quantitative and qualitative peptide-MHC binding affinity measurements in the dataset, the above measured value contributes to the loss only if it violates the inequality.

[0081] The output above includes a numerical score representing the binding affinity of the peptide ligand to MHC class I or MHC class II.

[0082] The above neural network model includes predicting the probability that the tumor-specific neoantigen will be presented on the cell surface by an MHC class I or MHC class II protein. To predict the probability that the tumor-specific neoantigen will be presented on the cell surface by an MHC class I protein, the tumor-specific peptide sequence, the corresponding facies region, and the HLA allele pseudosequence are concatenated into a single numerical score.

[0083] One or more parameters (e.g., layers and / or functions) may be applied to the concatenated peptide sequence of the tumor-specific neoantigen, the corresponding adjacent region, and the HLA pseudosequence.

[0084] One or more fully bound dense layers may be applied. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more fully bound dense layers may be applied. To predict the probability that the tumor-specific neoantigen will be presented on the cell surface by MHC class I or class II proteins, it is preferable to apply at least three fully bound dense layers.

[0085] One or more activation functions may be assigned to a neuron or an entire layer of said neuron. Exemplary activation functions that can be applied are the ELU activation function or the reLU layer. Other activation layers known to those skilled in the art may be applied. These activation functions can convert aggregated, weighted inputs from nodes into activations of nodes or outputs. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more activation layers (functions) may be applied. Typically, about 1, 2, 3, 4, or 5 activation functions may be applied.

[0086] One or more dropout layers may be applied. Dropout layers are beneficial in reducing overfitting, thereby leading to better results. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more dropout layers may be applied. Typically, about 1, 2, 3, 4, or 5 dropout layers may be applied.

[0087] Using Focal Loss binary classification (see Scheme 3 above), the probability that the above tumor-specific neoantigen will be presented on the cell surface by an MHC class I or MHC class II protein can be predicted. The output above is the numerical probability that the peptide will be presented on the cell surface by an MHC class I or MHC class II protein.

[0088] The above neural network model may be further calibrated. If the model has not been calibrated, the neural network may over-predict or under-predict probabilities. Therefore, by calibrating the neural network described in this specification, the accuracy and reliability of the predicted probabilities can be improved. To calibrate the above neural network, probability calculations may be applied. In particular, probability calculations may be applied to one or more target HLA alleles. For example, 1, 2, 3, 4, 5, or 6 HLA alleles. Using probability calculations, the overall presentation probability of the target alleles can be inferred based on the model's predictions for each HLA allele. The above neural network can be further calibrated by calibrating the prediction of the presentation for the validation dataset. For example, a low-order polynomial may be applied to the calibration curve. The polynomial coefficients can be positively constrained to obtain a monotonically increasing function. Lasso linear regression may be used. The calibrated prediction of the presentation (e.g., the MHC class I binding affinity of the above tumor-specific neoantigen and the probability that the above tumor-specific neoantigen is presented on the cell surface by MHC class I protein) can be used as a surrogate value for immunogenicity.

[0089] The performance of the above neural network model can be evaluated using the immunogenicity data described in this specification. The predictions of the above neural network can be evaluated using one or more ranking metrics. Exemplary ranking metrics include top-k items, Precision@K, n DCG K Examples of ranking metrics include, but are not limited to, inverse ranking and positive predictive value metrics. All peptides corresponding to each allele in the immunogenicity dataset can be ranked using one or more ranking metrics based on the predicted probability of peptide presentation on the cell surface and / or the predicted binding affinity score. The ranking metrics can then be aggregated using weighted allele frequencies.

[0090] In one embodiment, self-monitoring pre-training may be performed. Exemplary training models include masked language modeling and next peptide prediction.

[0091] IV. Implementation of Computer Methods The methods disclosed herein can be carried out using a programmed or otherwise configured computer system. Such computer system may comprise a single computing device or multiple computing devices interconnected using one or more computing networks. The computer system can use the capabilities of the computer to execute the neural network models described herein.

[0092] The computer system described above may include a central processing unit, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The system may also include memory (such as random-access memory, read-only memory, or flash memory), electronic storage devices (e.g., a cloud platform), communication interfaces for communicating with one or more systems, and other peripheral devices such as data storage devices, other memory, and display adapters. The memory, storage devices, interfaces, and peripheral devices may communicate with the CPU via a communication bus. One or all of these components may communicate via a shared internal network or an external network, and the aggregate system may communicate with one or more user devices via the network. The network may be the Internet, an extranet, or an Internet / extranet communicating with the Internet. The network may include one or more computer servers, which may enable distributed computing such as cloud computing. The computer system may communicate with a processing system. The processing system may be configured to implement the methods disclosed herein.

[0093] Various examples of the computing devices described above include, but are not limited to, desktop computers, laptops, and mobile phones, tablet computers, personal computers, wearable computers, servers, personal digital assistants (PDAs), hybrid PDAs / mobile phones, mobile phones, e-book readers, set-top boxes, voice command devices, cameras, and digital media players. In some embodiments, the computer device may have one or more user interfaces, command-line interfaces (CLIs), application programming interfaces (APIs), and / or other programming interfaces for sending training requests, deployment requests, and / or execution requests. In some embodiments, the computer device may run a standalone application that interacts with the neural network model.

[0094] In some embodiments, the network may be any wired network, wireless network, or a combination thereof. For example, the network may be a personal area network, local area network, wide area network, wireless broadcast network (e.g., for radio or television), cable network, satellite network, cellular network, or a combination thereof. As a further example, the network may be a network of publicly accessible connected networks operated by various separate parties, such as the Internet. In some embodiments, the network may be a private or semi-private network, such as an intranet of a company or university. The network may include one or more wireless networks, such as a Global System for Mobile Communications (GSM) network, a Code Division Multiple Access (CDMA) network, a Long Term Evolution (LTE) network, or any other form of wireless network. The network may use protocols and components for communication over the Internet or any of the other forms of networks described above. For example, protocols used by the network may include HTTP, HTTP Secure (HTTPS), Message Queue Telemetry Transport (MQTT), Constrained Application Protocol (CoAP), and the like. Protocols and components for communication over the Internet or any of the aforementioned types of communication networks are well known to those skilled in the art, and therefore no further detailed description is provided herein.

[0095] Figure 9 shows examples of provider network (or “service provider system”) environments according to several embodiments. The provider network 900 may provide resource virtualization to customers through one or more virtualization services 910, enabling customers to purchase, rent, or otherwise acquire instances 912 of virtualized resources, including, but not limited to, compute and storage resources, that run on devices within the provider network or in one or more data centers. A local Internet Protocol (IP) address 916 may be associated with the resource instance 912, and the local IP address is the internal network address of the resource instance 912 on the provider network 900. In some embodiments, the provider network 900 may also provide public IP addresses 914 and / or public IP address ranges (e.g., Internet Protocol version 4 (IPv4) or Internet Protocol version 6 (IPv6) addresses) that customers can acquire from the provider 900.

[0096] Conventionally, the provider network 900 may enable a service provider's customer (for example, a customer operating one or more client networks 950A-950C, including one or more customer devices 952) to dynamically associate at least several public IP addresses 914 assigned or allocated to the customer with a specific resource instance 912 allocated to the customer via the virtualization service 910. The provider network 900 may also enable the customer to remap a public IP address 914 previously mapped to one virtualized computing resource instance 912 allocated to the customer to another virtualized computing resource instance 912 similarly allocated to the customer. Using the virtualized computing resource instances 912 and public IP addresses 914 provided by the service provider, a service provider's customer, such as an operator of customer networks 950A-950C, may, for example, run customer-specific applications and serve customer applications over an intermediate network 940, such as the internet. Next, other network entities 920 on the intermediate network 940 may generate traffic to a destination public IP address 914 exposed by customer networks 950A-950C, which is routed to the service provider's data center, where it is routed via the network infrastructure to the local IP address 916 of the virtualized compute resource instance 912 that is currently mapped to the destination public IP address 914. Similarly, response traffic from the virtualized compute resource instance 912 may be routed back to the source entity 920 on the intermediate network 940 via the network infrastructure.

[0097] In this specification, a local IP address refers, for example, to an internal or “private” network address of a resource instance within a provider network. A local IP address may reside within an address block reserved by Internet Engineering Task Force (IETF) Request for Comments (RFC) 1918 and / or within an address block in an address format specified by IETF RFC 4193, and may be modifiable within the provider network. Network traffic originating from outside the provider network is not directly routed to a local IP address; instead, such traffic uses a public IP address mapped to the resource instance's local IP address. The provider network may include networking devices or appliances that provide network address translation (NAT) or similar functionality, performing mapping from public IP addresses to local IP addresses and vice versa.

[0098] A public IP address is a changeable network address on the internet that is assigned to a resource instance by either the service provider or the customer. Traffic routed to a public IP address is translated, for example, via a 1:1 NAT and forwarded to the respective local IP address of the resource instance.

[0099] Some public IP addresses may be assigned to specific resource instances by the provider network infrastructure, and these public IP addresses may be referred to as standard public IP addresses, or simply standard IP addresses. In some embodiments, mapping standard IP addresses to the resource instance's local IP address is the default startup configuration for all forms of resource instances.

[0100] At least some public IP addresses may be allocated to or acquired by customers of Provider Network 900, and the customer may then allocate the allocated public IP addresses to specific resource instances allocated to that customer. These public IP addresses may be referred to as customer public IP addresses, or simply customer IP addresses. Instead of being allocated to resource instances by Provider Network 900 as standard IP addresses are, customer IP addresses may be allocated to resource instances by the customer, for example, via an API provided by the service provider. Unlike standard IP addresses, customer IP addresses are allocated to customer accounts and may be remapped to other resource instances by each customer as needed or desired. Customer IP addresses are associated with customer accounts, not specific resource instances, and the customer manages those IP addresses until the customer chooses to relinquish them. Customer IP addresses, unlike traditional static IP addresses, allow customers to mask resource instance or availability zone failures by remapping their public IP addresses to any resource instance associated with their customer account. For example, by remapping the customer's IP address to an alternative resource instance, the customer can avoid problems with their resource instance or software.

[0101] Figure 10 is a block diagram of an exemplary provider network that provides storage services and hardware virtualization services to customers according to several embodiments. The hardware virtualization service 1020 provides customers with a plurality of computing resources 1024 (e.g., VMs). The computing resources 1024 may be rented or leased to customers of the provider network 1000 (e.g., customers running customer network 1050). Each computing resource 1024 may be assigned one or more local IP addresses. The provider network 1000 may be configured to route packets from the local IP addresses of the computing resources 1024 to destinations on the public internet, and from public internet sources to the local IP addresses of the computing resources 1024.

[0102] The provider network 1000 may provide the ability to run a customer network 1050 connected to an intermediate network 1040 via, for example, a local network 1056, and a virtual computing system 1092 connected to the intermediate network 1040 and the provider network 1000 via a hardware virtualization service 1020. In some embodiments, the hardware virtualization service 1020 may provide one or more APIs 1002, such as a web service interface, through which the customer network 1050 can access the functions provided by the hardware virtualization service 1020, for example, via a console 1094 (e.g., a web application, a standalone application, a mobile application, etc.). In some embodiments, the provider network 1000 may allow each virtual computing system 1092 of the customer network 1050 to correspond to computing resources 1024 leased, rented, or otherwise provided to the customer network 1050.

[0103] From an instance of the virtual computing system 1092 and / or another customer device 1090 (for example, via a console 1094), a customer can access the functions of the storage service 1010, for example via one or more APIs 1002, to access and store data in the storage resources 1018A-1018N (e.g., folders or "buckets", virtualized volumes, databases, etc.) of the virtual datastore 1016 provided by the provider network 1000. In some embodiments, a virtualized datastore gateway (not shown) may be provided to the customer network 1050, which may locally cache at least some data, for example, frequently accessed or important data, and may communicate with the storage service 1010 via one or more communication channels for uploading new or modified data from the local cache so that the primary store of data (virtualized datastore 1016) is maintained. In some embodiments, a user can mount volumes of a virtual datastore 1016 via a virtual computing system 1092 and / or on another customer device 1090, and access the volumes of 1016 via a storage service 1010 that functions as a storage virtualization service, and these volumes may appear to the user as local (virtualized) storage 1098.

[0104] Although not shown in Figure 10, the virtualization service(s) can also be accessed from resource instances within the provider network 1000 via API(s) 1002. For example, a customer, equipment service provider, or other entity can access the virtualization service via API 1002 from within their respective virtual network on the provider network 1000 and request the allocation of one or more resource instances within that virtual network or another virtual network.

[0105] In some embodiments, a system implementing some or all of the techniques described herein may include a general-purpose computer system having, or configured to access, one or more computer-accessible media, such as the computer system 1100 shown in Figure 11. In the illustrated embodiment, the computer system 1100 includes one or more processors 1110 connected to system memory 1120 via an input / output (I / O) interface 1130. The computer system 1100 further includes a network interface 1140 connected to the I / O interface 1130. Although Figure 11 shows the computer system 1100 as a single computing device, in various embodiments, the computer system 1100 may include one computing device or any number of computing devices configured to cooperate as a single computer system 1100.

[0106] In various embodiments, the computer system 1100 may be a uniprocessor system having one processor 1110, or a multiprocessor system having several processors 1110 (e.g., two, four, eight, or another appropriate number). The processor 1110 may be any appropriate processor capable of executing instructions. For example, in various embodiments, the processor 1110 may be a general-purpose or embedded processor implementing any of the following instruction set architectures (ISAs), such as x86, ARM, PowerPC, SPARC, or MIPS ISA, or any other appropriate ISA. In a multiprocessor system, each of the processors 1110 may, but does not necessarily, implement the same ISA.

[0107] The system memory 1120 can store instructions and data accessible by the processor(s) 1110. In various embodiments, the system memory 1120 can be implemented using any and appropriate memory technology, such as random access memory (RAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash type memory, or any other type of memory. In the illustrated embodiment, it is shown that program instructions and data that perform one or more desired functions, such as the methods, techniques, and data described above, are stored in the system memory 1120.

[0108] In one embodiment, the I / O interface 1130 may be configured to coordinate I / O traffic between any peripheral devices in the device, including the processor 1110, system memory 1120, and network interface 1140 or other peripheral interfaces. In some embodiments, the I / O interface 1130 may perform any necessary protocols, timing, or other data conversions to convert data signals from one component (e.g., system memory 1120) into a format suitable for use by another component (e.g., processor 1110). In some embodiments, the I / O interface 1130 may include support for devices connected via various types of peripheral buses, such as variations of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard. In some embodiments, the functionality of the I / O interface 1130 may be divided into two or more separate components, such as a northbridge and a southbridge. Also, in some embodiments, some or all of the functionality of the I / O interface 1130, such as the interface to system memory 1120, may be directly incorporated into the processor 1110.

[0109] The network interface 1140 may be configured to enable data exchange between the computer system 1100 and other devices 1160 connected to a network such as other computer systems or devices as shown in Figure 1, or to multiple networks 1150. In various embodiments, the network interface 1140 can support communication over any and appropriate wired or wireless general data network, such as in the form of an Ethernet network. Furthermore, the network interface 1140 may support communication over telecommunications / telephone networks such as analog voice networks or digital fiber optic communication networks, over storage area networks (SANs) such as Fiber Channel SANs, or over any and appropriate other forms of I / O networks and / or protocols.

[0110] In some embodiments, the computer system 1100 includes one or more offload cards 1170 (including one or more processors 1175 and possibly one or more network interfaces 1140) coupled using an I / O interface 1130 (for example, a version of the Peripheral Component Interconnect Express (PCI-E) standard, or a bus implementing another interconnect such as QuickPath interconnect (QPI) or UltraPath interconnect (UPI)). For example, in some embodiments, the computer system 1100 may function as a host electronic device hosting compute instances (for example, operating as part of a hardware virtualization service), and one or more offload cards 1170 run a virtualization manager that can manage compute instances running on the host electronic device. As an example, in some embodiments, the offload card(s) 1170 may perform compute instance management operations such as pausing and / or unpausing compute instances, starting and / or terminating compute instances, and performing memory transfer / copy operations. These management operations may, in some embodiments, be performed by the offload card(s) 1170 in cooperation with a hypervisor run by other processors 1110A-1110N of the computer system 1100 (for example, in response to a request from the hypervisor). However, in some embodiments, the virtualization manager implemented by the offload card(s) 1170 may be able to respond to requests from other entities (for example, from the compute instance itself) and may not cooperate with (or service) any separate hypervisor.

[0111] In some embodiments, the system memory 1120 may be an embodiment of a computer-accessible medium configured to store program instructions and data as described above. However, in other embodiments, program instructions and / or data may be received, transmitted, or stored on a computer-accessible medium of various types. Generally speaking, computer-accessible media may include magnetic or optical media, such as disks or DVDs / CDs, that are connected to the computer system 1100 via the I / O interface 1130. Non-temporary computer-accessible media may also include any volatile or non-volatile media such as RAM (e.g., SDRAM, double data rate (DDR) SDRAM, SRAM, etc.) or read-only memory (ROM), which may be provided in some embodiments of the computer system 1100 as system memory 1120 or another type of memory. Furthermore, computer-accessible media may include transmission media or signals such as electrical signals, electromagnetic signals, or digital signals that are transmitted via communication media such as networks and / or wireless links, which may be performed via the network interface 1140.

[0112] The various embodiments discussed or suggested herein can run in a wide variety of operating environments and may include one or more user computers, computing devices, or processing devices that can be used to run any of a number of applications. The user or customer device may include a number of general-purpose personal computers, such as desktop or laptop computers running standard operating systems, as well as cellular, wireless, and handheld devices that run mobile software and support a number of networking and messaging protocols. Such a system may also include a number of workstations running various commercial operating systems or other known applications for purposes such as development and database management. These devices may also include other electronic devices, such as dummy terminals, thin clients, game systems, and / or other devices capable of communicating over a network.

[0113] Most embodiments utilize at least one network well known to those skilled in the art to support communications using one of a variety of widely available protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), File Transfer Protocol (FTP), Universal Plug and Play (UPnP), Network File System (NFS), Common Internet File System (CIFS), Extensible Messaging and Presence Protocol (XMPP), and AppleTalk. Examples of such networks include local area networks (LANs), wide area networks (WANs), virtual private networks (VPNs), the Internet, intranets, extranets, public switched telephone networks (PSTNs), infrared networks, wireless networks, and combinations thereof.

[0114] In embodiments utilizing a web server, the web server may run any of the following servers or mid-tier applications: an HTTP server, a File Transfer Protocol (FTP) server, a Common Gateway Interface (CGI) server, a data server, a Java server, a business application server, and so on. The server(s) may also have the ability to execute programs or scripts in response to requests from user devices, such as by running one or more web applications that can be implemented as one or more scripts or programs written in any programming language, such as Java®, C, C# or C++, or Perl, Python, PHP or TCL, or combinations thereof. The server(s) may also include, but are not limited to, database servers commercially available from Oracle®, Microsoft®, Sybase®, IBM®, and others. The database server may be a relational or non-relational ("NoSQL," etc.) database server, a distributed or non-distributed database server, and so on.

[0115] The environments disclosed herein may comprise various datastores and other memory and storage media, as described above. These may reside in various locations, such as on storage media local to (and / or inherent in) one or more computers, or on storage media remote to any or all computers in the entire network. In a particular set of embodiments, information may reside in a storage area network (SAN) as is well known to those skilled in the art. Similarly, any files necessary to perform functions originating from computers, servers, or other network devices may be appropriately stored locally and / or remotely. Where the system comprises computerized devices, each such device may comprise hardware components that are electrically connected via a bus, such as, for example, at least one central processing unit (CPU), at least one input device (e.g., mouse, keyboard, controller, touchscreen, or keypad), and / or at least one output device (e.g., display device, printer, or speaker). Such a system may also include a disk drive, an optical storage device, and a solid-state storage device such as random-access memory (RAM) or read-only memory (ROM), as well as one or more removable storage devices such as a media device, memory card, or flash card.

[0116] Such a device may also include a computer-readable storage medium reader, a communication device (e.g., a modem, a network card (wireless or wired), an infrared communication device, etc.), and the aforementioned working memory. The computer-readable storage medium reader may be connected to, or configured to receive, computer-readable storage medium representing remote, local, fixed, and / or removable storage devices, and storage media for storing, storing, transmitting, and retrieving computer-readable information temporarily and / or more permanently. The above system and various devices will also typically include an operating system and a number of software applications, modules, services, or other components located in at least one working memory device, including application programs such as client applications or web browsers. It should be understood that alternative embodiments may have many variations from the embodiments described above. For example, customized hardware may be used, and / or specific components may be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connections to other computing devices, such as network input / output devices, may be used.

[0117] Storage media and computer-readable media for storing code or parts of code include, but are not limited to, any and appropriate media known or used in the art, including storage media and communication media, such as volatile and non-volatile, removable and non-removable storage media, and including RAM, ROM, electrically erasable and rewritable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other media that can be used to store desired information and can be accessed by a system device. Those skilled in the art will understand other means and / or methods for carrying out various embodiments based on the disclosures and teachings provided herein.

[0118] The above description includes various embodiments. For explanatory purposes, specific configurations and details are described to enable a complete understanding of the embodiments. However, it will be apparent to those skilled in the art that the above embodiments can be carried out even without specific details. Furthermore, well-known features may be omitted or simplified in order to avoid complicating the described embodiments.

[0119] In this specification, blocks enclosed in parentheses and with dashed borders (e.g., long dashes, short dashes, dotted lines, and dotted lines) are used to indicate optional operations that add further features to some embodiments. However, such notation should not be construed as meaning that these are the only options or optional operations, and / or that in certain embodiments, blocks with solid borders are not optional.

[0120] V. Immunogenic composition The present invention further relates to personalized (i.e., target-specific) immunogenic compositions (e.g., cancer vaccines) comprising one or more tumor-specific antigens selected using the methods described herein. Such immunogenic compositions can be formulated according to standard procedures of the Art. These immunogenic compositions have the ability to elicit a specific immune response.

[0121] This immunogenic composition can be formulated so that the selection and number of tumor-specific neoantigens are individually tailored to the specific cancer being targeted. For example, the selection of tumor-specific neoantigens may depend on the specific type of cancer, the stage of the cancer, the immune status of the target, and the MHC type of the target.

[0122] This immunogenic composition may contain at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty This immunogenic composition may contain about 10 to 20 tumor-specific neoantigens, about 10 to 30 tumor-specific neoantigens, about 10 to 40 tumor-specific neoantigens, about 10 to 50 tumor-specific neoantigens, about 10 to 60 tumor-specific neoantigens, about 10 to 70 tumor-specific neoantigens, about 10 to 80 tumor-specific neoantigens, about 10 to 90 tumor-specific neoantigens, or about 10 to 100 tumor-specific neoantigens. Preferably, this immunogenic composition contains at least about 10 tumor-specific neoantigens, or at least about 20 tumor-specific neoantigens.

[0123] This immunogenic composition may further contain natural or synthetic antigens. These natural or synthetic antigens may enhance the immune response. Examples of natural or synthetic antigens include, but are not limited to, pan-DR epitopes (PADREs) and tetanus toxin antigens.

[0124] The immunogenic composition may be in any form, such as a synthetic long-chain peptide, RNA, DNA, cell, dendritic cell, nucleotide sequence, polypeptide sequence, plasmid, or vector.

[0125] Tumor-specific neoantigens can also be incorporated into viral vector-based vaccine platforms, including, but not limited to, lentiviruses such as vaccinia, fowlpox, self-replicating alphavirus, marabavirus, adenovirus (see Tatsis et al., Molecular Therapy, 10:616-629 (2004)), or second-generation, third-generation, or hybrid second / third-generation lentiviruses, and recombinant lentiviruses of any generation designed to target specific cell types or receptors (see, e.g., Hu et al., Immunol Rev., 239(1): 45-61 (2011); Sakma et al, Biochem J., 443(3):603-18 (2012)). Depending on the packaging capacity of the viral vector-based vaccine platform described above, this method can deliver one or more nucleotide sequences encoding one or more tumor-specific neoantigen peptides. The above sequence may be adjacent to a non-mutant sequence, separated by a linker, or preceded by one or more sequences targeting intracellular compartments (see, for example, Gros et al., Nat Med., 22 (4):433-8 (2016); Stronen et al., Science., 352(6291): 1337-1341 (2016); Lu et al., Clin Cancer Res., 20(13):3401-3410 (2014)). Upon introduction into a host, infected cells express one or more tumor-specific neoantigens, thereby inducing a host immune response (e.g., CD8+ or CD4+) to these neoantigens. Useful vaccinia vectors and methods for immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacille Calmette Guerin). The BCG vector is described in Stover et al. (Nature 351:456-460 (1991)).As will be apparent to those skilled in the art from the description herein, a wide variety of other vaccine vectors useful for the therapeutic administration or immunization of neoantigens can also be used.

[0126] This immunogenic composition may contain individualized components to meet the specific needs of the target individual.

[0127] The immunogenic compositions described herein may further contain an adjuvant. An adjuvant is any substance, when mixed with the immunogenic composition, that increases, or otherwise enhances and / or boosts, the immune response to tumor-specific neoantigens, but which, when administered alone, does not produce an immune response to tumor-specific neoantigens. Preferably, the adjuvant produces an immune response to the neoantigens but does not cause allergies or other adverse reactions. In this specification, the immunogenic compositions may be administered before, with, concurrently with, or after the administration of the immunogenic composition.

[0128] Adjuvants can enhance the immune response through several mechanisms, including, for example, the recruitment of lymphocytes, the stimulation of B cells and / or T cells, and the stimulation of macrophages. When the immunogenic composition of the present invention contains an adjuvant or is administered together with one or more adjuvants, the adjuvants that can be used include, but are not limited to, mineral salt adjuvants or mineral salt gel adjuvants, particle adjuvants, microparticle adjuvants, mucosal adjuvants, and immunostimulatory adjuvants. Examples of adjuvants include aluminum salts (alum) (aluminum hydroxide, aluminum phosphate, and aluminum sulfate, etc.), 3-O-deacylated monophosphoryl lipid A (MPL) (see GB2220211), MF59 (Novartis), AS03 (Glaxo SmithKline), AS04 (Glaxo SmithKline), polysorbate 80 (Tween 80; ICL Americas, Inc.), imidazopyridine compounds (international application PCT / US2007 / 064857, published as international publication WO2007 / 109812), imidazoquinoxaline compounds (see international application PCT / US2007 / 064858, published as international publication WO2007 / 109813), and saponins such as QS21 (Kensil et al, in Vaccine Design: The Subunit and Examples of adjuvant approaches include, but are not limited to, the Adjuvant Approach (eds. Powell & Newman, Plenum Press, NY, 1995; see U.S. Patent No. 5,057,540). In some embodiments, the adjuvant is a Freund's adjuvant (complete or incomplete). Other suitable adjuvants are oil-in-water emulsions (such as squalene or peanut oil) optionally combined with immunostimulants such as monophosphoryl lipid A (see Stoute et al., N. Engl. J. Med. 336, 86-91 (1997)).

[0129] It has also been reported that CpG immunostimulatory oligonucleotides enhance the adjuvant effect in the vaccine environment. RNAs that bind to other TLRs, such as TLR7, TLR8, and / or TLR9, can also be used.

[0130] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), Poly ICLC, non-CpG bacterial DNA or RNA, as well as immunoactive small molecules and antibodies that may act therapeutically and / or as adjuvants, such as cyclophosphamide, sunitinib, bevacizumab, Celebrex (celecoxib), NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175. In embodiments, Poly ICLC is a preferred adjuvant.

[0131] The immunogenic composition may contain one or more tumor-specific neoantigens described herein, either alone or with a pharmaceutically acceptable carrier. A suspension or dispersion of one or more tumor-specific neoantigens, particularly an isotonic aqueous suspension, dispersion, or amphiphilic solvent, may be used. The immunogenic composition may be sterile and / or contain excipients, such as preservatives, stabilizers, wetting agents and / or emulsifiers, solubilizers, salts and / or buffers for adjusting osmotic pressure, and may be prepared by methods known in themselves, for example, by conventional dispersion and suspension processes. In certain embodiments, such dispersions or suspensions may contain viscosity modifiers. The suspensions or dispersions should be maintained at a temperature of about 2°C to 8°C, or, for longer-term storage, it is preferable to freeze them and thaw them immediately before use. For injection, the vaccine or immunogenic formulation may be formulated in an aqueous solution, preferably a physiologically suitable buffer such as Hanks' solution, Ringer's solution, or physiological saline buffer. The above liquid may contain compounding agents such as suspending agents, stabilizers, and / or dispersants.

[0132] In certain embodiments, the compositions described herein further comprise a preservative, such as the mercury derivative thimerosal. In certain embodiments, the pharmaceutical compositions described herein comprise 0.001% to 0.01% thimerosal. In other embodiments, the pharmaceutical compositions described herein do not contain a preservative.

[0133] Excipients may exist independently of the adjuvant. The function of the excipients may be, for example, to increase the molecular weight of the immunogenic composition, to increase its activity or immunogenicity, to confer stability, to increase its biological activity, or to increase its serum half-life. Excipients may be used to assist in the presentation of one or more tumor-specific neoantigens to T cells (e.g., CD4+ or CD8+ T cells). The excipients may be carrier proteins such as keyhole limpet hemocyanin; serum proteins such as transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, immunoglobulins; or hormones such as insulin or palmitic acid. In the case of human immunization, the carriers are generally physiologically acceptable carriers that are tolerable and safe for humans. Alternatively, the carriers may be dextran, for example, cepharose.

[0134] Cytotoxic T cells recognize antigens not as intact foreign antigens themselves, but in the form of peptides bound to MHC molecules. These MHC molecules are present on the cell surface of antigen-presenting cells. Therefore, activation of cytotoxic T cells is possible when a trimer complex of peptide antigen, MHC molecule, and antigen-presenting cell (APC) is present. The immune response may be enhanced if, in addition to using one or more tumor-specific antigens to activate cytotoxic T cells, further APCs containing each of the MHC molecules are added. Therefore, in some embodiments, the immunogenic composition further comprises at least one APC.

[0135] The immunogenic composition may contain an acceptable carrier (e.g., an aqueous carrier). Various aqueous carriers can be used, such as water, buffer water, 0.9% physiological saline, 0.3% glycine, hyaluronic acid, etc. These compositions may be sterilized by conventional, well-known sterilization techniques or by sterile filtration. The resulting aqueous solution may be packaged for immediate use or lyophilized and mixed with a sterile solution before administration. The above compositions may contain pharmaceutically acceptable auxiliary substances necessary to approximate physiological conditions, such as pH adjusters and buffers, tonicity adjusters, and wetting agents, such as sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, and triethanolamine oleate.

[0136] Neoantigens may also be administered via liposomes, which target specific cell tissues such as lymphoid tissue. Liposomes are also useful for increasing half-life. Examples of liposomes include emulsions, foamy substances, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, and lamellar layers. The delivered neoantigens are incorporated into these preparations as part of liposomes, either alone or together with molecules that bind to receptors widely present in lymphoid cells, such as monoclonal antibodies that bind to the CD45 antigen, or with other therapeutic or immunogenic compositions. In this way, liposomes filled with the desired neoantigens can migrate to a site in lymphoid cells, which then deliver the selected immunogenic composition to that site. Liposomes can be formed from standard vesicle-forming lipids, which generally include neutral and negatively charged phospholipids and sterols such as cholesterol. Lipid selection is generally guided by considerations such as liposome size, acid instability, and stability of the liposomes in the bloodstream. Various methods are available for preparing liposomes, such as those described in Szoka et al., An. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Patents No. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.

[0137] To target immune cells, ligands incorporated into liposomes may include, for example, antibodies or antibody fragments specific to the cell surface determinants of desired immune system cells. The liposome suspension may be administered intravenously, locally, or topically, in doses that vary depending on the method of administration, the peptides delivered, and the stage of the disease being treated.

[0138] Alternative methods for targeting components of this immunogenic composition, such as immune cells, antigens (i.e., tumor-specific neoantigens), ligands, or adjuvants (e.g., TLRs), may be incorporated into lactic acid-glycolic acid copolymer microspheres. These lactic acid-glycolic acid copolymer microspheres can encapsulate components of this immunogenic composition as endosome delivery devices.

[0139] Nucleic acids encoding tumor-specific neoantigens as described herein may be administered to a target for therapeutic or immunization purposes. Many convenient methods are available for delivering the nucleic acids to a target. For example, the nucleic acids may be delivered directly as "naked DNA." This method is described, for example, in Wolff et al., Science 247: 1465-1468 (1990), and in U.S. Patents 5,580,859 and 5,589,466. The nucleic acids may also be delivered using ballistic delivery as described, for example, in U.S. Patent 5,204,253. Particles consisting solely of DNA may be delivered, or the DNA may be attached to particles such as gold particles. Methods for delivering nucleic acid sequences include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation. The nucleic acids may also be delivered in complex with cationic compounds such as cationic lipids.

[0140] The immunogenic compositions provided herein may be administered to a subject by routes including, but not limited to, oral, intradermal, intratumoral, intramuscular, intraperitoneal, intravenous, topical, subcutaneous, percutaneous, nasal, and inhalation routes, as well as by random cutting (for example, by scratching through the surface of the skin using a forked needle). The immunogenic compositions may be administered to a tumor site to induce a local immune response against the tumor.

[0141] The dosage of one or more tumor-specific neoantigens described above may depend on the type of composition, as well as the subject's age, weight, body surface area, individual condition, individual pharmacokinetic data, and method of administration.

[0142] This specification also discloses a method for producing an immunogenic composition comprising one or more tumor-specific neoantigens selected by carrying out the steps of the method disclosed herein. The immunogenic compositions described herein can be produced using methods known in the art. For example, a method for producing a tumor-specific neoantigen or vector disclosed herein (e.g., a vector comprising at least one sequence encoding one or more tumor-specific neoantigens) may include culturing host cells containing at least one polynucleotide encoding the neoantigen or vector under conditions suitable for expressing the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatography, electrophoresis, immunological techniques, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.

[0143] Examples of host cells include Chinese hamster ovary (CHO) cells, NS0 cells, yeast, or HEK293 cells. The host cells may be transformed with one or more polynucleotides containing at least one nucleic acid sequence encoding one or more tumor-specific neoantigens or vectors disclosed herein. In certain embodiments, the isolated polynucleotides may be cDNAs. [Examples]

[0144] Example 1.1: Training Data The model was trained to predict the probability of peptide-MHC binding and endogenous peptide presentation on MHC class I cells. These were treated as surrogates for CD8+ T cell immunogenicity. Selected peptide-MHC binding affinity data from MHCflurry ("curated_training_data.no_mass_spec.csv"1) was used, which includes data from IEDB[1] and Kim et al.[2]. The only processing step performed on this selected dataset was to extract adjacent regions and add the peptide source protein(s) for use in negative sampling. After this processing step, the final dataset used for training is named "curated_training_data.no_mass_spec.multiple_context.blast.v2.csv.". From this dataset, only entries containing HLA-A / B / C alleles and peptides of length 8-15 without post-translational modifications were retained. The target for these samples is either quantitative ("=") or qualitative ("<" / ">"), as proposed in MHCflurry-1.2[3]. Qualitative entries in the selected MHCflurry datasets represent qualitative values ​​of positive-high (<100nm), positive-medium (<1000nm), positive-low (<5000nm), or negative (>5000nm). In addition, we utilized a peptide presentation dataset to the cell surface, consisting of peptides presented on the cell surface as measured by peptide elution and mass spectrometry from several sources (Sarkizova et al.[4] dataset, which involved profiling of over 185,000 peptides eluted from 95 HLA-A, B, or C and G single-allele cell lines using mass spectrometry). In the MHCflurry dataset, when mass spectrometry is used to identify a sample, the mass spectrometry is considered accurate, and the relevant sample in the dataset is identified by the presence of a "mass spectrometry" value in the "measurement_source" column.This includes 226,684 ligands identified by MS, deposited in the IEDB[1] or the SysteMHC Atlas[5], or reported by Abelin et al.[6].

[0145] Furthermore, cell surface presentation data obtained by stably transfecting HEK293 cells at Fred Hutchinson Cancer Center to express secreted HLA molecules covalently bound to β-2-microglobulin was used. Subsequently, the cell supernatant was prepared to capture the secreted MHC peptide complex and analyzed by mass spectrometry. Reports had been made regarding peptides derived from the purified complex. The data file is available in the compressed file "RolandPeptidePresentationData.zip". All of these data sources consist of peptide sequences known to be presented on MHC class I molecules associated with specific HLA alleles. For our final presentation dataset ("mass_spec_data.multiple_context.blast.allele_supertypes.v3.csv"), samples of all mentioned sources, including HLA-A / B / C alleles and peptide source proteins(s), were extracted from adjacent regions and combined for use in a negative sampling method. Samples possessing identical alleles and peptides, with only a single additional amino acid at either edge, were considered duplicates due to inaccuracies in mass spectrometry measurements; therefore, "extended" samples that were close to duplicates were also filtered out. Samples (peptide-MHC pairs) found in the immunogenicity assessment data were also filtered out from both the affinity and presentation datasets to ensure that the assessment sets were completely "hidden" during the training phase and not shown to the model at all.

[0146] These datasets (affinity and presentation) were randomly split into training and validation splits before training the new model. For each allele, N was selected from the binding affinity samples of the i-th allele. bai Peptides were randomly sampled, and from the presentation sample of the i-th allele, N pi Peptides are randomly sampled. Here, N bai =min(0.25*|specific affinity peptide|i,100) and N pi =min(0.25*|unique presentation peptide| i ,100)

[0147] For each allele, these sampled peptides (a combination of peptides sampled from both datasets) are considered a hidden validation set, and all samples containing these peptides are removed from the training set and used solely for validation. The above validation set is used to determine early termination of model training if the validation loss does not improve in more than N consecutive epochs. N=20 was set for all experiments. When training multiple similar models, different training-validation splits are used for ensemble / model selection purposes.

[0148] Human rhinovirus (HR) data were included in our data as negative samples from ICS applied to a 1600 HRV 15-mer constructed from HRV-1A, HRV-B, and HRV-C mosaic (as described in Fischer et al., (2007), Nature Medicine, 13, 100-106). Specifically, negative samples randomly sampled from the ICS results were used for immunogenicity evaluation and tested as part of the immunogenicity evaluation data. The remaining HRV samples were treated as negative presentation samples and randomly sampled during training.

[0149] Example 1.2: Immunogenicity assessment data To investigate the extent to which BigMHC 1.0's predictions of peptide-MHC binding affinity and peptide presentation on the cell surface can be adequately translated into predictions of T cell immunogenicity, a T cell immunogenicity dataset was created, and the BigMHC machine learning model was tested and validated. Previously reported peptide-MHC pairs were used, derived from the summary tables of CTL / CD8+ epitopes in the HIV Molecular Immunology Database and the previously reported summary tables of CTL epitopes in the HCV Immunology Database. These tables provide experimentally validated HIV / HCV CTL / CD8+ epitopes.

[0150] All of these samples are in the class of immunogenicity positive. Similarly, to obtain immunogenicity negative peptide-MHC samples, experiments that produced positive pairs were manually reviewed. No negative findings were reported, and some of them were found to be reconstructible from positive findings. In detail, in many experiments, all possible peptide-MHC pair combinations were tested for a given set of peptides and a given set of HLA alleles, a method known as the matrix method. By extracting previously reported peptides and alleles from the positive findings, it could be concluded that at least (there may be further negative samples that the inventors have overlooked and that have not been reported, such as peptide / alleus being negative among all possible combinations) all possible peptide-MHC pair combinations within the above peptides and alleles were tested in the above experiments. Considering this list of tested peptide-MHC pairs, it can be concluded that any pair on this list that was not reported as positive is actually negative. To verify whether the matrix method was used, we reviewed 31 of the largest experiments (using the largest amount of test sample, |unique allele|x|unique peptide|), and found that this method was used in 18 of them. No negative samples were inferred from experiments using other methods.

[0151] For the remaining smaller experiments, a matrix method was used, and it was assumed that all negative samples were extracted unless the sample had been reported positive in another experiment. If a sample had been reported positive in another experiment, the inventors assumed that the sample was positive. Only HLA-A / B / C alleles and peptides of length 8–15 were retained.

[0152] In addition to the above data, the inventors add immunogenicity-positive samples reported in the IEDB (Vita et al., (2019), Nucleic acids research, 47, D339-D343) and additional randomly sampled HRV-negative samples. For alleles with a relatively low positive sample:negative sample ratio, HRV-negative samples are sampled. The inventors consider the ideal ratio to be 1:100 and, if possible (not all alleles have HRV samples), HRV-negative samples are sampled until approximately this ratio is reached. Following this balancing procedure, all samples for alleles with a ratio less than 1:5 are filtered out.

[0153] Following this procedure, pairs of 2,985 positive and 68,469 negative samples were obtained, covering 110 HLA alleles and 1,416 unique peptides.

[0154] This immunogenicity dataset was divided into two sets: a validation set 6 for tuning the hyperparameters of the inventors' model and selecting further configurations, and a test set 7 for which the inventors performed final benchmarking.

[0155] Example 1.3: Data Analysis To better visualize and understand the distribution of the above data, the following figure was plotted.

[0156] Peptide length: Figure 2B shows the distribution of peptide lengths across all three datasets.

[0157] Peptide diversity: The similarity between two peptides is called the number of overlapping amino acids in the best possible alignment. In Figure 2A, for each similarity threshold, the percentage of peptides that did not have a partner peptide with similarity exceeding a given threshold was calculated. In this analysis, only unique peptides were considered, and overlapping peptides within each dataset were ignored. In detail, the affinity dataset consisted of 35,467 unique peptides out of a total of 158,001 samples (22.45%), the presentation dataset consisted of 265,236 unique peptides out of a total of 384,812 samples (68.93%), and the immunogenicity dataset consisted of 1,416 unique peptides out of a total of 71,474 samples (1.98%).

[0158] Target Distribution: The distribution of targets and the imbalances between them differed across different datasets. Specifically, the binding affinity - peptide-MHC binding affinity dataset consisted of a mixture of quantitative and qualitative targets. Figure 3 shows a high level of qualitative distribution of binding affinity targets. Presentation - The peptide presentation dataset to the cell surface was binary, and our dataset consisted only of positive samples. During training, we applied negative sample mining at the start of each epoch to generate a corresponding negative sample for each positive sample. Immunogenicity - The T-cell immunogenicity dataset was binary but highly imbalanced, consisting of 2,985 positive peptides and 68,469 negative peptide-MHC pairs.

[0159] Sample Distribution by Supertype - Visualizing each allele individually is difficult. To better visualize and understand the underlying allele distribution in the dataset, Figure 4 shows the distribution of HLA allele supertypes in our dataset samples. We applied the HLA supertype classification determined by Sidney et al.[7].

[0160] Sample distribution by HLA allele - In addition to the distribution of HLA allele supertypes, Figure 4 also shows the distribution of HLA alleles in each peptide-MHC sample of the dataset.

[0161] Example 1.4: Negative mining data during training The peptide presentation data on the cell surface consists only of "positive" samples, but these positive samples do not provide the negative samples (which cannot be presented on the cell surface) necessary for training the binary presentation classifier. Therefore, to train such a classifier, the following strategy for stochastic negative mining during training was employed.

[0162] Given a positive sample consisting of an HLA allele shuffle peptide and its corresponding HLA allele, the given allele was replaced by randomly sampling a different allele that did not belong to the supertype(s) of the positive allele. The HLA supertype classification determined by Sidney et al.[7] (which assigns each HLA allele to one or more HLA supertypes) was applied. In this classification, some HLA alleles remained unclassified. These unclassified alleles were mapped to three further supertype classes, "Unclassified-A," "Unclassified-B," and "Unclassified-C," according to the corresponding HLA-A / B / C groups, and these groups were treated similarly to the other supertype classes.

[0163] Peptide Shuffle: Given a positive sample consisting of a peptide and its corresponding HLA allele, the given peptide was replaced with an amino acid subsequence of the same length randomly sampled from the peptide's source protein. Furthermore, the affinity training dataset was expanded to include random peptides sampled from the amino acid data distribution, with a qualitative weak affinity target (>20,000 nM), according to the method of MHCflurry-1.6[8]. The lengths of these random peptides were measured for each allele, by performing the same number of unbound data points for each peptide length.

[0164] HRV-negative sampling involves randomly sampling from negative HRV data (excluding samples used for immunogenicity assessment). The basic assumption here is that, in most cases, negative immunogenicity will result from negative presentation on the cell surface, and therefore this can be anticipated during training (although there are cases where presentation is positive and immunogenicity is negative).

[0165] Example 1.5: Estimating the goal of a cross-task For joint multitask training, all training samples from both the binding affinity dataset and the presentation dataset were used. However, for each sample, only one known target existed, and the corresponding single known target (not two tasks) in the corresponding dataset from which that target originated was also known. To mitigate this problem (and to make multitask training more effective), we inferred the target of one task from the target of the other task by assuming that samples presented on the cell surface (presentation positive) would also have high binding affinity values, and samples with low binding affinity values ​​would not be presented (presentation negative). In detail, for all presentation positive samples, a qualitative high affinity target (<500 nM) was inferred, and for samples with low binding affinity measurements (>5000 nM), a presentation negative target was inferred. The remaining "missing targets" (targets that could not be inferred) were simply ignored (masked) during training by assigning them a sample weight of zero (only for tasks where no target was found).

[0166] Example 1.6: Autodistillation Using the BigMHC predictor, we extracted estimates of binding affinity and presentation for various samples, and added these samples with the corresponding "weak" labels to the training dataset. This self-distillation process was carried out in the two scenarios described below.

[0167] 1. Mass Spectrometry Data of Multiple Alleles: The MULTI-ALLELIC OLD dataset described in MHCflurry-2.0 was used. This dataset contains mass spectrometry hits of over 200,000 positive samples. BigMHC-1.3.1 prediction was used to determine which allele was the allele responsible for the hit. First, hits of multiple alleles with known positive presenters were filtered out from the inventors' presentation training data. Next, the allele with the highest probability of presentation was selected, and if the probability of presentation was higher than a certain threshold (0.5) and the binding affinity was below a certain threshold (5000 nM), only that sample was retained.

[0168] 2. Positive Presenters For all positive presenters in the presenting training data whose binding affinity was unknown (after performing the steps above), the binding affinity was estimated based on BigMHC-1.3.1 prediction. All samples with a predicted binding affinity of less than 5000 nM were added to the binding affinity training data.

[0169] Example 2.1: Displaying an array Each HLA allele was represented by a 49-amino acid pseudosequence used in MHCflurry-1.4. This pseudosequence coding uses amino acids at 49 selected positions determined by multiple sequence alignments of numerous MHC class I alleles across species. The display of HLA allele pseudosequences is available in the "allele_sequences.csv" file. Furthermore, peptides of amino acid lengths 8–15 were displayed using fixed-length coding designed to preserve the positionality of residues that make the most important MHC stabilizing contacts, utilizing the peptide padding and coding method of O'Donnell et al.[3]. These "anchor positions" arise toward the beginning or end of the peptide for most alleles. The peptide is represented as a 15-length sequence, in which missing residues are filled with the letter "X", effectively the 21st amino acid. The first and last four residues of the above peptide are mapped to the first and last four positions in the above display. The middle seven residues are filled as needed. In the octamer, all intermediate positions remain as X, while in the 15mer, all positions are filled. In this way, the positions most likely to contain anchor residues are consistently mapped to the same positions in the above representation. Considering the 5 amino acids on each side linked at the edge of the encoded peptide, the adjacent regions of the peptide were also encoded.

[0170] In contrast to O'Donnell et al.[3], which uses immobilized amino acid embeddings based on the BLOSUM62 substitution matrix, our invention uses a trainable embedding layer that is trained end-to-end in conjunction with the remnants of the neural network. This embedding layer encodes every amino acid, both in the encoded peptide or in the allele pseudosequence, into a 16-dimensional vector.

[0171] Example 2.2: Adjacent Regions For each peptide sequence, we identified all instances in which this sequence is a subset of longer protein sequences present in the UniProt dataset[9]. We searched three UniProt files, namely (1) the UniProt human proteome dataset, "UP000005640_9606.fasta", (2) the complete UniProtKB / Swiss-Prot dataset, "uniprot_sprot.fasta", (3) the additional sequences from the UniProtKB / Swiss-Prot dataset, "uniprot_sprot_varsplic.fasta", representing all annotated splice variants, and (4) downloaded both the CD8 ("CD8_epitopes_netMHCpan.fas") and CD4 ("CD4_epitopes_netMHCIIpan.fsa") from the netMHCpan-4.0 immunogenicity dataset.

[0172] Each of the sequences longer than the above is called a "parent sequence". Each peptide is associated with one or more "adjacent regions" of length 10, where the adjacency regions are the five amino acids immediately preceding the peptide and the five amino acids immediately following the peptide sequence in each of the parent sequences of the peptide. All unique combinations of adjacency regions were stored in a file, and a weight was defined for each of them, inversely proportional to the number of unique sequences for each peptide. These weights were used as sample weights during training to ensure that peptides with many variations were not given too much weight, while allowing the network to learn all possible variations. For peptides that did not exactly match, BLAST10 was used to find the best-matching peptide, and its corresponding "parent sequence" was used to extract the relevant adjacency regions.

[0173] Example 2.3: Self-monitoring pre-training Using a large-scale protein database, we pre-trained our model to learn appropriate initial sequence representations. Using a subset of 25 million proteins from the Uniparc database, and inspired by BERT pre-training, we trained the peptide transformer model on the following two tasks.

[0174] 1. Masked Language Modeling We randomly selected 0.15 tokens from those considered "masked" and trained a token classification head (based on all other tokens) to attempt to predict the original tokens using cross-entropy loss. For the model input, the "masked" tokens could be randomly replaced (10%), remain unchanged (10%), or replaced with masking tokens (80%).

[0175] 2. Prediction of the next peptide In the pre-training phase, our input sequence was a concatenation of two peptide sequences (in contrast to the concatenation of a peptide sequence and an allele sequence in the main training phase). The sequence was given a special isolation token ( <sep>) are separated via and have different segment indices and embeddings (a segment sequence is an additional input to the network, simply indicating whether each token belongs to the first sequence, the second sequence, or a special token. Then, the embedding of the segment index is added to the token, and the embedding is placed). <cls>A classifier was trained on the token output to predict whether the second peptide was the next peptide produced in the protein (after the first peptide). The peptides were either derived from two consecutive peptides of the same length from a human protein, or randomly sampled from different proteins.

[0176] Example 2.3 Training Objective and Multitask Loss Our neural network was jointly trained to predict both peptide-MHC binding affinity and peptide presentation on the cell surface. We utilized a modified version of the mean squared error (MSE) loss function, as performed by O'Donnell et al.[3], where measurements with inequalities (>) or (<) contribute to the loss only if they violate the inequality, with respect to the processing of both quantitative and qualitative peptide-MHC binding affinity measurements in the dataset. The exact formulas are outlined in Scheme 1 and Scheme 2.

[0177] For the binary classification task of peptide presentation on the cell surface, we utilize the Focal Loss

[10] shown by LP-FL, which is a weighted extension of the standard binary cross-entropy loss that gives greater emphasis to samples that are not well classified. The exact formula is outlined in Scheme 3 (reproduced below).

number

[0178] Here, γ is a real parameter and is set to 1. In the binary case,

number

number

[10] have shown that Focal Loss is effective in addressing data imbalance, and a recent study by Mukhoti et al.

[11] has also shown that Focal Loss provides better network calibration compared to standard cross-entropy loss. The overall objective function is a linear combination of a variation of MSE for peptide-MHC binding affinity and a binary Focal Loss for peptide presentation to the cell surface, L= aL BA-MSE +(1-a)L p-FL That is the case.

[0179] During training, the negative sampling method described in Example 1.4 was also applied. These were applied at the start of each epoch for both the presentation task and the binding affinity task. Negative samples from the validation set were generated once at the start of the first epoch and fixed throughout the entire training process. Where possible, the goals of the other task inferred from each task were also used. A sample weighting mechanism was also applied for each loss term, assigning different weights to samples based on the amount of adjacent region variation and masking samples for which no goals were found (only samples for specific tasks where the goal was unknown were masked). Samples for which the goal was inferred, as described in Example 1.5, were given the same sample weights for both tasks.

[0180] To predict peptide presentation on the cell surface, a binary classifier is trained using binary cross-entropy loss. P-BCE The presentation objective represented by is

number

[0181] The overall objective function is a linear combination of the above training objectives, using a predefined weight loss coefficient:

number

[0182] During training, a negative sampling method was also applied to each randomly sampled negative sample, determining which of the four methods described above to use. The N ratio, a hyperparameter for the negative:positive ratio, was also defined and used to determine the number of negative samples to sample for each positive presentation. A sample weight of 1 / N ratio was assigned to each sampled negative sample, effectively ensuring that the LP-BCE objective is trained on balanced data (equal overall sample weights for positive and negative training samples).

[0183] Furthermore, during the main training phase, an auxiliary objective of masked language modeling was utilized, in which the masked tokens were randomly selected only from the adjacent regions of the peptide in question. This objective is applied only to natural peptide sequences and is therefore ignored for certain types of sampled negative samples where either the peptide (e.g., randomly sampled amino acid sequences) or its context (e.g., HRV-negative samples) is "synthetic". This auxiliary loss was L A-MLM It is represented as follows.

[0184] Example 2.4 Calibration Recent research by Guo et al.

[12] has revealed that modern neural networks are not properly calibrated. Calibration properties are critical to us because our ranking logic pipeline uses probability calculations and makes decisions based on the predicted probability of presentation. Specifically, in our vaccine design pipeline, we considered six control HLA alleles and applied the following probability calculations to estimate the overall presentation probability for each target allele based on the model's predictions for each single HLA allele.

number

[0185] If our predictors are properly calibrated, this type of calculation will improve. Therefore, we further calibrated the network's presentation predictions on the validation set by fitting a low-order polynomial to the calibration curve. To obtain a monotonically increasing function, we constrained all polynomial coefficients to be positive. We used Lasso linear regression for this task. This step was important to obtain well-calibrated presentation probabilities, which were later used in our inference pipeline. However, since this calibration step is monotonic, it does not affect the ranking of peptides for a single allele. In our vaccine design pipeline, we used the above-calibrated presentation predictions as our best surrogate for immunogenicity probability estimations.

[0186] Example 2.5 Model Architecture A. Model Architecture 1 The architecture of this model consisted of three main components: sequence processing, peptide-MHC binding affinity prediction, and subsequent peptide presentation prediction on the cell surface. The input to our model was a peptide primary sequence containing the aforementioned adjacency region and a 49-amino acid HLA allele pseudosequence. First, the peptide sequence was encoded into a fixed-length vector, and then all amino acid sequences were encoded using a shared d-dimensional amino acid embedding layer. Next, the embedded sequences were flattened to generate vector representations of each peptide, allele, and adjacency region. For the peptide-MHC binding affinity prediction component, the representations of the peptide and HLA allele were concatenated, and two dense layers of sizes 512 and 256, followed by Exponential Linear Unit (ELU) activation and a dropout layer with a dropout probability p=0.5, were applied to each. The output of this component is called the "affinity representation." Next, an additional dense layer was used to output the predicted binding affinity logit. Finally, for the peptide presentation prediction component on the cell surface, the representations of the peptide, adjacent region, and HLA allele were first concatenated into a single vector. Next, two fully bound layers with a similar structure to those in the peptide-MHC binding affinity prediction component were used, but with the output to the "affinity representation" also concatenated. The output of this component is called the "presentation representation." Subsequently, an additional linear dense layer was added to predict the probability logit of peptide presentation on the cell surface.

[0187] B. Model Architecture 1 The above model consists of three components: 1. array embedding, 2. Self-attention Transformer layer, and 3. prediction head.

[0188] 1. Sequence Embedding - The above model received two sequences as input: 1) an amino acid sequence (a linkage of an allele pseudosequence and a peptide containing adjacent regions); and 2) a segment identifier sequence (allele, peptide, context, and special tokens) that "notifies" the model which segment each amino acid belongs to. Figure 12 shows the structure of the token and segment input sequences. Each token in the above sequence is represented by an aa embedding, a learned position embedding, and additional segment embeddings.

number

[0189] 2. Self-attention Transformer Layers - The embedded array was processed with 12 consecutive Transformer layers, each containing a Multi-Head Self-attention module followed by a feedforward module (consisting of two linear layers with GELU activation in between), as shown in Figure 13. Layer normalization was applied at the beginning of each module, and residual dropout with a ratio p=0.1 was applied at the end of each component before residual connections.

[0190] 3. Predictive Heads - The above model included the following three predictive heads:

[0191] Masked language modeling: The final representation at each position in the sequence was fed into an LM predictive head consisting of a linear layer (including GELU activation and subsequent layer normalization) and an additional linear projection (to the token vocabulary size) with weights coupled to a token embedding matrix with additional learning biases.

[0192] Binding affinity: <cls>The final representation at token position was fed into a prediction head consisting of Linear + GELU + LayerNorm + Dropout + Linear.

[0193] Presentation: Linked to a single binding affinity logit, <cls>The final representation at token position was fed into a prediction head consisting of Linear + GELU + LayerNorm + Dropout + Linear.

[0194] Example 2.6 Model Ensemble The neural network architecture described above was used for training two types of models: (1) a general allele model containing all training data, and (2) an allele-specific model where the training data is divided by HLA allele. Each HLA allele-specific model was trained with sufficient training data for each HLA allele (using the criterion that the dataset contained at least 1000 binding affinity samples and 1000 presentation samples).

[0195] Both types have strictly identical architectures and undergo similar training. The only difference between the two models is the initialization of the training data, i.e., the list of alleles supported during inference and their weights. The general model is initialized randomly, while the allele-specific models are fine-tuned from the best-trained general model. During inference, an ensemble of trained models is used, which averages the predictions for each sample across all models supporting the sample's alleles.

[0196] Example 2.7 Model Selection Early experiments have shown that in many cases, a subset of models, or even a single model, performs better than an ensemble of many models. Therefore, we developed the following model selection procedure: 1. For any given model configuration (general models and allele-specific models for each allele with sufficient training data), train 10 models in a strict training setup using only different folds (training-validation folds). After training is complete, select the single best-performing model from the 10 trained models for each model configuration, based on our evaluation protocol using the validation immunogenicity data fold portion. Apply hierarchical model selection per allele, in which, for any allele (with immunogenicity data for performance validation), select the most optimal configuration possible from the following options: a) use only allele-specific models; b) use only general models; c) use an ensemble of general models + allele-specific models and average their predictions.

[0197] Example 2.8 Evaluation To evaluate the model's performance, T-cell immunogenicity data were used to determine how well a given trained model managed to rank positive immunogenic peptide-HLA allele pairs higher than negative non-immunogenic pairs. In detail, we focused on peptides ranked up to approximately the top 20, as this roughly represents the desired amount of peptide to produce a given vaccine. Given an immunogenicity validation / test set, model predictions could be extracted for all samples. With this in mind, we utilized three common ranking indices focusing on the top K-rank items: Precision@K, nDCGK, and Reciprocal Rank. We also utilized positive predictive value indices used in previous studies, such as O'Donnell et al.

[13] . For each allele, all corresponding peptides were individually ranked in the immunogenicity validation / test set based on their predicted probability of peptide presentation on the cell surface or their predicted binding affinity score. Precision@K, nDCGK, Reciprocal Rank, and positive predictive value index were then calculated for each allele.

[0198] After calculating these indices separately for each allele, they were aggregated across all HLA alleles by applying a weighted average. In this case, each allele was weighted by its frequency in the U.S. population. Allele frequencies from the four highest ethical standards in the U.S. were used.

[11] The frequencies derived from these files were further weighted by the frequencies of each ethnic group in the U.S. population.

[12] Specifically, 0.54 was used for European Caucasians, 0.22 for Hispanics, 0.17 for African Americans, and 0.07 for Asians. The motivation for weighting by allele frequencies in the population is that the above indices capture (or at least better correlate) the proportion of subjects for which this method could theoretically be useful. The frequency weights of HLA alleles in the immunogenicity assessment set, based on population, are plotted in Figure 7. The final metrics, Weighted Precision@K (WP@K), Weighted nDCGK (WnDCGK), Weighted-Reciprocal-Rank (WRR), and Weighted Positive Predictive Value (WPPV), are expressed by the following formulas:

number

number

[0199] In the formula, rank i This is the rank of the first positive sample in the ranked peptide of the i-th allele.

number

[0200] In the formula, Ni is the number of positive samples for the i-th allele. For all indicators,

number

[0201] While the PPV index is useful, it does not focus on the top few peptides, which is the area of ​​greatest interest to the inventors. The Reciprocal Rank index uses the rank of the first positive item among the ranked peptides, giving a higher score for higher ranks. However, this index ignores the presence or absence of further positive peptides in the higher-ranked items. This was not ideal, as the inventors wanted to ensure that they obtained as many potentially positive samples as possible in the vaccine they designed (since not all presented peptides are immunogenic and do not produce the desired immune response). The Precision@K index simply captures the proportion of positive samples among the top K ranked items. However, it does not consider the actual ranks within the top K (i.e., one positive item ranked 1st and one positive item ranked Kth would receive the exact same score, which is not ideal). The nDCGK index considers both the proportion of positive samples within the top K and their corresponding ranks. The DCG normalization coefficient for ideal rankings reduces the need to limit K when the number of positive samples is less than 20, and also produces values ​​within the [0,1] range. However, these scores are still somewhat less interpretable / intuitive compared to the precision@K index.

[0202] Example 3.1 Results Unless otherwise specified, all models were trained using the ADAM optimizer with a learning rate of 0.001 and a batch size of 256. The loss term coefficient was set to α=0.5, and both loss terms were given equal weight. Early stopping was applied if the inventors' validation loss, evaluated on a hidden validation set, did not improve after 20 epochs. Both a small LI regularization coefficient of 1e-9 and dropout with a drop rate of 0.5 were applied. Negative sampling methods and target estimation were applied to all models. In the case of allele-specific models, if it was undesirable to replace the allele of a positive sample with an allele of a different supertype (because we were only interested in samples of a specific allele), we did not start the repository of positive samples with alleles of other supertypes (other than the supertype(s) of the specific allele being trained), and positive samples randomly sampled from this set, by reversing the order of execution (until all other allele data had been filtered), but instead replaced the "external" allele with the currently trained allele. For the final predictor, we trained the following set of model types: 1. Pan models; and 2. Allele-specific models. These models were trained only on subsets of alleles with sufficient training data. Specifically, we trained only alleles with at least 1000 binding affinity training samples and 1000 presentation training samples.

[0203] For each type, ten similar models were trained on different training validation splits, and the model with the best performance on the immunogenicity validation set was selected. Hierarchical model selection was applied to determine which model(s) to use during prediction for each allele.

[0204] To verify whether multitask training is beneficial in a semi-inventor setting, the effect of the loss weight coefficient α was investigated. Three times the number of experiments were conducted on several general models with identical settings except for this hyperparameter, and the average WP@K for the immunogenicity test segment was reported. Results obtained by both affinity-prediction and presentation-prediction rankings were plotted to capture all effects. Figure 7 clearly shows that α=0 and α=1 (corresponding to presentation-only and affinity-only training, respectively) yield inferior results compared to intermediate region values ​​corresponding to joint training using weighted combinations of both objectives.

[0205] To test our hypothesis that the probability of presentation is a good surrogate for immunogenicity estimation, we plotted histograms of binned presentation predictions against the proportion of immunogenicity-positive samples in each bin. The positive slope seen in Figure 8A indicates that dP 免疫原性 / dP 提示 It was confirmed that as the probability of peptide presentation increases to >0, i.e., on average, the peptide is more likely to be immunogenic. For comparison, similar behavior regarding binding affinity prediction was also investigated, and a similar pattern was observed in Figure 8B. A performance comparison was also made between our model and other state-of-the-art predictors.

[0206] In detail, we compared the performance with O'Donnell et al.

[13] and its earlier version, MHCflurry-2.0. The above comparison was performed using our evaluation protocol and indices, and the reported values ​​were calculated on the immunogenicity test partition. As is evident from Table 1, our best predictor performed significantly better than the MHCflurry-2.0 predictor and all variations of its earlier versions. We were also able to demonstrate that our single pan-model performed better than a collection of allele-specific models (which is expected since the above collection does not support predictions for all alleles), and further performance was improved by combining the two, with the selection of hierarchical models. [Table 1]

[0207] Furthermore, we compared the allele-specific performance (for all alleles with immunogenicity data to be evaluated) between our best predictor, the "general-allele-specific presentation predictor," and the best predictor of MHCflurry-2.0, the "general(+ms) affinity predictor." The results are shown in Table 2. [Table 2-1] [Table 2-2] [Table 2-3]

[0208] References cited in the examples 1. Vita, R.;Mahajan, S.;Overton, J. A.;Dhanda, S.K.;Martini, S.;Cantrell, J.R.;Wheeler, D.K.;Sette, A.;Peters, B. The immune epitope database (IEDB): 2018 update. Nucleic acids research 2019, 47, D339-D343. 2. Kim, Y.;Sidney, J.;Buus, S.;Sette, A.;Nielsen, M.;Peters, B. Dataset size and composition impact the reliability of performance benchmarks for peptide-MHC binding predictions. BMC bioinformatics 2014, 15, 241. 3. O’Donnell, T.J.;Rubinsteyn, A.;Bonsack, M.;Riemer, A.B.;Laserson, U.;Hammerbacher, J. MHCflurry: open-source class I MHC binding affinity prediction. Cell systems 2018, 7, 129-132. 4. Sarkizova, S.;Klaeger, S.;Le, P.M.;Li, L.W.;Oliveira, G.;Keshishian, H.;Hartigan, C.R.;Zhang, W.;Braun, D.A.;Ligon, K.L.;others. A large peptidome dataset improves HLA class I epitope prediction across most of the human population. Nature Biotechnology 2020, 38, 199-209. 5. Shao, W.;Pedrioli, P.G.;Wolski, W.;Scurtescu, C.;Schmid, E.;Vizcaino, J. A.;Courcelles, M.;Schuster, H.;Kowalewski, D.;Marino, F.;others. The SysteMHC atlas project. Nucleic acids research 2018, 46, D1237-D1247. 6. Abelin, J.G.;Harjanto, D.;Malloy, M.;Suri, P.;Colson, T.;Goulding, S.P.;Creech, A.L.;Serrano, L.R.;Nasir, G.;Nasrallah, Y.;others. Defining HLA-II ligand processing and binding rules with mass spectrometry enhances cancer epitope prediction. Immunity 2019, 51, 766-779. 7. Sidney, J.;Peters, B.;Frahm, N.;Brander, C.;Sette, A. HLA class I supertypes: a revised and updated classification. BMC immunology 2008, 9, 1. 8. O’Donnell, T.;Rubinsteyn, A.;Laserson, U. Improved predictive models for peptide presentation on MHC I. BioRxiv 2020. 9. Consortium, U. UniProt: a worldwide hub of protein knowledge. Nucleic acids research 2019, 47, D506-D515. 10. Lin, T.Y.;Goyal, P.;Girshick, R.;He, K.;Dollar, P. Focal loss for dense object detection. Proceedings of the IEEE international conference on computer vision, 2017, pp. 2980-2988. 11. Mukhoti, J.;Kulharia, V.;Sanyal, A.;Golodetz, S.;Torr, P.H.;Dokania, P.K. Calibrating Deep Neural Networks using Focal Loss. arXiv preprint arXiv:2002.09437 2020. 12. Guo, C.;Pleiss, G.;Sun, Y.;Weinberger, K.Q. On calibration of modem neural networks. arXiv preprint arXiv: 1706.045992017. 13. O’Donnell, T.J.;Rubinsteyn, A.;Laserson, U. MHCflurry 2.0: Improved Pan-Allele Prediction of MHC Class I-Presented Peptides by Incorporating Antigen Processing. Cell Systems 2020. 14. Fischer, W.;Perkins, S.;Theiler, J.;Bhattacharya, T.;Yusim, K.;Funkhouser, R.;Kuiken, C.;Haynes, B.;Letvin, N.L.;Walker, B.D.;others. Polyvalent vaccines for optimal coverage of potential T-cell epitopes in global HIV-1 variants. Nature medicine 2007, 13, 100-106. 15. Jurtz, V.;Paul, S.;Andreatta, M.;Marcatili, P.;Peters, B.;Nielsen, M. NetMHCpan-4.0: improved peptide-MHC class I interaction predictions integrating eluted ligand and peptide binding affinity data. The Journal of Immunology 2017, 199, 3360-3368. 16. Consortium, U. UniProt: a worldwide hub of protein knowledge. Nucleic acids research 2019, 47, D506-D515.

[0209] Equivalents It will be readily apparent to those skilled in the art that other suitable modifications and adaptations of the methods of the invention described herein are obvious and can be made using suitable equivalents without departing from the scope of this disclosure or these embodiments. While specific compositions and methods have been described in detail so far, a clearer understanding of these compositions and methods will be gained by referring to the examples provided. These examples are presented for illustrative purposes only and are not intended to limit the scope.< / cls> < / cls> < / cls> < / sep> < / eos> < / cls>

Claims

1. A method for predicting the MHC class I immunogenicity of tumor-specific neoantigens, a) Obtaining the peptide sequence of a tumor-specific neoantigen and the corresponding adjacent region of the peptide sequence; encoding the peptide sequence and the adjacent region into a number vector, wherein each number vector includes the amino acid residues encoding the peptide of the tumor-specific neoantigen and the amino acid residues of the adjacent region, as well as the positions of the amino acid residues. b) Obtaining an HLA allele pseudosequence, wherein the HLA allele pseudosequence represents an HLA allele; and encoding the HLA allele sequence into a corresponding numerical vector. c) Using a neural network model, predict both the MHC class I binding affinity of the tumor-specific neoantigen and the numerical probability that the corresponding peptide for each tumor-specific neoantigen will be presented on the cell surface by the MHC class I protein, wherein the neural network model is (i) optimizing the performance of the neural network model by training it on a training dataset, wherein the training dataset includes a peptide-MHC class I affinity measurement dataset and a cell surface peptide presentation dataset, the training dataset includes positive training data and negative training data, the negative training data is generated by replacing the HLA alleles of the positive training data with different HLA alleles that do not belong to the supertype of the positive training data HLA alleles, and the optimization is performed accordingly. (ii) an input layer comprising the number vector including the peptide sequence and adjacent region of the tumor-specific neoantigen, and the number vector including the HLA allele pseudosequence layer, (iii) Encoding the numerical vectors including the peptide sequence and adjacent region of the tumor-specific neoantigen, and the numerical vectors including the HLA allele pseudosequence, into an amino acid embedding layer, (iv) Flattening the amino acid embedding layer to create a number vector representation of each peptide sequence of the tumor-specific neoantigen, the adjacent regions of the peptide sequences, and the pseudo-sequences of the HLA alleles, (v) Predicting the MHC class I binding affinity of the tumor-specific neoantigen by concatenating the peptide sequence of the tumor-specific neoantigen with the HLA allele pseudosequence and applying one or more layers and / or one or more activation functions, wherein the output is a numerical score representing the MHC class I binding affinity of the tumor-specific neoantigen, and the above prediction, (vi) predicting the probability that the tumor-specific neoantigen will be presented on the cell surface by an MHC class I protein by concatenating the target peptide sequence, the adjacent region of the peptide sequence, and the HLA allele pseudosequence in a single number vector and applying one or more layers and / or one or more activation functions, wherein the output is a numerical probability that the peptide will be presented on the cell surface by an MHC class I protein, the prediction includes, the prediction includes, The prediction method wherein the MHC class I binding affinity of the tumor-specific neoantigen and the probability expressed by the numerical value that the tumor-specific neoantigen will be presented on the cell surface by an MHC class I protein are surrogate values ​​for the MHC class I immunogenicity of the tumor-specific neoantigen.

2. (i) Applying one or more ranking indicators to the immunogenicity validation dataset, (ii) Ranking the peptides with respect to each allele in the immunogenicity validation dataset based on the predicted MHC class I binding affinity of the peptide and the probability expressed by the aforementioned numerical value that the peptide will be presented on the cell surface by an MHC class I protein, (iii) To aggregate one or more of the above ranking indicators for all alleles The method according to claim 1, further comprising verifying the neural network model by means of...

3. The method according to claim 2, wherein the one or more ranking indicators are aggregated using weighted allele frequencies.

4. The method according to any one of claims 1 to 3, wherein the neural network model is a pan-allele model, an allele-specific model, a super-type-specific model, or a combination thereof.

5. The method according to any one of claims 1 to 4, wherein the length of the HLA allele pseudosequence is 30 to 60 amino acids.

6. The method according to any one of claims 1 to 5, wherein the peptide sequence of the tumor-specific neoantigen is 8 amino acids to 15 amino acids in length.

7. The method according to any one of claims 1 to 6, wherein the adjacent region is immediately to the left of the peptide sequence of the tumor-specific neoantigen and / or immediately to the right of the peptide sequence of the tumor-specific neoantigen, and / or the length of the adjacent region is 10 amino acids, and / or the length of the adjacent region immediately to the left of the tumor-specific neoantigen is 5 amino acids, or the length of the adjacent region immediately to the right of the tumor-specific neoantigen is 5 amino acids.

8. The method according to any one of claims 1 to 7, further comprising performing calibration of the neural network model.

9. The method according to any one of claims 1 to 8, wherein the negative training data further comprises human rhinovirus (HRV) negative sampling, wherein the negative training data is a peptide.

10. The method according to any one of claims 1 to 9, wherein the negative training data includes a peptide that does not have MHC class I binding affinity for tumor-specific neoantigens and / or is not presented on the cell surface by MHC class I proteins.

11. The method according to any one of claims 1 to 10, wherein the HLA allele is HLA-A, HLA-B, or HLA-C.

12. The method according to any one of claims 1 to 11, wherein one or more tumor-specific neoantigens are selected for an immunogenic composition, wherein the MHC class I immunogenicity of the tumor-specific neoantigen is CD8+ T cell immunogenic, and / or is predicted to be MHC class I immunogenic, and 20 to 100 tumor-specific neoantigens are selected for the immunogenic composition.

13. The method according to any one of claims 1 to 12, wherein one or more of the layers are fully connected layers, and / or one or more of the layers are dropout layers.

14. The method according to any one of claims 1 to 13, wherein the one or more layers and / or the one or more activation functions include applying one or more fully connected layers, applying a dropout layer, and applying an activation function.

15. The method according to claim 8, wherein the neural network model is calibrated by probability calculation, and the probability calculation estimates the overall probability of presentation of the target allele.