Method and system for prediction of HLA epitopes

A machine-learning model predicts HLA peptide presentation to identify cancer-specific antigens, enhancing the development of immune-based therapies by achieving a positive predictive value of 0.2 in selecting effective peptide sequences for cancer treatment.

JP2025531662APending Publication Date: 2025-09-25BIONTECH US INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025507884
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-12
Filing Date
2023-08-11
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Current methods are inadequate for accurately predicting which cancer-specific antigens are likely to elicit T cell responses, hindering the development of effective immune-based therapies.

Method used

A method using a trained machine-learning HLA peptide presentation prediction model to identify and select peptide sequences presented by HLA alleles, based on amino acid sequence information and epitope presentation quantitative data, for use in cancer treatment.

Benefits of technology

The model achieves a positive predictive value of at least 0.2 in predicting peptide sequences presented by HLA proteins, enabling the selection of effective peptide sequences for cancer treatment, including pharmaceutical compositions and T cell therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531662000001_ABST
    Figure 2025531662000001_ABST
Patent Text Reader

Abstract

Methods for preparing personalized cancer vaccines and methods for training machine learning HLA peptide prediction models.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 397,669, filed August 12, 2022, which is incorporated herein by reference in its entirety. [Background technology]

[0002]

[0002] The major histocompatibility complex (MHC) is a gene complex that encodes human leukocyte antigen (HLA) genes. HLA genes are expressed on the surface of human cells as protein heterodimers that are presented to circulating T cells. HLA genes are highly polymorphic, allowing them to fine-tune the adaptive immune system. The adaptive immune response relies, in part, on the ability of T cells to identify and eliminate cells that display disease-associated peptide antigens bound to human leukocyte antigen (HLA) heterodimers.

[0003]

[0003] In humans, endogenous and exogenous proteins are processed into peptides by the proteasome and by cytosolic and endosomal / lysosomal proteases and peptidases, which can then be presented by two classes of cell surface proteins encoded by MHC genes. These cell surface proteins are called human leukocyte antigens (HLA class I and class II), and the groups of peptides that bind to them and elicit an immune response are named HLA epitopes. HLA epitopes are important components that enable the immune system to detect danger signals, such as pathogen infection and its own transformation. Typically, CD8+ T cells recognize MHC class I epitopes presented on antigen-presenting cells (APCs), such as dendritic cells and macrophages, while CD4+ T cells recognize class II MHC (HLA-DR, HLA-DQ, and HLA-DP) epitopes presented on APCs. The endogenous processing and presentation of HLA epitopes is a complex procedure, involving a diverse subset of chaperones and enzymes. HLA peptide presentation can activate cytotoxic and helper T cells, subsequently promoting B cell differentiation and antibody production as well as CTL responses.

[0004]

[0004] Understanding the peptide binding selectivity of all HLA class I or class II molecules is key to better predicting which cancer- or tumor-specific antigens are likely to elicit cancer- or tumor-specific T cell responses. Methods are needed to identify and isolate specific HLA class I- or class II-associated peptides (e.g., neoantigenic peptides). Such methods and isolated molecules are useful, for example, for the development of therapeutics, including, but not limited to, immune-based therapies. Summary of the Invention [Problem to be solved by the invention]

[0005]

[0005] Provided herein is a method for identifying peptide sequences presented by at least one of one or more proteins encoded by HLA alleles of cells of a subject, comprising: (a) inputting, using a computer processor, amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine-learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction of the plurality of presentation predictions indicates a likelihood that a peptide sequence of the set of candidate peptide sequences will be presented by an MHC protein of the single human subject; and wherein the trained machine-learning HLA peptide presentation prediction model: (i) generates a plurality of presentation predictions based on training data from training cells expressing the MHC protein. The method includes the steps of: (i) receiving a plurality of parameters, wherein the training data comprises a plurality of training peptide sequences and epitope presentation quantitative information, and the epitope presentation quantitative information comprises one or more amounts of the plurality of training peptide sequences presented by an MHC protein; and (ii) a function representing an association between amino acid sequence information received as input and a presentation probability generated as output based on the amino acid sequence information and the plurality of parameters; and (b) identifying, based at least on the plurality of presentation predictions, a peptide sequence of the plurality of peptide sequences of the set of candidate peptide sequences that is presented by at least one of the one or more proteins encoded by HLA alleles of the subject's cell.

[0006]

[0006] Provided herein is a method for selecting peptide sequences, the method comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, each presentation prediction of the plurality of presentation predictions indicating a presentation likelihood of a peptide sequence of the set of candidate peptide sequences being presented by an MHC protein of the single human subject; the trained machine learning HLA peptide presentation prediction model comprising: (i) a plurality of parameters based on training data from training cells expressing MHC proteins, the training data including a plurality of training peptide sequences and epitope presentation quantitative information, the epitope presentation quantitative information including one or more amounts of the plurality of training peptide sequences presented by the MHC protein; and (ii) a function representing the association between the amino acid sequence information received as input and the presentation likelihood generated as output based on the amino acid sequence information and the plurality of parameters; and (b) selecting a subset of peptide sequences of the set of candidate peptide sequences to generate a set of selected peptide sequences based at least on the plurality of presentation predictions.

[0007]

[0007] Provided herein is a method of treating cancer in a human subject in need thereof, comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine-learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, each presentation prediction indicating a likelihood that a peptide sequence of the set of candidate peptide sequences will be presented by an MHC protein of the single human subject; and the trained machine-learning HLA peptide presentation prediction model comprising: (i) a plurality of parameters based on training data from training cells expressing MHC proteins, the training data comprising a plurality of training peptide sequences and epitope presentation quantitative information, the epitope presentation quantitative information comprising one or more amounts of the plurality of training peptide sequences presented by the MHC protein. and (ii) a function representing an association between amino acid sequence information received as input and presentation probabilities generated as output based on the amino acid sequence information and a plurality of parameters; (b) selecting or identifying a subset of peptide sequences of the set of candidate peptide sequences to generate a set of selected or identified peptide sequences based at least on the plurality of presentation predictions; and (c) administering to a single human subject a pharmaceutical composition comprising: (i) a polypeptide having one or more of the selected peptide sequences, (ii) a polynucleotide encoding the polypeptide of (i); (iii) an APC comprising (i) or (ii), or (iv) a T cell comprising a T cell receptor (TCR) specific for an MHC protein of the single human subject in a complex with one or more of the peptide sequences selected or identified in (b).

[0008]

[0008] In some embodiments, the plurality of parameters is based on training data from training cells expressing MHC proteins of a single human subject.

[0009] In some embodiments, each training peptide sequence of the plurality associates with an MHC protein.

[0009]

[0010] In some embodiments, the training data comprises the identity of the MHC protein associated with each of the plurality of training peptide sequences.

[0011] In some embodiments, the training data comprises mass spectrometry observations that one or more of a plurality of training peptide sequences were presented by an MHC protein.

[0010]

[0012] In some embodiments, the MHC protein of the single human subject is a class I MHC protein.

[0013] In some embodiments, a plurality of candidate peptide sequences expressed by cancer cells of a single human subject are identified by comparing whole genome or whole exome sequence information from cancer cells of a single human subject with whole genome or whole exome sequence information from non-cancerous cells of a single human subject, and identifying nucleic acid sequences that are unique to the cancer cells and absent from the non-cancerous cells.

[0011]

[0014] In some embodiments, each candidate sequence of the plurality of candidate peptide sequences comprises a cancer-specific mutation.

[0015] In some embodiments, the trained machine learning HLA peptide presentation predictive model has a peptide presentation predictive value (PPV) of at least 0.2 according to a presentation PPV determination method.

[0012]

[0016] In some embodiments, the presentation PPV determination method includes inputting amino acid sequence information of a plurality of test peptide sequences into a trained machine learning HLA peptide presentation prediction model to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by HLA alleles may present a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences includes at least 500 test peptide sequences, including: (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained within proteins encoded by the genome of an organism, wherein the organism and the subject are of the same species.

[0013]

[0017] In some embodiments, the plurality of test peptide sequences has a 1:499 ratio of at least one hit peptide sequence to at least 499 decoy peptide sequences, and the top 0.2% of the plurality of test peptide sequences are predicted to be presented by HLA proteins expressed in cells by a trained machine learning HLA peptide presentation prediction model.

[0014]

[0018] In some embodiments, (i) the at least one hit peptide sequence comprises at least 10 hit peptide sequences, and (ii) the at least 499 decoy peptide sequences comprise at least 4,990 decoy peptide sequences.

[0015]

[0019] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the number of copies of one or more of the plurality of training peptide sequences presented by the MHC protein.

[0016]

[0020] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the number of one or more copies per cell of the plurality of training peptide sequences presented by the MHC protein.

[0017]

[0021] In some embodiments, the amount of one or more of the plurality of training peptide sequences presented by the MHC protein comprises the absolute amount, number of molecules, density, concentration, absolute amount per cell, number of molecules per cell, density per cell, or concentration in a cell of one or more of the plurality of training peptide sequences presented by the MHC protein.

[0018]

[0022] In some embodiments, the quantity of one or more of the plurality of training peptide sequences presented by the MHC proteins is based on a number of mass spectrometry observations, spectral counts, area under the curve (AUC), intensity-based absolute quantification (iBAQ), label-free quantification (LFQ), isotope dilution mass spectrometry, isobaric mass tagging, stable isotope labeling, and / or mass spectrometry peak intensities.

[0019]

[0023] In some embodiments, the quantity of one or more of the plurality of training peptide sequences presented by the MHC proteins is obtained from quantitative mass spectrometry.

[0024] In some embodiments, epitope presentation quantitative information is obtained from internal standard parallel reaction monitoring (IS-PRM) mass spectrometry.

[0020]

[0025] In some embodiments, epitope presentation quantitative information is obtained from xenograft samples.

[0026] In some embodiments, the xenograft sample is a patient-derived xenograft (PDX) sample.

[0021]

[0027] Also provided herein is a method for selecting peptide sequences, comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine-learning HLA-peptide antigen-specific T cell prediction model to generate a plurality of antigen-specific T cell predictions, wherein each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence from the set of candidate peptide sequences will stimulate T cells to be specific for a peptide sequence from the set of candidate peptide sequences; and wherein the trained machine-learning HLA-peptide cytotoxic T cell prediction model: (i) inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine-learning HLA-peptide antigen-specific T cell prediction model to generate a plurality of antigen-specific T cell predictions, wherein each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence from the set of candidate peptide sequences will stimulate T cells to be specific for a peptide sequence from the set of candidate peptide sequences; and (b) selecting, based at least on the plurality of antigen-specific T cell predictions, a subset of peptide sequences from the set of candidate peptide sequences to generate a set of selected peptide sequences.

[0022]

[0028] In some embodiments, each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of a single human subject and a peptide sequence from the set of candidate peptide sequences will stimulate a T cell to be specific for a neo-antigenic peptide sequence from the set of candidate peptide sequences.

[0023]

[0029] In some embodiments, the function is a function that represents the association between amino acid sequence information received as input and the likelihood that T cells specific to a neo-antigenic peptide sequence of the set of candidate peptide sequences will be generated as output based on the amino acid sequence information and multiple parameters.

[0024]

[0030] In some embodiments, each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of a single human subject and a peptide sequence from the set of candidate peptide sequences will stimulate a T cell to be cytotoxic.

[0025]

[0031] In some embodiments, the function is a function that represents the association between amino acid sequence information received as input and the likelihood that cytotoxic T cells will be generated as output based on the amino acid sequence information and multiple parameters.

[0026] Incorporation by Reference

[0032] All publications, patents, and patent applications mentioned herein are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained in the specification, the specification supersedes and / or takes precedence over any such conflicting material.

[0027]

[0033] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "FIG."). The drawings are as follows: [Brief explanation of the drawings]

[0028] [Figure 1A]

[0034] FIG. 1A shows data illustrating the results of an evaluation using a holdout partition of a monoallelic dataset. [Figure 1B]

[0035] FIG. 1B shows data showing the performance of predictors in ovarian tumors profiled by MS. [Figure 1C]

[0036] FIG. 1C shows the RECON presentation score. [Figure 2A]

[0037] FIG. 2A shows a diagram illustrating an exemplary workflow for generating patient-derived xenografts for targeted MS. [Figure 2B]

[0038] Figure 2B shows a diagram showing the sequence overlap of non-synonymous mutations. [Figure 3A]

[0039] FIG. 3A shows an exemplary workflow of a method for validation of predicted neoantigens using internal standard triggered parallel reaction monitoring. [Figure 3B]

[0040] Figure 3B shows data from validation of each of the predicted neoantigens by parallel reaction monitoring. [Figure 4]

[0041] FIG. 4 shows the RECON® presentation scores for epitopes targeted by MS. [Figure 5A]

[0042] Figure 5A shows an exemplary workflow for quantification of predicted neoantigens. [Figure 5B]

[0043] Figure 5B shows the quantification of predicted neoantigens. [Figure 5C]

[0044] Figure 5C shows the quantification of predicted neoantigens. [Figure 6A]

[0045] FIG. 6A shows the observed T cell responses to neoantigens. [Figure 6B]

[0046] FIG. 6B shows immune monitoring correlation data. [Figure 7]

[0047] FIG. 7 shows the binding affinity correlation data. [Figure 8]

[0048] Figure 8 shows the clinical outcome data. DETAILED DESCRIPTION OF THE INVENTION

[0029]

[0049] All terms are intended to be understood as they are understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0030]

[0050] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0051] Although various features of the present disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described herein for clarity in the context of separate embodiments, the disclosure may also be practiced in a single embodiment.

[0031]

[0052] The method and composition described herein can be used in a wide range of applications.For example, the method and composition described herein can be used to identify immunogenic antigen peptide, develop drugs (for example, personalized medicines), and isolate and characterize antigen-specific T cells.

[0032]

[0053] The method disclosed herein may include generating LC-MS / MS allele data for training an allele-specific machine learning method for epitope prediction.Such method may include: improving LC-MS / MS data quality by utilizing a series of quality metrics that strictly eliminate false positives and improve the performance of prediction models; identifying allele-specific HLA class I or class II binding cores from HLA-ligandome LC-MS / MS data sets; utilizing machine learning algorithms to improve HLA class I or class II-ligand and epitope prediction; and / or identifying biological variables that affect HLA class I or class II-ligand presentation and improve HLA class I or class II epitope prediction, such as gene expression, cleavage, gene bias, cellular localization and secondary structure.

[0033]

[0054] Provided herein are methods including: (a) processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each candidate peptide sequence of the plurality of candidate peptide sequences is encoded by a subject's genome or exome, the plurality of presentation predictions including an HLA presentation prediction for each of the plurality of candidate peptide sequences, each HLA presentation prediction indicating a likelihood that one or more proteins encoded by class I or class II HLA alleles of a cell of the subject may present a given candidate peptide sequence of the plurality of candidate peptide sequences, the machine learning HLA peptide presentation prediction model being trained using training data including sequence information for sequences of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in the training cells; and (b) identifying, based on at least the plurality of presentation predictions, a peptide sequence of the plurality of peptide sequences that is presented by at least one of the one or more proteins encoded by class I or class II HLA alleles of the cell of the subject, wherein the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 according to a presentation PPV determination method.

[0034]

[0055] Provided herein are methods including: (a) processing amino acid information for a plurality of peptide sequences encoded by a subject's genome or exome using a machine-learning HLA peptide binding prediction model to generate a plurality of binding predictions, wherein the plurality of binding predictions include an HLA binding prediction for each of a plurality of candidate peptide sequences, each binding prediction indicating the likelihood that one or more proteins encoded by class II HLA alleles of the subject's cells will bind to a given candidate peptide sequence of the plurality of candidate peptide sequences, and wherein the machine-learning HLA peptide binding prediction model is trained using training data including sequence information for sequences of peptides identified to bind to HLA class I or class II proteins or HLA class I or class II protein analogs; and (b) identifying, based at least on the plurality of binding predictions, peptide sequences of the plurality of peptide sequences that have a higher probability of binding to at least one of the one or more proteins encoded by class I or class II HLA alleles of the subject's cells than a binding prediction probability value threshold, wherein the machine-learning HLA peptide binding prediction model has a positive predictive value (PPV) of at least 0.1 according to a binding PPV determination method.

[0035]

[0056] In some embodiments, the machine learning HLA peptide presentation prediction model is trained using training data that includes sequence information for sequences of training peptides identified by mass spectrometry as being presented by HLA proteins expressed on training cells.

[0036]

[0057] In some embodiments, the method includes a step of ranking at least two peptides identified as being presented by at least one of the one or more proteins encoded by class I or class II HLA alleles of the subject's cells based on presentation prediction.

[0037]

[0058] In some embodiments, the method comprises selecting one or more peptides of the two or more ranked peptides.

[0059] In some embodiments, the method includes selecting one or more of a plurality of peptides identified as being presented by at least one of the one or more proteins encoded by class I or class II HLA alleles of the subject's cells.

[0038]

[0060] In some embodiments, the method comprises selecting one or more peptides from two or more peptides ranked based on the presentation prediction.

[0061] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 when amino acid information for a plurality of test peptide sequences is processed to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by class I or class II HLA alleles in a cell of a subject may present a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences includes at least 500 test peptide sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 499 decoy peptide sequences contained within a protein encoded by the genome of an organism, wherein the organism and the subject are the same species, and wherein the plurality of test peptide sequences has a 1:499 ratio of the at least one hit peptide sequence to the at least 499 decoy peptide sequences, and wherein the top percentage of the plurality of test peptide sequences are predicted by the machine learning HLA peptide presentation prediction model to be presented by an HLA protein expressed in the cell.

[0039]

[0062] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.1 when amino acid information for a plurality of test peptide sequences is processed to generate a plurality of test binding predictions, each test binding prediction indicating the likelihood that one or more proteins encoded by class I or class II HLA alleles of a cell of the subject will bind to a given test peptide sequence in the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, e.g., at least 20 test peptide sequences comprising at least 19 decoy peptide sequences contained within a protein comprising a single HLA protein expressed in the cell (e.g., a single-allelic cell), wherein the plurality of test peptide sequences comprises a 1:19 ratio of at least 19 decoy peptide sequences to at least one hit peptide sequence, and wherein the top percentage of the plurality of test peptide sequences are predicted by the machine learning HLA peptide presentation prediction model to bind to an HLA protein expressed in the cell.

[0040]

[0063] In some embodiments, no amino acid sequence overlap exists within at least one of the hit peptide sequence and the decoy peptide sequence.

[0064] In some embodiments, the machine learning HLA peptide presentation prediction model has a mean value of at least 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99.

[0041]

[0065] In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences.

[0042]

[0066] In some embodiments, the at least 499 decoy peptide sequences are at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 480 0, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 730 0, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 980 0, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 2900 0, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 500 00, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 1 25000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000,The hit:decoy ratio may be varied to vary the PPV.

[0043]

[0067] In some embodiments, the at least 500 test peptide sequences are at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, , 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400 , 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900 , 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 3000 0, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 525 00, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000,Contains 900,000 or 1,000,000 test peptide sequences.

[0044]

[0068] In some embodiments, the top percentage is the top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8. 0%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.0 0%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%, 9.90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%.

[0045]

[0069] In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences.

[0046]

[0070] In some embodiments, the at least 19 decoy peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 120 0, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 370 0, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 620 0, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 87 00, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20 000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 4 1000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500,The decoy peptide sequences include 80,000, 82,500, 85,000, 87,500, 90,000, 92,500, 95,000, 97,500, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 325,000, 350,000, 375,000, 400,000, 425,000, 450,000, 475,000, 500,000, 600,000, 700,000, 800,000, 900,000 or 1,000,000 decoy peptide sequences.

[0047]

[0071] In some embodiments, the at least 20 test peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5 , 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300 , 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17 000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 3 8000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000,At least 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000 test peptide sequences.

[0048]

[0072] In some embodiments, the top percentage is the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39% or 40%.

[0049]

[0073] In some embodiments, the subject is a single subject.

[0074] In some embodiments, the subject is a mammal.

[0075] In some embodiments, the subject is a human.

[0050]

[0076] In some embodiments, the training cells are cells that express a single protein encoded by a class I or class II HLA allele of the subject's cells.

[0077] In some embodiments, the training cells are monoallelic HLA cells or cells expressing an HLA allele that includes an affinity tag.

[0051]

[0078] In some embodiments, the cells of the subject comprise cancer cells.

[0079] In some embodiments, the method is for identifying a peptide sequence.

[0080] In some embodiments, the method is for selecting a peptide sequence.

[0052]

[0081] In some embodiments, the method is for preparing a cancer treatment.

[0082] In some embodiments, the method is for preparing a target-specific cancer treatment.

[0083] In some embodiments, the method is for preparing a cancer cell-specific cancer therapy.

[0053]

[0084] In some embodiments, each peptide sequence of the plurality of peptide sequences is associated with cancer.

[0085] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is overexpressed by cancer cells of the subject.

[0054]

[0086] In some embodiments, each peptide sequence of the plurality of peptide sequences is overexpressed by cancer cells of the subject.

[0087] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is a cancer cell-specific peptide.

[0055]

[0088] In some embodiments, each peptide sequence of the plurality of peptide sequences is a cancer cell-specific peptide.

[0089] In some embodiments, each peptide sequence of the plurality of peptide sequences is expressed by a cancer cell of the subject.

[0056]

[0090] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is not encoded by non-cancer cells of the subject.

[0091] In some embodiments, each peptide sequence of the plurality of peptide sequences is not encoded by non-cancer cells of the subject.

[0057]

[0092] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is not expressed by non-cancer cells of the subject.

[0093] In some embodiments, each peptide sequence of the plurality of peptide sequences is not expressed by non-cancer cells of the subject.

[0058]

[0094] In some embodiments, the method comprises obtaining a plurality of peptide sequences of interest.

[0095] In some embodiments, the method comprises obtaining a plurality of polynucleotide sequences of the subject.

[0059]

[0096] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences for a subject that encode a plurality of peptide sequences encoded by the genome or exome of the subject, or by a pathogen or virus in the subject.

[0060]

[0097] In some embodiments, the method includes obtaining, by a computer processor, a plurality of polynucleotide sequences for the subject that encode a plurality of peptide sequences encoded by the genome or exome of the subject.

[0061]

[0098] In some embodiments, the method comprises obtaining a plurality of polynucleotide sequences of the subject by genomic or exome sequencing.

[0099] In some embodiments, the method comprises obtaining a plurality of polynucleotide sequences of the subject by whole genome sequencing or by whole exome sequencing.

[0062]

[0100] In some embodiments, the processing step comprises processing by a computer processor.

[0101] In some embodiments, the processing step includes generating a plurality of predictor variables based at least on amino acid information of the plurality of peptide sequences.

[0063]

[0102] In some embodiments, processing a plurality of predictor variables using a machine learning HLA peptide presentation predictive model.

[0103] In some embodiments, the one or more proteins encoded by class I or class II HLA alleles in the subject's cells are one or more proteins encoded by class I or class II HLA alleles expressed by the subject.

[0064]

[0104] In some embodiments, the one or more proteins encoded by class I or class II HLA alleles of the subject's cells are one or more proteins encoded by class I or class II HLA alleles expressed by the subject's cancer cells.

[0065]

[0105] In some embodiments, the one or more proteins encoded by class I or class II HLA alleles of the subject's cells is a single protein encoded by class I or class II HLA alleles of the subject's cells.

[0066]

[0106] In some embodiments, the one or more proteins encoded by class II HLA alleles of the subject's cells are two, three, four, five, or more proteins encoded by class I or class II HLA alleles of the subject's cells.

[0067]

[0107] In some embodiments, the one or more proteins encoded by class I or class II HLA alleles of the subject's cells is each protein encoded by class I or class II HLA alleles of the subject's cells.

[0068]

[0108] In some embodiments, the method further comprises administering to the subject a composition comprising one or more selected subsets of peptide sequences.

[0109] In some embodiments, identifying a plurality of peptide sequences comprises comparing DNA, RNA, or protein sequences from cancer cells of the subject to DNA, RNA, or protein sequences from normal cells of the subject, wherein each of the plurality of peptides comprises at least one mutation that is present in the cancer cells of the subject and not present in the normal cells of the subject.

[0069]

[0110] In some embodiments, the machine learning HLA peptide presentation prediction model includes a plurality of predictor variables identified based at least on training data, the training data including training peptide sequence information including amino acid position information, the training peptide sequence information associated with HLA proteins expressed in cells; and a function representing the association between the amino acid position information and presentation probability generated as output based on the amino acid position information and the plurality of predictor variables.

[0070]

[0111] In some embodiments, the identifying step comprises identifying, based on at least the plurality of presentation predictions, a peptide sequence of the plurality of peptide sequences that has a probability higher than a presentation prediction probability value threshold to be presented by at least one of the one or more proteins encoded by class I or class II HLA alleles of the subject's cell.

[0071]

[0112] In some embodiments, one or more of 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model have a probability higher than a presentation prediction probability value threshold to be presented by at least one of the one or more proteins encoded by class I or class II HLA alleles of the subject's cells.

[0072]

[0113] In some embodiments, 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model each have a probability greater than a presentation prediction probability value threshold of being presented by at least one of the one or more proteins encoded by class I or class II HLA alleles of the subject's cells.

[0073]

[0114] In some embodiments, the number of positives is constrained to be equal to the number of hits.

[0115] In some embodiments, the mass spectrometry is single allele mass spectrometry.

[0116] In some embodiments, the peptides are presented by HLA proteins expressed in cells through autophagy.

[0074]

[0117] In some embodiments, the peptides are presented by HLA proteins expressed on cells through phagocytosis.

[0118] In some embodiments, the plurality of predictor variables comprises an expression level predictor of a source protein comprising a peptide.

[0075]

[0119] In some embodiments, the plurality of predictor variables comprises a stability predictor of a source protein that includes the peptide.

[0120] In some embodiments, the plurality of predictor variables comprises a degradation rate predictor of a source protein that includes the peptide.

[0076]

[0121] In some embodiments, the plurality of predictor variables comprises a proteolytic cleavability predictor of a source protein that comprises the peptide.

[0122] In some embodiments, the plurality of predictor variables comprises a cellular or tissue localization predictor of a source protein comprising a peptide.

[0077]

[0123] In some embodiments, the plurality of predictor variables comprises predictors for the intracellular processing mode of a source protein comprising a peptide, wherein the processing mode of the source protein comprises predictors for whether the source protein is subject to, among other things, autophagy, phagocytosis, and intracellular trafficking.

[0078]

[0124] In some embodiments, the quality of the training data is improved by using multiple quality metrics.

[0125] In some embodiments, the multiple quality metrics include removal of common contaminating peptides, high scored peak intensity, high score, and high mass accuracy.

[0079]

[0126] In some embodiments, the scored peak intensity is at least 50%.

[0127] In some embodiments, the scored peak intensity is at least 60%.

[0128] In some embodiments, the score is at least 7.

[0080]

[0129] In some embodiments, the mass accuracy is up to 5 ppm.

[0130] In some embodiments, the peptides presented by an HLA protein expressed on a cell are peptides presented by a single immunoprecipitated HLA protein expressed on a cell.

[0081]

[0131] In some embodiments, the peptides presented by an HLA protein expressed on the cell are peptides presented by a single exogenous HLA protein expressed on the cell.

[0082]

[0132] In some embodiments, the peptides presented by an HLA protein expressed on a cell are peptides presented by a single recombinant HLA protein expressed on a cell.

[0083]

[0133] In some embodiments, the plurality of predictors comprises a peptide-HLA affinity predictor.

[0134] In some embodiments, peptides presented by HLA proteins include peptides identified by searching non-enzyme specific peptide databases that do not have the modification.

[0084]

[0135] In some embodiments, peptides presented by HLA proteins include peptides identified by searching peptide databases using a reverse database search strategy.

[0085]

[0136] In some embodiments, peptides presented by HLA proteins include peptides identified by comparing the MS / MS spectrum of an HLA peptide with the MS / MS spectrum of one or more peptides or proteins in a peptide or protein database.

[0086]

[0137] In some embodiments, the mutation is selected from the group consisting of a point mutation, a splice site mutation, a frameshift mutation, a readthrough mutation, and a gene fusion mutation.

[0138] In some embodiments, the peptides presented by the HLA proteins have a length of 8 to 12 or 15 to 40 amino acids.

[0087]

[0139] In some embodiments, peptides presented by HLA proteins include peptides identified by identifying peptides presented by HLA proteins by comparing the MS / MS spectrum of the HLA peptide with the MS / MS spectrum of one or more peptides or proteins in a peptide or protein database.

[0088]

[0140] In some embodiments, the personalized cancer treatment further comprises an adjuvant.

[0141] In some embodiments, the personalized cancer treatment further comprises an immune checkpoint inhibitor.

[0089]

[0142] In some embodiments, the training data includes structured data, time series data, unstructured data, relational data, or any combination thereof.

[0143] In some embodiments, the unstructured data includes image data.

[0090]

[0144] In some embodiments, the relational data includes data from customer systems, enterprise systems, operational systems, websites, web-accessible application program interfaces (APIs), or any combination thereof.

[0091]

[0145] In some embodiments, the training data is uploaded to a cloud-based database.

[0146] In some embodiments, the training is performed using a convolutional neural network.

[0092]

[0147] In some embodiments, the convolutional neural network includes at least two convolutional layers.

[0148] In some embodiments, the convolutional neural network includes at least one batch normalization step.

[0093]

[0149] In some embodiments, the convolutional neural network includes at least one spatial dropout step.

[0150] In some embodiments, the convolutional neural network includes at least one global max pooling step.

[0094]

[0151] In some embodiments, the convolutional neural network includes at least one densely connected layer.

[0152] In some embodiments, identifying peptide sequences comprises identifying peptide sequences having mutations that are expressed in cancer cells of the subject.

[0095]

[0153] In some embodiments, identifying a peptide sequence comprises identifying a peptide sequence that is not expressed in normal cells of the subject.

[0154] In some embodiments, identifying a peptide sequence comprises identifying a viral peptide sequence.

[0096]

[0155] In some embodiments, identifying peptide sequences comprises identifying overexpressed peptide sequences.

[0156] Provided herein is a method for identifying HLA class I or class II specific peptides for immunotherapy of a subject, the method comprising: obtaining, by a computer processor, candidate peptides comprising an epitope and a plurality of peptide sequences each comprising the epitope; processing, by the computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences to an immune cell, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by HLA class I or class II alleles will be able to present a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is presented by HLA proteins expressed in cells and identified by mass spectrometry. The method includes the steps of: training using training data including sequence information of peptide sequences; selecting proteins predicted by the machine learning HLA peptide presentation prediction model to bind to a candidate peptide from one or more proteins encoded by HLA class I or class II alleles of the subject's cells, wherein the proteins have a probability of presenting the candidate peptide to immune cells higher than a presentation prediction probability value threshold; contacting the candidate peptide with the selected protein so that the candidate peptide competes with a placeholder peptide that associates with the selected protein; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder.

[0097]

[0157] In some embodiments, the obtaining step comprises identifying candidate peptides, wherein identifying candidate peptides comprises comparing DNA, RNA or protein sequences from the subject's cancer cells with DNA, RNA or protein sequences from the subject's normal cells.

[0098]

[0158] In some embodiments, the processing step comprises identifying a plurality of predictor variables based on at least amino acid information of the plurality of peptide sequences, and processing the plurality of predictor variables using a machine learning HLA peptide presentation prediction model.

[0099]

[0159] In some embodiments, the machine learning HLA peptide presentation prediction model comprises a plurality of predictor variables identified based at least on training data, the training data comprising: training peptide sequence information including amino acid position information, the training peptide sequence information associated with HLA proteins expressed in cells; and a function representing the association between the amino acid position information and presentation probability generated as an output based on the amino acid position information and the plurality of predictor variables.

[0100]

[0160] In some embodiments, the number of positives is constrained to be equal to the number of hits.

[0161] In some embodiments, the mass spectrometry is single allele mass spectrometry.

[0162] In some embodiments, the plurality of predictor variables comprises any one or more of: an expression level predictor of the source protein comprising the peptide, a stability predictor, a degradation rate predictor, a cleavability predictor, a cellular or tissue localization predictor, and an intracellular processing mode including autophagy, phagocytosis, and intracellular trafficking predictor.

[0101]

[0163] In some embodiments, the quality of the training data is improved by using multiple quality metrics.

[0164] In some embodiments, the multiple quality metrics include removal of common contaminating peptides, high scored peak intensity, high score, and high mass accuracy.

[0102]

[0165] In some embodiments, the scored peak intensity is at least 50%.

[0166] In some embodiments, the scored peak intensity is at least 60%.

[0167] In some embodiments, the placeholder peptide is a CLIP peptide.

[0103]

[0168] In some embodiments, the placeholder peptide is a CMV peptide.

[0169] In some embodiments, the method comprises: determining the IC of replacement of a placeholder peptide with a target peptide; 50 The method further includes measuring:

[0104]

[0170] In some embodiments, the IC of the replacement of the placeholder peptide with the target peptide 50 is less than 500 nM.

[0171] In some embodiments, the target peptide is further identified by mass spectrometry.

[0105]

[0172] In some embodiments, at least one protein encoded by an HLA class I or class II allele of the subject's cells is a recombinant protein.

[0173] In some embodiments, at least one protein encoded by an HLA class I or class II allele of the subject's cell is expressed in a eukaryotic cell.

[0106]

[0174] In some embodiments, the peptides are presented by HLA proteins expressed in cells through autophagy.

[0175] In some embodiments, the peptides are presented by HLA proteins expressed on cells through phagocytosis.

[0107]

[0176] In some embodiments, the peptides presented by an HLA protein expressed on a cell are peptides presented by a single immunoprecipitated HLA protein expressed on a cell.

[0108]

[0177] In some embodiments, the peptides presented by an HLA protein expressed on the cell are peptides presented by a single exogenous HLA protein expressed on the cell.

[0109]

[0178] In some embodiments, the peptides presented by an HLA protein expressed on a cell are peptides presented by a single recombinant HLA protein expressed on a cell.

[0110]

[0179] In some embodiments, the plurality of predictors comprises a peptide-HLA affinity predictor.

[0180] In some embodiments, peptides presented by HLA proteins include peptides identified by searching non-enzyme specific peptide databases that do not have the modification.

[0111]

[0181] In some embodiments, peptides presented by HLA proteins include peptides identified by searching peptide databases using a reverse database search strategy.

[0112]

[0182] In some embodiments, the immunotherapy is cancer immunotherapy.

[0183] In some embodiments, the epitope is a cancer-specific epitope.

[0184] In some embodiments, the identity of the peptide is known.

[0113]

[0185] In some embodiments, the identity of the peptide is unknown.

[0186] In some embodiments, the identity of the peptide is determined by mass spectrometry.

[0187] In some embodiments, the peptide exchange assay comprises detection of a peptide fluorescent probe or tag.

[0114]

[0188] In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide has the amino acid sequence PVSKMRMATPLLMQA (SEQ ID NO: 1).

[0115]

[0189] In some embodiments, the polynucleic acid construct comprises an expression vector, further comprising one or more of: a promoter, a secretion signal, a dimerization factor, a ribosome skipping sequence, one or more tags for purification and / or detection.

[0116]

[0190] In some embodiments, the placeholder peptide sequence is encoded by a nucleic acid sequence in a vector.

[0191] In some embodiments, the sequence encoding the cleavable domain is positioned between the sequence encoding the placeholder peptide and the HLA beta 1 peptide.

[0117]

[0192] Provided herein is a method for assaying the immunogenicity of an MHC class I or class II-binding peptide, the method comprising: selecting proteins encoded by HLA class I or class II alleles that are predicted to bind to the MHC class I or class II-binding peptide by a machine-learning HLA peptide presentation prediction model, the machine-learning HLA peptide presentation prediction model being configured to generate a presentation prediction for a given peptide sequence, the presentation prediction indicating the likelihood that one or more proteins encoded by the HLA class II alleles can present the given peptide sequence, the proteins having a probability for presenting the MHC class I or class II-binding peptide that is higher than a presentation prediction probability value threshold; contacting the peptide with the selected protein such that the peptide competes with the placeholder peptide for association with the selected protein, displacing the placeholder peptide, thereby forming a complex comprising the HLA class I or class II protein and the MHC class I or class II-binding peptide; contacting the complex with CD4+ T cells and assaying one or more activation parameters of the CD4+ T cells selected from the group consisting of cytokine induction, chemokine induction, and cell surface marker expression.

[0118]

[0193] Provided herein is a method for inducing CD4+ T cell activation in a subject for cancer immunotherapy, the method comprising: identifying a peptide sequence associated with cancer and comprising a cancer mutation, wherein identifying the peptide sequence comprises comparing a DNA, RNA, or protein sequence from the subject's cancer cells with a DNA, RNA, or protein sequence from the subject's normal cells; selecting a protein encoded by an HLA class I or HLA class II allele that is normally expressed by the subject's cells and predicted by a machine learning HLA peptide presentation prediction model to bind to the peptide, wherein the prediction model has a positive predictive value of at least 0.1, with a recall of at least 0.1%, 0.1% to 50%, or up to 50%, and the protein has a probability higher than a presentation prediction probability value threshold for presenting the identified peptide sequence; and selecting a placeholder peptide that associates with the selected protein encoded by the HLA class I or HLA class II allele with an IC of less than 500 nM to replace the placeholder peptide. 50 The method includes contacting the identified peptide with a selected protein encoded by an HLA class I or HLA class II allele to verify whether the identified peptide competes with the selected protein; optionally purifying the identified peptide; and administering to a subject an effective amount of a polypeptide comprising the sequence of the identified peptide or a polynucleotide encoding the polypeptide.

[0119]

[0194] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by the computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information associated with HLA proteins expressed in the cell; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic to the subject; and administering to the subject a composition comprising the drug.

[0120]

[0195] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by the computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of peptide sequences presented by HLA proteins expressed in the cell and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence will be immunogenic to the subject.

[0121]

[0196] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor, inputting amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction represents the probability that one or more proteins encoded by class I or II MHC alleles of the subject's cells will present an epitope sequence of the given peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, the training data including: sequence information of sequences of peptides presented by HLA proteins expressed in the cell and identified by mass spectrometry, training peptide sequence information including amino acid position information, the training peptide sequence information associated with the HLA proteins expressed in the cell, and a function representing the association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables; determining or predicting that each of the peptide sequences of the polypeptide sequence will not be immunogenic to the subject based on the set of presentation predictions; and administering a composition comprising the drug to the subject.

[0122]

[0197] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor, inputting amino acid information of peptide sequences of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction represents the probability that one or more proteins encoded by class I or II MHC alleles of the subject's cells will present an epitope sequence of the given peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, wherein the training data includes: sequence information of sequences of peptides presented by HLA proteins expressed in the cell and identified by mass spectrometry, training peptide sequence information including amino acid position information, the training peptide sequence information associated with the HLA proteins expressed in the cell, and a function representing the association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables; and determining or predicting that at least one of the peptide sequences of the polypeptide sequence will be immunogenic to the subject based on the set of presentation predictions.

[0123]

[0198] Provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by the computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information associated with HLA proteins expressed in the cell; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic to the subject; and administering to the subject a composition comprising the drug.

[0124]

[0199] In some embodiments, the method further comprises deciding not to administer the drug to the subject.

[0200] In some embodiments, the drug comprises an antibody or a binding fragment thereof.

[0125]

[0201] In some embodiments, the peptide sequence of the polypeptide sequence has a length of 8, 9, 10, 11, or 12 amino acids, and wherein the protein encoded by a class I or II MHC allele of the subject's cell is a protein encoded by a class I MHC allele of the subject's cell.

[0126]

[0202] In some embodiments, the peptide sequence of the polypeptide sequence has a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, and wherein the protein encoded by a class I or II MHC allele of the subject's cell is a protein encoded by a class II MHC allele of the subject's cell.

[0127]

[0203] Provided herein are methods of treating a subject having an autoimmune disease or condition, the methods comprising: (a) identifying or predicting an epitope of an expressed protein that is presented by class I or II MHC of a cell of the subject, wherein a complex comprising the identified or predicted epitope and class I or II MHC is targeted by CD8 or CD4 T cells of the subject; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells from the subject or allogeneic regulatory T cells; and (d) administering the TCR-expressing regulatory T cells to the subject.

[0128]

[0204] Provided herein are methods of treating a subject having an autoimmune disease or condition, the methods comprising administering to the subject regulatory T cells that express a T cell receptor (TCR) that binds to a complex comprising: (i) an epitope of an expressed protein identified or predicted to be presented by class I or II MHC on cells of the subject, and (ii) class I or II MHC, wherein the complex is targeted by CD8 or CD4 T cells of the subject.

[0129]

[0205] Provided herein is a computer system for identifying peptide sequences for personalized cancer treatment of a subject, comprising: a database configured to store a plurality of peptide sequences of the subject; and one or more computer processors operably linked to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class I or class II MHC alleles of a cell of the subject can present a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of sequences of peptides presented by HLA proteins expressed in the cell and identified by mass spectrometry; and selecting a subset of the plurality of peptide sequences for personalized cancer treatment of the subject based on at least the plurality of presentation predictions.

[0130]

[0206] Provided herein is a computer system for identifying HLA class I or HLA class II specific peptides for immunotherapy for a subject, comprising: a database configured to store candidate peptides comprising an epitope and a plurality of peptide sequences each comprising the epitope; and one or more computer processors operably linked to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences to immune cells, each presentation prediction indicating the likelihood that one or more proteins encoded by an HLA class I or HLA class II allele may present a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is programmed to generate presentation predictions for each of the plurality of peptide sequences to immune cells, each presentation prediction indicating the likelihood that one or more proteins encoded by an HLA class I or HLA class II allele may present a given peptide sequence of the plurality of peptide sequences, The machine-learned HLA peptide presentation prediction model is trained using training data including sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; selecting a protein from one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cells and predicted to bind to a candidate peptide by the machine-learned HLA peptide presentation prediction model, wherein the protein has a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide is contacted with the selected protein so that the candidate peptide competes with the placeholder peptide associated with the selected protein.

[0131]

[0207] Provided herein is a computer system for screening drugs comprising polypeptide sequences for immunogenicity in a subject, comprising: a database configured to store a plurality of peptide sequences of the polypeptide sequence; and one or more computer processors operably linked to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information associated with HLA proteins expressed in the cell; and determining or predicting, based on the set of the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic to the subject, wherein a composition comprising a drug is administered to the subject.

[0132]

[0208] Provided herein is a computer system for screening drugs comprising polypeptide sequences for immunogenicity in a subject, comprising: a database configured to store a plurality of peptide sequences of the polypeptide sequence; and one or more computer processors operably linked to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of peptide sequences presented by HLA proteins expressed in the cell and identified by mass spectrometry; and determining or predicting that at least one of the plurality of peptide sequences of the polypeptide sequence will be immunogenic to the subject based on the plurality of presentation predictions.

[0133]

[0209] Provided herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, performs a method for identifying peptide sequences for personalized cancer treatment in a subject, the method comprising: obtaining a plurality of peptide sequences for the subject; processing amino acid information of the plurality of peptide sequences using a machine-learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or class II MHC alleles of a cell of the subject may present a given peptide sequence of the plurality of peptide sequences, the machine-learning HLA peptide presentation prediction model being trained using training data comprising sequence information of sequences of peptides presented by HLA proteins expressed in the cells and identified by mass spectrometry; and selecting a subset of the plurality of peptide sequences for personalized cancer treatment in the subject based on at least the plurality of presentation predictions.

[0134]

[0210] Provided herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method for identifying HLA class II-specific peptides for immunotherapy of a subject, the method comprising: obtaining candidate peptides comprising an epitope and a plurality of peptide sequences, each comprising the epitope; processing amino acid information of the plurality of peptide sequences using a machine-learning HLA peptide presentation prediction model to generate presentation predictions to immune cells for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by HLA class I or HLA class II alleles may present a given peptide sequence of the plurality of peptide sequences, and wherein the machine-learning HLA peptide presentation prediction model is expressed in a cell. the machine learning HLA peptide presentation prediction model is trained using training data containing sequence information of peptide sequences presented by HLA proteins identified by mass spectrometry; selecting a protein from one or more proteins encoded by HLA class II alleles of the subject's cells and predicted to bind to a candidate peptide by the machine learning HLA peptide presentation prediction model, the protein having a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide is contacted with the selected protein so that the candidate peptide competes with the placeholder peptide.

[0135]

[0211] Provided herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method of screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine-learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, the machine-learning HLA peptide presentation prediction model being trained using training data comprising sequence information associated with HLA proteins expressed in the cell; and determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic to the subject, wherein a composition comprising the drug is administered to the subject.

[0136]

[0212] Provided herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method for screening drugs comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine-learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or II MHC alleles of a cell of the subject may present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, wherein the machine-learning HLA peptide presentation prediction model is trained using training data comprising sequence information of peptide sequences presented by HLA proteins expressed in the cell and identified by mass spectrometry; and determining or predicting that at least one of the plurality of peptide sequences of the polypeptide sequence will be immunogenic to the subject based on the plurality of presentation predictions.

[0137]

[0213] The present invention provides a method for identifying a peptide sequence of the plurality of peptide sequences that has a higher probability than a presentation prediction probability value threshold to be presented by at least one of the one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cells, the method comprising: processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each of the plurality of candidate peptide sequences is encoded by a genome or exome of a subject, the plurality of presentation predictions comprising an HLA presentation prediction for each of the plurality of candidate peptide sequences, each presentation prediction indicating a likelihood that one or more proteins encoded by HLA class I or HLA class II alleles of a cell of the subject may present the plurality of given candidate peptide sequences, the machine learning HLA peptide presentation prediction model being trained using training data comprising sequence information of sequences of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in the training cells; and identifying, based at least on the plurality of presentation predictions, a peptide sequence of the plurality of peptide sequences that has a higher probability than a presentation prediction probability value threshold to be presented by at least one of the one or more proteins encoded by HLA class I or HLA class II alleles of the cell of the subject, the machine learning HLA peptide presentation prediction model being trained using training data comprising sequence information of sequences of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in the training cells, the method comprising: The machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 when amino acid information of a plurality of test peptide sequences is processed to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by HLA class I or HLA class II alleles of a cell of a subject may present a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 500 test peptide sequences, the test peptide sequences comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell and (ii) at least 499 decoy peptide sequences contained within a protein encoded by the genome of an organism, wherein the organism and the subject are the same species, and the plurality of test peptide sequences comprises a ratio of at least one hit peptide sequence to at least 499 decoy peptide sequences of 1:499, wherein 0.2% of the plurality of test peptide sequences are predicted by the machine learning HLA peptide presentation prediction model to be presented by an HLA protein expressed in the cell.

[0138]

[0214] The present invention provides a method for identifying a peptide sequence of a plurality of candidate peptide sequences encoded by a genome or exome of a subject, the method comprising: processing amino acid information of a plurality of candidate peptide sequences encoded by the genome or exome of the subject using a machine learning HLA peptide binding prediction model to generate a plurality of binding predictions, the plurality of binding predictions including an HLA binding prediction for each of the plurality of candidate peptide sequences, each binding prediction indicating a likelihood that one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cell will bind to a given candidate peptide sequence of the plurality of candidate peptide sequences, the machine learning HLA peptide binding prediction model being trained using training data including sequence information of sequences of peptides identified to bind to HLA class I or HLA class II proteins or HLA class I or HLA class II protein analogs; and identifying, based at least on the plurality of binding predictions, a peptide sequence of the plurality of test peptide sequences that has a probability higher than a binding prediction probability value threshold of binding to at least one of the one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cell, the amino acid information of the plurality of test peptide sequences being generated. The method further includes the step of: when the amino acid information is processed to generate a plurality of test binding predictions, the machine learning HLA peptide binding prediction model has a positive predictive value (PPV) of at least 0.1, each test binding prediction indicating the likelihood that one or more proteins encoded by HLA class I or HLA class II of the subject's cells will bind to a given test peptide sequence of the plurality of test peptide sequences; the plurality of test peptide sequences includes at least 50 test peptide sequences, the test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed by the cell, and (ii) at least 19 decoy peptide sequences contained within a protein that includes the peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed by the cell; the organism and the subject are of the same species; the plurality of test peptide sequences includes a 1:19 ratio of the at least one hit peptide sequence to the at least 19 decoy peptide sequences; and 5% of the plurality of test peptide sequences are predicted by the machine learning HLA peptide presentation prediction model to bind to an HLA protein expressed by the cell.

[0139]

[0215] In some embodiments, the machine learning HLA peptide presentation prediction model is trained using training data that includes sequence information for sequences of training peptides identified by mass spectrometry as being presented by HLA proteins expressed on training cells.

[0140]

[0216] In some embodiments, one or more of 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model have a probability higher than a presentation prediction probability value threshold to be presented by at least one of the one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cells.

[0141]

[0217] In some embodiments, 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model each have a probability higher than a presentation prediction probability value threshold of being presented by at least one of the one or more proteins encoded by HLA class I or HLA class II alleles of the subject cell.

[0142]

[0218] Provided herein is a method for preparing a personalized cancer therapy, the method comprising: identifying a peptide sequence associated with cancer, the identifying comprising comparing a DNA, RNA, or protein sequence from a cancer cell of the subject with a DNA, RNA, or protein sequence from a normal cell of the subject; using a computer processor, inputting amino acid position information of the identified peptide sequence into a machine learning HLA peptide presentation prediction model to generate a set of presentation predictions for the identified peptide sequence, each presentation prediction indicating a probability that one or more proteins encoded by an HLA class I or HLA class II allele of the cell of the subject will present a given sequence of the identified peptide sequence, The presented prediction model includes: a plurality of predictor variables identified based at least on training data, where the training data includes: sequence information of peptide sequences expressed in cells and presented by HLA proteins identified by mass spectrometry, training peptide sequence information including amino acid position information, where the training peptide sequence information is associated with the HLA proteins expressed in the cells, and a function representing an association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables; and selecting a subset of the identified peptide sequences based on the set of presentation predictions for preparing a personalized cancer treatment, wherein the prediction model has a positive predictive value of at least 0.1 with a recall of at least 0.1%, 0.1% to 50%, or up to 50%.

[0143]

[0219] Provided herein is a method comprising training a machine-learned HLA peptide presentation prediction model, the training step comprising inputting, using a computer processor, amino acid positional information sequences of HLA peptides isolated from one or more HLA-peptide complexes derived from cells expressing HLA class II alleles into the HLA peptide presentation prediction model, wherein the machine-learned HLA peptide presentation prediction model comprises: sequence information of sequences of peptides presented by HLA proteins expressed in the cells and identified by mass spectrometry; training peptide sequence information comprising amino acid positional information of training peptides associated with the HLA proteins expressed in the cells; and a plurality of predictor variables identified based at least on the training data, the training peptide sequence information comprising a function representing an association between the amino acid positional information received as input and a presentation probability generated as an output based on the amino acid positional information and the predictor variables.

[0144]

[0220] In some embodiments, the presented model has a positive predictive value of at least 0.25 at a recall of at least 0.1%, between 0.1% and 50%, or up to 50%.

[0221] In some embodiments, the presented model has a positive predictive value of at least 0.4 at a recall of at least 0.1%, between 0.1% and 50%, or up to 50%.

[0145]

[0222] In some embodiments, the presented model has a positive predictive value of at least 0.6 at a recall of at least 0.1%, between 0.1% and 50%, or up to 50%.

[0223] In some embodiments, the mass spectrometry is single allele mass spectrometry.

[0146]

[0224] In some embodiments, the peptides are presented by HLA proteins expressed in cells through autophagy.

[0225] In some embodiments, the peptides are presented by HLA proteins expressed on cells through phagocytosis.

[0147]

[0226] In some embodiments, the quality of the training data is improved by using multiple quality metrics.

[0227] In some embodiments, the multiple quality metrics include removal of common contaminating peptides, high scored peak intensity, high score, and high mass accuracy.

[0148]

[0228] In some embodiments, the scored peak intensity is at least 50%.

[0229] In some embodiments, the scored peak intensity is at least 60%.

[0230] In some embodiments, the score is at least 7.

[0149]

[0231] In some embodiments, the mass accuracy is up to 5 ppm.

[0232] In some embodiments, the mass accuracy is up to 2 ppm.

[0233] In some embodiments, the backbone cleavage score is at least 5.

[0150]

[0234] In some embodiments, the backbone cleavage score is at least 8.

[0235] In some embodiments, the peptides presented by an HLA protein expressed on a cell are peptides presented by a single immunoprecipitated HLA protein expressed on a cell.

[0151]

[0236] In some embodiments, the peptides presented by an HLA protein expressed on the cell are peptides presented by a single exogenous HLA protein expressed on the cell.

[0152]

[0237] In some embodiments, the peptides presented by an HLA protein expressed on a cell are peptides presented by a single recombinant HLA protein expressed on a cell.

[0153]

[0238] In some embodiments, the plurality of predictors comprises a peptide-HLA affinity predictor.

[0239] In some embodiments, the plurality of predictor variables comprises a source protein expression level predictor variable.

[0154]

[0240] In some embodiments, the plurality of predictors comprises a peptide cleavability predictor.

[0241] In some embodiments, the training peptide sequence information includes sequences derived from peptides presented by HLA proteins, including peptides identified by searching non-enzyme specific peptide databases that lack modifications. In some embodiments, the peptides presented by HLA proteins include peptides identified by searching de novo peptide sequencing tools.

[0155]

[0242] In some embodiments, peptides presented by HLA proteins include peptides identified by searching peptide databases using a reverse database search strategy.

[0156]

[0243] In some embodiments, the HLA protein comprises an HLA-DR and an HLA-DP or HLA-DQ protein. In some embodiments, the HLA protein comprises an HLA-DR protein selected from the group consisting of an HLA-DR and an HLA-DP or HLA-DQ protein.Also known as HLA-DPB1*01:01 / HLA-DPA1*01:03 HLA- DPB1*02:01 / HLA-DPA1*01:03、HLA-DPB1*03:01 / HLA-DPA1*01:03 、HLA-DPB1*04:01 / HLA-DPA1*01:03、HLA-DPB1*04:02 / HLA-DPA1*01:03、HLA-DPB1*06:01 / HLA-DPA1*01:03、HLA-02:02A / DQB1* -DQA1*05:01、HLA-DQB1*02:02 / HLA-DQA1*02:01、HLA-DQB1*06:02 / HLA-DQA1*01:02、HLA-DQB1*06:04 / HLA-1:DQA0-2* 1*01:01、HLA-DRB1*01:02、HLA-DRB1*03:01、HLA-DRB1*03:02、HLA-DRB1*04:01、HLA-DRB1*04:02、HLA-DRB1*04:04: 04、HLA-DRB1*04:05、HLA-DRB1*04:07、HLA-DRB1*07:01、HLA-DRB1*08:01、HLA-DRB1*08:02、HLA-DRB1*08:03DRB1、HL LA-DRB1*09:01、HLA-DRB1*10:01、HLA-DRB1*11:01、HLA-DRB1*11:02、HLA-DRB1*11:04、HLA-DRB1*12:01、D、HLA-12:0-DRB1* RB1*13:01、HLA-DRB1*13:02、HLA-DRB1*13:03、HLA-DRB1*14:01、HLA-DRB1*15:01、HLA-DRB1*15:02、HLA-0B3DRB1*15: 16:01、HLA-DRB3*01:01、HLA-DRB3*02:02、HLA-DRB3*03:01、HLA- DRB4*01:01 HLA-DRB5*01:01 must be activated when HLA-DR is activated.

[0157]

[0244] In some embodiments, peptides presented by HLA proteins include peptides identified by comparing the MS / MS spectrum of an HLA peptide with the MS / MS spectra of one or more HLA peptides in a peptide database.

[0158]

[0245] In some embodiments, the mutation is selected from the group consisting of a point mutation, a splice site mutation, a frameshift mutation, a readthrough mutation, and a gene fusion mutation.

[0246] In some embodiments, peptides presented by HLA proteins have a length of 8 to 12 or 15 to 40 amino acids.

[0159]

[0247] In some embodiments, the peptides presented by HLA proteins include peptides identified by the steps of: (a) isolating one or more HLA complexes from a cell line expressing a single HLA class I or HLA class II allele; (b) isolating one or more HLA peptides from the one or more isolated HLA complexes; (c) obtaining MS / MS spectra for the one or more isolated HLA peptides; and (d) obtaining peptide sequences from a peptide database corresponding to the MS / MS spectra of the one or more isolated HLA peptides, wherein the one or more sequences obtained from step (d) identify the sequences of the one or more isolated HLA peptides.

[0160]

[0248] In some embodiments, the personalized cancer treatment further comprises an adjuvant.

[0249] In some embodiments, the personalized cancer treatment further comprises an immune checkpoint inhibitor.

[0161]

[0250] In some embodiments, the training data includes structured data, time series data, unstructured data, relational data, or any combination thereof.

[0251] In some embodiments, the unstructured data includes image data.

[0162]

[0252] In some embodiments, the relational data includes data from customer systems, enterprise systems, operational systems, websites, web-accessible application program interfaces (APIs), or any combination thereof.

[0163]

[0253] In some embodiments, the training data is uploaded to a cloud-based database.

[0254] In some embodiments, the training is performed using a convolutional neural network.

[0164]

[0255] In some embodiments, the convolutional neural network includes at least two convolutional layers.

[0256] In some embodiments, the convolutional neural network (CNN) includes at least one batch normalization step.

[0165]

[0257] In some embodiments, the convolutional neural network includes at least one spatial dropout step.

[0258] In some embodiments, the convolutional neural network includes at least one global max pooling step.

[0166]

[0259] In some embodiments, the convolutional neural network includes at least one densely connected layer.

[0260] In some embodiments, identifying peptide sequences comprises identifying peptide sequences having mutations that are expressed in cancer cells of the subject.

[0167]

[0261] In some embodiments, identifying a peptide sequence comprises identifying a peptide sequence that is not expressed in normal cells of the subject.

[0262] In some embodiments, identifying peptide sequences comprises identifying overexpressed peptide sequences.

[0168]

[0263] In some embodiments, identifying the peptide sequence comprises identifying a viral peptide sequence. In one aspect, a method for identifying an HLA class I or HLA class II-specific peptide for specific immunotherapy for a subject is provided, the method comprising: identifying candidate peptides comprising the epitope; inputting, using a computer processor, amino acid information of a plurality of peptide sequences, each comprising the epitope, into a machine learning HLA peptide presentation prediction model to generate a set of HLA presentation predictions for the peptide sequences to immune cells, wherein each presentation prediction indicates a probability that one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cells will present a given peptide sequence comprising the epitope, and the prediction model has a positive predictive value of at least 0.1 with a recall of at least 0.1%, between 0.1% and 50%, or up to 50%; The method includes the steps of: selecting a protein predicted by the prediction model to bind to the candidate peptide from one or more proteins encoded by HLA class I or HLA class II alleles, wherein the protein has a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; contacting the candidate peptide with a protein encoded by an HLA class I or HLA class II allele so that the candidate peptide competes with a placeholder peptide that associates with the protein encoded by the HLA class I or HLA class II allele; and identifying the candidate peptide as a peptide for immunotherapy specific for the protein encoded by the HLA class II allele based on whether the candidate peptide replaces the placeholder peptide.

[0169]

[0264] In some embodiments, the immunotherapy is cancer immunotherapy.

[0265] In some embodiments, the identifying step comprises comparing a DNA, RNA, or protein sequence from the subject's cancer cells with a DNA, RNA, or protein sequence from the subject's normal cells. In some embodiments, the epitope is a cancer-specific epitope.

[0170]

[0266] In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide is a CMV peptide. In some embodiments, the method comprises determining IC of replacement of the placeholder peptide with the target peptide. 50 In some embodiments, the IC of the replacement of the placeholder peptide with the target peptide is measured. 50 In some embodiments, the target peptide is further identified by mass spectrometry. In some embodiments, the at least one protein encoded by the HLA class I or HLA class II allele of the target cell is a recombinant protein. In some embodiments, the at least one protein encoded by the HLA class I or HLA class II allele of the target cell is expressed in a eukaryotic cell.

[0171]

[0267] In one aspect, an assay method for verifying the specificity of a candidate peptide for binding to an HLA class I or HLA class II protein is provided, the method comprising: expressing in a eukaryotic cell a polynucleic acid construct comprising a nucleic acid sequence encoding an HLA class I or HLA class II protein comprising an alpha chain and a beta chain, or a fragment thereof, capable of binding to a peptide comprising an MHC-binding epitope, wherein the expressed HLA class I or HLA class II protein, or fragment thereof, remains associated with a placeholder peptide; isolating the HLA class I or HLA class II protein, or portion thereof, expressed in the eukaryotic cell; (a) adding increasing amounts of the candidate peptide to determine whether the candidate peptide displaces the placeholder peptide associated with the HLA class I or HLA class II protein, or portion thereof; and (b) measuring the IC of the displacement reaction to determine the affinity of the candidate peptide for the HLA class I or HLA class II protein, or portion thereof, relative to the placeholder peptide. 50 performing a peptide exchange assay by calculating the specificity of the candidate peptide for binding to an HLA class I or HLA class II protein, thereby verifying the specificity of the candidate peptide for binding to an HLA class I or HLA class II protein.

[0172]

[0268] In some embodiments, the identity of the peptide is known. In some embodiments, the identity of the peptide is unknown. In some embodiments, the identity of the peptide is determined by mass spectrometry.

[0173]

[0269] In some embodiments, the peptide exchange assay comprises detection of a peptide fluorescent probe or tag. In some embodiments, the placeholder peptide is a CLIP peptide.

[0174]

[0270] In some embodiments, the polynucleic acid construct comprises an expression vector further comprising one or more of: a promoter, a linker, one or more protease cleavage sites, a secretion signal, a dimerization factor, a ribosomal skipping sequence, one or more tags for purification and or detection.

[0175]

[0271] In one aspect, provided herein is a method for assaying the immunogenicity of an MHC class II-binding peptide, the method comprising: selecting a protein encoded by an HLA class II allele predicted by a machine learning HLA peptide presentation prediction model to bind to the peptide, wherein the prediction model has a positive predictive value of at least 0.1 with a recall of at least 0.1%, 0.1% to 50%, or up to 50%, and the protein has a probability higher than a presentation prediction probability value threshold for presenting the identified peptide sequence; contacting the peptide with the selected protein encoded by the HLA class II allele such that the peptide competes with a placeholder peptide for association with the selected protein encoded by the HLA class II allele; displacing the placeholder peptide, thereby forming a complex comprising the HLA class II protein and the identified peptide; contacting the HLA class II protein and identified peptide complex with CD4+ T cells and assaying for one or more activation parameters of the CD4+ T cells selected from cytokine induction, chemokine induction, and cell surface marker expression.

[0176]

[0272] In one aspect, provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: (a) using a computer processor, inputting amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction indicates a probability that one or more proteins encoded by HLA class I or II alleles of the subject's cells will present an epitope sequence of the given peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, the training data including: sequence information of sequences of peptides presented by HLA proteins expressed in the cell and identified by mass spectrometry, training peptide sequence information including amino acid position information, the training peptide sequence information associated with the HLA proteins expressed in the cell, and a function representing an association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables; (b) determining or predicting that each of the peptide sequences of the polypeptide sequence will not be immunogenic to the subject based on the set of presentation predictions; and (c) administering a composition comprising the drug to the subject.

[0177]

[0273] In one aspect, provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: (a) using a computer processor, inputting amino acid information of peptide sequences of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction indicates a probability that one or more proteins encoded by HLA class I or II alleles of the subject's cells will present an epitope sequence of the given peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, the training data including: sequence information of sequences of peptides presented by HLA proteins expressed in the cell and identified by mass spectrometry, training peptide sequence information including amino acid position information, the training peptide sequence information associated with the HLA proteins expressed in the cell, and a function representing an association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables; (b) determining or predicting that at least one peptide sequence of the polypeptide sequence will be immunogenic to the subject based on the set of presentation predictions.

[0178]

[0274] In one embodiment, the method further comprises deciding not to administer the drug to the subject.

[0275] In one embodiment, the drug comprises an antibody or a binding fragment thereof.

[0179]

[0276] In one embodiment, the peptide sequences of the polypeptide sequence comprise each contiguous peptide sequence of the polypeptide sequence having a length of 8, 9, 10, 11 or 12 amino acids, and wherein the protein encoded by an HLA class I or II allele of the subject's cell is a protein encoded by an HLA class I allele of the subject's cell.

[0180]

[0277] In one embodiment, the peptide sequences of the polypeptide sequence comprise each contiguous peptide sequence of the polypeptide sequence having a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, wherein the protein encoded by an HLA class I or II allele of the subject's cell is a protein encoded by an HLA class II allele of the subject's cell.

[0181]

[0278] In one aspect, provided herein is a method of treating a subject having an autoimmune disease or condition, comprising: (a) identifying or predicting an epitope of a protein presented and expressed by HLA class I or II on cells of the subject, wherein a complex comprising the identified or predicted epitope and HLA class I or II is targeted by CD8 or CD4 T cells of the subject; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells from the subject or allogeneic regulatory T cells; and (d) administering the TCR-expressing regulatory T cells to the subject.

[0182]

[0279] In one embodiment, the autoimmune disease or condition is diabetes.

[0280] In one embodiment, the cells are islet cells.

[0281] In one aspect, provided herein is a method of treating a subject having an autoimmune disease or condition, comprising administering to the subject regulatory T cells that express a T cell receptor (TCR) that binds to a complex comprising (i) an epitope of an expressed protein identified or predicted to be presented by HLA class I or II on cells of the subject, and (ii) HLA class I or II, wherein the complex is targeted by CD8 or CD4 T cells of the subject.

[0183]

[0282]

[0013] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein merely illustrative embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and descriptions are to be regarded as illustrative in nature and not restrictive.

[0184]

[0283] In one aspect, provided herein is a method for treating cancer in a subject, the method comprising: identifying a peptide sequence, wherein the peptide sequence is associated with the cancer, wherein the identifying comprises comparing a DNA, RNA, or protein sequence from cancer cells of the subject with a DNA, RNA, or protein sequence from normal cells of the subject; using a computer processor, inputting amino acid position information of the identified peptide sequence into a machine learning HLA peptide presentation prediction model to generate a set of presentation predictions for the identified peptide sequence, wherein each presentation prediction indicates a probability that one or more proteins encoded by HLA class I or HLA class II alleles of the subject's cells will present a given sequence of the identified peptide sequence, wherein the machine learning HLA peptide presentation prediction model: the training data includes: sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, training peptide sequence information including amino acid position information, the training peptide sequence information associated with HLA proteins expressed in cells, and a function representing an association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables; and selecting a subset of the identified peptide sequences based on the set of presentation predictions for preparing a personalized cancer treatment; and administering a composition comprising one or more peptides to a subject, wherein the predictive model has a positive predictive value of at least 0.1 with a recall of at least 0.1%, between 0.1% and 50%, or up to 50%.

[0185]

[0284] In some embodiments, the machine learning HLA peptide presentation predictive model comprises sequence information of sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry after performing reverse-phase offline fractionation.

[0186]

[0285] In some embodiments, the predictive model exhibits a 1.1x to 100x fold improvement compared to NetMHCIIpan or NetMHCI. In some embodiments, the predictive model exhibits a 1.1, 2, 3, 4, 5, 6, 7, 7.4, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 ,40,41,42,43,44,45,50,55,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,8,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,showing over 100x improvement.

[0187]

[0286] In one aspect, the present disclosure provides a method for predicting peptides that can accurately pair with or bind to specific HLA class I or class II molecules, such that high-fidelity binding of the peptide to HLA class I or class II proteins (including alpha and beta chain heterodimers) ensures presentation of the specific peptide to T lymphocytes, thereby eliciting a specific immune response and avoiding any cross-reactivity or immune disruption. Several recent studies have shown that CD8+ or CD4+ T cells also recognize HLA class I or class II-presented ligands, contributing to tumor management. Cancer vaccines and other immunotherapies ideally utilize directing CD8+ or CD4+ T cell responses, but current efforts have completely halted HLA class I or class II antigen prediction due to the insufficient accuracy of current prediction tools.

[0188]

[0287] In one aspect, the present disclosure provides methods for predicting peptides that can accurately bind to specific HLA class I or class II proteins such that, when the peptides are administered therapeutically to subjects expressing the specific cognate HLA class I or class II proteins, a sustained and robust immune response can be activated with the peptides by means of the HLA class I or class II proteins' ability to stimulate CD8+ or CD4+ T cell activation and immunological memory. In some embodiments, the methods provided herein demonstrate an improvement in specific HLA class I or class II protein prediction over currently available predictors. In some embodiments, the methods provided herein demonstrate at least about a 1.1-fold improvement in specific HLA class I or class II protein prediction over currently available predictors. In some embodiments, the methods provided herein demonstrate at least about a 2-fold improvement in specific HLA class I or class II protein prediction over currently available predictors. In some embodiments, the methods provided herein demonstrate at least about a 3-fold improvement in specific HLA class I or class II protein prediction over currently available predictors. In some embodiments, the methods provided herein exhibit at least about a 4-fold improvement over currently available predictors in predicting certain HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 5-fold improvement over currently available predictors in predicting certain HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 6-fold improvement over currently available predictors in predicting certain HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 7-fold improvement over currently available predictors in predicting certain HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about an 8-fold improvement over currently available predictors in predicting certain HLA class I or class II proteins.In some embodiments, the methods provided herein exhibit at least about a 9-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 10-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 15-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 20-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 30-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 40-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 50-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins. In some embodiments, the methods provided herein exhibit at least about a 60-fold improvement over currently available predictors in predicting specific HLA class I or class II proteins.

[0189]

[0288] In one aspect, the present invention provides a method for immunotherapy that is tailored or personalized for a specific subject.Every subject or patient expresses a specific array of HLA class I and HLA class II proteins.HLA typing is a known technique that allows the specific repertoire of HLA proteins expressed by a subject to be determined.Once the HLA heterodimers expressed by a specific subject are understood, as described herein, an improved, refined and reliable method for predicting the peptides that can bind with high fidelity to specific HLA class I or class II molecules or complexes can ensure that a specific immune response can be generated specifically for a specific subject.

[0190]

[0289] In this application, the use of the singular includes the plural unless specifically stated otherwise. It should be noted that, as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless stated otherwise. Furthermore, the use of the term "including" and other forms, such as "include," "includes," and "included," is not limiting. The term "one or more" or "at least one," for example, one or more, or at least one member(s) of a group of members, will itself be clear with further illustration; in particular, the term encompasses reference to any one of said members, or any two or more of said members, for example, >3, >4, >5, >6, or >7, etc., of any of said members, and up to all said members.

[0191]

[0290] References herein to "some embodiments," "an embodiment," "one embodiment," or "other embodiments" mean that a feature, structure, or characteristic described in connection with an embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments.

[0192]

[0291] As used in the specification and claim(s), the words "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include"), or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed herein can be implemented in connection with any method or composition of the disclosure, and vice versa. Furthermore, the compositions of the disclosure can be used to achieve the methods of the disclosure.

[0193]

[0292] As used herein, the terms "about" or "approximately," when referring to a measurable value such as a parameter, amount, time period, etc., are meant to encompass a variation of no more than + / - 20%, no more than + / - 10%, no more than + / - 5%, or no more than + / - 1% of the specified value, insofar as such variations are appropriate to practice the present disclosure. It is understood that the value to which the modifier "about" or "approximately" refers is itself expressly disclosed.

[0194]

[0293] The term "immune response" includes T cell-mediated and / or B cell-mediated immune responses that are affected by the regulation of T cell costimulation. Exemplary immune responses include T cell responses, such as cytokine production and cytotoxic activity. In addition, the term immune response includes immune responses that are indirectly affected by T cell activation, such as antibody production (humoral response) and activation of cytokine-responsive cells, such as macrophages.

[0195]

[0294] "Receptor" is understood to mean a biological molecule or group of molecules that can bind to a ligand. Receptors serve to transmit information in cells, cell formations, or organisms. A receptor includes at least one receptor unit and can contain two or more receptor units, where each receptor unit can be a protein molecule, for example, a glycoprotein molecule. A receptor has a structure that complements the structure of a ligand and can complex with the ligand as a binding partner. Signaling information can be transmitted by a conformational change of the receptor following binding with the ligand on the cell surface. According to the present disclosure, a receptor can refer to MHC class I and II proteins that can form a receptor / ligand complex with a ligand, for example, a peptide or peptide fragment of a suitable length. Class I and class II MHC peptides encoded by HLA class I and class II alleles are often referred to herein as HLA class I and HLA class II peptides, or HLA class I and HLA class II peptides, or HLA class I and class II proteins, or HLA class I and HLA class II proteins, or HLA class I and class II molecules, or common variants thereof, as will be well understood in the context of discussion by those skilled in the art.

[0196]

[0295] A "ligand" is a molecule capable of forming a complex with a receptor. In accordance with the present disclosure, a ligand is understood to mean, for example, a peptide or peptide fragment having a suitable length and a suitable binding motif in its amino acid sequence so that the peptide or peptide fragment can bind to an MHC class I or MHC class II protein (i.e., an HLA class I or HLA class II protein) and form a complex.

[0197]

[0296] An "antigen" is a molecule capable of stimulating an immune response and may be produced by cancer cells, infectious agents, or autoimmune diseases. Antigens recognized by either T cells, helper T lymphocytes (helper T (TH) cells), or cytotoxic T lymphocytes (CTLs) are not recognized as intact proteins, but rather as small peptides associated with HLA class I or class II proteins on the cell surface. During the course of a natural immune response, antigens recognized in association with HLA class II molecules on antigen-presenting cells (APCs) are obtained extracellularly, internalized, and processed into small peptides that associate with HLA class II molecules. APCs can also cross-present peptide antigens by processing exogenous antigens and presenting the processed antigens on HLA class I molecules. Antigens that give rise to peptides recognized in association with HLA class I MHC molecules are generally peptides produced intracellularly, which are processed and associated with class I MHC molecules. It is now understood that peptides that associate with a given HLA class I or class II molecule are characterized as having a common binding motif, and binding motifs for a number of different HLA class I and II molecules have been determined. Synthetic peptides corresponding to the amino acid sequence of a given antigen and containing a binding motif for a given HLA class I or II molecule can also be synthesized. These peptides can then be added to appropriate APCs, which can be used in vitro or in vivo to stimulate helper T cell or CTL responses. Binding motifs, methods for synthesizing peptides, and methods for stimulating helper T cell or CTL responses are all well known and readily available to those skilled in the art.

[0198]

[0297] The term "peptide" is used interchangeably herein with "mutant peptide" and "neoantigenic peptide." Similarly, the term "polypeptide" is used interchangeably herein with "mutant polypeptide" and "neoantigenic polypeptide." By "neoantigen" or "neoepitope" is meant a class of tumor antigens or tumor epitopes that arise from tumor-specific mutations in expressed proteins. The present disclosure further includes peptides containing tumor-specific mutations, peptides containing known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods of the present disclosure. These peptides and polypeptides are referred to herein as "neoantigenic peptides" or "neoantigenic polypeptides." The polypeptides or peptides may be of various lengths, in either their neutral (uncharged) form or in salt form, and with or without modifications such as glycosylation, side chain oxidation, phosphorylation, or post-translational modifications, subject to conditions under which the modifications do not destroy the biological activity of the polypeptides described herein. In some embodiments, neo-antigenic peptides of the present disclosure can comprise: for HLA class I, 22 residues or less in length, e.g., about 8 to about 22 residues, about 8 to about 15 residues, or 9 or 10 residues; for HLA class II, 40 residues or less in length, e.g., about 8 to about 40 residues in length, about 8 to about 24 residues in length, about 12 to about 19 residues, or about 14 to about 18 residues in length. In some embodiments, the neo-antigenic peptide or neo-antigenic polypeptide comprises a neo-epitope.

[0199]

[0298] The term "epitope" includes any protein determinant capable of specific binding to an antibody, antibody peptide, and / or antibody-like molecule (including, but not limited to, a T-cell receptor) as defined herein. Epitope determinants typically consist of chemically active surface groupings of molecules such as amino acids or sugar side chains and generally have specific three-dimensional structural characteristics, as well as specific charge characteristics.

[0200]

[0299] A "T cell epitope" is a peptide sequence that can be bound by class I or II MHC molecules in the form of peptide-presenting MHC molecules or MHC complexes, and in this form is recognized and bound by cytotoxic T lymphocytes or helper T cells, respectively.

[0201]

[0300] As used herein, the term "antibody" refers to whole antibodies, including IgG (including IgG1, IgG2, IgG3, and IgG4), IgA (including IgA1 and IgA2), IgD, IgE, IgM, and IgY, including single-chain whole antibodies and their antigen-binding (Fab) fragments. Antigen-binding antibody fragments include, but are not limited to, Fab, Fab', and F(ab'), Fd (consisting of VH and CH1), single-chain variable fragments (scFv), single-chain antibodies, disulfide-linked variable fragments (dsFv), and fragments containing either the VL or VH domain. Antibodies can be derived from any animal. Antigen-binding antibody fragments, including single-chain antibodies, can contain the variable region(s) alone or in combination with all or part of the following: hinge region, CH1, CH2, and CH3 domains. Also included are any combinations of the variable region(s) with the hinge region, CH1, CH2, and CH3 domains. The antibodies may be, for example, monoclonal, polyclonal, chimeric, humanized, and human monoclonal and polyclonal antibodies that specifically bind to HLA-associated polypeptides or HLA-HLA binding peptide (HLA peptide) complexes. Those skilled in the art will recognize that various immunoaffinity techniques are suitable for enriching soluble proteins, such as soluble HLA-peptide complexes, or membrane-bound HLA-associated polypeptides, such as those proteolytically cleaved from the membrane. These include techniques in which (1) one or more antibodies capable of specifically binding to soluble proteins are immobilized on a fixed or mobile substrate (e.g., plastic wells or resin, latex, or paramagnetic beads), and (2) a solution containing soluble proteins from a biological sample is passed over the antibody-coated substrate, allowing the soluble proteins to bind to the antibodies. The substrate containing the antibodies and bound soluble proteins is separated from the solution, and optionally, the antibodies and soluble proteins are dissociated, for example, by varying the pH and / or ionic strength and / or ionic composition of the solution in which the antibodies are immersed. Alternatively, immunoprecipitation techniques can be used in which antibodies and soluble proteins are combined to form macromolecular aggregates, which can be separated from the solution by size exclusion techniques or by centrifugation.

[0202]

[0301] The term "immunopurification (IP)" (or immunoaffinity purification or immunoprecipitation) is a process known in the art and widely used for the isolation of a desired antigen from a sample. Generally, the process involves contacting a sample containing the desired antigen with an affinity matrix in which an antibody against the antigen is covalently bound to a solid phase. The antigen in the sample becomes bound to the affinity matrix through immunochemical binding. The affinity matrix is ​​then washed to remove any unbound molecular species. The antigen is removed from the affinity matrix by changing the chemical composition of the solution in contact with the affinity matrix. Immunopurification may be performed on a column containing the affinity matrix, in which case the solution is the eluate. Alternatively, immunopurification may be performed in a batch process where the affinity matrix is ​​maintained in suspension in solution. A critical step in the process is the removal of the antigen from the matrix. This is generally achieved by increasing the ionic strength of the solution in contact with the affinity matrix, for example, by adding inorganic salts. Changing the pH may also be effective in dissociating the immunochemical bond between the antigen and the affinity matrix.

[0203]

[0302] An "agent" is any small molecule chemical, antibody, nucleic acid molecule or polypeptide or fragment thereof.

[0303] A "change" or "variation" is an increase or decrease. The change may be as little as 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, or up to 40%, 50%, 60%, or even as much as 70%, 75%, 80%, 90%, or 100%.

[0204]

[0304] A "biological sample" is any tissue, cell, body fluid, or other substance derived from an organism. As used herein, the term "sample" includes biological samples, such as any tissue, cell, body fluid, or other substance derived from an organism. "Specifically binds" refers to a compound (e.g., a peptide) that recognizes and binds to a molecule (e.g., a polypeptide) but does not substantially recognize and bind to other molecules in a sample, e.g., a biological sample.

[0205]

[0305] A "capture reagent" refers to a reagent that specifically binds to a molecule (eg, a nucleic acid molecule or a polypeptide) in order to select or isolate the molecule (eg, a nucleic acid molecule or a polypeptide).

[0206]

[0306] As used herein, the terms "determine," "assess," "assay," "measure," "detect," and their grammatical equivalents refer to both quantitative and qualitative determinations, and thus the term "determine" is used interchangeably herein with "assay," "measure," etc. When a quantitative determination is intended, the phrase "determine the amount" of an analyte, etc. is used. When a qualitative and / or quantitative determination is intended, the phrase "determine the level" of an analyte or "determining" an analyte is used.

[0207]

[0307] A "fragment" is a portion of a protein or nucleic acid that is substantially identical to a reference protein or nucleic acid. In some embodiments, the portion retains at least 50%, 75%, or 80%, or 90%, 95%, or even 99% of the biological activity of the reference protein or nucleic acid described herein.

[0208]

[0308] The terms "isolated," "purified," "biologically pure," and their grammatical equivalents refer to material that is free, to varying degrees, from components normally associated with it when found in its native state. "Isolated" describes a degree of separation from the original source or environment. "Purified" describes a degree of separation greater than isolation. A "purified" or "biologically pure" protein is substantially free of other substances, such that all impurities do not substantially affect the biological properties of the protein or produce other adverse effects. That is, a nucleic acid or peptide of the present disclosure is purified when it is substantially free of cellular material, viral material, or culture medium, if produced by recombinant DNA technology, or chemical precursors or other chemicals, if chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" may describe a nucleic acid or protein that yields essentially one band in electrophoresis. For proteins that may be subject to modifications, such as phosphorylation or glycosylation, different modifications may result in different isolated proteins that can be separately purified.

[0209]

[0309] An "isolated" polypeptide (e.g., a peptide derived from an HLA peptide complex) or polypeptide complex (e.g., an HLA peptide complex) is a polypeptide or polypeptide complex of the present disclosure that has been separated from components that naturally accompany it. Typically, a polypeptide or polypeptide complex is isolated when it is at least 60% by weight free from the proteins and naturally occurring organic molecules that naturally accompany it. A preparation may be at least 75%, at least 90%, or at least 99% by weight of a polypeptide or polypeptide complex of the present disclosure. An isolated polypeptide or polypeptide complex of the present disclosure may be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding one or more components of such a polypeptide or polypeptide complex, or by chemically synthesizing one or more components of the polypeptide or polypeptide complex. Purity may be measured by any appropriate method, for example, column chromatography, polyacrylamide electrophoresis, or HPLC analysis. In some cases, HLA allele-encoded MHC class II protein (i.e., MHC class II peptide) is referred to interchangeably herein as HLA class II protein (or HLA class II peptide).

[0210]

[0310] The term "vector" refers to a nucleic acid molecule capable of mediating the transport or expression of heterologous nucleic acids. Plasmids are a type of molecular species encompassed by the term "vector." Typically, a vector refers to a nucleic acid sequence containing an origin of replication and other entities necessary for replication and / or maintenance in a host cell. Vectors capable of directing the expression of operably linked genes and / or nucleic acid sequences are referred to herein as "expression vectors." Generally, useful expression vectors are often in the form of "plasmids," which refer to circular double-stranded DNA molecules that, in their vector form, are not bound to a chromosome and typically contain entities for stable or transient expression of the encoded DNA. Other expression vectors that can be used in the methods disclosed herein include, but are not limited to, plasmids, episomes, bacterial artificial chromosomes, yeast artificial chromosomes, bacteriophage, or viral vectors; such vectors can integrate into the host genome or replicate autonomously in the cell. Vectors can be DNA or RNA vectors. Other forms of expression vectors known to those skilled in the art which serve equivalent functions, e.g., self-replicating extrachromosomal vectors or vectors capable of integrating into a host genome, can also be used. Exemplary vectors are those capable of autonomous replication and / or expression of nucleic acids to which they are linked.

[0211]

[0311] When used in reference to a fusion protein, the term "spacer" or "linker" refers to a peptide that links proteins comprising the fusion protein. Spacers generally have no specific biological activity other than linking protein or RNA sequences or maintaining a certain minimum distance or other spatial relationship. However, in some embodiments, the constituent amino acids of a spacer may be selected to influence certain properties, such as molecular folding, net charge, or hydrophobicity. Linkers suitable for use in embodiments of the present disclosure are known to those of skill in the art and include, but are not limited to, linear or branched carbon linkers, heterocyclic carbon linkers, or peptide linkers. In some embodiments, a linker is used to separate two antigenic peptides by a distance sufficient to ensure proper folding of each antigenic peptide. Exemplary peptide linker sequences adopt flexible, extended conformations and do not exhibit a tendency to form a prescribed secondary structure. Typical amino acids in flexible protein regions include Gly, Asn, and Ser. Virtually any permutation of an amino acid sequence containing Gly, Asn, and Ser is expected to meet the above criteria for a linker sequence. Other near-neutral amino acids, such as Thr and Ala, can also be used in the linker sequence. Further amino acid sequences that can be used as linkers are disclosed in Maratea et al. (1985), Gene 40:39-46; Murphy et al. (1986) Proc. Nat'l. Acad. Sci. USA 83:8258-62; U.S. Patent No. 4,935,233; and U.S. Patent No. 4,751,180.

[0212]

[0312] The term "neoplasm" refers to any disease that is caused by or results from an inappropriately high level of cell division, an inappropriately low level of apoptosis, or both. Glioblastoma is a non-limiting example of neoplasm or cancer. The term "cancer" or "tumor" or "hyperproliferative disorder" refers to the presence of cells that have the typical characteristics of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological properties. Cancer cells are often in the form of tumors, but such cells may exist alone in animals, or may be non-tumorigenic cancer cells, such as leukemia cells. Cancers include, but are not limited to, B-cell cancers (e.g., multiple myeloma, Waldenstrom's hypergammaglobulinemia), heavy chain diseases (e.g., alpha chain diseases, gamma chain diseases, and mu chain diseases), benign gammopathy and immune cell amyloidosis, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer (e.g., metastatic, hormone-refractory prostate cancer), pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral cavity or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or appendix cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancers of the blood tissue, and the like.Other non-limiting examples of cancer types amenable to methods encompassed by the present disclosure include human sarcomas and carcinomas, such as fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, liver carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, pulmonary arterial ... cancer, Wilms' tumor, cervical cancer, bone cancer, brain tumor, testicular cancer, lung cancer, small cell lung cancer, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, retinoblastoma; leukemias, e.g., acute lymphocytic leukemia and acute myeloid leukemia (myeloblastic, promyelocytic, myelomonocytic, monocytic, and erythroleukemia); chronic leukemias (chronic myeloid (granulocytic) leukemia and chronic lymphocytic leukemia); and polycythemia vera, lymphoma (Hodgkin's disease and non-Hodgkin's disease), multiple myeloma, Waldenstrom's hypergammaglobulinemia, and heavy chain disease. In some embodiments, the cancer is an epithelial cancer, such as, but not limited to, bladder cancer, breast cancer, cervical cancer, colon cancer, gynecological cancer, renal cancer, laryngeal cancer, lung cancer, oral cancer, head and neck cancer, ovarian cancer, pancreatic cancer, prostate cancer, or skin cancer. In other embodiments, the cancer is breast cancer, prostate cancer, lung cancer, or colon cancer. In still other embodiments, the epithelial cancer is non-small cell lung cancer, non-papillary renal cell carcinoma, cervical cancer, ovarian cancer (e.g., serous ovarian cancer), or breast cancer. Epithelial cancers may also be characterized in various other ways, including, but not limited to, serous, endometrioid, mucinous, clear cell, Brenner, or anaplastic. In some embodiments, the present disclosure is used in the treatment, diagnosis, and / or prognosis of lymphoma or its subtypes, including, but not limited to, mantle cell lymphoma. Lymphoproliferative disorders are also considered to be proliferative diseases.

[0213]

[0313] The term "vaccine" is understood to mean a composition for generating immunity for the prevention and / or treatment of a disease (e.g., neoplasm / tumor / infectious pathogen / autoimmune disease). Thus, a vaccine is a pharmaceutical that contains an antigen and is intended for use in humans or animals to generate specific defenses and protective substances by vaccination. A "vaccine composition" may include a pharmaceutically acceptable excipient, carrier, or diluent. Aspects of the present disclosure relate to the use of the technology in preparing antigen-based vaccines. In these embodiments, vaccine is meant to refer to one or more disease-specific antigen peptides (or corresponding nucleic acids that encode them). In some embodiments, the antigen-based vaccine contains at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30 or more antigenic peptides.In some embodiments, the antigen-based vaccine comprises a vaccine comprising between 2 and 100 species, between 2 and 75 species, between 2 and 50 species, between 2 and 25 species, between 2 and 20 species, between 2 and 19 species, between 2 and 18 species, between 2 and 17 species, between 2 and 16 species, between 2 and 15 species, between 2 and 14 species, between 2 and 13 species, between 2 and 12 species, between 2 and 10 species, between 2 and 9 species, between 2 and 8 species, between 2 and 7 species, between 2 and 6 species, between 2 and 5 species, between 2 and 4 species, between 3 and 100 species, between 3 and 75 species, between 3 and 50 species, between 3 and 25 species, between 3 and 20 species, between 3 and 19 species, between 3 and 18 species, between 3 and 17 species, between 3 and 16 species, between 3 and 15 species, between 3 and 14 species, between 3 and 13 species, between 3 and 12 species, between 3 and 10 species, between 3 and 9 species, between 3 and 8 species, between 3 and 4 species, between 3 and 5 species, between 3 and 6 species, between 3 and 5 ... The antigen peptides may contain 7 to 7, 3 to 6, 3 to 5, 4 to 100, 4 to 75, 4 to 50, 4 to 25, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 5 to 100, 5 to 75, 5 to 50, 5 to 25, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 10, 5 to 9, 5 to 8, or 5 to 7 antigen peptides. In some embodiments, the antigen-based vaccine contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 antigenic peptides. In some cases, the antigenic peptides are neo-antigenic peptides. In some cases, the antigenic peptides include one or more neo-epitopes.

[0214]

[0314] The term "pharmaceutically acceptable" refers to a substance approved or available for approval by a federal or state government regulatory agency or listed in the United States Pharmacopoeia or other generally recognized pharmacopeia for use in animals, including humans. A "pharmaceutically acceptable excipient, carrier, or diluent" refers to an excipient, carrier, or diluent that can be administered to a subject with an agent, which does not destroy the pharmacological activity of the agent when administered in a dose sufficient to deliver a therapeutic amount of the agent, and which is non-toxic. The "pharmaceutically acceptable salts" of the pooled disease-specific antigens listed herein may be acid or base salts generally considered in the art to be suitable for use in contact with human or animal tissues without excessive toxicity, irritation, allergic response, or other problems or complications. Such salts include inorganic and organic acid salts of basic residues, such as amines, and alkali or organic salts of acidic residues, such as carboxylic acids. Illustrative pharmaceutical salts include, but are not limited to, salts of acids such as hydrochloric acid, phosphoric acid, hydrobromic acid, malic acid, glycolic acid, fumaric acid, sulfuric acid, sulfamic acid, sulfanilic acid, formic acid, toluenesulfonic acid, methanesulfonic acid, benzenesulfonic acid, ethanedisulfonic acid, 2-hydroxyethylsulfonic acid, nitric acid, benzoic acid, 2-acetoxybenzoic acid, citric acid, tartaric acid, lactic acid, stearic acid, salicylic acid, glutamic acid, ascorbic acid, pamoic acid, succinic acid, fumaric acid, maleic acid, propionic acid, hydroxymaleic acid, hydroiodic acid, phenylacetic acid, alkanoic acids such as acetic acid, HOOC-(CH)-COOH, where n is 0 to 4. Similarly, pharmaceutically acceptable cations include, but are not limited to, sodium, potassium, calcium, aluminum, lithium, and ammonium. Those skilled in the art will recognize from this disclosure and knowledge in the art additional pharmaceutically acceptable salts for the pooled disease-specific antigens provided herein, including those listed by Remington's Pharmaceutical Sciences, 17th ed., Mack Publishing Company, Easton, PA, page 1418 (1985).In general, pharmaceutically acceptable acid or base salts can be synthesized from a parent compound that contains a basic or acidic moiety by any conventional chemical method. Briefly, such salts can be prepared by reacting the free acid or base form of these compounds with a stoichiometric amount of the appropriate base or acid in a suitable solvent.

[0215]

[0315] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide or fragment thereof of the present disclosure. Such nucleic acid molecules do not need to be 100% identical to an endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide that has substantial identity to an endogenous sequence can typically hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridization" refers to the situation where a pair of nucleic acid molecules forms a double-stranded molecule with a complementary polynucleotide sequence or a portion thereof under various stringency conditions (see, for example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507). For example, a stringent salt concentration can typically be less than about 750 mM NaCl and 75 mM trisodium citrate, less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents, such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, or at least about 50% formamide. Stringent temperature conditions typically include temperatures of at least about 30°C, at least about 37°C, or at least about 42°C. Varying additional parameters, such as hybridization time, detergent concentration, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are known to those skilled in the art. Various levels of stringency can be achieved by combining these various conditions as needed. In an exemplary embodiment, hybridization may occur at 30° C. in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another exemplary embodiment, hybridization may occur at 37° C. in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA).In another exemplary embodiment, hybridization may occur at 42°C, 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations in these conditions will be readily apparent to those of skill in the art. For most applications, post-hybridization wash steps may also vary in stringency. Wash stringency conditions may be defined by salt concentration and by temperature. As noted above, wash stringency may be increased by decreasing salt concentration or increasing temperature. For example, stringent salt concentrations for wash steps may be less than about 30 mM NaCl and 3 mM trisodium citrate or less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for wash steps may include temperatures of at least about 25°C, at least about 42°C, or at least about 68°C. In an exemplary embodiment, the wash steps may be performed at 25° C., 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another exemplary embodiment, the wash steps may be performed at 42° C., 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In another exemplary embodiment, the wash steps may be performed at 68° C., 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations on these conditions will be readily apparent to one of ordinary skill in the art.Hybridization techniques are known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0216]

[0316] "Substantially identical" refers to a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). Such sequences can be at least 60%, 80%, or 85%, 90%, 95%, 96%, 97%, 98%, or even 99% or more identical at the amino acid level or at the nucleic acid level to the sequence used for comparison. Sequence identity is typically measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Typically, conservative substitutions include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach to determining the degree of identity, the BLAST program uses a 3 to 5 sequence identity scale to determine the degree of identity. 〇 A "reference" is a standard of comparison.

[0217]

[0317] The term "subject" or "patient" refers to an animal that has been the object of treatment, observation, or experiment. By way of example only, a subject includes, but is not limited to, a mammal, including, but not limited to, a human or a non-human mammal, such as a non-human primate, mouse, cow, horse, dog, sheep, or cat.

[0218]

[0318] The terms "treat," "treated," "treating," "treatment," and the like are meant to refer to reducing, preventing, or ameliorating a disorder and / or its associated symptoms (e.g., a neoplasm or tumor, or an infectious agent, or an autoimmune disease). "Treating" can refer to administering a therapy to a subject after the onset or suspected onset of a disease (e.g., cancer, or infection by an infectious agent, or an autoimmune disease). "Treating" includes the concept of "alleviating," which refers to reducing the frequency of onset or recurrence, or the severity, of any symptoms associated with the disease or other adverse effects of the illness and / or side effects associated with treatment. The term "treating" also encompasses the concept of "managing," which refers to reducing the severity of a disease or disorder in a patient, for example, extending the lifespan or chance of survival of a patient with the disease, or delaying its recurrence, for example, extending the period of remission in a patient suffering from the disease. It is understood, although not excluded, that treating a disorder or condition does not require the complete elimination of the disorder, condition, or symptoms associated therewith.

[0219]

[0319] As used herein, the terms "prevent," "preventing," "prevention," and their grammatical equivalents mean to avert or delay the onset of symptoms associated with a disease or condition in a subject who is not symptomatic at the time administration of an agent or compound is initiated.

[0220]

[0320] The term "therapeutic effect" refers to the relief to some extent of one or more symptoms of a disorder (e.g., a neoplasm, tumor, or infection by an infectious agent or autoimmune disease) or its associated pathology. As used herein, a "therapeutically effective amount" refers to an amount of an agent administered to a cell or subject, in single or multiple doses, that is effective in extending the survival chances of a patient with such a disorder beyond that expected in the absence of such treatment, reducing, preventing, or delaying one or more signs or symptoms of the disorder, etc. A "therapeutically effective amount" is intended to qualify the amount necessary to achieve a therapeutic effect. A physician or veterinarian having ordinary skill in the art can readily determine and prescribe the "therapeutically effective amount" (e.g., ED50) of the required pharmaceutical composition. For example, a physician or veterinarian can start doses of the disclosed compounds used in the pharmaceutical composition at levels lower than those required to achieve the desired therapeutic effect, and gradually increase the dosage until the desired effect is achieved. Disease, condition, and disorder are used interchangeably herein.

[0221]

[0321] Those skilled in the art will recognize that the terms "peptide tag," "affinity tag," "epitope tag," or "affinity acceptor tag" are used interchangeably herein. As used herein, the term "affinity acceptor tag" refers to an amino acid sequence that allows the tagged protein to be easily detected or purified, for example, by affinity purification. The affinity acceptor tag is generally (but not necessarily) located at or near the N- or C-terminus of the HLA allele. A variety of peptide tags are known in the art. Non-limiting examples include a polyhistidine tag (e.g., 4 to 15 consecutive His residues (SEQ ID NO: 4), e.g., 8 consecutive His residues (SEQ ID NO: 5)); a poly-histidine-glycine tag; an HA tag (e.g., Field et al., Mol. Cell. Biol., 8:2159, 1988); a c-myc tag (e.g., Evans et al., Mol. Cell. Biol. 5:3610, 1985); a herpes simplex virus glycoprotein D (gD) tag (e.g., Paborsky et al., Protein Engineering, 3:547, 1990); FLAG tag (e.g., Hopp et al., BioTechnology, 6:1204, 1988; U.S. Patent Nos. 4,703,004 and 4,851,341); KT3 epitope tag (e.g., Martine et al., Science, 255:192, 1992); tubulin epitope tag (e.g., Skinner, Biol. Chem., 266:15173, 1991); T7 gene 10 protein peptide tag (e.g., Lutz-Freyemuth et al., Proc. Natl. Acad. Sci. USA, 87:6393, 1990); streptavidin tag (StrepTag.™ or StrepTagII.™; see, e.g., Schmidt et al., J. Mol. Biol., 255(5):753-766, 1996 or U.S. Pat. No. 5,506,121; commercially available from Sigma-Genosys); or the VSV-G epitope tag derived from the vesicular stomatitis virus glycoprotein; or the V5 tag derived from a small epitope (Pk) found in the P and V proteins of the paramyxovirus simian virus 5 (SV5).In some embodiments, the affinity acceptor tag is an "epitope tag," a type of peptide tag that adds a recognizable epitope (antibody binding site) to an HLA protein to provide binding of the corresponding antibody, thereby enabling identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are Protein A or Protein G, which bind to IgG. In some embodiments, the matrix of IgG Sepharose 6 Fast Flow chromatography resin is covalently bound to human IgG. This resin allows for high flow rates for rapid and convenient purification of proteins bearing a Protein A tag. Numerous other tag components are known to those of skill in the art and are envisioned and contemplated herein. Any peptide tag can be used as long as it can be expressed as an element of an affinity acceptor-tagged HLA peptide complex.

[0222]

[0322] As used herein, the term "affinity molecule" refers to a molecule or ligand that binds to an affinity acceptor peptide with chemical specificity. Chemical specificity is the ability of a protein's binding site to bind to a specific ligand. The fewer ligands a protein can bind, the higher its specificity. Specificity describes the strength of binding between a given protein and a ligand. This relationship can be described by the dissociation constant (KD), which characterizes the equilibrium between binding and non-binding in a protein-ligand system.

[0223]

[0323] The term "affinity acceptor-tagged HLA peptide complex" refers to a complex comprising an HLA class I or class II associated peptide or a portion thereof that specifically binds to a monoallelic recombinant HLA class I or class II peptide, including an affinity acceptor peptide.

[0224]

[0324] The terms "specific binding" or "specifically binds," when used in reference to the interaction of an affinity molecule with an affinity acceptor tag, or an epitope with an HLA peptide, mean that the interaction is dependent on the presence of a particular structure (e.g., an antigenic determinant or epitope) on the protein; in other words, the affinity molecule recognizes and binds to a particular affinity acceptor peptide structure rather than the protein generally.

[0225]

[0325] As used herein, the term "affinity" refers to a measure of the strength of binding between two members of a binding pair, e.g., an "affinity acceptor tag" and an "affinity molecule," and an HLA-binding peptide and an HLA class I or II molecule. KD is the dissociation constant and has units of molar. The affinity constant is the reciprocal of the dissociation constant. Affinity constant is sometimes used as a generic term to describe the chemical entity. It is a direct measure of the energy of binding. Affinity can be determined experimentally, for example, by surface plasmon resonance (SPR) using a commercially available Biacore SPR instrument. Affinity is measured as the concentration at which 50% of the peptide is displaced, the inhibitory concentration 50 (IC 50 ) can also be expressed as lnIC 50 is IC 50 Koff refers to the natural logarithm of, for example, the off-rate constant for dissociation of an affinity molecule from an affinity acceptor-tagged HLA peptide complex.

[0226]

[0326] In some embodiments, affinity acceptor-tagged HLA peptide complexes containing a biotin acceptor peptide (BAP) are immunopurified from complex cell mixtures using streptavidin / NeutrAvidin beads. Biotin-avidin / streptavidin binding is the strongest known natural non-covalent interaction. This property is exploited as a biological tool for a wide range of applications, such as immunopurification of proteins to which biotin is covalently bound. In an exemplary embodiment, a nucleic acid sequence encoding an HLA allele incorporates a biotin acceptor peptide (BAP) as an affinity acceptor tag for immunopurification. The BAP can be specifically biotinylated at a single lysine residue within the tag in vivo or in vitro (e.g., U.S. Patent Nos. 5,723,584; 5,874,239; and 5,932,433; and UK Patent No. GB2370039). BAPs are typically 15 amino acids long and contain a single lysine as a biotin acceptor residue. In some embodiments, the BAP is located at or near the N- or C-terminus of a monoallelic HLA peptide. In some embodiments, the BAP is located between the heavy chain domain and the β2 microglobulin domain of an HLA class I peptide. In some embodiments, the BAP is located between the β chain domain and the α chain domain of an HLA class II peptide. In some embodiments, the BAP is located in the loop region between the α1, α2, and α3 domains of the heavy chain of HLA class I, or between the α1 and α2 and β1 and β2 domains of the α chain and β chain of HLA class II, respectively.

[0227]

[0327] As used herein, the term "biotin" refers to the compound biotin itself and its analogs, derivatives, and variants. Thus, the term "biotin" includes biotin (cis-hexahydro-2-oxo-1H-thieno[3,4]imidazole-4-pentanoic acid) and any derivatives and analogs thereof, including biotin-like compounds. Such compounds include, for example, biotin-eN-lysine, biocytin hydrazide, 2-iminobiotin and amino or sulfhydryl derivatives of biotinyl-E-aminocaproic acid-N-hydroxysuccinimide ester, sulfosuccinimidoiminoviotin, biotin bromoacetylhydrazide, p-diazobenzoylbiocytin, 3-(N-maleimidopropionyl)biocytin, desthiobiotin, etc. The term "biotin" also includes biotin variants that are capable of specifically binding to one or more rhizavidin, avidin, streptavidin, tamavidin moieties or other avidin-like peptides.

[0228]

[0328] As used herein, "PPV determination method" can refer to a presentation PPV determination method. For example, the "PPV determination method" includes the following steps: (a) using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, to process amino acid information of a plurality of test peptide sequences to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by the class II HLA alleles of a cell, for example, the class II HLA alleles of a cell of a subject, can present a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences includes (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 500 test peptide sequences, including at least 499 decoy peptide sequences contained in proteins encoded by the genome of an organism, wherein the organism is the same species as the subject, and the plurality of test peptide sequences has a ratio of the number of hit peptide sequences to the number of decoy peptide sequences of less than 1, for example, a ratio of at least one hit peptide sequence to at least 499 decoy peptide sequences of 1:499; (b) using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by the class II HLA alleles of a cell, for example, the class II HLA alleles of a cell of a subject, can present a given test peptide sequence of the plurality of test peptide sequences; (c) identifying or calling the top percentage of a plurality of test peptide sequences, for example, the top 0.2% of a plurality of test peptide sequences, as presented by HLA alleles; and (c) calculating the PPV of an HLA peptide presentation prediction model, wherein the PPV is the percentage of the plurality of test peptide sequences identified or called as presented by the Class II HLA alleles of the cell that are peptides observed by mass spectrometry to be presented by the Class II HLA alleles of the cell. In some embodiments, the decoy peptides are the same length, i.e., contain the same number of amino acids as the hit peptides. In some embodiments, the decoy peptides may contain one more or one less amino acid compared to the hit peptides. In some embodiments, the decoy peptides are peptides that are endogenous peptides. In some embodiments, the decoy peptides are synthetic peptides. In some embodiments, the decoy peptides areThe decoy peptide is an endogenous peptide identified by mass spectrometry as binding to a first MHC class I or class II protein, wherein the first MHC class I or class II protein is different from a second MHC class I or class II protein that binds to the hit peptide. In some embodiments, the decoy peptide may be a scrambled peptide, e.g., the decoy peptide may include an amino acid sequence in which amino acid positions within the peptide length are rearranged compared to those of the hit peptide. In some embodiments, the PPV determination method may be a display PPV determination method. In some embodiments, the ratio of the number of hit peptide sequences to the number of decoy peptide sequences is about 1:10, 1:20, 1:50, 1:100, 1:250, 1:500, 1:1000, 1:1500, 1:2000, 1:2500, 1:5000, 1:7500, 1:10000, 1:25000, 1:50000, or 1:100000. In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences. In some embodiments, the at least 499 decoy peptide sequences are at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700,4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 87 00, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 320 00, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 875 The decoy peptide sequences include 00, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1,000,000 decoy peptide sequences. In some embodiments, the at least 500 test peptide sequences are at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000,4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8 200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 290 00, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500 , 85,000, 87,500, 90,000, 92,500, 95,000, 97,500, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 325,000, 350,000, 375,000, 400,000, 425,000, 450,000, 475,000, 500,000, 600,000, 700,000, 800,000, 900,000 or 1,000,000 test peptide sequences. In some embodiments, identifying or calling the top percentage of the plurality of test peptide sequences as presented by class II HLA alleles of the cell includes identifying or calling the top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.00%, 7.10%, 7.20%, 7.20%, 7.30%, 7.40%, 7.50%, 7.50%, 7.60%, 7.70%,1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.7 0%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.6 In some embodiments, the cells are monoallelic cells.

[0229]

[0329] As used herein, "PPV determination method" can refer to a binding PPV determination method. For example, "PPV determination method" can refer to a step of (a) processing amino acid information of a plurality of test peptide sequences using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, to generate a plurality of test presentation predictions, wherein each test presentation prediction is associated with a class I or class II HLA allele of a cell, for example, a class I or class II HLA allele of a cell of a subject. and (c) calculating a PPV of the HLA peptide binding prediction model, the PPV being a percentage of the HLA peptide binding prediction model that indicates the likelihood that one or more proteins encoded by HLA alleles can bind to a given test peptide sequence of a plurality of test peptide sequences, the plurality of test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 20 test peptide sequences including at least 19 decoy peptide sequences contained within a protein comprising at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, the plurality of test peptide sequences having a ratio of the number of hit peptide sequences to the number of decoy peptide sequences of less than 1, e.g., a ratio of at least one hit peptide sequence to at least 19 decoy peptide sequences of 1:19; (b) identifying or calling a top percentage of the plurality of test peptide sequences as binding to the HLA protein, e.g., the top 5% of the plurality of test peptide sequences; and (c) calculating a PPV of the HLA peptide binding prediction model, the PPV being a percentage of the class I or class II HLA alleles of the cell that are peptides observed by mass spectrometry as being presented by a class I or class II HLA allele of the cell. This may refer to a method in which the ratio of a plurality of test peptide sequences identified or called as binding to an HLA allele is about 1:2, 1:3, 1:4, 1:5, 1:10, 1:20, 1:25, 1:30, 1:40, 1:50, 1:75, 1:100, 1:200, 1:250, 1:500, or 1:1000. In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95,12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130 Contains 9, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences. In some embodiments, the at least 19 decoy peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 310 0, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 56 00, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8 100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000,15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000 , 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 6500 The decoy peptide sequence may include 0, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000, or 1,000,000 decoy peptide sequences. In some embodiments, the at least 20 test peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200,6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 980 0, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 3800 0, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 950 00, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1 million test peptide sequences. In some embodiments, identifying or calling the top percentage of the plurality of test peptide sequences as presented by class II HLA alleles of the cell comprises identifying or calling the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% as presented by class II HLA alleles of the cell.The cells are monoallelic cells.

[0230] Human leukocyte antigen (HLA) system

[0330] The immune system can be divided into two functional subsystems: the innate and adaptive immune systems. The innate immune system is the first line of defense against infection; most potential pathogens are rapidly neutralized by this system before they can cause significant infection. The adaptive immune system reacts with molecular structures called antigens of invading organisms. Unlike the innate immune system, the adaptive immune system is highly pathogen-specific. Acquired immunity can also provide long-term protection; for example, individuals who recover from measles are protected from measles for the rest of their lives. There are two types of adaptive immune responses: humoral and cell-mediated. In the humoral immune response, antibodies secreted into bodily fluids by B cells bind to pathogen-derived antigens and lead to their elimination through various mechanisms, such as complement-mediated lysis. In the cell-mediated immune response, T cells, which can destroy other cells, are activated. For example, if disease-related proteins are present in cells, they are proteolytically fragmented into peptides within the cell. Certain cellular proteins then attach themselves to the antigens or peptides formed in this way and transport them to the surface of the cell where they are presented to the body's molecular defense mechanisms in T cells. Cytotoxic T cells recognize these antigens and kill the antigen-bearing cells.

[0231]

[0331] The terms "major histocompatibility complex (MHC)," "MHC molecule," or "MHC protein" refer to proteins that can bind to peptides resulting from proteolytic cleavage of protein antigens, present potential T cell epitopes, transport them to the cell surface, and present the peptides on specific cells, such as cytotoxic T lymphocytes or helper T cells. Human MHC is also called the HLA complex. Thus, the terms "human leukocyte antigen (HLA) system," "HLA molecule," or "HLA protein" refer to the gene complex that encodes MHC proteins in humans. The term MHC is referred to as the "H-2" complex in mouse species. Those skilled in the art will recognize that the terms "major histocompatibility complex (MHC)," "MHC molecule," "MHC protein," and "human leukocyte antigen (HLA) system," "HLA molecule," and "HLA protein" are used interchangeably herein.

[0232]

[0332] HLA proteins are classified into two types, designated HLA class I and HLA class II. Although the structures of the two HLA class proteins are very similar, they have very different functions. HLA class I proteins are present on the surface of most cells in the body, including most tumor cells. HLA class I proteins are loaded with antigens, usually derived from endogenous proteins or pathogens present inside the cell, and then presented to naive or cytotoxic T lymphocytes (CTLs). HLA class II proteins are present on antigen-presenting cells (APCs), including, but not limited to, dendritic cells, B cells, and macrophages. They primarily present processed peptides from external, e.g., extracellular, antigen sources to helper T cells. Most peptides bound by HLA class I proteins are derived from cytoplasmic proteins produced in the organism's own healthy host cells and do not normally stimulate an immune response.

[0233]

[0333] HLA class I molecules consist of two non-covalently linked polypeptide chains: the HLA-encoded α chain (heavy chain, 44 to 47 kD) and a non-HLA-encoded subunit called β2-microglobulin (or β2m) (12 kD). The α chain has three extracellular domains, α1, α2, and α3, and a transmembrane region, and the α1 and α2 regions can bind peptides of approximately 7 to 13 amino acids (e.g., approximately 8 to 11 amino acids, or 9 or 10 amino acids). HLA class I molecules bind peptides with a suitable binding motif and present them to cytotoxic T lymphocytes. The HLA class I heavy chain can be the protein product of an HLA-A allele (also called an HLA-A monomer), or the protein product of an HLA-B allele (also called an HLA-B monomer), or the protein product of an HLA-C allele (also called an HLA-C monomer), each of which is complexed with β2-microglobulin. α1 is present on the non-HLA protein β2m; β2m is encoded by the beta-2-microglobulin gene located on human chromosome 15. The α3 domain is connected to the transmembrane region, anchoring the HLA class I molecule to the cell membrane. The peptide to be presented is held in the floor of the peptide-binding groove in the central region of the α1 / α2 heterodimer (a molecule composed of two non-identical subunits). HLA class IA, HLA class IB, and HLA class IC are highly polymorphic. Each of the HLA class 1-A gene (HLA-A gene), HLA class 1-B gene (HLA-B gene), and HLA class 1-C gene (HLA-C gene) contains eight exons: exon 1 encodes the leader peptide, exons 2 and 3 encode the α1 and α2 domains, exon 5 encodes the transmembrane region, and exons 6 and 7 encode the cytoplasmic tail. Polymorphisms in exons 2 and 3 are responsible for the peptide binding specificity of each class 1 molecule. HLA class IB genes (HLA-B) have a large number of possible variations, expression patterns, and antigens presented.This group is subdivided into those encoded within the HLA locus, e.g., HLA-E, HLA-F, HLA-G, and non-HLA stress ligands, e.g., ULBPs, Rae1, and H60. The antigens / ligands for many of these molecules remain unknown, but they can interact with CD8+ T cells, NKT cells, and NK cells, respectively.

[0234]

[0334] In some embodiments, the present disclosure utilizes non-classical HLA class IE alleles. HLA-E molecules are recognized by natural killer (NK) cells and CD8+ T cells. HLA-E is expressed in almost all tissues, including lung, liver, skin, and placental cells. HLA-E expression has also been detected in solid tumors (e.g., osteosarcoma and melanoma). HLA-E molecules bind to TCRs expressed on CD8+ T cells, resulting in T cell activation. HLA-E is also known to bind to the CD94 / NKG2 receptor expressed on NK cells and CD8+ T cells. CD94 can pair with several different isoforms of NKG2 to form receptors that can either inhibit (NKG2A, NKG2B) or promote (NKG2C) cell activation. HLA-E can bind peptides derived from amino acid residues 3-11 of the leader sequences of most HLA-A, -B, -C, and -G molecules, but cannot bind its own leader peptide. HLA-E has also been shown to present peptides derived from endogenous proteins similar to HLA-A, -B, or -C alleles. Under physiological conditions, the association of CD94 / NKG2A with HLA-E loaded with peptides derived from HLA class I leader sequences typically induces inhibitory signals. Cytomegalovirus (CMV) utilizes a mechanism for evading NK cell immune surveillance through the expression of the UL40 glycoprotein, which mimics the HLA-A leader. However, it has also been reported that CD8+ T cells can recognize HLA-E loaded with the UL40 peptide from the Toledo strain of CMV and play a role in defense against CMV. Numerous studies have revealed several important functions of HLA-E in infectious diseases and cancer.

[0235]

[0335] Peptide antigens attach themselves to HLA class I molecules through competitive affinity binding in the endoplasmic reticulum before they are presented on the cell surface. Herein, the affinity of each peptide antigen is directly related to its amino acid sequence and the presence of specific binding motifs at specific positions within the amino acid sequence. If the sequence of such peptides is known, it is possible to use, for example, peptide vaccines to manipulate the immune system against diseased cells.

[0236]

[0336] MHC molecules are highly polymorphic, meaning there are numerous MHC variants. Each variant is encoded by a variation in the protein-coding gene, and each such variant gene is called an allele. In humans, MHC is known as the human leukocyte antigen (HLA) molecule, which includes three types of HLA class II molecules: DP, DQ, and DR. HLA class II peptides (Figure 1) have two chains, α and β, each with two domains—α1 and α2, and β1 and β2—and each chain has a transmembrane domain, α2 and β2, respectively, that anchor the HLA class II molecule to the cell membrane. The peptide-binding groove is formed by a heterodimer of α1 and β1. The most widely studied HLA-DR molecule contains DRA and DRB, which correspond to the α and β domains, respectively. DRB is diverse, while DRA is nearly identical. As a result, the binding specificity of a DRB allele is that of the corresponding HLA-DR. Each MHC protein has its own binding specificity, meaning that the set of peptides binding to one MHC molecule may differ from that of another. Classical molecules present peptides to CD4+ lymphocytes. Non-classical molecules with intracellular functions, accessories that are not exposed to the plasma membrane but reside in the internal membrane of lysosomes, usually load antigenic peptides onto classical HLA class II molecules.

[0237]

[0337] In the HLA class II system, phagocytes such as macrophages and immature dendritic cells endocytose proteins into phagosomes—although B cells typically endocytose into endosomes—and fuse with lysosomes, where acidic enzymes cleave the endocytosed proteins into numerous, diverse peptides. Autophagy is another source of HLA class II peptides. Through physicochemical interactions with host-derived HLA class II variants encoded in the host genome, specific peptides exhibit immunodominance and are loaded onto HLA class II molecules, which are transported to the cell surface and externalized. The most studied subclasses of HLA class II genes are: HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRA, and HLA-DRB1.

[0238]

[0338] Peptide presentation to CD4+ helper T cells by HLA class II molecules is necessary for immune responses to foreign antigens (Roche and Furuta, 2015). Upon activation, CD4+ T cells promote B cell differentiation and antibody production, as well as CD8+ T cell (CTL) responses. CD4+ T cells also secrete cytokines and chemokines that activate and induce differentiation of other immune cells. HLA class II molecules are heterodimers of α- and β-chains that interact to form a peptide-binding groove that is more open than the HLA class I peptide-binding groove (Unanue et al., 2016). Peptides that bind to HLA class II molecules are thought to have a nine-amino acid binding core that protrudes from the groove, including adjacent residues on either the N- or C-terminus (Jardetzky et al., 1996; Stern et al., 1994). These peptides are typically 12–16 amino acids in length and often contain three to four anchor residues at positions P1, P4, P6 / 7, and P9 of the binding register ( Rossjohn et al., 2015 ).

[0239]

[0339] HLA alleles are expressed in a codominant manner, meaning that alleles (variants) inherited from both parents are equally expressed. For example, if each individual possesses two alleles of each of the three class I genes (HLA-A, HLA-B, and HLA-C), six different types of HLA class II may be expressed. At the HLA class II locus, each individual inherits a pair of HLA-DP genes (DPA1 and DPB1, encoding the α and β chains), HLA-DQ (DQA1 and DQB1, encoding the α and β chains), one gene HLA-DRα (DRA1), and one or more genes HLA-DRβ (DRB1 and DRB3, -4, or -5). HLA-DRB1, for example, has approximately more than 400 known alleles. This means that a heterozygous individual may inherit six or eight functional HLA class II alleles: three or more from each parent. As a result, HLA genes are highly polymorphic; many different alleles exist in different individuals within a population. Genes encoding HLA proteins have a large number of possible variations, allowing each individual's immune system to respond to a wide range of foreign invaders. Some HLA genes have hundreds of identified versions (alleles), each of which is given a specific number. In some embodiments, HLA class I alleles are HLA-A*02:01, HLA-B*14:02, HLA-A*23:01, and HLA-E*01:01 (non-classical). In some embodiments, HLA class II alleles are HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*11:01, HLA-DRB*15:01, and HLA-DRB*07:01.

[0240]

[0340] The subject-specific HLA alleles or the subject's HLA genotype can be determined by any method known in the art. In an exemplary embodiment, the HLA genotype is determined by any method described in International Patent Application No. PCT / US2014 / 068746, filed June 11, 2015 as WO2015085147, which is incorporated herein by reference in its entirety. Briefly, the method includes determining a polymorphism genotype, which may include generating an alignment of reads extracted from a sequencing dataset to a genetic reference set comprising allelic variants of the polymorphic gene; determining a first posterior probability or posterior probability-derived score for each allelic variant in the alignment; identifying an allelic variant having the maximum first posterior probability or posterior probability-derived score as the first allelic variant; identifying one or more overlapping reads aligned with the first allelic variant and one or more other allelic variants; determining a second posterior probability or posterior probability-derived score for the one or more other allelic variants using a weighting factor; identifying a second allelic variant by selecting the allelic variant having the maximum second posterior probability or posterior probability-derived score, wherein the first and second allelic variants define a genotype for the polymorphic gene; and providing an output of the first and second allelic variants.

[0241]

[0341] In some embodiments, the MHC class II peptide:antigen peptide binding and presentation prediction method described herein has the ability to predict binders from a wide repertoire of MHC class II peptides encoded by an individual's HLA alleles. In some embodiments, the MAPTAC technology is trained using a large database of mass spectrometry-verified HLA-matched peptides. In some embodiments, the large database of mass spectrometry-verified HLA-matched peptides contains more than 1.2 x 10 such HLA-matched peptides. In some embodiments, the large database of mass spectrometry-verified HLA-matched peptides covers more than 150 HLA alleles, including both MHC class I and class II allele subtypes. In some embodiments, the database covers at least 95% of the US population for HLA-I and HLA-II (DR subtypes).

[0242]

[0342] As described herein, there is extensive evidence in both animals and humans that mutated epitopes are effective in inducing immune responses, that cases of spontaneous tumor regression or long-term survival correlate with CD8+ T cell responses to mutated epitopes, and that "immune editing" can be traced to changes in the expression of overt mutated antigens in mice and humans.

[0243]

[0343] Sequencing technology has revealed that each tumor contains multiple patient-specific mutations that alter the protein-coding content of genes. Such mutations create altered proteins ranging from single amino acid changes (caused by missense mutations) to the addition of long stretches of novel amino acid sequences through frameshifts, stop codon readthrough, or translation of intronic regions (novel open reading frame mutations; neoORFs). These mutant proteins are valuable targets for the host immune response to tumors, distinct from native proteins; they are not subject to the immune-damaging effects of self-tolerance. Therefore, mutant proteins are more likely to be immunogenic and more specific to tumor cells compared with the patient's normal cells. Essentially, short peptides (8–24 amino acids long) containing cancer-associated mutations are candidates for cancer immunotherapy.

[0244]

[0344] In some embodiments, the algorithms driving the prediction methods may be further utilized to call mutations in peptides, hi some embodiments, the prediction methods may be used to determine driver mutation status, and / or RNA expression status, and / or intrapeptide cleavage predictions.

[0245]

[0345] The term "T cells" includes CD4+ T cells and CD8+ T cells. The term T cells also includes both type 1 and type 2 helper T cells. As used herein, T cells are generally classified into two main classes: helper T (TH) cells and cytotoxic T lymphocytes (CTLs) according to function and cell surface antigens (cluster differentiation antigens or CDs) that also facilitate binding of the T cell receptor to antigen.

[0246]

[0346] Mature helper T (TH) cells express the surface protein CD4 and are called CD4+ T cells. Following T cell development, mature, naive T cells leave the thymus and begin to disseminate throughout the body, including lymph nodes. Naive T cells are T cells that have never been exposed to the antigen to which they are programmed. Like all T cells, they express the T cell receptor-CD3 complex. The T cell receptor (TCR) consists of both constant and variable regions. The variable region determines which antigens the T cell can respond to. CD4+ T cells have TCRs with affinity for MHC class II proteins, and CD4 is involved in determining MHC affinity during maturation in the thymus. MHC class II proteins are found only on the surface of specialized antigen-presenting cells (APCs). Specialized APCs are primarily dendritic cells, macrophages, and B cells, but dendritic cells are the only cell population that continuously (at all times) express MHC class II. Some APCs also bind native (or unprocessed) antigens to their surface, e.g., follicular dendritic cells, but unprocessed antigens do not interact with T cells and are not involved in their activation. Peptide antigens that bind to HLA class I proteins are typically shorter than peptide antigens that bind to HLA class II proteins.

[0247]

[0347] Cytotoxic T lymphocytes (CTLs), also known as cytotoxic T cells, cytolytic T cells, CD8+ T cells, or killer T cells, are lymphocytes that induce apoptosis in target cells. CTLs form antigen-specific conjugates with target cells through the interaction of their TCRs with processed antigens (Ag) on ​​the target cell surface, resulting in apoptosis of the target cell. The apoptotic bodies are removed by macrophages. The term "CTL response" refers to the primary immune response mediated by CTL cells. Cytotoxic T lymphocytes possess both T cell receptors (TCRs) and CD8 molecules on their surface. T cell receptors can recognize and bind peptides complexed with HLA class I molecules. Each cytotoxic T lymphocyte expresses a unique T cell receptor that can bind to a specific MHC / peptide complex. Most cytotoxic T cells express a T cell receptor (TCR) that can recognize a specific antigen. For TCRs to bind to HLA class I molecules, they must be accompanied by a glycoprotein called CD8, which binds to the constant portion of the HLA class I molecule. Therefore, these T cells are called CD8+ T cells. The affinity between CD8 and MHC molecules keeps T cells tightly bound to target cells during antigen-specific activation. Once activated, CD8+ T cells are recognized as T cells and are generally classified as having a predetermined cytotoxic role within the immune system. However, CD8+ T cells also have the ability to produce some cytokines.

[0248]

[0348] A "T cell receptor (TCR)" is a cell surface receptor that participates in T cell activation in response to antigen presentation. TCRs are generally composed of two chains, alpha and beta, which assemble to form heterodimers and associate with CD3-transducing subunits to form the T cell receptor complex present on the cell surface. Each of the alpha and beta chains of the TCR consists of immunoglobulin-like N-terminal variable (V) and constant (C) regions, a hydrophobic transmembrane domain, and a short intracytoplasmic region. The variable regions of the alpha and beta chains, as with immunoglobulin molecules, are generated by V(D)J recombination, creating a wide diversity of antigen specificities within a population of T cells. However, in contrast to immunoglobulins that recognize intact antigens, T cells are activated by processed peptide fragments associated with MHC molecules, introducing another aspect of antigen recognition by T cells, known as MHC restriction. Recognition of MHC differences between donor and recipient through the T cell receptor leads to T cell proliferation and the potential development of graft-versus-host disease (GVHD). Normal surface expression of the TCR has been shown to depend on the coordinate synthesis and assembly of all seven components of the complex (Ashwell and Klusner 1990). Inactivation of TCRα or TCRβ can result in removal of the TCR from the surface of the T cell, preventing recognition of alloantigen and thereby GVHD. However, TCR disruption generally results in removal of the CD3 signaling component, altering the means for further T cell expansion.

[0249]

[0349] The term "HLA peptidome" refers to a pool of peptides that specifically interact with a particular HLA class and can encompass thousands of different sequences. The HLA peptidome includes peptide diversity and is derived from both normal and abnormal proteins expressed in cells. Thus, the HLA peptidome can be studied to identify cancer-specific peptides for tumor immunotherapy and as a source of information on protein synthesis and degradation schemes within cancer cells. In some embodiments, the HLA peptidome is a pool of soluble HLA peptides (sHLA). In some embodiments, the HLA peptidome is a pool of membrane-associated HLA (mHLA).

[0250]

[0350] "Antigen-presenting cells" or "APCs" include professional antigen-presenting cells (e.g., B lymphocytes, macrophages, monocytes, dendritic cells, Langerhans cells) and other antigen-presenting cells (e.g., keratinocytes, endothelial cells, astrocytes, fibroblasts, oligodendrocytes, thymic epithelial cells, thyroid epithelial cells, glial cells (brain), pancreatic beta cells, and vascular endothelial cells). "Antigen-presenting cells" or "APCs" are cells that express major histocompatibility complex (MHC) molecules and can display foreign antigens complexed with MHC on their surface.

[0251] Monoallelic HLA cell lines

[0351] Single-allele cell lines expressing either a single HLA class I allele, a single pair of HLA class II alleles, or a single pair of HLA class I and class II alleles can be generated by transducing or transfecting an appropriate cell population with a polynucleic acid, e.g., a vector encoding a single HLA allele. Suitable cell populations include, for example, HLA class I-deficient cell lines in which a single HLA class I allele is exogenously expressed, HLA class II-deficient cell lines in which a single exogenous pair of HLA class II alleles is expressed, or class I and class II-deficient cells in which a single pair of HLA class I and / or class II alleles is exogenously expressed. In one exemplary embodiment, the HLA class I-deficient B cell line is B721.221. However, it will be apparent to those skilled in the art that other cell populations that are HLA class I- and / or class II-deficient can be generated. An exemplary method for deleting / inactivating endogenous HLA class I or HLA class II genes includes, for example, CRISPR-Cas9-mediated genome editing in THP-1 cells. In some embodiments, the population of cells is professional antigen-presenting cells, such as macrophages, B cells, and dendritic cells. The cells can be B cells or dendritic cells. In some embodiments, the cells are tumor cells or cells derived from a tumor cell line. In some embodiments, the cells are isolated from a patient. In some embodiments, the cells contain an infectious agent or a portion thereof. In some embodiments, the population of cells comprises at least 10 cells. 7In some embodiments, the population of cells is further modified by increasing or decreasing the expression and / or activity of at least one gene. In some embodiments, the gene encodes a member of the immunoproteasome. The immunoproteasome is known to be involved in processing HLA class I-bound peptides and includes the LMP2 (β1i), MECL-1 (β2i), and LMP7 (β5i) subunits. The immunoproteasome can also be induced by interferon gamma. Thus, in some embodiments, the population of cells can be contacted with one or more cytokines, growth factors, or other proteins. The cells can be stimulated with inflammatory cytokines, such as interferon gamma, IL-10, IL-6, and / or TNF-α. The population of cells can also be subjected to various environmental conditions, such as stress (heat stress, oxygen deprivation, glucose starvation, DNA-damaging agents, etc.). In some embodiments, the cells are contacted with one or more of a chemotherapeutic agent, radiation, targeted therapy, or immunotherapy. Therefore, the methods disclosed herein can be used to study the effects of various genes or conditions on HLA peptide processing and presentation. In some embodiments, the conditions used are selected to match the condition of the patient in which the population of HLA peptides is identified.

[0252]

[0352] The single HLA alleles of the present disclosure can be encoded and expressed using a virus-based system (e.g., an adenovirus system, an adeno-associated virus (AAV) vector, a poxvirus, or a lentivirus). Plasmids that can be used for adeno-associated virus, adenovirus, and lentivirus delivery have been previously described (see, e.g., U.S. Patent Nos. 6,955,808 and 6,943,019 and U.S. Patent Application No. 20080254008, which are incorporated herein by reference). Of the vectors that can be used in practicing the present disclosure, integration into the cellular host genome is possible using retroviral gene transfer methods, often resulting in long-term expression of the inserted transgene. In an exemplary embodiment, the retrovirus is a lentivirus. Additionally, high transduction efficiencies have been observed in a number of different T cell types and target tissues. The tropism of retroviruses can be altered by incorporating foreign envelope proteins, expanding the potential target population of cells. Retroviruses can be engineered to allow conditional expression of inserted transgenes, so that only certain cell types are infected by lentiviruses. Cell-type specific promoters can be used to target expression in specific cell types. Lentiviral vectors are retroviral vectors (and for this reason, both lentiviral and retroviral vectors can be used in the implementation of the present disclosure). Furthermore, lentiviral vectors can transduce or infect non-dividing cells and typically produce high viral titers.

[0253]

[0353] The choice of retroviral gene transfer system can depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats that have the capacity to package up to 6-10 kb of foreign sequence. The minimal cis-acting LTRs are sufficient for vector replication and packaging, and are then used to integrate the desired nucleic acid into target cells, resulting in persistent expression. Widely used retroviral vectors that can be used in the practice of the present disclosure include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., (1992) J. Virol. 66:2731-2739; Johann et al., (1992) J. Virol. 66:1635-1640; Sommnerfelt et al., (1990) Virol. 176:58-59; Wilson et al., (1998) J. Virol. 63:2374-2378; Miller et al., (1991) J. Virol. 65:2220-2224; PCT / US94 / 05700). Also useful in the practice of the present disclosure are minimal non-primate lentiviral vectors, such as lentiviral vectors based on equine infectious anemia virus (EIAV) (see, e.g., Balagaan, (2006) J Gene Med;8:275-285, Published online 21 November 2005 in Wiley InterScience DOI:10.1002 / jgm.845). The vector may have a cytomegalovirus (CMV) promoter driving expression of the target gene. Accordingly, the present disclosure contemplates viral vectors, including retroviral and lentiviral vectors, among the vector(s) useful in the practice of the present disclosure.

[0254]

[0354] Any HLA allele can be expressed in the cell population. In an exemplary embodiment, the HLA allele is an HLA class I allele. In some embodiments, the HLA class I allele is an HLA-A allele or an HLA-B allele. In some embodiments, the HLA allele is an HLA class II allele. The sequences of HLA class I and class II alleles can be found in the IPD-IMGT / HLA database. Exemplary HLA alleles include, but are not limited to, HLA-A*02:01, HLA-B*14:02, HLA-A*23:01, HLA-E*01:01, HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*11:01, HLA-DRB*15:01, and HLA-DRB*07:01.

[0255]

[0355] In some embodiments, HLA allele is selected to correspond to the genotype of interest.In some embodiments, HLA allele is a variant HLA allele, and can be an allele that does not naturally occur in affected patients or an allele that naturally occurs.The method disclosed herein also has the advantage of identifying the HLA binding peptide for the HLA allele associated with various disorders and the allele that exists at low frequency.Therefore, in some embodiments, the method provided herein can identify the HLA allele even if it exists at a frequency of less than 1% in a population, such as a Caucasian population.

[0256]

[0356] In some embodiments, the nucleic acid sequence encoding the HLA allele further comprises an affinity acceptor tag that can be used to immunopurify the HLA protein. Suitable tags are known in the art.In some embodiments, the affinity acceptor tag is a polyhistidine tag, polyhistidine glycine tag, polyarginine tag, polyaspartic acid tag, polycysteine ​​tag, polyphenylalanine, c-myc tag, herpes simplex virus glycoprotein D (gD) tag, FLAG tag, KT3 epitope tag, tubulin epitope tag, T7 gene 10 protein peptide tag, streptavidin tag, streptavidin binding peptide (SPB) tag, Strep tag, Strep tag II, albumin binding protein (ABP) tag, alkaline phosphatase (AP) tag, bluetongue virus tag (B tag), calmodulin binding peptide (CBP) tag, chloramphenicol acetyltransferase (ATP) tag, or a combination thereof. Catalytic enzyme (CAT) tag, choline-binding domain (CBD) tag, chitin-binding domain (CBD) tag, cellulose-binding domain (CBP) tag, dihydrofolate reductase (DHFR) tag, galactose-binding protein (GBP) tag, maltose-binding protein (MBP), glutathione-S-transferase (GST), Glu-Glu (EE) tag, human influenza hemagglutinin (HA) tag, horseradish peroxidase (HRP) tag, NE tag, HSV tag, ketosteroid isomerase (KSI) tag, KT3 tag, LacZ tag, luciferase tag, NusA tag, PDZ domain tag, AviTag, calmodulin tag, E tag, S tag, SBP tag, SoftAg These include Softag 3, TC tag, VSV tag, Xpress tag, Isopeptag, SpyTag, SnoopTag, Profinity eXact tag, Protein C tag, S1 tag, S tag, biotin carboxy carrier protein (BCCP) tag, green fluorescent protein (GFP) tag, small ubiquitin-like modifier (SUMO) tag, tandem affinity purification (TAP) tag, HaloTag, Nus tag, thioredoxin tag, Fc tag, CYD tag, HPC tag, TrpE tag, ubiquitin tag, VSV-G epitope tag derived from the vesicular stomatitis virus glycoprotein, and V5 tag derived from a small epitope (Pk) found in the P and V proteins of the paramyxovirus simian virus 5 (SV5).In some embodiments, the affinity acceptor tag is an "epitope tag," a type of peptide tag that adds a recognizable epitope (antibody binding site) to an HLA protein to result in binding of the corresponding antibody, thereby enabling identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are protein A or protein G, which bind to IgG. In some embodiments, the affinity acceptor tag comprises a biotin acceptor peptide (BAP) or a human influenza hemagglutinin (HA) peptide sequence. Numerous other tag moieties are known to those skilled in the art and can be envisioned and are contemplated herein. Any peptide tag can be used as long as it can be expressed as an element of an affinity acceptor-tagged HLA peptide complex.

[0257]

[0357] The methods provided herein include isolating HLA peptide complexes from transferase or transduced cells by affinity pull-down of HLA constructs. In some embodiments, the complexes can be isolated using standard immunoprecipitation techniques well known in the art using commercially available antibodies. Cells can first be lysed. HLA class I-peptide complexes can be isolated using an HLA class I-specific antibody, such as the W6 / 32 antibody, while HLA class II-peptide complexes can be isolated using an HLA class II-specific antibody, such as the M5 / 114.15.2 monoclonal antibody. In some embodiments, a single (or pair of) HLA alleles is expressed as a fusion protein with a peptide tag, and the HLA peptide complex is isolated using a binding molecule that recognizes the peptide tag.

[0258]

[0358] The method further comprises isolating peptides from the HLA-peptide complex and sequencing the peptides. The peptides are isolated from the complex by any method known to those skilled in the art, such as acid elution. Meanwhile, any sequencing method can be used, and in some embodiments, a method using mass spectrometry, such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or alternatively HPLC-MS or HPLC-MS / MS), is utilized. These sequencing methods are known to those skilled in the art and are reviewed in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb;34(1):43-63.

[0259]

[0359] In some embodiments, the population of cells expresses one or more endogenous HLA alleles. In some embodiments, the population of cells is a population of engineered cells lacking one or more endogenous HLA class I alleles. In some embodiments, the population of cells is a population of engineered cells lacking endogenous HLA class I alleles. In some embodiments, the population of cells is a population of engineered cells lacking one or more endogenous HLA class II alleles. In some embodiments, the population of cells is a population of engineered cells lacking endogenous HLA class II alleles or a population of engineered cells lacking endogenous HLA class I alleles and endogenous HLA class II alleles. In some embodiments, the population of cells comprises cells enriched or sorted, such as by fluorescence-activated cell sorting (FACS). In some embodiments, fluorescence-activated cell sorting (FACS) is used to sort the population of cells. In some embodiments, the population of cells has been previously FACS-sorted for cell surface expression of either HLA class I or class II, or both HLA class I and class II. For example, FACS can be used to sort a population of cells for cell surface expression of HLA class I alleles, HLA class II alleles, or a combination thereof.

[0260] Methods for preparing personalized cancer vaccines

[0360] When a cancer-specific mutation is identified, such that the mutation is present in the DNA of cancer cells but not in normal cells of the same human subject, the mutation results in a change in one or more amino acids in the protein encoded by the DNA, and the mutation can become a target of the host immune response. The innate immune response is directed against the mutated protein, leading to the destruction of the cancer cells expressing the protein. Due to the natural tolerance response and immunocompromised environment in cancerous tissues, immunotherapy is a clinical pathway that attempts to enhance such immune responses to override the body's tolerance and immunosuppressive effects. Therefore, proteins or peptides containing the above-described mutations are suitable candidates for immunotherapy.

[0261]

[0361] The mutant proteins are phagocyte-mediated endocytic cleavage and digestion by professional phagocytes acting as antigen-presenting cells (APCs). They are then presented on the cell surface as antigens in antigen-presenting complexes containing major histocompatibility complex (MHC) proteins for T cell activation. Human MHC proteins are called human leukocyte antigens (HLA). MHC proteins can be MHC class I or class II proteins, and several functional differences result from peptide presentation by either class I or class II MHC proteins (HLA class I and HLA class II proteins). One notable difference is that HLA class I-peptide complexes present antigens to cytotoxic CD8+ T cells, while HLA class II peptide complexes can also activate CD4+ T cells, resulting in long-term immune responses. CD8+ T cells are essential for the cell-by-cell elimination of diseased cells, such as infected or tumor cells. CD4+ T cells have a more sustained effect on activation, the most important of which is the generation of immunological memory. CD4 subsets are recruited differently depending on the type of immunological threat, and multiple subsets with overlapping or disparate functions may be recruited simultaneously. This helps balance the immunological response to pathogen threats. In this context, HLA class I or class II peptide-mediated antigen presentation influences a sustained and tailored immune response. On the other hand, HLA class I or class II binding to peptides can be promiscuous, resulting in nonspecific peptide binding and presentation to the immune system, leading to aberrant immune responses such as autoimmunity.

[0262]

[0362] In one aspect, the present disclosure provides a method for predicting peptides that can precisely pair or bind to specific HLA class I or class II molecules, such that high fidelity binding of the peptide to an HLA class I or class II protein ensures presentation of the specific peptide to T lymphocytes, thereby eliciting a specific immune response and avoiding all cross-reactivity and immune disruption conditions.

[0263]

[0363] In one aspect, the present disclosure provides a method for predicting peptides that can accurately bind to a specific HLA class I or class II protein, such that when the peptide is administered therapeutically to a subject expressing the specific cognate HLA class I or class II protein, the peptide's ability to activate CD4+ T cells and stimulate immunological memory can activate a more sustained and robust immune response. In some embodiments, a given peptide predicted to bind to an HLA class I or class II protein with high specificity is a peptide containing a mutation, where the mutation is prevalent in the subject's cancer or tumor cells; whereas the same HLA class I or class II protein predicted to bind to the mutant peptide either (a) does not bind or (b) binds with significantly lower affinity to the corresponding unmutated wild-type peptide compared to its affinity for binding to the subject's mutant peptide. Preferential binding of HLA to mutant peptides is advantageous in the development of immunotherapies, as cells expressing the wild-type peptide are spared immune attack by T cells reactive to HLA-presented peptides. In some embodiments, the predicted peptides that specifically bind to HLA class I or class II proteins are peptides that have post-translational modifications. Exemplary post-translational modifications include, but are not limited to, phosphorylation, ubiquitination, dephosphorylation, glycosylation, methylation, or acetylation. In some embodiments, the predicted peptides are subjected to post-translational modifications before use in immunotherapy.

[0264]

[0364] In some embodiments, the immunotherapeutic methods and strategies disclosed herein may also be applicable to suppressing unwanted immune activation, for example, in autoimmune reactions. Specifically, peptides identified as potential binders for specific HLA subtypes may be tailored to bind to specific HLA molecules, inducing tolerance rather than generating an immunogenic response.

[0265]

[0365] In one aspect, the present invention provides a method for targeted or personalized immunotherapy for specific patients.Every subject or patient expresses a specific array of HLA class I and HLA class II proteins.HLA typing is a well-known technique that allows the specific repertoire of HLA proteins expressed by a subject to be determined.Once the HLA heterodimers expressed by a specific subject are understood, as described herein, an improved, refined and reliable method for predicting peptides that can bind to specific HLA class I or class II complexes with high fidelity can ensure that specific immune responses can be generated specifically for the subject.

[0266]

[0366] Genes encoding HLA heterodimers are highly polymorphic, with over 4,000 HLA class II allele variants identified in the human population. From maternal and paternal HLA haplotypes, individuals may inherit different alleles at each HLA class II locus, and each HLA class II heterodimer is composed of an α- and β-chain. Due to the large number of α- and β-chain pairing combinations, particularly for HLA-DP and HLA-DQ alleles, the population of possible HLA heterodimers is highly complex. HLA class II heterodimers are translated in the endoplasmic reticulum (ER) and assemble into stable complexes with the invariant chain (Ii) derived from the protein CD74. Ii stabilizes the class II complex by enabling correct protein folding and allows transport of HLA class II heterodimers to endosomal / lysosomal compartments. Within these HLA class II loading compartments, Ii is proteolytically cleaved by cathepsins into a placeholder peptide called CLIP. CLIP is then exchanged for higher affinity peptides in a low pH environment by the chaperone HLA-DM, a non-classical HLA class II heterodimer. The high-affinity peptide-loaded HLA class II complex then resides in the trans-Golgi and finally on the cell surface for presentation to CD4+ T cells.

[0267]

[0367] Each HLA heterodimer is estimated to bind thousands of peptides with allele-specific binding selectivity. In fact, it is estimated that each HLA allele binds and presents approximately 1,000–10,000 unique peptides to T cells. Given such diversity in HLA binding, accurate prediction of whether a peptide is likely to bind to a particular HLA allele is extremely difficult. Due to heterogeneity in α- and β-chain pairing, data complexity that limits the ability to confidently assign core binding epitopes, and the lack of immunoprecipitation-grade allele-specific antibodies necessary for high-resolution biochemical analysis, little is known about the allele-specific peptide binding characteristics of HLA class II molecules. Furthermore, analyzing peptide epitopes from a given HLA allele can lead to ambiguity when multiple HLA alleles are presented on the cell surface.

[0268]

[0368] Disclosed herein is a method for preparing personalized cancer vaccine.The method for preparing personalized cancer vaccine can include the following steps: Identify the peptide sequence that contains mutations expressed in cancer cells of target;Use a computer processor to input the amino acid position information of identified peptide sequence into a machine learning HLA peptide presentation prediction model, to generate a series of presentation predictions for identified peptide sequence, wherein each presentation prediction indicates the probability that one or more proteins that are encoded by the class I or class II MHC allele of cancer cells of target will present a given sequence of identified peptide sequence;And select a subset of identified peptide sequence based on the series of presentation predictions to prepare personalized cancer vaccine.

[0269]

[0369] In some embodiments, one or more results obtained from the methods described herein can provide one or more quantitative values ​​indicative of one or more of the following: likelihood of diagnostic accuracy, likelihood of the presence of a condition in a subject, likelihood of a subject developing a condition, likelihood of success of a particular treatment, or any combination thereof. In some embodiments, the methods described herein can predict the risk or likelihood of developing a condition. In some embodiments, the methods described herein can be an indication of early diagnosis of developing a condition. In some embodiments, the methods described herein can confirm the diagnosis or existence of a condition. In some embodiments, the methods described herein can monitor the progression of a condition. In some embodiments, the methods described herein can monitor the effectiveness of a treatment for a condition in a subject.

[0270] Methods for the identification of MHC-presented peptides

[0370] In one aspect, a method for identifying one or more peptides presented by an MHC protein for immune activation is provided. In some embodiments, the one or more peptides comprise an epitope. In some embodiments, the method comprises a computer prediction of the likelihood that a specific epitope will be presented by an MHC protein. In some embodiments, the method comprises a computer prediction of the specificity of the epitope for MHC presentation. In some embodiments, the computer prediction method comprises evaluating peptide-MHC interactions. In some embodiments, the computer prediction method comprises predicting the allele specificity of a peptide for antigen presentation.

[0271]

[0371] In some embodiments, the computational prediction method incorporates bioinformatics information, such as nucleotide sequences, structural motifs of biomolecules, protein-protein interaction properties, and functional strengths such as immunogenicity. In some embodiments, the computational prediction method involves machine learning. Numerous immunoinformatics methods for predicting peptide-MHC interactions have been developed for both MHC class I and II based on machine learning approaches, such as simple pattern motifs, support vector machines (SVMs), hidden Markov models (HMMs), neural network (NN) models, quantitative structure-activity relationship (QSAR) analysis, structure-based methods, and biophysical methods. These methods can be classified into two categories: intra-allelic (allele-specific) and trans-allelic (pan-specific) methods. Intra-allelic methods are trained on a specific MHC molecule with a limited set of experimental peptide binding data and are applied to predict peptides that bind to that molecule. Due to the extreme polymorphism of MHC molecules, the existence of thousands of allelic variants, combined with the lack of sufficient experimental binding data, makes it impossible to build a predictive model for each allele. Therefore, transallelic and general-purpose methods, such as NetMHCIIpan (Karosiene E et al., NetMHCIIpan-3.0, a common pan-specific MHC class II prediction method including all three human MHC class II isotypes, HLA-DR, HLA-DP, and HLA-DQ. Immunogenetics (2013) 65(10):711-24) and TEPITOPEpan (Zhang L, et al., TEPITOPEpan: extending TEPITOPE for peptide binding prediction covering over 700 HLA-DR molecules. PLoS One (2012) 7(2):e30483), have been developed using peptide binding data that span multiple alleles or species. Similar methods for MHC-I are also available, such as NetMHCpan and KISS.

[0272]

[0372] In some embodiments, the peptide sequence may not be expressed in normal cells of the subject. In some embodiments, each and every cell of the subject may not be a cancer cell. Cancer cells include, but are not limited to, thyroid cancer, adrenocortical carcinoma, anal cancer, aplastic anemia, bile duct cancer, bladder cancer, bone cancer, bone metastasis, central nervous system (CNS) cancer, peripheral nervous system (PNS) cancer, breast cancer, Castleman's disease, cervical cancer, childhood non-Hodgkin's lymphoma, lymphoma, colon and rectal cancer, endometrial cancer, esophageal cancer, Ewing's family of tumors (e.g., Ewing's sarcoma), eye cancer, gallbladder cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, gestational trophoblastic disease, hairy cell leukemia, Hodgkin's disease, Kaposi's sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, acute lymphocytic leukemia, acute myeloid leukemia, childhood leukemia, chronic lymphocytic leukemia, and the like. The antibodies may be produced through a variety of cancers, including myeloid leukemia, chronic myeloid leukemia, liver cancer, lung cancer, pulmonary carcinoid tumors, non-Hodgkin's lymphoma, male breast cancer, malignant mesothelioma, multiple myeloma, myelodysplastic syndromes, myeloproliferative disorders, nasal cavity and sinus cancer, nasopharyngeal cancer, neuroblastoma, oral cavity and oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, pituitary tumors, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma (adult soft tissue cancer), melanoma skin cancer, non-melanoma skin cancer, stomach cancer, testicular cancer, thymic cancer, uterine cancer (e.g., uterine sarcoma), vaginal cancer, vulvar cancer, or Waldenstrom's hypergammaglobulinemia.

[0273]

[0373] Identifying comprises comparing the DNA, RNA or protein sequence from a subject's cancer cells with the DNA, RNA or protein sequence from a subject's normal cells.The DNA, RNA or protein sequence from a subject's cancer cells may be different from the DNA, RNA or protein sequence from a subject's normal cells.Identifying can identify nucleic acid variants with high sensitivity.

[0274]

[0374] The machine learning HLA peptide presentation prediction model may include a plurality of predictor variables identified based on at least training data, including sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry, training peptide sequence information including amino acid position information associated with the HLA proteins expressed in cells, and a function representing the association between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and the predictor variables.

[0275]

[0375] In some embodiments, the training data includes structured data, time-series data, unstructured data, and relational data. Unstructured data may include audio data, image data, video, mechanical data, electrical data, chemical data, and combinations thereof for use in accurately simulating or training robotics or simulations. Time-series data may include data from one or more of smart meters, smart appliances, smart devices, monitoring systems, telemetry devices, or sensors. Relational data includes data from customer systems, enterprise systems, operational systems, websites, web-accessible application program interfaces (APIs), or any combination thereof. This may be done by a user through any method of inputting files or other data formats into the software or system.

[0276]

[0376] In some embodiments, the training data may be stored in a database. The database may be stored in a computer-readable format. A computer processor may be configured to access the data stored in a computer-readable memory. In some embodiments, the computer system may be used to analyze the data to obtain results. The results may be stored remotely or internally on a storage medium and communicated to personnel such as a medical professional. In some embodiments, the computer system may be operatively coupled to components for communicating the results. Components for communication may include wired and wireless components. Examples of wired communication components may include a Universal Serial Bus (USB) connection, a coaxial cable connection, an Ethernet cable, e.g., a Cat5 or Cat6 cable, a fiber optic cable, or a telephone line. Examples of wireless communication components may include a Wi-Fi receiver, a component for accessing a mobile data standard such as a 3G or 4G LTE data signal, or a Bluetooth receiver. In some embodiments, all of this data in a storage medium is collected and archived to build a data warehouse.

[0277]

[0377] In some embodiments, the database comprises an external database, such as, but not limited to, a medical database, such as, Adverse Drug Effects Database, AHFS Supplemental File, Allergen Picklist File, Average WAC Pricing File, Brand Probability File, Canadian Drug File v2, Comprehensive Price History, Controlled Substances File, Drug Allergy Cross-Reference File, Drug Application File, Drug Dosing & Administration Database, Drug Image Database v2.0 / Drug Imprint Database v2.0, Drug Inactive Date File, Drug Indications Database, Drug Lab Conflict Database, Drug Therapy Monitoring System (DTMS) v2.2 / DTMS Consumer Monographs、Duplicate Therapy Database、Federal Government Pricing File、Healthcare Common Procedure Coding System Codes(HCPCS)Database、ICD-10 Mapping Files、Immunization Cross-Reference File、Integrated A to Z Drug Facts Module、Integrated Patient Education、Master Parameters Database、Medi-Span Electronic Drug File(MED-File)v2、Medicaid Rebate File、Medicare Plans File、Medical Condition Picklist File、Medical Conditions Master Database、Medication Order Management Database(MOMD)、Parameters to Monitor Database、Patient Safety Programs File、Payment Allowance Limit-Part B(PAL-B)v2.0、Precautions Database、RxNorm Cross-Reference File、Standard Drug Identifiers Database、Substitution Groups File、Supplemental Names File、Uniform System of Classification Cross-Reference FileまたはWarning Label Databaseであってよい。.

[0278]

[0378] In some embodiments, training data may also be obtained through other data sources. Data sources may include sensors or smart devices, such as appliances, smart meters, wearables, monitoring systems, data stores, customer systems, billing systems, financial systems, cloud-sourced data, weather data, social networks, or any other sensors, enterprise systems, or data stores. Examples of smart meters or sensors may include meters or sensors located at customer premises, or meters or centers located between customers and generation or source locations. By incorporating data from a wide range of sources, the system can perform complex and detailed analyses. In some embodiments, data sources may include, without limitation, sensors or databases for other medical platforms.

[0279]

[0379] HLA typing is conventionally performed either by serological methods using antibodies or by PCR-based methods, such as sequence-specific oligonucleotide probes (SSOP) or sequence-based typing (SBT). The former is hampered by a potentially high degree of cross-reactivity and limited resolution capabilities, while the latter suffers from difficulties related to PCR efficiency due to the very limited possibilities for primer positioning due to the location of the polymorphisms.

[0280]

[0380] In some embodiments, the sequence information is identified by either a sequencing method or a method using mass spectrometry, such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or alternatively, HPLC-MS or HPLC-MS / MS). These sequencing methods are known to those skilled in the art and are reviewed in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb;34(1):43-63. In some embodiments, the mass spectrometry is single-allele mass spectrometry. In some embodiments, the mass spectrometry can be MS analysis, MS / MS analysis, LC-MS / MS analysis, or a combination thereof. In some embodiments, MS analysis can be used to determine the mass of an intact peptide. For example, determining can include determining the mass of an intact peptide (e.g., MS analysis). In some embodiments, MS / MS analysis can be used to determine the mass of a peptide fragment. For example, determining can include determining the mass of a peptide fragment, which can be used to determine the amino acid sequence of the peptide or portion thereof (e.g., MS / MS analysis). In some embodiments, the mass of the peptide fragments can be used to determine the sequence of amino acids within the peptide. In some embodiments, LC-MS / MS analysis can be used to separate a complex peptide mixture. For example, determining can include separating the complex peptide mixture, for example, by liquid chromatography, and determining the mass of the intact peptide, the mass of the peptide fragment, or a combination thereof (e.g., LC-MS / MS analysis). This data can be used, for example, for peptide sequencing.

[0281]

[0381] In some embodiments, the training peptide sequence information includes amino acid position information of the training peptides. In some embodiments, the training peptide sequence information includes up to about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or less of the sequence information of the sequences of peptides expressed in cells and presented by HLA proteins identified by mass spectrometry. In some embodiments, the training peptide sequence information may include at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the sequence information of the sequences of peptides expressed in cells and presented by HLA proteins identified by mass spectrometry.

[0282]

[0382] Any information and data may be paired with the subject that is the source of the information and data. The subject or a medical professional can retrieve the information and data from a storage or server using the subject's identity. The subject's identity may include a patient's photograph, name, address, Social Security number, date of birth, telephone number, zip code, or any combination thereof. The subject's identity may be encrypted and coded into a visual graphical code. The visual graphical code may be a one-time barcode uniquely associated with the subject's identity. The barcode may be a UPC barcode, EAN barcode, Code39 barcode, Code128 barcode, ITF barcode, CodaBar barcode, GS1 DataBar barcode, MSI Plessey barcode, QR barcode, Datamatrix code, PDF417 code, or Aztec barcode. The visual graphical code may be configured to be displayed on a display screen. The barcode may include a QR that is optically captured and machine-readable. A barcode can define elements such as the version, format, position, alignment, or timing of the barcode to enable it to be read and decoded. A barcode can encode various types of information in any type of suitable format, for example, binary or alphanumeric information. A QR Code® can have various symbol sizes, as long as the QR Code® can be scanned from a reasonable distance by an imaging device. A QR Code® can be in any image format (e.g., EPS or SVG vector graphics, PNG, TIF, GIF, or JPEG raster graphics formats).

[0283]

[0383] In some embodiments, the function representing the relationship between the amino acid position information received as input and the likelihood of presentation generated as output based on the amino acid position information and the predictor variable includes a linear or nonlinear function, such as a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions, such as saturated hyperbolic tangent, identity map, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sinusoid, sinc, Gaussian, or sigmoid function, or any combination thereof.

[0284]

[0384] In some embodiments, a linear function can be obtained through linear regression. In some embodiments, linear regression is a method of predicting a target variable by fitting a best linear relationship between a dependent variable and independent variables. A best fit can mean that the sum of all distances between the shape at each point and the actual observation is minimized. Linear regression can include simple linear regression or multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting a best linear relationship. A non-linear function can be obtained using nonlinear regression. Nonlinear regression can be a form of regression analysis in which observed data is modeled by a function that is a nonlinear combination of model parameters and depends on one or more independent variables. Nonlinear regression can include step functions, piecewise functions, splines, and generalized additive models.

[0285]

[0385] In some embodiments, the likelihood of presentation is represented by a one-dimensional value (e.g., probability). In some embodiments, the probability is configured to measure the likelihood that an event may occur. In some embodiments, the probability is in the range of about 0 and 1, 0.1 to 0.9, 0.2 to 0.8, 0.3 to 0.7, or 0.4 to 0.6. The higher the probability of an event, the higher the likelihood that the event will occur. In some embodiments, the event may include, by way of non-limiting example, any type of situation, including whether an HLA peptide presents several peptides with certain amino acid position information, and whether an individual will become ill based on the amino acid position information. In some embodiments, the likelihood may be represented by a multidimensional value. The multidimensional value may be represented by a multidimensional space, a heat map, or a spreadsheet.

[0286]

[0386] In one embodiment, selecting a subset of peptide sequences identified based on a set of presentation predictions is configured to prepare a personalized cancer vaccine. In some embodiments, the subset includes up to about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or less of the peptide sequences identified based on a set of presentation predictions. In other cases, the subset may include at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more of the peptide sequences identified based on a set of presentation predictions. The cancer vaccine may be a vaccine that either treats existing cancer or prevents the onset of cancer. The vaccine may be prepared from a sample collected from a patient and may be specific to that patient.

[0287]

[0387] In some embodiments, poxviruses are used in disease (e.g., cancer) vaccines or immunogenic compositions. These include orthopoxvirus, avian pox, vaccinia, MVA, NYVAC, canarypox, ALVAC, fowlpox, TROVAC, etc. Advantages of vectors may include simple construction, the ability to accommodate large amounts of foreign DNA, and high expression levels. Information regarding poxviruses that can be used in the practice of the present disclosure, such as chordopoxvirinae poxviruses (vertebrate poxviruses), e.g., orthopoxviruses and avian poxviruses, such as vaccinia virus (e.g., Wyeth strain, WR strain (e.g., ATCC® VR-1354), Copenhagen strain, NYVAC, NYVAC.1, NYVAC.2, MVA, MVA-BN), canarypox virus (e.g., Wheatley C93 strain, ALVAC), fowlpox virus (e.g., FP9 strain, Webster strain, TROVAC), dovepox, pigeonpox, quailpox, and raccoonpox, among others, synthetic or non-naturally occurring recombinant forms thereof, their uses, and methods for making and using such recombinants, can be found in the scientific and patent literature.

[0288]

[0388] In some embodiments, vaccinia viruses are used in disease vaccines or immunogenic compositions to express antigens, and recombinant vaccinia viruses are capable of replicating within the cytoplasm of infected host cells, allowing the polypeptide of interest to induce an immune response.

[0289]

[0389] In some embodiments, ALVAC is used as a vector in a disease vaccine or immunogenic composition. ALVAC may be modified to express a foreign transgene and may be a canarypox virus that has been used as a method for vaccination against both prokaryotic and eukaryotic antigens.

[0290]

[0390] In some embodiments, modified vaccinia Ankara (MVA) virus is used as a viral vector for antigen vaccines or immunogenic compositions. MVA may be a member of the Orthopoxvirus family and was generated by serial passage of the Ankara strain of vaccinia virus (CVA) approximately 570 times in chicken embryo fibroblast cells. As a result of these passages, the resulting MVA virus may contain 31 kilobases less genomic information than CVA and may be highly host-cell restricted. MVA may be characterized by its extreme attenuation, i.e., reduced pathogenicity or infectious capacity while still retaining excellent immunogenicity. When tested in various animal models, MVA may prove nonpathogenic even in immunosuppressed individuals. Furthermore, MVA-BN®-HER2 may be a candidate immunotherapy designed for the treatment of HER-2-positive breast cancer and is currently undergoing clinical testing.

[0291]

[0391] In some embodiments, positive predictive value (PPV) is used as part of a predictive model. PPV, also known as an accuracy measure, is the probability that an individual diagnosed with a disease or condition, for example, using a test or model, actually has the disease or condition. It can be calculated by dividing the number of true positive results by the total number of positive results (including false positives). PPV = true positives / (true positives + false positives). For example, in a set of 100 patients, if a model identifies 50 patients as positive results, and 25 of those are true positives, the PPV is 25 / 50 = 0.5. A PPV closer to 1 represents a more accurate diagnostic method, such as a test or model. PPV can be used to determine the accuracy of a predictive model. PPV can be used to adjust a predictive model to accommodate for false positive results that may be generated by the model.

[0292]

[0392] Recall can be used as part of a predictive model. Recall is considered the percentage of true positive results out of the total number of positives in a sample set. Recall = true positives / (true positives + false negatives). For example, in a set of 100 patients, if a model identifies positive results in 50 patients, 25 of which are true positives, for a total of 75 positives in the patient set, then the recall is {25 / (25+25)}x100=50%. Recall can be used to determine the accuracy of a predictive model. Recall can be used to adjust a predictive model to account for false positive or false negative results generated by the model.

[0293]

[0393] In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall rate of 0.1% to 10%. In some embodiments, the predictive model may have a positive predictive value of up to 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less at a recall rate of 0.1% to 10%. The predictive model may have a positive predictive value of at least 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall rate of less than 0.1%. In some embodiments, the predictive model may have a positive predictive value of up to 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less at a recall rate of less than 0.1%. The predictive model may have a positive predictive value of at least 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or more at a recall rate of greater than 10%. In some embodiments, the predictive model may have a positive predictive value of up to 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less at a recall rate of greater than 10%.

[0294]

[0394] In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 0.1% to 10%. In some embodiments, the predictive model has a positive predictive value of at least 0.1% to 0.5%, 0.1% to 1%, 0.1% to 2%, 0.1% to 3%, 0.1% to 4%, 0.1% to 5%, 0.1% to 6%, 0.1% to 7%, 0.1% to 8%, 0.1% to 9%, 0.1% to 10%, 0.5% to 1%, 0.5 to 2%, 0.5 to 3%, 0.5 to 4%, 0.5% to 5%, 0.5% to 6%, 0.5% to 7%, 0.5% to 8%, 0.5% to 9%, 0.5% to 10%, 1% to 2%, 1% to 3%, 1% to 4%, 1% to 5%, 1% to 6%, 1% to 7%, 1% to 8%, 1% to 9%, 1% to 10%, 2% to 3%, 2% to 4%, 2% to 5%, 2% to 6%, 2% or to 7%, 2% to 8%, 2% to 9%, 2% to 10%, 3% to 4%, 3% to 5%, 3% to 6%, 3% to 7%, 3% to 8%, 3% to 9%, 3% to 10%, 4% to 5%, 4% to 6%, 4% to 7%, 4% to 8%, 4% to 9%, 4% to 10%, 5% to 6%, 5% to 7%, 5% to 8%, 5% to 9%, 5% to 1 and a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 0%, 6% to 7%, 6% to 8%, 6% to 9%, 6% to 10%, 7% to 8%, 7% to 9%, 7% to 10%, 8% to 9%, 8% to 10%, or 9% to 10%. In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%. In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater with a recall of at least 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, or 9%.In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater with a recall of up to 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%.

[0295]

[0395] In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 10% to 20%. In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 10% to 20%. In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 10% to 20%. to 16%, 11% to 17%, 11% to 18%, 11% to 19%, 11% to 20%, 12% to 13%, 12% to 14%, 12% to 15%, 12% to 16%, 12% to 17%, 12% to 18%, 12% to 19%, 12% to 20%, 13% to 14%, 13% to 15%, 13% to 16%, 13% to 17% %, 13% to 18%, 13% to 19%, 13% to 20%, 14% to 15%, 14% to 16%, 14% to 17%, 14% to 18%, 14% to 19%, 14% to 20%, 15% to 16%, 15% to 17%, 15% to 18%, 15% to 19%, 15% to 20%, 16% to 17%, 16% to 18%, 1 and a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 6% to 19%, 16% to 20%, 17% to 18%, 17% to 19%, 17% to 20%, 18% to 19%, 18% to 20%, or 19% to 20%. In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall of 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%. In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater with a recall of at least 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, or 19%.In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater at a recall rate of 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%.

[0296]

[0396] In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or greater at a recall rate of at least 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%. For example, the predictive model has a positive predictive value of at least 0.1 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.2 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.3 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.5 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.6 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.7 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.8 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.9 at a recall rate of at least 10%. For example, the predictive model may have a positive predictive value of at least 0.1 at a recall rate of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.2 at a recall rate of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.3 at a recall rate of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall rate of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.5 at a recall rate of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.6 with a recall of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.7 with a recall of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.8 with a recall of at least 5%. For example, the predictive model may have a positive predictive value of at least 0.9 with a recall of at least 5%.For example, the predictive model may have a positive predictive value of at least 0.1 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.2 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.3 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.5 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.6 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.7 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.8 at a recall rate of at least 20%. For example, the predictive model may have a positive predictive value of at least 0.9 at a recall rate of at least 20%.

[0297]

[0397] In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or greater at a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%. For example, the predictive model may have a positive predictive value of at least 0.1 at a recall of about 10%. For example, the predictive model may have a positive predictive value of at least 0.2 at a recall of about 10%. For example, the predictive model may have a positive predictive value of at least 0.3 at a recall of about 10%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall of about 10%. For example, the predictive model may have a positive predictive value of at least 0.5 at about 10% recall. For example, the predictive model may have a positive predictive value of at least 0.6 at about 10% recall. For example, the predictive model may have a positive predictive value of at least 0.7 at about 10% recall. For example, the predictive model may have a positive predictive value of at least 0.8 at about 10% recall. For example, the predictive model may have a positive predictive value of at least 0.9 at about 10% recall. For example, the predictive model may have a positive predictive value of at least 0.1 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.2 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.3 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.4 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.5 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.6 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.7 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.8 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.9 at about 5% recall. For example, the predictive model may have a positive predictive value of at least 0.1 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.2 at about 20% recall.For example, the predictive model may have a positive predictive value of at least 0.3 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.4 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.5 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.6 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.7 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.8 at about 20% recall. For example, the predictive model may have a positive predictive value of at least 0.9 at about 20% recall.

[0298]

[0398] In some embodiments, the predictive model has a positive predictive value of at least 0.05, 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or greater at a recall rate of less than 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%. For example, the predictive model may have a positive predictive value of at least 0.1 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.2 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.3 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.5 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.6 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.7 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.8 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.9 at a recall rate of up to 10%. For example, the predictive model may have a positive predictive value of at least 0.1 at a recall rate of up to 5%. For example, the predictive model may have a positive predictive value of at least 0.2 at a recall rate of up to 5%. For example, the predictive model may have a positive predictive value of at least 0.3 at a recall rate of up to 5%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall rate of up to 5%. For example, the predictive model may have a positive predictive value of at least 0.5 at a recall rate of up to 5%. For example, the predictive model may have a positive predictive value of at least 0.6 at up to 5% recall. For example, the predictive model may have a positive predictive value of at least 0.7 at up to 5% recall. For example, the predictive model may have a positive predictive value of at least 0.8 at up to 5% recall. For example, the predictive model may have a positive predictive value of at least 0.9 at up to 5% recall. For example, the predictive model may have a positive predictive value of at least 0.1 at up to 20% recall. For example, the predictive model may have a positive predictive value of at least 0.2 at up to 20% recall.For example, the predictive model may have a positive predictive value of at least 0.3 at a recall rate of up to 20%. For example, the predictive model may have a positive predictive value of at least 0.4 at a recall rate of up to 20%. For example, the predictive model may have a positive predictive value of at least 0.5 at a recall rate of up to 20%. For example, the predictive model may have a positive predictive value of at least 0.6 at a recall rate of up to 20%. For example, the predictive model may have a positive predictive value of at least 0.7 at a recall rate of up to 20%. For example, the predictive model may have a positive predictive value of at least 0.8 at a recall rate of up to 20%. For example, the predictive model may have a positive predictive value of at least 0.9 at a recall rate of up to 20%.

[0299]

[0399] In some embodiments, the predictive model has a positive predictive value of 0.05% to 0.6% at a recall rate of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%.With a recall of approximately 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%, the predictive model is 0.05% to 0.1%, 0.05% to 0.15%, 0.05% to 0.2%, 0.05% to 0.25%, 0.05% to 0.3%, 0.05% to 0.35%, 0.05% to 0.4%, 0.05% to 0.45%, 0.05% to 0.5%, 0.05% to 0.55%, 0.05% to 0.65% to 0.6%, 0.1% to 0.15%, 0.1% to 0.2%, 0.1% to 0.25%, 0.1% to 0.3%, 0.1% to 0.35%, 0.1% to 0.4%, 0.1% to 0.45%, 0.1% to 0.5%, 0.1% to 0.55%, 0.1% to 0.6%, 0.15% to 0.2%, 0.15% to 0.25%, 0.15% to 0.3%, 0.15% to 0.35%, 0.15% to 0.4%, 0.15% to 0.45%, 0.15% to 0.5%, 0.15% to 0.55%, 0.1 5% to 0.6%, 0.2% to 0.25%, 0.2% to 0.3%, 0.2% to 0.35%, 0.2% to 0.4%, 0.2% to 0.45%, 0.2% to 0.5%, 0.2% to 0.55%, 0.2% to 0.6%, 0.25% to 0.3%, 0.25% to 0.35%, 0.25% to 0.4%, 0.25% to 0.45%, 0.25% to 0.5%, 0.25% to 0.55%, 0.25% to 0.6%, 0.3% to 0.35%, 0.3% to 0.4%, 0.3% to 0.45%, 0. It may have a positive predictive value of 3% to 0.5%, 0.3% to 0.55%, 0.3% to 0.6%, 0.35% to 0.4%, 0.35% to 0.45%, 0.35% to 0.5%, 0.35% to 0.55%, 0.35% to 0.6%, 0.4% to 0.45%, 0.4% to 0.5%, 0.4% to 0.55%, 0.4% to 0.6%, 0.45% to 0.5%, 0.45% to 0.55%, 0.45% to 0.6%, 0.5% to 0.55%, 0.5% to 0.6% or 0.55% to 0.6%.At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, 0.55%, or 0.6%. At a recall rate of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of at least 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5% or 0.55%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of up to 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, 0.55%, or 0.6%.

[0300]

[0400] At a recall rate of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of 0.45% to 0.98%.With a recall of approximately 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%, the predictive model performed well from 0.45% to 0.5%, 0.45% to 0.55%, 0.45% to 0.6%, 0.45% to 0.65%, 0.45% to 0.7%, 0.45% to 0.75%, 0.45% to 0.8%, 0.45% to 0.85%, 0.45% to 0.9%, 0.45% to 0.96%, 0.45% to 0.96%. 0.98%, 0.5% to 0.55%, 0.5% to 0.6%, 0.5% to 0.65%, 0.5% to 0.7%, 0.5% to 0.75%, 0.5% to 0.8%, 0.5% to 0.85%, 0.5% to 0.9%, 0.5% to 0.96%, 0.5% to 0.98%, 0.55% to 0.6%, 0.55% to 0.65%, 0.55% to 0.7%, 0.55% to 0.75%, 0.55% to 0.8%, 0.55% to 0.85%, 0.55% to 0.9%, 0.55% to 0.96%, 0.55% or to 0.98%, 0.6% to 0.65%, 0.6% to 0.7%, 0.6% to 0.75%, 0.6% to 0.8%, 0.6% to 0.85%, 0.6% to 0.9%, 0.6% to 0.96%, 0.6% to 0.98%, 0.65% to 0.7%, 0.65% to 0.75%, 0.65% to 0.8%, 0.65% to 0.85%, 0.65% to 0.9%, 0.65% to 0.96%, 0.65% to 0.98%, 0.7% to 0.75%, 0.7% to 0.8%, 0.7% to 0.85%, 0.7% to 0.9%, 0.7% to 0.96%, 0.7% to 0.98%, 0.75% to 0.8%, 0.75% to 0.85%, 0.75% to 0.9%, 0.75% to 0.96%, 0.75% to 0.98%, 0.8% to 0.85%, 0.8% to 0.9%, 0.8% to 0.96%, 0.8% to 0.98%, 0.85% to 0.9%, 0.85% to 0.96%, 0.85% to 0.98%, 0.9% to 0.96%, 0.9% to 0.98%, or 0.96% to 0.98%.At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of 0.45%, 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, 0.96%, or 0.98%. At a recall rate of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of at least 0.45%, 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, or 0.96%. At a recall of about 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, the predictive model may have a positive predictive value of 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, 0.96%, or 0.98%.

[0301] A method for training a machine learning HLA peptide presentation prediction model

[0401] In one aspect, a method of training a machine learning HLA peptide presentation prediction model may include inputting, using a computer processor, amino acid positional information sequences of HLA peptides isolated from one or more HLA peptide complexes from cells expressing HLA class I or class II alleles into the HLA peptide presentation prediction model; training the machine learning HLA peptide presentation prediction model may include adjusting weights at nodes of a neural network to best match the provided training data.

[0302]

[0402] The training data includes sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information of training peptides associated with HLA proteins expressed in cells; and a function representing the association between the amino acid position information received as input and presentation probability generated as output based on the amino acid position information and predictor variables. The training data, training peptide sequence information, function, and presentation probability are disclosed elsewhere herein.

[0303]

[0403] A trained algorithm may include one or more neural networks. A neural network may be a type of computer system based on a graph of several connected neurons (or nodes) in a series of layers. A neural network may include an input layer, to which data is presented; one or more internal and / or "hidden" layers; and an output layer, from which results are presented. A neural network can learn the relationship between an input data set and a target data set by adjusting a series of connection weights. Neurons may be connected to neurons in other layers via connections with weights that control the strength of the connections. The number of neurons in each layer may be determined by the complexity of the problem to be solved. The minimum number of neurons required in a layer may be determined by the complexity of the problem, and the maximum number may be limited by the ability of the neuron neural network to generalize. Input neurons may receive the presented data and then transmit the data to nodes in the first hidden layer through connection weights, which are modified during training. A result node may sum the products of all pairs of inputs and their associated weights. The weighted sum may be offset using a bias to adjust the value of the result node. The output of a node or neuron can be gated using a threshold or an activation function. The activation function can be a linear or non-linear function. The activation function can be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions, such as a saturated hyperbolic tangent, identity map, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sinusoid, sinc, Gaussian, or sigmoid function, or any combination thereof.

[0304]

[0404] Hidden layers in neural networks process data and communicate the results to the next layer through a second set of weighted connections. Each subsequent layer may "pool" the results from the previous layer into more complex relationships. Neural networks can be trained using a known sample set of training data (data collected from one or more sensors) by allowing them to modify themselves during (and after) training to provide a desired output, such as an output value, from a given set of inputs. Trained algorithms can include convolutional neural networks, recurrent neural networks, augmented convolutional neural networks, fully connected neural networks, deep generative models, and Boltzmann machines.

[0305]

[0405] The weight coefficients, bias values, and thresholds or other computational parameters of a neural network may be "taught" or "learned" in a training phase using one or more sets of training data. For example, the parameters may be trained using input data from a training data set and gradient descent or backpropagation techniques so that the output value(s) from the neural network match the examples contained in the training data set.

[0306]

[0406] The number of nodes used in the input layer of the neural network may be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000 or more. In other cases, the number of nodes used in the input layer may be up to about 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9000, 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, or 10 or fewer. In some cases, the total number of layers used in a neural network (including input and output layers) may be at least about 3, 4, 5, 10, 15, 20, or more. In other cases, the total number of layers may be up to about 20, 15, 10, 5, 4, 3 or fewer.

[0307]

[0407] In some cases, the total number of learnable or trainable parameters, e.g., weighting coefficients, biases, or thresholds, used in a neural network may be at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more. In other cases, the number of learnable parameters may be up to about 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9000, 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, or 10 or less.

[0308]

[0408] The neural network may include a convolutional neural network. The convolutional neural network may include one or more convolutional layers, extension layers, or fully connected layers. The number of convolutional layers may be between 1 and 10, and the number of extension layers may be between 0 and 10. The total number of convolutional layers (including input and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more, and the total number of extension layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more. The total number of convolutional layers may be up to about 20, 15, 10, 5, 4, 3, or less, and the total number of extension layers may be up to about 20, 15, 10, 5, 4, 3, or less. In some embodiments, the number of convolutional layers is between 1 and 10, and the number of fully connected layers is between 0 and 10. The total number of convolutional layers (including input and output layers) may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more, and the total number of fully connected layers may be at least about 1, 2, 3, 4, 5, 10, 15, 20, or more. The total number of convolutional layers may be up to about 20, 15, 10, 5, 4, 3, or less, and the total number of fully connected layers may be up to about 20, 15, 10, 5, 4, 3, or less.

[0309]

[0409] A convolutional neural network (CNN) may be a deep and feed-forward artificial neural network. CNNs may be applicable to analyzing visual images. CNNs may include an input layer, an output layer, and multiple hidden layers. The hidden layers of a CNN may include a convolutional layer, a pooling layer, a fully connected layer, and a normalization layer. Layers may be constructed in three dimensions: width, height, and depth.

[0310]

[0410] A convolutional layer can apply a convolutional operation to the input and pass the results of the convolutional operation to the next layer. For processing images, convolutional operations can reduce the number of free parameters, allowing the network to become deeper with fewer parameters. In a convolutional layer, neurons can receive input from only a limited subarea of ​​the previous layer. The parameters of a convolutional layer can include a set of learnable filters (or kernels). Learnable filters have small receptive fields that can extend through the full depth of the input capacity. During a forward pass, each filter can be convolved across the width and height of the input capacity, computing the dot product between the filter's entries and the input, producing a two-dimensional activation map for that filter. As a result, the network can learn filters that activate when a particular type of feature is detected at a certain spatial location in the input.

[0311]

[0411] The pooling layer may include a global pooling layer. A global pooling layer may combine the outputs of neuron clusters in one layer into a single neuron in the next layer. For example, a max pooling layer may use the maximum value from each cluster of neurons in the previous layer; an average pooling layer may use the average value from each cluster of neurons in the previous layer. A fully connected layer may connect all neurons in one layer to all neurons in another layer. In a fully connected layer, each neuron may receive input from all elements in the previous layer. The normalization layer may be a batch normalization layer. A batch normalization layer can improve the performance and stability of a neural network. A batch normalization layer can provide any layer in a neural network with inputs that are zero mean / unit variance. Advantages of using a batch normalization layer include faster trained networks, higher learning speed, easier weight initialization, more feasible activation functions, and a simpler process for creating deep networks.

[0312]

[0412] The neural network may include a recurrent neural network. The recurrent neural network may be configured to receive sequential data, e.g., continuous data input, as input, and the recurrent neural network software module may update its internal state at every time step. The recurrent neural network may use its internal state (memory) to process the sequence of inputs. Recurrent neural networks may be applied to tasks such as handwriting or speech recognition, next word prediction, music composition, image captioning, time series anomaly detection, machine translation, scene labeling, and stock market prediction. The recurrent neural network may include a fully recurrent neural network, an independent recurrent neural network, an Elman network, a Jordan network, an echo state, a neural history compressor, a long-short-term memory, a gated recurrent unit, a multiple timescale model, a neural Turing machine, a differentiable neural computer, a neural network pushdown automaton, or any combination thereof.

[0313]

[0413] The trained algorithm may include supervised or unsupervised learning methods, such as SVM, random forest, clustering algorithms (or software modules), gradient boosting, logistic regression, and / or decision trees. A supervised learning algorithm may be an algorithm that relies on the use of a set of labeled and paired training data examples to infer associations between input data and output data. An unsupervised learning algorithm may be an algorithm used to derive output data from a training dataset. An unsupervised learning algorithm may include cluster analysis, which may be used for exploratory data analysis to find hidden patterns or groupings in process data. An example of an unsupervised learning method may include principal component analysis, which may include dimensionality reduction of one or more variables. The dimensionality of a given variable may be at least 1, 5, 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800 or more. The dimensionality of a given variable may be up to 1800, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10 or less.

[0314]

[0414] The training algorithm may be obtained through statistical techniques, which in some embodiments may include linear regression, classification, resampling methods, subset selection, shrinkage, dimensionality reduction, non-linear models, tree-based methods, support vector machines, unsupervised learning, or any combination thereof.

[0315]

[0415] Linear regression is a method for predicting a target variable by fitting the best linear relationship between the dependent and independent variables. The best fit can mean that the sum of all distances between the shape at each point and the actual observations is minimal. Linear regression can include simple linear regression and multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting the best linear relationship.

[0316]

[0416] Classification can be a data mining technique that assigns categories to a collection of data to achieve accurate predictions and analysis. Classification techniques can include logistic regression and discriminant analysis. Logistic regression can be used when the dependent variable is dichotomous (binary). Logistic regression can be used to discover and describe the relationship between one dependent binary variable and one or more nominal, ordinal, interval, or ratio-level independent variables. Resampling can be a method that involves drawing replicate samples from the original data sample. Resampling may not involve the use of general distribution tables to calculate approximate probability values. Resampling can generate a unique sampling distribution based on actual data. In some embodiments, resampling may use experimental rather than analytical methods to generate a unique sampling distribution. Resampling techniques can include bootstrapping and cross-validation. Bootstrapping can be performed by sampling with replacement from the original data and taking "unselected" data points as test cases. Cross-validation can be performed by dividing the training data into multiple parts.

[0317]

[0417] Subset selection can identify a subset of predictors related to the response. Subset selection can include best subset selection, forward stepwise selection, backward stepwise selection, hybrid methods, or any combination thereof. In some embodiments, shrinkage fits a model involving all predictors, but the estimated coefficients are shrunk towards zero compared to the least squares method. This shrinkage can reduce variance. Shrinkage can include ridge regression and lasso. Dimension reduction can reduce the problem of estimating n + 1 coefficients to the easier problem of m + 1 coefficients (where m < n). This can be done by computing n different linear combinations of the variables or by projection. These n projections are then used as predictors to fit a linear regression model by the least squares method. Dimension reduction can include principal component regression and partial least squares. Principal component regression can be used to extract a low-dimensional set of features from a large set of variables. The principal components used in principal component regression can then capture the maximum variance in the data using linear combinations of the data in orthogonal directions. Since partial least squares uses response variance to identify new features, partial least squares can be used instead of principal component regression.

[0318]

[0418] Nonlinear regression can be a form of regression analysis in which the observed data are modeled by a function that is a nonlinear combination of model parameters and depends on one or more independent variables. Nonlinear regression can include step functions, piecewise functions, splines, and generalized additive models, or any combination thereof.

[0319]

[0419] Tree-based methods can be used for both regression and classification problems. Regression and classification problems can involve stratifying or segmenting the predictor space into multiple simple regions. Tree-based methods can include bagging, boosting, random forests, or any combination of these. Bagging can reduce the variability of predictions by generating additional data for training from the original data using a combination of iterations to produce multiple steps of the same carnality / size as the original data. Boosting can calculate the output using several different models and then average the results using a weighted average approach. Random forest algorithms can derive random bootstraps of the training set. Support vector machines are classification techniques. Support vector machines can involve finding a hyperplane that best separates two classes of points with the largest margin. Support vector machines can constrain the optimization problem so that the margin is maximized subject to the constraint of perfectly classifying the data.

[0320]

[0420] An unsupervised method can be a method for drawing inferences from a dataset that includes input data that does not include labeled responses. Unsupervised methods can include clustering, principal component analysis, k-means clustering, hierarchical clustering, or any combination thereof.

[0321]

[0421] The mass spectrometry may be single-allele mass spectrometry. In some embodiments, the mass spectrometry may be MS analysis, MS / MS analysis, LC-MS / MS analysis, or a combination thereof. In some embodiments, MS analysis may be used to determine the mass of an intact peptide. For example, determining may include determining the mass of an intact peptide (e.g., MS analysis). In some embodiments, MS / MS analysis may be used to determine the mass of a peptide fragment. For example, determining may include determining the mass of a peptide fragment, which may be used to determine the amino acid sequence of the peptide or portion thereof (e.g., MS / MS analysis). In some embodiments, the mass of a peptide fragment may be used to determine the sequence of amino acids within the peptide. In some embodiments, LC-MS / MS analysis may be used to separate a complex peptide mixture. For example, determining may include separating a complex peptide mixture, e.g., by liquid chromatography, and determining the mass of an intact peptide, the mass of a peptide fragment, or a combination thereof (e.g., LC-MS / MS analysis). This data may be used, for example, for peptide sequencing.

[0322]

[0422] Peptides can be presented by HLA proteins expressed in cells through autophagy, which allows for the orderly degradation and recycling of cellular components. Autophagy can include macroautophagy, microautophagy, and chaperone-mediated autophagy. Peptides can be presented by HLA proteins expressed in cells through phagocytosis, which may be the primary mechanism used to remove pathogens and cellular debris. For example, when macrophages ingest pathogenic microorganisms, the pathogens are captured in phagosomes, which then fuse with lysosomes to form phagolysosomes. In the case of HLA class II, phagocytic cells, such as macrophages and immature dendritic cells, can ingest entities by phagocytosis into phagosomes—although B cells exhibit a more general endocytosis into endosomes—which fuse with lysosomes, where their acidic enzymes cleave the ingested proteins into numerous different peptides.

[0323]

[0423] The quality of the training data can be improved by using multiple quality metrics. The multiple quality metrics can include removal of common contaminating peptides, high scoring peak intensity, high score, and high mass accuracy. The scoring peak intensity can be used before performing scoring. The MS / MS search first screens MS / MS spectra for candidate sequences using a simple filter. This filter can be a minimum scoring peak intensity. Using the scoring peak intensity can enhance the speed of the search by ensuring that candidate sequences are quickly and immediately rejected if a sufficient number of spectral peaks are considered and found not to meet the threshold established by the filter. The scoring peak intensity can be at least 50%. The scoring peak intensity can be at least 70%. The scoring peak intensity can be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more. In some cases, the scored peak intensity may be up to 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or less. The score may be at least 7. The score may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more. In some cases, the score may be up to about 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or less. The mass accuracy may be up to 5 ppm. The mass accuracy may be up to 10 ppm, 9 ppm, 8 ppm, 7 ppm, 6 ppm, 5 ppm, 4 ppm, 3 ppm, 2 ppm, 1 ppm, or less. The mass accuracy may be at least 1 ppm, 2 ppm, 3 ppm, 4 ppm, 5 ppm, 6 ppm, 7 ppm, 8 ppm, 9 ppm, 10 ppm, or more.

[0324]

[0424] In some embodiments, the mass accuracy is up to 2 ppm. In some embodiments, the backbone cleavage score is at least 5. In some embodiments, the backbone cleavage score is at least 8.

[0325]

[0425] The peptide presented by the HLA protein expressed on the cell can be a peptide presented by a single immunoprecipitated HLA protein expressed on the cell. Immunoprecipitation (IP) can be a technique in which a protein antigen is precipitated from solution using an antibody that specifically binds to a particular protein. This process can be used to isolate and concentrate a specific protein from a sample containing thousands of different proteins. Immunoprecipitation may require that the antibody be coupled to a solid substrate at some point during the procedure.

[0326]

[0426] The peptides presented by HLA proteins expressed in cells can be peptides presented by a single exogenous HLA protein expressed in cells. The single exogenous HLA protein can be generated by introducing one or more exogenous peptides into a population of cells. In some embodiments, the introducing comprises contacting the population of cells with one or more exogenous peptides or expressing one or more exogenous peptides in the population of cells. In some embodiments, the introducing comprises contacting the population of cells with one or more nucleic acids encoding one or more exogenous peptides. In some embodiments, the one or more nucleic acids encoding one or more peptides are DNA. In some embodiments, the one or more nucleic acids encoding one or more peptides are RNA, optionally wherein the RNA is mRNA. In some embodiments, the enrichment does not involve the use of a tetramer (or multimer) reagent.

[0327]

[0427] The peptide presented by the HLA protein expressed in the cell can be a peptide presented by a single recombinant HLA protein expressed in the cell. The recombinant HLA protein can be encoded by a recombinant HLA class I or HLA class II allele. The HLA class I can be selected from the group consisting of HLA-A, HLA-B, and HLA-C. The HLA class I can be a non-classical class-Ib group. The HLA class I can be selected from the group consisting of HLA-E, HLA-F, and HLA-G. The HLA class I can be a non-classical class-Ib group selected from the group consisting of HLA-E, HLA-F, and HLA-G. In some embodiments, the HLA class II comprises an HLA class II α chain, an HLA class II β chain, or a combination thereof.

[0328]

[0428] The plurality of predictor variables may include a peptide-HLA affinity predictor variable. The plurality of predictor variables may include a source protein expression level predictor variable. The source protein expression level may be the expression level of the source protein of the peptide in the cell. In some embodiments, the expression level may be determined by measuring the amount of the source protein or the amount of RNA encoding the source protein. The plurality of predictor variables may include peptide sequence, amino acid properties, peptide properties, the expression level of the source protein of the peptide in the cell, protein stability, protein translation rate, ubiquitination sites, protein degradation rate, translation efficiency from ribosome profiling, protein cleavability, protein localization, motifs in host proteins that promote TAP transport, host proteins subject to autophagy, motifs that support ribosome stalling (e.g., polyproline or polylysine stretches), and protein features that support NMD (e.g., long 3'UTR, stop codon > 50 nt upstream of the last exon:exon junction, and peptide cleavability).

[0329]

[0429] The plurality of predictors may include a peptide cleavability predictor. Peptide cleavability may be associated with a cleavage linker or cleavage sequence. In some embodiments, the cleavage linker is a ribosome skipping site or an internal ribosome entry site (IRES) element. In some embodiments, the ribosome skipping site or IRES is cleaved when expressed in a cell. In some embodiments, the ribosome skipping site is selected from the group consisting of F2A, T2A, P2A, and E2A. In some embodiments, the IRES element is selected from common cellular or viral IRES sequences. The cleavage sequence, e.g., F2A, or an internal ribosome entry site (IRES), may be located between the alpha chain and beta2-microglobulin (HLA class I) or between the alpha chain and beta chain (HLA class II). In some embodiments, the single HLA class I allele is HLA-A*02:01, HLA-A*23:01 and HLA-B*14:02, or HLA-E*01:01, and the HLA class II allele is HLA-DRB*01:01, HLA-DRB*01:02 and HLA-DRB*11:01, HLA-DRB*15:01, or HLA-DRB*07:01. In some embodiments, the cleavage sequence is a T2A, P2A, E2A, or F2A sequence. For example, the cleavage sequence can be EGRGSLTCGDVENPGP (SEQ ID NO: 6) (T2A), ATNFSLKQAGDVENPGP (SEQ ID NO: 7) (P2A), QCTNYALKLAGDVESNPGP (SEQ ID NO: 8) (E2A), or VKQTLNFDLKLAGDVESNPGP (SEQ ID NO: 9) (F2A).

[0330]

[0430] In some embodiments, the cleavage sequence may be the thrombin cleavage site CLIP.

[0431] The peptides presented by HLA proteins may include peptides identified by searching a non-enzyme-specific peptide database without modifications. The peptide database may be a non-enzyme-specific peptide database, for example, a database without modifications or a database with modifications (e.g., phosphorylation or cysteinylation). In some embodiments, the peptide database is a polypeptide database. In some embodiments, the polypeptide database may be a protein database. In some embodiments, the method further includes searching the peptide database using a reverse database search strategy. In some embodiments, the method further includes searching a protein database using a reverse database search strategy. In some embodiments, a de novo search is performed, for example, to discover novel peptides not included in normal peptide or protein databases. The peptide database may be generated by providing first and second populations of cells, each comprising one or more cells containing sequence affinity acceptor-tagged HLA peptides comprising various recombinant polypeptides encoded by various HLA alleles operably linked to affinity acceptor peptides; enriching for affinity acceptor-tagged HLA peptide complexes; characterizing peptides or portions thereof that bind to the affinity acceptor-tagged HLA peptide complexes from the enrichment; and generating an HLA allele-specific peptide database.

[0331]

[0432] Peptides presented by an HLA protein can include peptides identified by comparing the MS / MS spectrum of an HLA peptide to the MS / MS spectra of one or more HLA peptides in a peptide database.

[0332]

[0433] Mutations may occur in either the peptide or the nucleic acid encoding the peptide. The mutations may be selected from the group consisting of point mutations, splice site mutations, frameshift mutations, read-through mutations, and gene fusion mutations. Point mutations may be genetic mutations in which a single nucleotide base is changed, inserted, or deleted from the DNA or RNA sequence. Splice site mutations may be genetic mutations in which multiple nucleotides are inserted, deleted, or changed at a specific site where splicing occurs during processing of precursor messenger RNA into mature messenger RNA. Frameshift mutations may be genetic mutations caused by indels (insertions or deletions) of multiple nucleotides in the DNA sequence that are not divisible by three. Mutations may also include insertions, deletions, substitution mutations, gene duplications, chromosomal translocations, and chromosomal inversions.

[0333]

[0434] In some embodiments, the HLA class II protein comprises an HLA-DR protein.

[0435] In some embodiments, the HLA class II protein comprises an HLA-DP protein.

[0334]

[0436] In some embodiments, the HLA class II protein comprises an HLA-DQ protein.

[0437] In some embodiments, the HLA class II protein may be selected from the group consisting of HLA-DR and HLA-DP or HLA-DQ proteins. In some embodiments, the HLA proteins are: HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA -DQA1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB 1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04: 04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, H LA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-D The peptide presented by the HLA protein may be an HLA class II protein selected from the group consisting of HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, and HLA-DRB5*01:01. The peptide presented by the HLA protein may be 15 to 40 amino acids in length.Peptides presented by HLA proteins can have a length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or more amino acids. In some embodiments, peptides presented by HLA proteins can have a length of up to 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11 or fewer amino acids.

[0335]

[0438] Peptides presented by HLA proteins may include peptides identified by the steps of: (a) isolating one or more HLA complexes from a cell line expressing a single HLA class II allele; (b) isolating one or more HLA peptides from the one or more isolated HLA complexes; (c) obtaining MS / MS spectra for the one or more isolated HLA peptides; and (d) obtaining peptide sequences from a peptide database corresponding to the MS / MS spectra of the one or more isolated HLA peptides; wherein the one or more sequences obtained from steps (a, b, c) and (d) identify the sequences of the one or more isolated HLA peptides.

[0336]

[0439] Isolating may include isolating HLA-peptide complexes from cells transfected or transduced with affinity-tagged HLA constructs. In some embodiments, the complexes may be isolated using standard immunoprecipitation techniques well known in the art using commercially available antibodies. Cells may first be lysed. HLA class II-peptide complexes may be isolated using an HLA class II-specific antibody, such as the M5 / 114.15.2 monoclonal antibody. In some embodiments, a single (or pair of) HLA alleles is expressed as a fusion protein with a peptide tag, and the HLA-peptide complex is isolated using a binding molecule that recognizes the peptide tag.

[0337]

[0440] Isolating may further include isolating peptides from HLA-peptide complexes and sequencing the peptides. Peptides are isolated from the complexes by any method known to those skilled in the art, such as acid elution. Any sequencing method may be used, and in some embodiments, a method using mass spectrometry, such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or alternatively HPLC-MS or HPLC-MS / MS), is utilized. These sequencing methods are known to those skilled in the art and are generally described in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb;34(1):43-63.

[0338]

[0441] Additional candidate components and molecules suitable for isolation or purification can include binding molecules such as biotin (biotin-avidin specific binding pair), antibodies, receptors, ligands, lectins, or molecules containing solid supports, including, for example, plastic or polystyrene beads, plates or beads, magnetic beads, test strips, and membranes. Purification methods such as cation exchange chromatography can be used to separate conjugates based on charge differences, which effectively separates conjugates into their various molecular weights. The contents of the fractions obtained by cation exchange chromatography can be identified by molecular weight using conventional methods, such as mass spectrometry, SDS-PAGE, or other well-known methods for separating molecular entities by molecular weight.

[0339]

[0442] In some embodiments, the method further comprises isolating the peptide from the affinity acceptor-tagged HLA peptide complex prior to characterization. In some embodiments, the HLA peptide complex is isolated using an anti-HLA antibody. In some cases, the HLA peptide complex with or without an affinity tag is isolated using an anti-HLA antibody. In some cases, soluble HLA (sHLA) with or without an affinity tag is isolated from the cell culture medium. In some cases, soluble HLA (sHLA) with or without an affinity tag is isolated using an anti-HLA antibody. For example, HLA, e.g., soluble HLA (sHLA) with or without an affinity tag, can be isolated using beads or a column containing an anti-HLA antibody. In some embodiments, the peptide is isolated using an anti-HLA antibody. In some cases, soluble HLA (sHLA) with or without an affinity tag is isolated using an anti-HLA antibody. In some cases, soluble HLA (sHLA) with or without an affinity tag is isolated using a column containing an anti-HLA antibody. In some embodiments, the method further comprises removing one or more amino acids from the termini of the peptides bound to the affinity acceptor-tagged HLA peptide complexes.

[0340]

[0443] A personalized cancer vaccine may further include an adjuvant. For example, poly-ICLC, an agonist of TLR3 and the RNA helicase domains of MDA5 and RIG3, has demonstrated several desirable properties for a vaccine adjuvant. These properties include inducing local and systemic activation of immune cells in vivo, producing stimulatory chemokines and cytokines, and stimulating antigen presentation by DCs. Furthermore, poly-ICLC can induce sustained CD4+ and CD8+ responses in humans. Importantly, striking similarities in the upregulation of transcriptional and signaling pathways were observed in subjects vaccinated with poly-ICLC and in volunteers who received a highly effective, replicative yellow fever vaccine. Fur...

Claims

1. 1. A method for identifying a peptide sequence presented by at least one of one or more proteins encoded by HLA alleles of a cell of a subject, comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction of the plurality of presentation predictions indicates a likelihood that a peptide sequence of the set of candidate peptide sequences will be presented by an MHC protein of the single human subject; and wherein the trained machine learning HLA peptide presentation prediction model: (i) a plurality of parameters based on training data from training cells expressing MHC proteins, the training data comprising a plurality of training peptide sequences and epitope presentation quantitative information, the epitope presentation quantitative information comprising one or more amounts of the plurality of training peptide sequences presented by the MHC proteins; and (ii) a function that represents the relationship between amino acid sequence information received as input and a representation possibility generated as output based on the amino acid sequence information and a plurality of parameters; and (b) identifying a peptide sequence of the plurality of peptide sequences of the set of candidate peptide sequences that is presented by at least one of the one or more proteins encoded by HLA alleles of the cell of the subject based at least on the plurality of presentation predictions. A method comprising:

2. 1. A method for selecting a peptide sequence, comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction of the plurality of presentation predictions indicates a likelihood that a peptide sequence of the set of candidate peptide sequences will be presented by an MHC protein of the single human subject; and wherein the trained machine learning HLA peptide presentation prediction model: (i) a plurality of parameters based on training data from training cells expressing MHC proteins, the training data comprising a plurality of training peptide sequences and epitope presentation quantitative information, the epitope presentation quantitative information comprising one or more amounts of the plurality of training peptide sequences presented by the MHC proteins; and (ii) a function that represents the relationship between amino acid sequence information received as input and a representation possibility generated as output based on the amino acid sequence information and a plurality of parameters; and (b) selecting a subset of peptide sequences of the set of candidate peptide sequences to generate a set of selected peptide sequences based at least on the plurality of presented predictions. A method comprising:

3. 1. A method of treating cancer in a human subject in need thereof, comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each presentation prediction of the plurality of presentation predictions indicates a likelihood that a peptide sequence of the set of candidate peptide sequences will be presented by an MHC protein of the single human subject; and wherein the trained machine learning HLA peptide presentation prediction model: (i) a plurality of parameters based on training data from training cells expressing MHC proteins, the training data comprising a plurality of training peptide sequences and epitope presentation quantitative information, the epitope presentation quantitative information comprising one or more amounts of the plurality of training peptide sequences presented by the MHC proteins; and (ii) a function that represents the relationship between amino acid sequence information received as input and a representation possibility generated as output based on the amino acid sequence information and a plurality of parameters; the steps of: (b) selecting or identifying a subset of peptide sequences of the set of candidate peptide sequences to generate a set of selected or identified peptide sequences based at least on the plurality of presented predictions; and (c)(i) a polypeptide having one or more of the selected peptide sequences; (ii) a polynucleotide encoding the polypeptide of (i); (iii) an APC comprising (i) or (ii); or (iv) T cells comprising a T cell receptor (TCR) specific for a single MHC protein of a human subject in complex with one or more of the peptide sequences selected or identified in (b). administering to a single human subject a pharmaceutical composition comprising A method comprising:

4. 4. The method of claim 1, wherein the plurality of parameters is based on training data from training cells expressing MHC proteins of a single human subject.

5. 5. The method of claim 4, wherein each of the plurality of training peptide sequences is associated with an MHC protein.

6. 6. The method of claim 5, wherein the training data comprises the identity of the MHC protein associated with each of the plurality of training peptide sequences.

7. 7. The method of claim 6, wherein the training data comprises mass spectrometry observations that one or more of the plurality of training peptide sequences are presented by MHC proteins.

8. The method of any one of claims 1 to 7, wherein the MHC protein of a single human subject is a class I MHC protein.

9. 9. The method of any one of claims 1 to 8, wherein the plurality of candidate peptide sequences expressed by cancer cells of a single human subject is identified by comparing whole genome or whole exome sequence information from cancer cells of the single human subject with whole genome or whole exome sequence information from non-cancerous cells of the single human subject, and identifying nucleic acid sequences that are unique to the cancer cells and absent from the non-cancerous cells.

10. The method according to any one of claims 1 to 9, wherein each candidate sequence of the plurality of candidate peptide sequences comprises a cancer-specific mutation.

11. 11. The method of any one of claims 1 to 10, wherein the trained machine learning HLA peptide presentation prediction model has a peptide presentation predictive value (PPV) of at least 0.2 according to a presentation PPV determination method.

12. A method for determining presentation PPV includes inputting amino acid sequence information of a plurality of test peptide sequences into a trained machine learning HLA peptide presentation prediction model to generate a plurality of test presentation predictions, each test presentation prediction indicating the likelihood that one or more proteins encoded by an HLA allele can present a given test peptide sequence of the plurality of test peptide sequences, wherein the plurality of test peptide sequences: (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell; and (ii) at least 499 decoy peptide sequences contained within proteins encoded by the genome of the organism; 12. The method of any one of claims 1 to 11, comprising at least 500 test peptide sequences comprising:

13. 13. The method of claim 12, wherein the plurality of test peptide sequences has a 1:499 ratio of at least one hit peptide sequence to at least 499 decoy peptide sequences, and the top 0.2% of the plurality of test peptide sequences are predicted to be presented by HLA proteins expressed in cells by a trained machine learning HLA peptide presentation prediction model.

14. The method of claim 13, wherein (i) the at least one hit peptide sequence comprises at least 10 hit peptide sequences, and (ii) the at least 499 decoy peptide sequences comprise at least 4,990 decoy peptide sequences.

15. 15. The method of any one of claims 1 to 14, wherein the one or more amounts of the plurality of training peptide sequences presented by the MHC proteins comprise the number of copies of the one or more of the plurality of training peptide sequences presented by the MHC proteins.

16. 16. The method of any one of claims 1 to 15, wherein the one or more amounts of the plurality of training peptide sequences presented by MHC proteins comprise one or more numbers of copies per cell of the plurality of training peptide sequences presented by MHC proteins.

17. 17. The method of any one of claims 1 to 16, wherein the amount of one or more of the plurality of training peptide sequences presented by the MHC proteins comprises an absolute amount, number of molecules, density, concentration, absolute amount per cell, number of molecules per cell, density per cell, or concentration in a cell of one or more of the plurality of training peptide sequences presented by the MHC proteins.

18. 18. The method of any one of claims 1 to 17, wherein the quantity of one or more of the plurality of training peptide sequences presented by the MHC proteins is based on a number of mass spectrometry observations, spectral counts, area under the curve (AUC), intensity-based absolute quantification (iBAQ), label-free quantification (LFQ), isotope dilution mass spectrometry, isobaric mass tagging, stable isotope labeling, and / or mass spectrometry peak intensities.

19. 19. The method of any one of claims 1 to 18, wherein the quantity of one or more of the plurality of training peptide sequences presented by the MHC proteins is obtained from quantitative mass spectrometry.

20. The method of any one of claims 1 to 19, wherein epitope presentation quantitative information is obtained from internal standard parallel reaction monitoring (IS-PRM) mass spectrometry.

21. The method of any one of claims 1 to 20, wherein epitope presentation quantitative information is obtained from xenograft samples.

22. 22. The method of claim 21, wherein the xenograft sample is a patient-derived xenograft (PDX) sample.

23. 1. A method for selecting a peptide sequence, comprising: (a) using a computer processor, inputting amino acid sequence information of a set of candidate peptide sequences expressed by cancer cells of a single human subject into a trained machine-learning HLA-peptide antigen-specific T cell prediction model to generate a plurality of antigen-specific T cell predictions, wherein each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of the single human subject and a peptide sequence of the set of candidate peptide sequences will stimulate a T cell to be specific for a peptide sequence of the set of candidate peptide sequences; and wherein the trained machine-learning HLA-peptide cytotoxic T cell prediction model: (i) a plurality of parameters based on training data from training cells expressing MHC proteins, the training data comprising a plurality of training peptide sequences and epitope presentation quantitative information, the epitope presentation quantitative information comprising one or more amounts of the plurality of training peptide sequences presented by the MHC proteins; and (ii) a function representing the relationship between amino acid sequence information received as input and the likelihood that T cells specific to a peptide sequence in the set of candidate peptide sequences will be generated as output based on the amino acid sequence information and multiple parameters; and (b) selecting a subset of peptide sequences of the set of candidate peptide sequences to generate a set of selected peptide sequences based at least on the plurality of antigen-specific T cell predictions. A method comprising:

24. 24. The method of claim 23, wherein each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of a single human subject and a peptide sequence of the set of candidate peptide sequences will stimulate a T cell to be specific for a neo-antigenic peptide sequence of the set of candidate peptide sequences.

25. 25. The method of claim 24, wherein the function represents the relationship between amino acid sequence information received as input and the likelihood that T cells specific to the neo-antigenic peptide sequences of the set of candidate peptide sequences will be generated as output based on the amino acid sequence information and multiple parameters.

26. 24. The method of claim 23, wherein each antigen-specific T cell prediction of the plurality of antigen-specific T cell predictions indicates the likelihood that an MHC complex comprising an MHC protein of a single human subject and a peptide sequence of the set of candidate peptide sequences will stimulate a T cell to be cytotoxic.

27. 27. The method of claim 26, wherein the function represents the relationship between amino acid sequence information received as input and the likelihood that cytotoxic T cells will be generated as output based on the amino acid sequence information and multiple parameters.