Methods and systems for predicting HLA class II-specific epitopes and characterizing CD4+ T cells

JP2025518569A5Pending Publication Date: 2026-05-26BIONTECH US INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
BIONTECH US INC
Filing Date
2023-05-18
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Current methods for predicting and identifying HLA class II-binding epitopes are limited by inaccurate prediction standards, such as NetMHCIIpan, and lack of reliable quantification of gene expression, enzymatic cleavage, and pathway/localization biases, which hampers the development of effective immunotherapy treatments.

Method used

The use of machine learning algorithms trained with high-quality LC-MS/MS single allele data to improve the prediction of HLA class II-ligand and epitope presentation, including the identification of allele-specific binding cores and biological variables affecting presentation.

Benefits of technology

This approach enhances the accuracy of predicting HLA class II-binding epitopes, improving the identification of truly presented epitopes and potentially increasing the therapeutic efficacy of immunotherapy by targeting specific tumor-associated antigens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods for preparing individualized cancer vaccines and methods for training machine learning HLA peptide presentation prediction models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference

[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 343,913, filed May 19, 2022, which is incorporated herein by reference in its entirety.

Background Art

[0002]

[0002] The major histocompatibility complex (MHC) is a gene complex that encodes the human leukocyte antigen (HLA) genes. The HLA genes are expressed as protein heterodimers that are presented to circulating T cells on the surface of human cells. The HLA genes are highly polymorphic, allowing them to finely tune the adaptive immune system. The adaptive immune response depends in part on the ability of T cells to identify and eliminate cells presenting disease-associated peptide antigens bound to human leukocyte antigen (HLA) heterodimers.

[0003]

[0003] In humans, endogenous and exogenous proteins are processed into peptides by the proteasome, as well as by cytosolic and endosomal / lysosomal proteases and peptidases, and can be presented by two classes of cell surface proteins encoded by the MHC genes. These cell surface proteins are referred to as human leukocyte antigens (HLA class I and class II), and the group of peptides that bind to them and elicit an immune response are named HLA epitopes. HLA epitopes are important components that enable the immune system to detect danger signals such as pathogen infections and self-transformations. CD4+ T cells recognize class II MHC (HLA-DR, HLA-DQ, and HLA-DP) epitopes presented on antigen-presenting cells (APCs) such as dendritic cells and macrophages. The endogenous processing and presentation of HLA class II-ligands is a complex procedure involving a subset of various chaperones and enzymes that are not all fully characterized. HLA class II-peptide presentation activates helper T cells, subsequently promoting B cell differentiation and antibody production as well as CTL responses. Activated helper T cells also activate other T cells and secrete cytokines and chemokines that induce differentiation.

[0004]

[0004] Understanding the peptide-binding selectivity of all HLA class II heterodimers is important for successfully predicting which cancers or tumor-specific antigens may be able to elicit cancer- or tumor-specific T cell responses. Methods are needed to identify and isolate specific HLA class II-related peptides (e.g., neoantigen peptides). Such methods and isolated molecules are useful, for example, but not limited to, the development of treatments such as immunotherapy-based treatments.

Summary of the Invention

Problems to be Solved by the Invention

[0005]

[0005] The methods and compositions described herein are used in a wide range of applications. For example, the methods and compositions described herein can be used to identify immunogenic antigenic peptides, to develop drugs, such as personalized pharmaceuticals, and to isolate and characterize antigen-specific T cells.

[0006]

[0006] CD4+ T cell responses can have antitumor activity. High rates of CD4+ T cell responses can be shown without using class II predictions (e.g., see 60% of SLP epitopes in the NeoVax study (49% in NT-001, Ott et al., Nature, 2017 Jul 13;547(7662):217-221) and 48% of mRNA epitopes in the Biontech study, Sahin et al., Nature, 2017 Jul 13;547(7662):222-226). It may not be clear whether these epitopes are naturally and typically presented (by tumors or by phagocytic DCs). It may be desirable to convert high CD4+ T response rates into therapeutic efficacy by improving the identification of truly presented HLA class II-binding epitopes.

[0007]

[0007] The roles of gene expression, enzymatic cleavage, and pathway / localization biases may not be reliably quantified. Most existing MS data can be presumed to be from autophagy, but it may be unclear which of autophagy (HLA class II presentation by tumor cells) or phagocytosis (HLA class II presentation of tumor epitopes by APCs) is the more appropriate pathway. NetMHCIIpan is the current prediction standard, but it may not be considered accurate. Of the three HLA class II gene loci (DR, DP, and DQ), data exist only for specific common alleles of HLA-DR.

[0008]

[0008] There are various data generation approaches for learning the rules of HLA class II presentation, including field standard and proposed approach. The field standard can include affinity measurements that can serve as the basis for NetMHCIIpan predictors, which result in low throughput, require radioactive reagents, and miss processing rules. The proposed approach can include mass spectrometry that can assist in determining processing rules related to autophagy from cell line / tissue / tumor-derived data, and single allele MS can enable the determination of allele-specific binding rules (multi-allele MS data is presumed to be overly complex for efficient learning (Bassani-Sternberg. MCP. 2018)).

[0009]

[0009] There are various methods for validating novel HLA class II predictors: validation with holdout MS data, which may be the initial setting; retrospective vaccine studies (e.g., NT-001), where immune monitoring data evaluates vaccine peptides loaded onto APCs rather than tumor presentation, and the data may be thinly spread across a large number of different alleles; biochemical affinity measurements, which can be configured to obtain measurements for peptides predicted inconsistently (for only 2 - 3 alleles); T cell induction, which can be configured to test the rate at which Neon-prioritized and NetMHCIIpan-prioritized epitopes induce ex vivo T cell responses.

[0010]

[0010] For validation through T cell induction, the initial approach may include evaluating neoORFs from TCGA predicted inconsistently, where the inducer may include healthy donor APCs and T cells, and induction and reading may use SLP (peptides approximately 15 amino acids long). Random peptides result in a high rate of response, and SLP may inadequately address processing. A possible solution may include induction via mRNA.

Means for Solving the Problems

[0011]

[0011] The methods disclosed herein may include the step of generating LC-MS / MS single allele data for training of an allele-specific machine learning method for epitope prediction. Such methods may include the step of improving LC-MS / MS data quality using a set of quality metrics that improve the performance of the prediction model and strictly remove false positives; the step of identifying allele-specific HLA class II binding cores from an HLA-ligandome LC-MS / MS data set; the step of using a machine learning algorithm to improve HLA class II-ligand and epitope prediction; and / or the step of identifying biological variables, such as gene expression, cleavage, gene bias, subcellular localization, and secondary structure, that affect HLA class II-ligand presentation and improve HLA class II epitope prediction.

[0012]

[0012] Herein: (a) processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each candidate peptide sequence of the plurality of candidate peptide sequences is encoded by the genome or exome of a subject, the plurality of presentation predictions includes HLA presentation predictions for each of the plurality of candidate peptide sequences, and each HLA presentation prediction indicates the likelihood that one or more proteins encoded by the class II HLA alleles of the subject's cells can present a given candidate peptide sequence of the plurality of candidate peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in training cells; and (b) identifying peptide sequences of a plurality of peptide sequences presented by at least one of one or more proteins encoded by the class II HLA alleles of the subject's cells based at least on the plurality of presentation predictions, wherein the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 according to a presentation PPV determination method, a method is provided.

[0013] In this specification: (a) a step of processing amino acid information of a plurality of peptide sequences encoded by a genome or exome of a subject using a machine learning HLA peptide binding prediction model to generate a plurality of binding predictions, wherein the plurality of binding predictions includes HLA binding predictions for each of the plurality of candidate peptide sequences, and each binding prediction indicates the possibility that one or more proteins encoded by a class II HLA allele of the subject's cells bind to a given candidate peptide sequence of the plurality of candidate peptide sequences, and the machine learning HLA peptide binding prediction model is trained using training data including sequence information of peptides identified as binding to an HLA class II protein or an HLA class II protein analog; and (b) identifying peptide sequences of a plurality of peptide sequences having a probability higher than a binding prediction probability value threshold for binding to at least one of one or more proteins encoded by a class II HLA allele of the subject's cells, based at least on the plurality of binding predictions, wherein the machine learning HLA peptide binding prediction model has a positive predictive value (PPV) of at least 0.1 according to a binding PPV determination method.

[0014] In some embodiments, the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in training cells.

[0015] In some embodiments, the method includes ranking at least two peptides identified as being presented by at least one of one or more proteins encoded by a class II HLA allele of the subject's cells, based on presentation predictions.

[0016] In some embodiments, the method includes selecting one or more peptides of two or more ranked peptides.

[0017] In some embodiments, the method includes selecting one or more of a plurality of peptides identified as being presented by at least one of one or more proteins encoded by a Class II HLA allele of a target cell.

[0017]

[0018] In some embodiments, the method includes selecting one or more peptides from two or more peptides ranked based on presentation prediction.

[0019] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 when amino acid information of a plurality of test peptide sequences is processed to generate a plurality of test presentation predictions, where each test presentation prediction indicates the likelihood that one or more proteins encoded by a Class II HLA allele of a target cell can present a given test peptide sequence of the plurality of test peptide sequences, where the plurality of test peptide sequences includes at least 500 test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained within proteins encoded by the genome of an organism, where the organism and the target are of the same species, where the plurality of test peptide sequences has a 1:499 ratio of at least 499 decoy peptide sequences to at least one hit peptide sequence, and where the top percentage of the plurality of test peptide sequences is predicted by the machine learning HLA peptide presentation prediction model to be presented by an HLA protein expressed in the cell.

[0018]

[0020] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.1 when the amino acid information of a plurality of test peptide sequences is processed to generate a plurality of test binding predictions, and each test binding prediction indicates the likelihood that one or more proteins encoded by a class II HLA allele of a subject cell will bind to a given test peptide sequence of the plurality of test peptide sequences, where the plurality of test peptide sequences includes (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, for example, at least 19 decoy peptide sequences contained within a protein that includes a single HLA protein expressed in a cell (e.g., a single allele cell), such that the plurality of test peptide sequences includes at least 20 test peptide sequences, where the plurality of test peptide sequences includes a 1:19 ratio of at least 19 decoy peptide sequences to at least one hit peptide sequence, and the top percentage of the plurality of test peptide sequences is predicted by the machine learning HLA peptide presentation prediction model to bind to an HLA protein expressed in a cell.

[0019]

[0021] In some embodiments, the amino acid sequence overlap is not present within at least one hit peptide sequence and decoy peptide sequence.

[0022] In some embodiments, the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98 or 0.99.

[0020]

[0023] In some embodiments, at least one hit peptide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences.

[0021]

[0024] In some embodiments, at least 499 decoy peptide sequences are at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000,It contains 800,000, 900,000 or 1,000,000 decoy peptide sequences. Those skilled in the art can recognize that changing the hit:decoy ratio changes the PPV.

[0022]

[0025] In some embodiments, at least 500 test peptide sequences are at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000,It contains 900,000 or 1,000,000 test peptide sequences.

[0023]

[0026] In some embodiments, the top percentage is the top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%, 9.90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%.

[0024]

[0027] In some embodiments, at least one hit peptide sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences.

[0025]

[0028] In some embodiments, at least 19 decoy peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500,It contains 80,000, 82,500, 85,000, 87,500, 90,000, 92,500, 95,000, 97,500, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 325,000, 350,000, 375,000, 400,000, 425,000, 450,000, 475,000, 500,000, 600,000, 700,000, 800,000, 900,000 or 1,000,000 decoy peptide arrays.,

[0026]

[0029] In some embodiments, at least 20 test peptide sequences are among at least 500 test peptide sequences that are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000,Comprising at least including 72,500, 75,000, 77,500, 80,000, 82,500, 85,000, 87,500, 90,000, 92,500, 95,000, 97,500, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 325,000, 350,000, 375,000, 400,000, 425,000, 450,000, 475,000, 500,000, 600,000, 700,000, 800,000, 900,000 or 1,000,000 test peptide sequences.,

[0027]

[0030] In some embodiments, the top percentage is the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39% or 40%.,

[0028]

[0031] In some embodiments, the PPV is greater than each of the PPVs in column 2 of Table 11 for the proteins encoded by the corresponding HLA alleles in Table 11. In some embodiments, the PPV is at least equal to each of the PPVs in column 3 of Table 11 for the proteins encoded by the corresponding HLA alleles in Table 11.,

[0029]

[0032] In some embodiments, the PPV is equal to or greater than each of the PPVs in column 2 of Table 12 for the proteins encoded by the HLA class II alleles.,

[0030]

[0033] In some embodiments, the PPV is greater than each of the PPVs in column 2 of Table 16 for the proteins encoded by the HLA class II alleles.,

[0034] In some embodiments, the subject is a single subject.,

[0031]

[0035] In some embodiments, the subject is a mammal.

[0036] In some embodiments, the subject is a human.

[0037] In some embodiments, the training cells are cells that express a single protein encoded by a class II HLA allele of the subject's cells.

[0032]

[0038] In some embodiments, the training cells are cells that express a single allele HLA cell or an HLA allele containing an affinity tag.

[0039] In some embodiments, the subject's cells include cancer cells.

[0033]

[0040] In some embodiments, the method is for identifying a peptide sequence.

[0041] In some embodiments, the method is for selecting a peptide sequence.

[0042] In some embodiments, the method is for preparing a cancer treatment.

[0034]

[0043] In some embodiments, the method is for preparing a subject-specific cancer treatment.

[0044] In some embodiments, the method is for preparing a cancer cell-specific cancer treatment.

[0035]

[0045] In some embodiments, each peptide sequence of the plurality of peptide sequences is associated with cancer.

[0046] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is overexpressed by the subject's cancer cells.

[0036]

[0047] In some embodiments, each peptide sequence of the plurality of peptide sequences is overexpressed by the subject's cancer cells.

[0048] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is a cancer cell-specific peptide.

[0037]

[0049] In some embodiments, each peptide sequence of the plurality of peptide sequences is a cancer cell-specific peptide.

[0050] In some embodiments, each peptide sequence of the plurality of peptide sequences is expressed by the cancer cells of the subject.

[0038]

[0051] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is not encoded by the non-cancer cells of the subject.

[0052] In some embodiments, each peptide sequence of the plurality of peptide sequences is not encoded by the non-cancer cells of the subject.

[0039]

[0053] In some embodiments, at least one peptide sequence of the plurality of peptide sequences is not expressed by the non-cancer cells of the subject.

[0054] In some embodiments, each peptide sequence of the plurality of peptide sequences is not expressed by the non-cancer cells of the subject.

[0040]

[0055] In some embodiments, the method includes obtaining a plurality of peptide sequences of the subject.

[0056] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences of the subject.

[0041]

[0057] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences of the subject that encode a plurality of peptide sequences encoded by the genome or exome of the subject, or by a pathogen or virus in the subject.

[0042]

[0058] In some embodiments, the method includes obtaining, by a computer processor, a plurality of polynucleotide sequences of the subject that encode a plurality of peptide sequences encoded by the genome or exome of the subject.

[0043]

[0059] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences of interest by genomic or exome sequencing.

[0060] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences of interest by whole genome sequencing or by whole exome sequencing.

[0044]

[0061] In some embodiments, the step of processing includes processing by a computer processor.

[0062] In some embodiments, the step of processing includes generating a plurality of predictor variables based at least on amino acid information of a plurality of peptide sequences.

[0045]

[0063] In some embodiments, the step of processing a plurality of predictor variables using a machine learning HLA peptide presentation prediction model.

[0064] In some embodiments, the one or more proteins encoded by the class II HLA alleles of the subject's cells are the one or more proteins encoded by the class II HLA alleles expressed by the subject.

[0046]

[0065] In some embodiments, the one or more proteins encoded by the class II HLA alleles of the subject's cells are the one or more proteins encoded by the class II HLA alleles expressed by the subject's cancer cells.

[0047]

[0066] In some embodiments, the one or more proteins encoded by the class II HLA alleles of the subject's cells are a single protein encoded by the class II HLA alleles of the subject's cells.

[0048]

[0067] In some embodiments, the one or more proteins encoded by the class II HLA alleles of the subject's cells are two, three, four, five, or six or more proteins encoded by the class II HLA alleles of the subject's cells.

[0049]

[0068] In some embodiments, the one or more proteins encoded by the class II HLA alleles of the subject's cells are each protein encoded by the class II HLA alleles of the subject's cells.

[0050]

[0069] In some embodiments, the method further comprises administering to the subject a composition comprising one or more selected subsets of peptide sequences.

[0070] In some embodiments, the step of identifying a plurality of peptide sequences comprises comparing a DNA, RNA, or protein sequence from the subject's cancer cells to a DNA, RNA, or protein sequence from the subject's normal cells, wherein each of the plurality of peptides comprises at least one mutation that is present in the subject's cancer cells and not present in the subject's normal cells.

[0051]

[0071] In some embodiments, the machine learning HLA peptide presentation prediction model comprises a plurality of predictor variables identified based at least on training data, the training data comprising training peptide sequence information comprising amino acid position information, the training peptide sequence information being associated with an HLA protein expressed in a cell; and a function representing the relationship between the amino acid position information and the presentability generated as an output based on the amino acid position information and the plurality of predictor variables.

[0052]

[0072] In some embodiments, the step of identifying comprises identifying, based at least on a plurality of presentation predictions, peptide sequences of a plurality of peptide sequences having a probability higher than a presentation prediction probability value threshold presented by at least one of the one or more proteins encoded by the class II HLA alleles of the subject's cells.

[0053]

[0073] In some embodiments, one or more of 0.2% of a plurality of test peptide sequences predicted to be presented by a machine learning HLA peptide presentation prediction model have a probability higher than a presentation prediction probability value threshold presented by at least one of one or more proteins encoded by the class II HLA alleles of the cells of the subject.

[0054]

[0074] In some embodiments, each of 0.2% of a plurality of test peptide sequences predicted to be presented by a machine learning HLA peptide presentation prediction model has a probability higher than a presentation prediction probability value threshold presented by at least one of one or more proteins encoded by the class II HLA alleles of the cells of the subject.

[0055]

[0075] In some embodiments, the number of positives is restricted to be equal to the number of hits.

[0076] In some embodiments, the mass spectrometry is single-allele mass spectrometry.

[0077] In some embodiments, the peptide is presented by an HLA protein expressed in the cell through autophagy.

[0056]

[0078] In some embodiments, the peptide is presented by an HLA protein expressed in the cell through phagocytosis.

[0079] In some embodiments, the plurality of predictor variables includes expression level predictors of a source protein containing the peptide.

[0057]

[0080] In some embodiments, the plurality of predictor variables includes stability predictors of a source protein containing the peptide.

[0081] In some embodiments, the plurality of predictor variables includes degradation rate predictors of a source protein containing the peptide.

[0058]

[0082] In some embodiments, the plurality of predictor variables includes protein cleavability predictors of a source protein containing the peptide.

[0083] In some embodiments, the plurality of predictor variables includes predictors of the cellular or tissue localization of the source protein containing the peptide.

[0059]

[0084] In some embodiments, the plurality of predictor variables includes predictors of the intracellular processing pattern of the source protein containing the peptide, where the processing pattern of the source protein includes predictors as to whether the source protein is subject to autophagy, phagocytosis, and intracellular trafficking, in particular.

[0060]

[0085] In some embodiments, the quality of the training data is improved by using a plurality of quality metrics.

[0086] In some embodiments, the plurality of quality metrics includes common contaminant peptide removal, high scored peak intensity, high score, and high mass accuracy.

[0061]

[0087] In some embodiments, the scored peak intensity is at least 50%.

[0088] In some embodiments, the scored peak intensity is at least 60%.

[0089] In some embodiments, the score is at least 7.

[0062]

[0090] In some embodiments, the mass accuracy is at most 5 ppm.

[0091] In some embodiments, the peptides presented by the HLA protein expressed in the cell are the peptides presented by a single immunoprecipitated HLA protein expressed in the cell.

[0063]

[0092] In some embodiments, the peptides presented by the HLA protein expressed in the cell are the peptides presented by a single exogenous HLA protein expressed in the cell.

[0064]

[0093] In some embodiments, the peptides presented by the HLA proteins expressed in the cell are peptides presented by a single recombinant HLA protein expressed in the cell.

[0065]

[0094] In some embodiments, the plurality of predictor variables includes peptide-HLA affinity predictor variables.

[0095] In some embodiments, the peptides presented by the HLA proteins include peptides identified by searching a non-enzymatic specificity peptide database without modifications.

[0066]

[0096] In some embodiments, the peptides presented by the HLA proteins include peptides identified by searching a peptide database using an inverse database search strategy.

[0067]

[0097] In some embodiments, the HLA proteins include HLA-DR, HLA-DQ or HLA-DP proteins.

[0098] In some embodiments, the HLA protein comprises an HLA class II protein selected from the group consisting of: HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, HLA-DRB5*01:01.

[0068]

[0099] In some embodiments, HLA-DR pairs with DRA*01:01.

[0100] In some embodiments, the HLA protein is an HLA class II protein selected from the group consisting of: DPA*01:03 / DPB*04:01, DRB1*01:01, DRB1*01:02, DRB1*03:01, DRB1*04:01, DRB1*04:02, DRB1*04:04, DRB1*04:05, DRB1*07:01, DRB1*08:01, DRB1*08:02, DRB1*08:03, DRB1*09:01, DRB1*11:01, DRB1*11:02, DRB1*11:04, DRB1*12:01, DRB1*13:01, DRB1*13:02, DRB1*13:03, DRB1*14:01, DRB1*15:01, DRB1*15:02, DRB1*15:03, DRB1*16:02, DRB3*01:01, DRB3*02:01, DRB3*02:02, DRB3*03:01, DRB4*01:01, DRB4*01:03, and DRB5*01:01.

[0069]

[0101] In some embodiments, the HLA-DR protein comprises DRA*01:01 as a dimer.

[0102] In some embodiments, the HLA protein comprises an HLA-DP protein selected from the group consisting of: DPB1*01:01, DPB1*02:01, DPB1*02:02, DPB1*03:01, DPB1*04:01, DPB1*04:02, DPB1*05:01, DPB1*06:01, DPB1*11:01, DPB1*13:01, DPB1*17:01.

[0070]

[0103] In some embodiments, the HLA-DP protein comprises DPA1*01:03 for pairing.

[0104] In some embodiments, the HLA protein comprises an HLA-DQ protein complex selected from the group consisting of: A1*01:01+B1*05:01, A1*01:02+B1*06:02, A1*01:02+B1*06:04, A1*01:03+B1*06:03, A1*02:01+B1*02:02, A1*02:01+B1*03:03, A1*03:01+B1*03:02, A1*03:03+B1*03:01, A1*05:01+B1*02:01 and A1*05:05+B1*03:01.

[0071]

[0105] In some embodiments, the peptides presented by the HLA protein comprise peptides identified by comparing the MS / MS spectrum of the HLA peptide with the MS / MS spectrum of one or more peptides or proteins in a peptide or protein database.

[0072]

[0106] In some embodiments, the mutation is selected from the group consisting of point mutations, splice site mutations, frameshift mutations, read-through mutations, and gene fusion mutations.

[0107] In some embodiments, the peptides presented by the HLA protein have a length of 15 to 40 amino acids.

[0073]

[0108] In some embodiments, the peptides presented by the HLA protein comprise peptides identified by identifying the peptides presented by the HLA protein by comparing the MS / MS spectrum of the HLA peptide with the MS / MS spectrum of one or more peptides or proteins in a peptide or protein database.

[0074]

[0109] In some embodiments, the personalized cancer treatment further comprises an adjuvant.

[0110] In some embodiments, the personalized cancer treatment further comprises an immune checkpoint inhibitor.

[0075]

[0111] In some embodiments, the training data includes structured data, time series data, unstructured data, relational data, or any combination thereof.

[0112] In some embodiments, the unstructured data includes image data.

[0076]

[0113] In some embodiments, the relational data includes data derived from a customer system, an enterprise system, an operations system, a website, a web-accessible application programming interface (API), or any combination thereof.

[0077]

[0114] In some embodiments, the training data is uploaded to a cloud-based database.

[0115] In some embodiments, the training is performed using a convolutional neural network.

[0078]

[0116] In some embodiments, the convolutional neural network includes at least two convolutional layers.

[0117] In some embodiments, the convolutional neural network includes at least one batch normalization step.

[0079]

[0118] In some embodiments, the convolutional neural network includes at least one spatial dropout step.

[0119] In some embodiments, the convolutional neural network includes at least one global max pooling step.

[0080]

[0120] In some embodiments, the convolutional neural network includes at least one dense connection layer.

[0121] In some embodiments, the step of identifying a peptide sequence includes the step of identifying a peptide sequence having a mutation expressed in a target cancer cell.

[0081]

[0122] In some embodiments, the step of identifying a peptide sequence includes the step of identifying a peptide sequence that is not expressed in normal cells of the subject.

[0123] In some embodiments, the step of identifying a peptide sequence includes the step of identifying a viral peptide sequence.

[0082]

[0124] In some embodiments, the step of identifying a peptide sequence includes the step of identifying an overexpressed peptide sequence.

[0125] Provided herein is a method for identifying an HLA class II-specific peptide for immunotherapy of a subject, the method comprising: obtaining, by a computer processor, a candidate peptide comprising an epitope and a plurality of peptide sequences each comprising an epitope; processing, by the computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences to immune cells, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by an HLA class II allele can present a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; selecting, from one or more proteins encoded by the HLA class II alleles of the subject's cells, a protein predicted by the machine learning HLA peptide presentation prediction model to bind to the candidate peptide, the protein having a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; contacting the candidate peptide with the selected protein such that the candidate peptide competes with a placeholder peptide that associates with the selected protein; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder.

[0083]

[0126] In some embodiments, the obtaining step includes the step of identifying a candidate peptide, where the step of identifying a candidate peptide includes the step of comparing a DNA, RNA, or protein sequence from a target cancer cell with a DNA, RNA, or protein sequence from a target normal cell.

[0084]

[0127] In some embodiments, the processing step includes the step of identifying a plurality of predictor variables based on at least amino acid information of a plurality of peptide sequences, and the step of processing the plurality of predictor variables using a machine learning HLA peptide presentation prediction model.

[0085]

[0128] In some embodiments, the machine learning HLA peptide presentation prediction model includes a plurality of predictor variables identified based at least on training data, where the training data is: training peptide sequence information including amino acid position information, which is associated with an HLA protein expressed in a cell; and a function representing the relationship between the amino acid position information and the presentability generated as an output based on the amino acid position information and the plurality of predictor variables.

[0086]

[0129] In some embodiments, the number of positives is constrained to be equal to the number of hits.

[0130] In some embodiments, the mass spectrometry is single allele mass spectrometry.

[0131] In some embodiments, the plurality of predictor variables include: an expression level predictor of a source protein containing a peptide, a stability predictor, a degradation rate predictor, a cleavability predictor, a cell or tissue localization predictor, and any one or more of intracellular processing modes including autophagy, phagocytosis, and intracellular transport predictors.

[0087]

[0132] In some embodiments, the quality of the training data is improved by using a plurality of quality measurement criteria.

[0133] In some embodiments, the plurality of quality measurement criteria include common contaminant peptide removal, high scored peak intensity, high score, and high mass accuracy.

[0088]

[0134] In some embodiments, the scored peak intensity is at least 50%.

[0135] In some embodiments, the scored peak intensity is at least 60%.

[0136] In some embodiments, the placeholder peptide is a CLIP peptide.

[0089]

[0137] In some embodiments, the placeholder peptide is a CMV peptide.

[0138] In some embodiments, the method further includes measuring the IC50 of the replacement of the placeholder peptide by the target peptide.

[0090]

[0139] In some embodiments, the IC50 of the replacement of the placeholder peptide by the target peptide is less than 500 nM.

[0140] In some embodiments, at least one protein derived from one or more proteins encoded by the HLA class II alleles of the target cells is an HLA class II tetramer or multimer.

[0091]

[0141] In some embodiments, the target peptide is further identified by mass spectrometry.

[0142] In some embodiments, at least one protein encoded by the HLA class II alleles of the target cells is a recombinant protein.

[0092]

[0143] In some embodiments, at least one protein encoded by the HLA class II alleles of the target cells is expressed in eukaryotic cells.

[0144] In some embodiments, the peptide is presented by HLA proteins expressed in cells through autophagy.

[0093]

[0145] In some embodiments, the peptide is presented by an HLA protein expressed in the cell through phagocytosis.

[0146] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single immunoprecipitated HLA protein expressed in the cell.

[0094]

[0147] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single exogenous HLA protein expressed in the cell.

[0095]

[0148] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single recombinant HLA protein expressed in the cell.

[0096]

[0149] In some embodiments, the plurality of predictor variables includes peptide-HLA affinity predictor variables.

[0150] In some embodiments, the peptides presented by the HLA protein include peptides identified by searching a non-enzymatic specificity peptide database without modifications.

[0097]

[0151] In some embodiments, the peptides presented by the HLA protein include peptides identified by searching a peptide database using an inverse database search strategy.

[0098]

[0152] In some embodiments, the HLA protein includes an HLA-DR, HLA-DQ, or HLA-DP protein.

[0153] In some embodiments, the immunotherapy is cancer immunotherapy.

[0099]

[0154] In some embodiments, the epitope is a cancer-specific epitope.

[0155] In some embodiments, at least one protein encoded by an HLA class II allele comprises at least the alpha 1 subunit and the beta 1 subunit of the HLA protein, which exist in a dimeric form.

[0100]

[0156] In some embodiments, the peptide identity is well-known.

[0157] In some embodiments, the peptide identity is unknown.

[0158] In some embodiments, the peptide identity is determined by mass spectrometry.

[0101]

[0159] In some embodiments, the peptide exchange assay involves detection of a peptide fluorescent probe or tag.

[0160] In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide has the amino acid sequence of PVSKMRMATPLLMQA (SEQ ID NO: 1).

[0102]

[0161] In some embodiments, the polynucleic acid construct comprises an expression vector, further comprising one or more of a promoter, a secretion signal, a dimerization factor, a ribosome skipping sequence, and one or more of one or more tags for purification and / or detection.

[0103]

[0162] In some embodiments, the placeholder peptide sequence is encoded by a nucleic acid sequence within the vector.

[0163] In some embodiments, the sequence encoding the cleavable domain is positioned between the sequences encoding the placeholder peptide and the HLA beta 1 peptide.

[0104]

[0164] Provided herein is a method for assaying the immunogenicity of an MHC class II-binding peptide, the method comprising: selecting, by a machine learning HLA peptide presentation prediction model, a protein encoded by an HLA class II allele predicted to bind to the MHC class II-binding peptide, wherein the machine learning HLA peptide presentation prediction model is configured to generate a presentation prediction for a given peptide sequence, the presentation prediction indicating the likelihood that one or more proteins encoded by an HLA class II allele can present the given peptide sequence, and the protein has a probability higher than a presentation prediction probability value threshold for presenting the MHC class II-binding peptide; contacting the peptide with the selected protein so as to compete with a placeholder peptide that associates with the selected protein, thereby replacing the placeholder peptide and forming a complex comprising an HLA class II protein and the MHC class II-binding peptide; and assaying one or more activation parameters of CD4+ T cells selected from the group consisting of induction of cytokines, induction of chemokines, and expression of cell surface markers, by contacting the complex with the CD4+ T cells.

[0105]

[0165] In some embodiments, the HLA class II allele is a tetramer or multimer.

[0166] In some embodiments, the cytokine is IL-2.

[0167] Provided herein is a method for inducing CD4+ T cell activation in a subject for cancer immunotherapy, the method comprising: identifying a peptide sequence associated with cancer and comprising a cancer mutation, the step of identifying the peptide sequence comprising comparing a DNA, RNA or protein sequence from the subject's cancer cells to a DNA, RNA or protein sequence from the subject's normal cells; selecting a protein encoded by an HLA class II allele that is normally expressed by the subject's cells and is predicted by a machine learning HLA peptide presentation prediction model to bind to the peptide, the prediction model having a positive predictive value of at least 0.1 with a reproducibility rate of at least 0.1%, 0.1% - 50% or up to 50%, the protein having a probability higher than a presentation prediction probability value threshold for presenting the identified peptide sequence; contacting the identified peptide with the selected protein encoded by the HLA class II allele to verify whether the identified peptide competes with a placeholder peptide that associates with the selected protein encoded by the HLA class II allele and has an IC50 value of less than 500 nM to replace the placeholder peptide; optionally, purifying the identified peptide; and administering to the subject an effective amount of a polypeptide comprising the sequence of the identified peptide or a polynucleotide encoding the polypeptide.

[0106]

[0168] This specification provides a method for screening a drug containing a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by a computer processor, the amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class I or II MHC alleles of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data containing sequence information related to HLA proteins expressed in cells; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic in the subject; and administering to the subject a composition containing the drug.

[0107]

[0169] This specification provides a method for producing HLA class II tetramers or multimers by conjugation of four individual HLA protein alpha1 and beta1 heterodimers, the method comprising: expressing, in a eukaryotic cell, a vector containing nucleic acid sequences encoding an alpha chain and a beta chain of an HLA protein, a secretion signal, a biotinylation motif, and at least one tag for identification or purification, such that each HLA protein alpha1 and beta1 heterodimer is secreted in a dimeric state, wherein the heterodimer is associated with a placeholder peptide; purifying the secreted heterodimer from the cell culture medium; verifying the peptide binding activity using a peptide exchange assay; adding streptavidin, thereby conjugating the heterodimer to a tetramer; and purifying the tetramer to have a yield greater than 1 mg / L. Multimers, such as pentamers, hexamers, or octamers, are equally contemplated herein and can be generated similarly.

[0108]

[0170] In some embodiments, the vector comprises a CMV promoter.

[0171] In some embodiments, the vector comprises an array encoding a placeholder peptide that is linked to the beta 1 chain via a cleavable site.

[0109]

[0172] In some embodiments, the peptide exchange assay includes pre-cleavage of the placeholder peptide from the beta chain.

[0173] In some embodiments, the cleavable site is a thrombin cleavage site.

[0110]

[0174] In some embodiments, the peptide exchange assay is a FRET assay.

[0175] In some embodiments, purification is by any one of: column chromatography, ion exchange chromatography, size exclusion chromatography, affinity chromatography, or LC-MS.

[0111]

[0176] Provided herein are HLA class II tetramers or multimers in which each heterodimer comprises an alpha and a beta chain of HLA-DR or HLA-DP or HLA-DQ heterodimer, where the heterodimer is purified and present at a concentration higher than 1 mg / L.

[0112]

[0177] In some embodiments, the HLA class II tetramer is selected from Tables 8A - 8C.

[0178] In some embodiments, the HLA class II tetramer comprises a heterodimer pair selected from the group consisting of HLA-DR, HLA-DP, and HLA-DQ proteins.

[0113]

[0179] In some embodiments, the HLA protein is an HLA class II protein selected from the group consisting of: HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, and HLA-DRB5*01:01.

[0114]

[0180] In some embodiments, the heterodimer pair is expressed in eukaryotic cells.

[0181] In some embodiments, the heterodimer pair is encoded by a vector.

[0182] Provided herein is a vector, where the vector encodes a nucleic acid sequence for the alpha and beta chains of the HLA proteins described herein, a secretion signal, a biotinylation motif and at least one tag for identification or for purification, such that each HLA protein alpha1 and beta1 heterodimer is secreted in a dimeric state, where the secreted heterodimer optionally associates with a placeholder peptide.

[0115]

[0183] Provided herein are cells comprising the vectors described herein.

[0184] In some embodiments, the HLA class II heterodimer is secreted from eukaryotic cells into the cell culture medium, which is further purified by any one of: column chromatography, ion exchange chromatography, size exclusion chromatography, affinity chromatography or LC-MS.

[0116]

[0185] The present specification provides a method for screening a drug containing a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by the computer processor, amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by a class I or II MHC allele of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence will be immunogenic in the subject.

[0117]

[0186] The present specification provides a method for screening a drug containing a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor to input amino acid information of a peptide sequence of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a series of presentation predictions for the peptide sequence, each presentation prediction representing the probability that one or more proteins encoded by class I or II MHC alleles of the subject's cells present an epitope sequence of a given peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictive variables identified based at least on training data, the training data being: training peptide sequence information including sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, amino acid position information, the training peptide sequence information being related to HLA proteins expressed in cells, and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictive variables; determining or predicting, based on the set of presentation predictions, that each of the peptide sequences of the polypeptide sequence will not be immunogenic in the subject; and administering to the subject a composition comprising the drug.

[0118]

[0187] The present specification provides a method for screening a drug containing a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor to input amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a series of presentation predictions for the peptide sequence, each presentation prediction representing the probability that one or more proteins encoded by class I or II MHC alleles of the subject's cells present an epitope sequence of a given peptide sequence, and the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, wherein the training data is training peptide sequence information comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, amino acid position information, the training peptide sequence information being associated with the HLA proteins expressed in the cells, and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictor variables; and determining or predicting, based on the set of presentation predictions, that at least one of the peptide sequences of the polypeptide sequence will be immunogenic in the subject.

[0119]

[0188] This specification provides a method for screening a drug containing a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining, by a computer processor, a plurality of peptide sequences of the polypeptide sequence; processing, by the computer processor, the amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by a class I or II MHC allele of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data containing sequence information related to HLA proteins expressed in cells; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic in the subject; and administering to the subject a composition comprising the drug.

[0120]

[0189] In some embodiments, the method further comprises determining not to administer the drug to the subject.

[0190] In some embodiments, the drug comprises an antibody or a binding fragment thereof.

[0121]

[0191] In some embodiments, the peptide sequences of the polypeptide sequence have a length of 8, 9, 10, 11 or 12 amino acids, wherein the protein encoded by a class I or II MHC allele of the subject's cells is the protein encoded by a class I MHC allele of the subject's cells.

[0122]

[0192] In some embodiments, the peptide sequences of the polypeptide sequence have a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids, wherein the protein encoded by a class I or II MHC allele of the subject's cells is the protein encoded by a class II MHC allele of the subject's cells.

[0123]

[0193] Disclosed herein is a method of treating a subject having an autoimmune disease or condition, the method comprising: (a) identifying or predicting an epitope of an expressed protein presented by class I or II MHC of the subject's cells, wherein the identified or predicted epitope and the complex comprising class I or II MHC is targeted by the subject's CD8 or CD4 T cells; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells from the subject or allogeneic regulatory T cells; and (d) administering to the subject the regulatory T cells expressing the TCR.

[0124]

[0194] In some embodiments, the autoimmune disease or condition is diabetes.

[0195] In some embodiments, the cells are islet cells.

[0196] Disclosed herein is a method of treating a subject having an autoimmune disease or condition, the method comprising administering to the subject regulatory T cells expressing: (i) an epitope of an expressed protein identified or predicted to be presented by class I or II MHC of the subject's cells, and (ii) a T cell receptor (TCR) that binds to a complex comprising class I or II MHC, wherein the complex is targeted by the subject's CD8 or CD4 T cells.

[0125]

[0197] The present specification provides a computer system for identifying peptide sequences for individualized cancer treatment of a subject, which includes: a database configured to store a plurality of peptide sequences of the subject; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually or collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class II MHC alleles of the subject's cells can present a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and selecting a subset of the plurality of peptide sequences for individualized cancer treatment of the subject based on at least the plurality of presentation predictions.

[0126]

[0198] The present specification provides a computer system for identifying HLA class II-specific peptides for immunotherapy for a subject, which includes: a database configured to store candidate peptides containing epitopes and a plurality of peptide sequences each containing an epitope; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences to immune cells, each presentation prediction indicating the likelihood that one or more proteins encoded by an HLA class II allele can present a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; selecting one protein from one or more proteins encoded by the HLA class II allele of the subject's cells and predicted to bind to the candidate peptide by the machine learning HLA peptide presentation prediction model, wherein the protein has a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide is contacted with the selected protein such that the candidate peptide competes with the placeholder peptide associated with the selected protein.

[0127]

[0199] The present specification provides a computer system for screening a drug containing a polypeptide sequence for immunogenicity in a subject, which includes: a database configured to store a plurality of peptide sequences of the polypeptide sequence; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by a class I or II MHC allele of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data including sequence information related to HLA proteins expressed in cells; based on a set of the plurality of presentation predictions, determining or predicting that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic in the subject, wherein a composition containing the drug is administered to the subject.

[0128]

[0200] This specification provides a computer system for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the system comprising: a database configured to store a plurality of peptide sequences of the polypeptide sequence; and one or more computer processors operably coupled to the database, wherein the one or more computer processors are individually and collectively programmed to process amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class I or II MHC alleles of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; based on the plurality of presentation predictions, determining or predicting that at least one of the plurality of peptide sequences of the polypeptide sequence will be immunogenic in the subject.

[0129]

[0201] A non - transitory computer - readable medium is provided that includes machine - executable code for performing, by one or more computer processors, a method for identifying peptide sequences for individualized cancer treatment of a subject, the method comprising: obtaining a plurality of peptide sequences of the subject; processing amino acid information of the plurality of peptide sequences using a machine - learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, wherein each presentation prediction indicates the likelihood that one or more proteins encoded by class II MHC alleles of the subject's cells can present a given peptide sequence of the plurality of peptide sequences, and the machine - learning HLA peptide presentation prediction model is trained using training data that includes sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and selecting a subset of the plurality of peptide sequences for individualized cancer treatment of the subject based on at least the plurality of presentation predictions.

[0130]

[0202] Provided is a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements a method for identifying HLA class II-specific peptides for immunotherapy of a subject, the method comprising: obtaining a candidate peptide comprising an epitope and a plurality of peptide sequences each comprising an epitope; processing amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences for presentation to immune cells, each presentation prediction indicating the likelihood that one or more proteins encoded by an HLA class II allele can present a given peptide sequence of the plurality of peptide sequences, the machine learning HLA peptide presentation prediction model being trained using training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; selecting a protein from one or more proteins encoded by the HLA class II allele of the subject's cells and predicted to bind to the candidate peptide by the machine learning HLA peptide presentation prediction model, the protein having a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein based on whether the candidate peptide replaces a placeholder peptide when contacting the candidate peptide with the selected protein selected such that the candidate peptide competes with the placeholder peptide.

[0131]

[0203] Provided herein is a non-transitory computer-readable medium comprising machine-executable code for performing, by one or more computer processors, a method of screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class I or II MHC alleles of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, the machine learning HLA peptide presentation prediction model being trained using training data comprising sequence information related to HLA proteins expressed in cells; and determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence will not be immunogenic in the subject, wherein the composition comprising the drug is administered to the subject.

[0132]

[0204] Provided herein is a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements a method of screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine learning HLA peptide presentation prediction model to generate a presentation prediction for each of the plurality of peptide sequences, each presentation prediction indicating the likelihood that one or more proteins encoded by class I or II MHC alleles of the subject's cells can present an epitope sequence of a given peptide sequence of the plurality of peptide sequences, wherein the machine learning HLA peptide presentation prediction model is trained using training data comprising sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence will be immunogenic in the subject.

[0133]

[0205] In this specification: a step of processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each of the plurality of candidate peptide sequences is encoded by a genome or exome of a subject, the plurality of presentation predictions includes HLA presentation predictions for each of the plurality of candidate peptide sequences, each presentation prediction indicates the possibility that one or more proteins encoded by class II HLA alleles of the subject's cells can present a plurality of given candidate peptide sequences, and the machine learning HLA peptide presentation prediction model is trained using training data including sequence information of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in training cells; and a step of identifying peptide sequences of a plurality of peptide sequences having a probability higher than a presentation prediction probability value threshold presented by at least one of one or more proteins encoded by class II HLA alleles of the subject's cells, based at least on the plurality of presentation predictions, wherein the machine learning HLA peptide presentation prediction model has a positive predictive value (PPV) of at least 0.07 when amino acid information of a plurality of test peptide sequences is processed to generate a plurality of test presentation predictions, each test presentation prediction indicates the possibility that one or more proteins encoded by class II HLA alleles of the subject's cells can present a given test peptide sequence of the plurality of test peptide sequences, the plurality of test peptide sequences includes at least 500 test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell and (ii) at least 499 decoy peptide sequences contained within a protein encoded by the genome of an organism, the organism and the subject are of the same species, the plurality of test peptide sequences includes a ratio of 1:499 of at least one hit peptide sequence to at least 499 decoy peptide sequences, and 0.2% of the plurality of test peptide sequences are predicted by the machine learning HLA peptide presentation prediction model to be presented by an HLA protein expressed in a cell, a method is provided.

[0134]

[0206] In this specification: a step of processing amino acid information of a plurality of candidate peptide sequences encoded by a genome or exome of a subject using a machine learning HLA peptide binding prediction model to generate a plurality of binding predictions, wherein the plurality of binding predictions includes HLA binding predictions for each of the plurality of candidate peptide sequences, and each binding prediction indicates the likelihood that one or more proteins encoded by class II HLA alleles of the subject's cells bind to a given candidate peptide sequence of the plurality of candidate peptide sequences, and the machine learning HLA peptide binding prediction model is trained using training data including sequence information of peptides identified as binding to HLA class II proteins or HLA class II protein analogs; and identifying peptide sequences of a plurality of peptide sequences having a probability higher than a binding prediction probability value threshold for binding to at least one of one or more proteins encoded by class II HLA alleles of the subject's cells, based at least on the plurality of binding predictions, wherein when amino acid information of the plurality of test peptide sequences is processed to generate a plurality of test binding predictions, the machine learning HLA peptide binding prediction model has a positive predictive value (PPV) of at least 0.1, and each test binding prediction indicates the likelihood that one or more proteins encoded by class II HLA alleles of the subject's cells bind to a given test peptide sequence of the plurality of test peptide sequences, the plurality of test peptide sequences includes at least 50 test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell and (ii) at least 19 decoy peptide sequences contained within a protein including the peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, the organism and the subject are of the same species, the plurality of test peptide sequences includes a 1:19 ratio of at least 19 decoy peptide sequences to at least one hit peptide sequence, and 5% of the plurality of test peptide sequences are predicted by the machine learning HLA peptide presentation prediction model to bind to an HLA protein expressed in the cell, a method is provided.

[0135]

[0207] In some embodiments, the machine learning HLA peptide presentation prediction model is trained using training data that includes sequence information of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in the training cells.

[0136]

[0208] In some embodiments, one or more of 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model have a probability higher than the presentation prediction probability value threshold presented by at least one of one or more proteins encoded by the class II HLA alleles of the subject cells.

[0137]

[0209] In some embodiments, each of 0.2% of the plurality of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model has a probability higher than the presentation prediction probability value threshold presented by at least one of one or more proteins encoded by the class II HLA alleles of the subject cells.

[0138]

[0210] In some embodiments, the PPV is greater than each of the PPVs in column 2 of Table 11 for the proteins encoded by the corresponding HLA alleles in Table 13. In some embodiments, the PPV is at least equal to each of the PPVs in column 3 of Table 11 for the proteins encoded by the corresponding HLA alleles in Table 11.

[0139]

[0211] In some embodiments, the PPV is greater than each of the PPVs in column 2 of Table 12 for the proteins encoded by the HLA class II alleles.

[0212] In some embodiments, the PPV is at least equal to each of the PPVs in column 2 of Table 16 for the proteins encoded by the corresponding HLA alleles in Table 16.

[0140]

[0213] The present specification provides a method for preparing an individualized cancer treatment, the method comprising: identifying a peptide sequence, the peptide sequence being associated with cancer, the identifying step comprising comparing a DNA, RNA, or protein sequence from a cancer cell of a subject with a DNA, RNA, or protein sequence from a normal cell of the subject; using a computer processor to input amino acid position information of the identified peptide sequence into a machine learning HLA peptide presentation prediction model to generate a series of presentation predictions for the identified peptide sequence, each presentation prediction indicating the probability that one or more proteins encoded by the HLA class II alleles of the subject's cells present a given sequence of the identified peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictive variables identified based at least on training data, the training data comprising: sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, training peptide sequence information comprising amino acid position information, the training peptide sequence information being associated with HLA proteins expressed in cells, and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictive variables; and selecting a subset of the peptide sequences identified based on the series of presentation predictions for preparing an individualized cancer treatment, the prediction model having a positive predictive value of at least 0.1%, from 0.1% to 50%, or up to 50% with a recall rate of at least 0.1.

[0141]

[0214] This specification provides a method including the step of training a machine learning HLA peptide presentation prediction model, and the training step includes inputting, using a computer processor, an amino acid position information sequence of an HLA peptide isolated from one or more HLA peptide complexes derived from cells expressing an HLA class II allele into the HLA peptide presentation prediction model. The machine learning HLA peptide presentation prediction model includes: sequence information of the sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information of training peptides, which is related to the HLA proteins expressed in cells; and a plurality of prediction variables identified based at least on training data including a function representing the relationship between the input amino acid position information and the presentability generated as an output based on the amino acid position information and the prediction variables.

[0142]

[0215] In some embodiments, the presentation model has a positive predictive value of at least 0.25 at a recall rate of at least 0.1%, 0.1% - 50% or up to 50%.

[0216] In some embodiments, the presentation model has a positive predictive value of at least 0.4 at a recall rate of at least 0.1%, 0.1% - 50% or up to 50%.

[0143]

[0217] In some embodiments, the presentation model has a positive predictive value of at least 0.6 at a recall rate of at least 0.1%, 0.1% - 50% or up to 50%.

[0218] In some embodiments, the mass spectrometry is single - allele mass spectrometry.

[0144]

[0219] In some embodiments, the peptides are presented by HLA proteins expressed in cells through autophagy.

[0220] In some embodiments, the peptides are presented by HLA proteins expressed in cells through phagocytosis.

[0145]

[0221] In some embodiments, the quality of the training data is improved by using multiple quality metrics.

[0222] In some embodiments, the multiple quality metrics include common contaminant peptide removal, high scored peak intensity, high score, and high mass accuracy.

[0146]

[0223] In some embodiments, the scored peak intensity is at least 50%.

[0224] In some embodiments, the scored peak intensity is at least 60%.

[0225] In some embodiments, the score is at least 7.

[0147]

[0226] In some embodiments, the mass accuracy is at most 5 ppm.

[0227] In some embodiments, the mass accuracy is at most 2 ppm.

[0228] In some embodiments, the backbone cleavage score is at least 5.

[0148]

[0229] In some embodiments, the backbone cleavage score is at least 8.

[0230] In some embodiments, the peptides presented by HLA proteins expressed in cells are peptides presented by a single immunoprecipitated HLA protein expressed in the cells.

[0149]

[0231] In some embodiments, the peptides presented by HLA proteins expressed in cells are peptides presented by a single exogenous HLA protein expressed in the cells.

[0150]

[0232] In some embodiments, the peptides presented by HLA proteins expressed in cells are peptides presented by a single recombinant HLA protein expressed in the cells.

[0151]

[0233] In some embodiments, the plurality of predictive variables includes a peptide-HLA affinity predictive variable.

[0234] In some embodiments, the plurality of predictive variables includes a source protein expression level predictive variable.

[0152]

[0235] In some embodiments, the plurality of predictive variables includes a peptide cleavability predictive variable.

[0236] In some embodiments, the training peptide sequence information includes sequences derived from peptides presented by HLA proteins, which includes peptides identified by searching a non-enzymatic specificity peptide database without modifications. In some embodiments, the peptides presented by HLA proteins include peptides identified by searching a novel peptide sequencing tool.

[0153]

[0237] In some embodiments, the peptides presented by HLA proteins include peptides identified by searching a peptide database using an inverse database search strategy.

[0154]

[0238] In some embodiments, the HLA protein comprises HLA-DR and HLA-DP or HLA-DQ proteins. In some embodiments, the HLA protein comprises an HLA-DR protein selected from the group consisting of HLA-DR and HLA-DP or HLA-DQ proteins.In some embodiments, the HLA protein comprises an HLA-DR protein selected from the group consisting of: HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, and HLA-DRB5*01:01.

[0155]

[0239] In some embodiments, the peptides presented by the HLA protein include peptides identified by comparing the MS / MS spectrum of the HLA peptide with the MS / MS spectra of one or more HLA peptides in a peptide database.

[0156]

[0240] In some embodiments, the mutation is selected from the group consisting of point mutations, splice site mutations, frameshift mutations, read-through mutations, and gene fusion mutations.

[0241] In some embodiments, the peptides presented by the HLA protein have a length of 15 to 40 amino acids.

[0157]

[0242] In some embodiments, the peptides presented by the HLA protein include peptides identified by: (a) isolating one or more HLA complexes from a cell line expressing a single HLA class II allele; (b) isolating one or more HLA peptides from the one or more isolated HLA complexes; (c) obtaining an MS / MS spectrum for the one or more isolated HLA peptides; and (d) obtaining from a peptide database a peptide sequence corresponding to the MS / MS spectrum of the one or more isolated HLA peptides, wherein the one or more sequences obtained from step (d) identify the sequences of the one or more isolated HLA peptides.

[0158]

[0243] In some embodiments, the personalized cancer treatment further includes an adjuvant.

[0244] In some embodiments, the personalized cancer treatment further includes an immune checkpoint inhibitor.

[0159]

[0245] In some embodiments, the training data includes structured data, time series data, unstructured data, relational data, or any combination thereof.

[0246] In some embodiments, the unstructured data includes image data.

[0160]

[0247] In some embodiments, the relational data includes data from a customer system, an enterprise system, an operations system, a website, an application programming interface (API) accessible from the web, or any combination thereof.

[0161]

[0248] In some embodiments, the training data is uploaded to a cloud-based database.

[0249] In some embodiments, the training is performed using a convolutional neural network.

[0162]

[0250] In some embodiments, the convolutional neural network includes at least two convolutional layers.

[0251] In some embodiments, the convolutional neural network (CNN) includes at least one batch normalization step.

[0163]

[0252] In some embodiments, the convolutional neural network includes at least one spatial dropout step.

[0253] In some embodiments, the convolutional neural network includes at least one global max pooling step.

[0164]

[0254] In some embodiments, the convolutional neural network includes at least one dense connection layer.

[0255] In some embodiments, the step of identifying a peptide sequence includes the step of identifying a peptide sequence having a mutation expressed in a target cancer cell.

[0165]

[0256] In some embodiments, the step of identifying a peptide sequence includes the step of identifying a peptide sequence not expressed in a target normal cell.

[0257] In some embodiments, the step of identifying a peptide sequence includes the step of identifying an overexpressed peptide sequence.

[0166]

[0258] In some embodiments, the step of identifying the peptide sequence includes the step of identifying a viral peptide sequence. In one aspect, a method for identifying HLA class II-specific peptides for specific immunotherapy of a subject is provided, the method comprising: identifying candidate peptides comprising an epitope; using a computer processor to input the amino acid information of a plurality of peptide sequences each comprising an epitope into a machine learning HLA peptide presentation prediction model to generate a series of HLA presentation predictions for the peptide sequences to immune cells, each presentation prediction indicating the probability that one or more proteins encoded by the HLA class II alleles of the subject's cells present a given peptide sequence comprising the epitope, the prediction model having a positive predictive value of at least 0.1%, 0.1% - 50% or at least 0.1 with a maximum of 50% in terms of recall; selecting a protein predicted by the prediction model to bind to the candidate peptide from one or more proteins encoded by the HLA class II alleles of the subject's cells, the protein having a probability higher than a presentation prediction probability value threshold for presenting the candidate peptide to immune cells; contacting the candidate peptide with a protein encoded by the HLA class II allele such that the candidate peptide competes with a placeholder peptide that associates with the protein encoded by the HLA class II allele; and identifying the candidate peptide as a peptide for specific immunotherapy for a protein encoded by the HLA class II allele based on whether the candidate peptide replaces the placeholder peptide.

[0167]

[0259] In some embodiments, the immunotherapy is cancer immunotherapy.

[0260] In some embodiments, the step of identifying includes comparing the DNA, RNA or protein sequence derived from the subject's cancer cells with the DNA, RNA or protein sequence derived from the subject's normal cells. In some embodiments, the epitope is a cancer-specific epitope.

[0168]

[0261] In some embodiments, at least one protein encoded by an HLA class II allele comprises at least the alpha1 subunit and the beta1 subunit of the HLA protein or fragments thereof and exists in a dimeric form. In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide is a CMV peptide. In some embodiments, the method further comprises measuring the IC50 of the replacement of the placeholder peptide by the target peptide. In some embodiments, the IC50 of the replacement of the placeholder peptide by the target peptide is less than 500 nM. In some embodiments, at least one protein derived from one or more proteins encoded by an HLA class II allele of a subject's cells is an HLA class II tetramer or multimer. In some embodiments, the target peptide is further identified by mass spectrometry. In some embodiments, at least one protein encoded by an HLA class II allele of a subject's cells is a recombinant protein. In some embodiments, at least one protein encoded by an HLA class II allele of a subject's cells is expressed in eukaryotic cells.

[0169]

[0262] In one aspect, an assay method is provided for verifying the specificity of a candidate peptide for binding to an HLA class II protein, the method comprising: expressing in a eukaryotic cell a polynucleotide construct comprising a nucleic acid sequence encoding an HLA class II protein comprising an alpha chain and a beta chain or a fragment thereof, which can bind to a peptide comprising an MHC-II binding epitope, wherein the expressed HLA class II protein or fragment thereof remains associated with a placeholder peptide; isolating the HLA class II protein or a portion thereof expressed in the eukaryotic cell; (a) increasing the amount of the candidate peptide and adding it to determine whether the candidate peptide replaces the placeholder peptide that associates with the HLA class II protein or a portion thereof; and (b) performing a peptide exchange assay by calculating the IC50 of the substitution reaction to determine the affinity of the candidate peptide for the HLA class II protein or a portion thereof relative to the placeholder peptide, thereby verifying the specificity of the candidate peptide for binding to the HLA class II protein.

[0170]

[0263] In some embodiments, the peptide identity is well-known. In some embodiments, the peptide identity is unknown. In some embodiments, the peptide identity is determined by mass spectrometry.

[0171]

[0264] In some embodiments, the peptide exchange assay comprises detection of a peptide fluorescent probe or tag. In some embodiments, the placeholder peptide is a CLIP peptide.

[0172]

[0265] In some embodiments, the polynucleotide construct further comprises an expression vector comprising one or more of a promoter, a linker, one or more protease cleavage sites, a secretion signal, a dimerization factor, a ribosome skipping sequence, and one or more tags for purification and / or detection.

[0173]

[0266] In one aspect, provided herein is a method for assaying the immunogenicity of an MHC class II-binding peptide, the method comprising: selecting a protein encoded by an HLA class II allele predicted by a machine learning HLA peptide presentation prediction model that binds to the peptide, wherein the prediction model has a positive predictive value of at least 0.1%, from 0.1% to 50% or at most 50% with a recall rate of at least 0.1, and the protein has a probability higher than a presentation prediction probability value threshold for presenting the identified peptide sequence; contacting the peptide with the selected protein encoded by the HLA class II allele such that the peptide competes with a placeholder peptide that associates with the selected protein encoded by the HLA class II allele; replacing the placeholder peptide, thereby forming a complex comprising the HLA class II protein and the identified peptide; and contacting the HLA class II protein and the identified peptide complex with CD4+ T cells and assaying one or more activation parameters of the CD4+ T cells selected from induction of cytokines, induction of chemokines, and expression of cell surface markers.

[0174]

[0267] In some embodiments, the HLA class II allele is a tetramer or multimer. In some embodiments, the cytokine is IL-2. In some embodiments, the cytokine is IFN-gamma.

[0175]

[0268] In one aspect, provided herein is a method for inducing CD4+ T cell activation in a subject for cancer immunotherapy, the method comprising: identifying a peptide sequence associated with cancer and comprising a cancer mutation, the identifying step comprising comparing a DNA, RNA or protein sequence from the subject's cancer cells to a DNA, RNA or protein sequence from the subject's normal cells; selecting a protein encoded by an HLA class II allele that is normally expressed by the subject's cells and is predicted by a machine learning HLA peptide presentation prediction model to bind to the peptide, the prediction model having a positive predictive value of at least 0.1 with a reproducibility of at least 0.1%, 0.1% - 50% or up to 50%, and the protein having a probability higher than a presentation prediction probability value threshold for presenting the identified peptide sequence; contacting the identified peptide with the selected protein encoded by the HLA class II allele to verify whether the identified peptide competes with a placeholder peptide that associates with the selected protein encoded by the HLA class II allele and competes with an IC50 value of less than 500 nM to replace the placeholder peptide; purifying the identified peptide; and administering an effective amount of the identified peptide to the subject.

[0176]

[0269] In one aspect, provided herein is a method for producing an HLA class II tetramer or multimer, the method comprising: expressing in a eukaryotic cell a vector comprising a nucleic acid sequence encoding an alpha chain and a beta chain of an HLA protein, a linker, one or more protease cleavage sites, a secretion signal, a biotinylation motif, and at least one tag for identification or purification, such that each HLA protein alpha1 and beta1 heterodimer is secreted in a dimeric state, the heterodimer associating with a placeholder peptide; purifying the secreted heterodimer from the cell culture medium; verifying peptide binding activity using a peptide exchange assay; adding streptavidin, thereby conjugating the heterodimer to a tetramer; and purifying the tetramer to have a yield greater than 1 mg / L.

[0177]

[0270] In some embodiments, the vector comprises a CMV promoter. In some embodiments, the vector comprises a sequence encoding a placeholder peptide linked to the beta1 chain via a cleavable site. In some embodiments, the peptide exchange assay involves pre-cleavage of the placeholder peptide from the beta chain. In some embodiments, the cleavable site is a thrombin cleavage site. In some embodiments, the peptide exchange assay is a FRET assay. In some embodiments, the purification is by any one of column chromatography, batch chromatography, ion exchange chromatography, size exclusion chromatography, affinity chromatography, or LC-MS.

[0178]

[0271] In one aspect, the present specification provides a composition comprising an HLA class II tetramer comprising any one of HLA-DR or HLA-DP or HLA-DQ heterodimers, each heterodimer comprising an alpha and a beta chain, purified and present at a concentration higher than 0.25 mg / L. In some embodiments, the HLA class II tetramer comprises: a heterodimer pair selected from the group consisting of proteins that can be selected from the group consisting of proteins where the protein is selected from the group consisting of HLA-DR and HLA-DP or HLA-DQ proteins.In some embodiments, the HLA protein is selected from the group consisting of: HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA-DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, HLA-DRB5*01:01.

[0179]

[0272] In some embodiments, the heterodimer pair is expressed in eukaryotic cells. In some embodiments, the heterodimer pair is encoded by a vector. In some embodiments, the vector comprises a nucleic acid sequence encoding an alpha chain and a beta chain of an HLA protein, a secretion signal, a biotinylation motif and at least one tag for identification or purification, such that each HLA protein alpha1 and beta1 heterodimer is secreted in a dimeric state and the secreted heterodimer associates with a placeholder peptide. In some embodiments, the vector comprises a nucleic acid sequence encoding an alpha chain and a beta chain of an HLA protein, a secretion signal, a biotinylation motif and at least one tag for identification or purification, such that each HLA protein alpha1 and beta1 heterodimer is secreted in a dimeric state and the secreted heterodimer associates with a placeholder peptide.

[0180]

[0273] In some embodiments, the HLA class II heterodimer is secreted from eukaryotic cells into the cell culture medium and purified by any one of column or batch chromatography, ion exchange chromatography, size exclusion chromatography, affinity chromatography or LC-MS.

[0181]

[0274] In one aspect, the present specification provides a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: using a computer processor to input amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a series of presentation predictions for the peptide sequence, each presentation prediction indicating the probability that one or more proteins encoded by HLA class I or II alleles of the subject's cells present an epitope sequence of a given peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, the training data comprising: sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, training peptide sequence information comprising amino acid position information, the training peptide sequence information being related to HLA proteins expressed in cells, and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictor variables; (b) determining or predicting, based on the set of presentation predictions, that each of the peptide sequences of the polypeptide sequence will not be immunogenic in the subject; and (c) administering a composition comprising the drug to the subject.

[0182]

[0275] In one aspect, provided herein is a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, the method comprising: (a) using a computer processor to input amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA peptide presentation prediction model to generate a series of presentation predictions for the peptide sequence, wherein each presentation prediction indicates the probability that one or more proteins encoded by the HLA class I or II alleles of the subject's cells present an epitope sequence of a given peptide sequence, and the machine learning HLA peptide presentation prediction model comprises: a plurality of predictor variables identified based at least on training data, the training data comprising: sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, training peptide sequence information comprising amino acid position information, the training peptide sequence information being associated with HLA proteins expressed in cells, and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictor variables; (b) determining or predicting, based on the set of presentation predictions, that at least one of the peptide sequences of the polypeptide sequence will be immunogenic in the subject.

[0183]

[0276] In one embodiment, the method further comprises determining not to administer the drug to the subject.

[0277] In one embodiment, the drug comprises an antibody or a binding fragment thereof.

[0184]

[0278] In one embodiment, the peptide sequence of the polypeptide sequence comprises each neighboring peptide sequence of a polypeptide sequence having a length of 8, 9, 10, 11 or 12 amino acids, wherein the protein encoded by the HLA class I or II allele of the subject's cells is the protein encoded by the HLA class I allele of the subject's cells.

[0185]

[0279] In one embodiment, the peptide sequences of the polypeptide sequences include each neighboring peptide sequence of a polypeptide sequence having a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids, where the protein encoded by the HLA class I or II allele of the subject cell is the protein encoded by the HLA class II allele of the subject cell.

[0186]

[0280] In one aspect, provided herein is a method of treating a subject having an autoimmune disease or condition, the method comprising: (a) identifying or predicting an epitope of a protein presented and expressed by HLA class I or II of the subject's cells, wherein the identified or predicted epitope and the complex comprising HLA class I or II are targeted by the subject's CD8 or CD4 T cells; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells derived from the subject or allogeneic regulatory T cells; and (d) administering to the subject the regulatory T cells expressing the TCR.

[0187]

[0281] In one embodiment, the autoimmune disease or condition is diabetes.

[0282] In one embodiment, the cells are islet cells.

[0283] In one aspect, provided herein is a method of treating a subject having an autoimmune disease or condition, the method comprising administering to the subject regulatory T cells expressing (i) an epitope of an expressed protein identified or predicted to be presented by HLA class I or II of the subject's cells, and (ii) a T cell receptor (TCR) that binds to a complex comprising HLA class I or II, wherein the complex is targeted by the subject's CD8 or CD4 T cells.

[0188]

[0284] Additional aspects and advantages of the present disclosure will be readily apparent to those of ordinary skill in the art from the following detailed description, where merely exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure can be other and different embodiments, and some of the details thereof can be variations in various obvious matters without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.

[0189]

[0285] The MAPTAC™ construct can be used for high-throughput peptide binding assays in which peptides bound to HLA class II are measured using LC-MS / MS at various time points after isolation using the MAPTAC™ construct to obtain an array of peptides having different stabilities and under various conditions such as heating at 37°C.

[0190]

[0286] In one aspect, provided herein is a method for treating cancer in a subject, the method comprising: identifying a peptide sequence, the peptide sequence being associated with cancer, the identifying step comprising comparing a DNA, RNA, or protein sequence from the subject's cancer cells to a DNA, RNA, or protein sequence from the subject's normal cells; using a computer processor to input amino acid position information of the identified peptide sequence into a machine learning HLA peptide presentation prediction model to generate a series of presentation predictions for the identified peptide sequence, each presentation prediction indicating the probability that one or more proteins encoded by the HLA class II alleles of the subject's cells will present a given sequence of the identified peptide sequence, the machine learning HLA peptide presentation prediction model comprising: a plurality of predictor variables identified based at least on training data, the training data comprising: sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry, amino acid position information, training peptide sequence information associated with HLA proteins expressed in cells, and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictor variables; selecting a subset of the identified peptide sequence based on the series of presentation predictions for preparing an individualized cancer treatment; and administering to the subject a composition comprising one or more peptides, wherein the prediction model has a positive predictive value of at least 0.1%, from 0.1% to 50%, or up to 50% at a recall rate of at least 0.1.

[0191]

[0287] In some embodiments, the machine learning HLA peptide presentation prediction model comprises sequence information of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry after performing reverse-phase offline fractionation.

[0192]

[0288] In some embodiments, the prediction model shows an improvement of 1.1x to 100x compared to NetMHCIIpan. In some embodiments, the prediction model shows an improvement of 1.1, 2, 3, 4, 5, 6, 7, 7.4, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 8, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100x or more compared to NetMHCIIpan.

[0193] Incorporation by reference

[0289] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the incorporated publications and patents or patent applications conflict with the disclosure contained in the specification, the specification supersedes and / or takes precedence over any such conflicting material.

[0194]

[0290] The novel features of the invention are set forth in detail in the appended claims. A better understanding of the features and advantages of the invention will be obtained from the following detailed description which sets forth illustrative embodiments in which the principles of the invention are utilized, and from the appended drawings (also referred to herein as "FIG."). The drawings are as follows.

Brief Description of the Drawings

[0195]

Figure 1A

[0291] The illustration of FIG. 1A represents a peptide docked on an MHC class I protein. The figure discloses SEQ ID NO: 36.

Figure 1B

[0292] Figure 1B shows an exemplary illustration depicting a peptide docked onto an MHC class II protein. The figure discloses SEQ ID NO: 37.

Figure 2

[0293] Figure 2 shows an exemplary experimental approach for generating data on monoallelic HLA class II-binding peptides. HLA class II peptides are introduced into any cell, including cells that do not express HLA class II, such that a specific HLA class II allele is expressed in the cells. The population of cells expressing the genetically engineered HLA is recovered, lysed, tagged (e.g., biotinylated) on its HLA peptide complex, and immunopurified (e.g., using the biotin-streptavidin interaction). HLA-associated peptides specific for a single HLA can be eluted from the tagged (e.g., biotinylated) complex and evaluated (e.g., sequenced using high-resolution LC-MS / MS).

Figure 3

[0294] Figure 3 shows an exemplary sequence logo representation of HLA class II-DRB1*11:01-associated peptides across Neon BAP, Expi293 cell line; Neon BAP, A375 cell line; IEDB, affinity <50 nM; and Pan-HLA class II Ab, homozygous LCL. Figure 3 shows an example where the MS-derived motif matches a known pattern, demonstrating consistency across the transfected cell lines.

Figure 4

[0295] Figure 4 is an exemplary depiction of the performance of HLA class II binding predictors. Figure 4 is a bar plot showing the performance of the binding predictors (neonmhc2) and NetMHCIIpan applied to a validation data set consisting of observed mass spectrometry of peptides and decoy peptides generated at a ratio of 1:19 (hit:decoy) by randomly shuffling the hit peptides. For the Neon binding predictor neonmhc2, a separate model is constructed for each MHC II allele shown. The height of the bar represents the positive predictive value (PPV) defined as the percentage of predicted binders in the validation set that were actually hit peptides. The alleles are sorted by the performance of the model when predicting for that allele.

Figure 5

[0296] Figure 5 shows an exemplary effect of the scored peak intensity (SPI) threshold on binding predictor validation. Figure 5 shows the performance of HLA class II binding predictors when trained / validated on a set of peptides with different scored peak intensity (SPI) cutoffs. For each allele-specific model trained, the performance of the model in three settings is shown: trained and evaluated on a data set using observed MS hit peptides greater than or equal to 70SPI, trained on peptides greater than or equal to 50SPI, and validated on peptides greater than or equal to 70SPI, and trained and validated on peptides greater than or equal to 50SPI.

Figure 6

[0297] Figure 6 shows an exemplary bar plot showing representative data from a number of peptides observed by allele profiling by LC-MS / MS greater than or equal to a 70 scored peak intensity (SPI) cutoff. Each bar represents the total number of observed peptides for an allele. Data exists for 35 HLA-DR alleles collected. The data for the 35 HLA-DR alleles collected has a population coverage of >95% of HLA-DR (USA allele frequency).

Figure 7A

[0298] Figure 7A shows the PPV of the model when applied to the test partition of the data for the indicated HLA class II alleles. The decoy peptides used were scrambled sequences of the positive (hit) peptide sequences at a 1:19 hit-to-decoy ratio. The PPV was determined by identifying the top 5% of the peptides scored in the test partition and determining the proportion of those that were positive with respect to binding to the proteins encoded by each HLA class II allele.

Figure 7B

[0299] Figures 7B - 7D show exemplary prediction performance as a function of the size of the training set (curves obtained by artificially downsampling the training set). Figures 7B - 7D generally show that for the 35 HLA - DR alleles collected, as the size of the training set increases, the value of the PPV increases.

Figure 7C

Figure 7D

Figure 8

[0300] FIG. 8 shows an exemplary graph demonstrating that processing-related variables can further improve prediction. Random sequences of peptides observed by MS selected from the exome encoding proteins can be distinguished. In the partition of the training data, logistic regression can be fitted to predict HLA class II presentation using binding strength (predictors of NetMHCIIpan or Neon) and processing features (terms of RNA-Seq expression and derived gene-level biases). In a separate evaluation partition, the positions of exons overlapping with MHC II peptides (“hits”) observed by MS can be scored alongside the positions of random exons not observed by MS (a ratio of 1:499). The top 0.2% (1 / 500) can be called positive, and the positive predictive value can be evaluated by this threshold.

Figure 9

[0301] Figure 9 shows an exemplary neural network architecture. The input peptide is represented as a 20mer, and shorter peptides are filled with "missing" characters. Each peptide has a 31-dimensional embedding, and thus the input to the neural network is a 20×31 matrix. Before being processed by the neural network, feature normalization of the 20×31 matrix is performed based on the mean and standard deviation of the feature values in the training set. The first convolutional layer has a 9-amino acid kernel and 50 filters (also called channels) with a rectified linear unit (ReLU) activation function. This is followed by batch normalization and then spatial dropout with a 20% dropout rate. This is followed by another convolutional layer with a 3-amino acid kernel and 20 filters with a ReLU activation function, and then again by batch normalization and spatial dropout with a 20% dropout rate. Then global max pooling is applied to obtain the neuron that is maximally activated in each of the 20 filters; these 20 values are then passed to a fully connected (densely connected) layer with a single neuron using the sigmoid activation function. The output of this layer is processed as a bound / unbound prediction. L2 regularization is applied to the weights of the first convolutional layer, the second convolutional layer, and the densely connected layer with weights of 0.05, 0.1, and 0.01, respectively. Additional models used are those with the number of convolutional layers and the kernel size of each layer changed.

Figure 10

[0302] Figure 10 shows an exemplary computer control system programmed to execute the methods provided herein or otherwise configured.

Figure 11A

[0303] Figure 11A shows an exemplary overview of the MAPTAC (trademark) experimental workflow. The figure discloses SEQ ID NO: 38.

Figure 11B

[0304] Figure 11B shows the exemplary number of peptides per allele integrated over replicates.

Figure 11C

[0305] Figure 11C shows the exemplary peptide length distribution of HLA class I and HLA class II alleles profiled by MAPTAC™.

Figure 11D

[0306] Figure 11D shows the exemplary cysteine frequency per residue observed with MAPTAC™ and IEDB (alleles DRB1*01:01, DRB1*03:01, DRB1*09:01, and DRB1*11:01), the human proteome, and multi-allelic MS data from previous publications.

Figure 12A

[0307] Figure 12A shows the Caucasian frequencies of HLA-DR, -DP, and -DQ alleles present in >1% of individuals and the number of peptides from the indicated sources measured as strong binders (<50 nM).

Figure 12B

[0308] Figure 12B shows the exemplary length distribution of IEDB peptides and the associated HLA class II affinity measurements.

Figure 12C

[0309] Figure 12C shows exemplary Western blots of (1) Expi293, (2) HeLa, and (3) A375 cell lines individually transfected with two HLA class I and two HLA class II alleles: HLA-A*02:01, HLA-B*45:01, HLA-DRB1*01:01, and HLA-DRB1*11:01. The membrane was blotted with an anti-biotin ligase epitope tag to visualize biotin acceptor peptide (BAP) and anti-beta-tubulin as a loading control. The lanes correspond to the following fractions collected during the MAPTAC™ protocol: lane 1 input, lane 2 biotinylated input, and lane 3 input after pull-down.

Figure 12D

[0310] Figure 12D shows the exemplary amino acid frequency per residue observed with MAPTAC™ and IEDB (alleles DRB1*01:01, DRB1*03:01, DRB1*09:01, and DRB1*11:01), the human proteome, and multi-allelic MS data from previous publications.

Figure 12E

[0311] Figure 12E shows the white frequencies of HLA-DR, -DP, and -DQ alleles present in >1% of the population and the number of peptides from the indicated sources measured as strong binders (<50 nM). This figure includes additional data related to Figure 12A. The additional data was taken from tools.iedb.org / main / datasets / .

Figure 12F

[0312] Figure 12F shows the amino acid frequencies for exemplary residues observed with MAPTAC™ (reduced and alkylated), MAPTAC™ (untreated), and IEDB (alleles DRB1*01:01, DRB1*03:01, DRB1*09:01, and DRB1*11:01), the human proteome, and multi-allele MS data from previous publications.

Figure 13

[0313] Figure 13 shows an exemplary representation of the core binding sequence logos for MHC II alleles by MAPTAC™ and IEDB. The sequence logo is a graphical representation where the height of each amino acid is proportional to its frequency of occurrence in peptides that bind to the MHC protein encoded by the allele. Positions with the lowest entropy are presented in color, and the colors correspond to amino acid properties. Peptides were derived from the indicated datasets and aligned according to a CNN-based predictor (Methods). The logo represents all peptides that did not exactly match the entire motif (e.g., there are no peptides sequestered in the “trash” cluster).

Figure 14A

[0314] Figure 14A shows an exemplary sequence logo of HLA-A*02:01-binding peptides (ligands) analyzed using various HLA-ligand profiling techniques such as binding assays, stability assays, soluble HLA (sHLA) mass spectrometry, mono-allelic mass spectrometry, and MAPTAC™ in two different cell lines (A375 and expi293).

Figure 14B

[0315] Figure 14B shows exemplary fractions of the MAPTAC™ peptide and displays 0, 1, 2, 3, and 4 of the anchors defined to aid discovery.

Figure 14C

[0316] Figure 14C shows an exemplary distribution of the predicted binding affinities of peptides observed with MAPTAC™ (20 peptides per allele, each having a nested set with SPI > 70 and size ≥ 2) and length-matched decoys sampled from the proteome by NetMHCIIpan.

Figure 15A

[0317] Figure 15A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish monoallelic MHC peptides from length-matched scrambled decoys. The schematic shows the embedding of amino acid features, the use of two convolutional layers with different filter sizes, and the use of global max pooling as input to the final logistic output node.

Figure 15B

[0318] Figure 15B is an exemplary result showing Kendall's tau statistics for the correlation between measured IEDB affinity and binding predictions from either neonmhc2 or NetMHCIIpan. The peptides evaluated include only those submitted to IEDB in the years after NetMHCIIpan was released.

Figure 16

[0319] Figure 16 is an exemplary depiction of the performance of neonmhc2 as a function of training dataset size.

Figure 17A

[0320] Figure 17A shows the exemplary cluster assignment of MAPTAC™ peptides (20 per allele) spiked into the pan-DR and pan-class II MHC MS datasets. The datasets were deconvolved using GibbsCluster. Each box represents one MAPTAC™ peptide. The color of the box indicates which cluster it was assigned to, and the gray bar indicates from which allele the peptide actually originated. The total number of clusters in the Gibbs cluster solution (right side) was selected using the mutual information (MI) metric. The MI score also determines how the samples are sorted; samples with high MI solutions are considered superior.

Figure 17B

[0321] Figure 17B shows exemplary core-binding sequence logos of multi-allelic MS data deconvolved by GibbsCluster. Each set of peptides corresponds to the cluster that aligned best with the MAPTAC™ spike-in.

Figure 17C

[0322] Figure 17C shows the representative performance of models using either MAPTAC™ data or deconvolved multi-allelic data to predict hold-out MAPTAC™ peptides. For each allele, the larger of the two data sources (usually MAPTAC™) was downsampled so that the predictors were based on an equal number of training examples. The performance of NetMHCIIpan is shown as an additional comparison.

Figure 17D

[0323] Figure 17D shows exemplary core-binding sequence logos obtained from multi-allelic MS data from the indicated source.

Figure 18A

[0324] Figure 18A shows an exemplary graph of the ratio of gene expression of the source (transcripts per million (TPM)) for peptides observed by MS and random proteome decoys (data replotted from Schuster et al., 2017).

Figure 18B

[0325] Figure 18B shows the observed and predicted number of class II peptides per exemplary gene determined by the co - analysis of datasets of colorectal cancer, melanoma, and ovarian cancer (Loffler et al., 2018, and Schuster et al., 2017). The predicted number is derived by multiplying the length of the gene by the expression level. The predicted and observed numbers were summed across the relevant samples. Genes known to be present in plasma are marked according to their concentrations (insert figure).

Figure 18C

[0326] Figure 18C shows an exemplary distribution of the enrichment scores of genes related to autophagy (the ratio of observed to predicted observations, similar to that described in Figure 18B).

Figure 18D

[0327] Figure 18D shows an exemplary distribution of the enrichment scores by the localization of each source gene. Source gene localization was determined using Uniprot (uniprot_sprot.dat).

Figure 18E

[0328] Figure 18E shows exemplary data representing a comparison of the predicted frequency versus the observed frequency of the proportion of the total number of peptides with MHC - II binding affinity, separated based on their cellular localization characteristics.

Figure 18F

[0329] Figure 18F shows exemplary representative data of the relative agreement of peptides in observations regarding two different gene expression profiles. For each sample, the number of gene - level peptides was modeled as a linear combination of the bulk tumor gene expression and the professional APC (macrophage) gene expression profile. The ratio of the coefficients determines the relative agreement between each expression profile and the peptide repertoire. Error bars correspond to the 95% confidence intervals computed computationally by bootstrap resampling.

Figure 19A

[0330] Figure 19A shows exemplary representative data of the expression levels of HLA-DRB1 in the study of five examples. Each dot represents the expression in an individual cell type of an individual patient averaged across all cells.

Figure 19B

[0331] Figure 19B shows exemplary representative data of tumor- and stroma-derived HLA-DRB1 expression when input from RNA-Seq of TCGA patients. The horizontal bars correspond to individual patients and are grouped by tumor type. Patients were included if they had a mutation in an HLA class II pathway gene (CIITA, CD74, or CTSSS) as determined by DNA-based variant calling. For each patient, the fraction of HLA-DRB1 expression attributable to the tumor was estimated as min(1, 2f) (where f is the fraction of RNA-Seq reads in CIITA, CD74, or CTSS that show a mutation).

Figure 19C

[0332] Figure 19C shows exemplary representative data of an additional single-cell RNA-Seq study including biopsies before and after checkpoint blockade immunotherapy.

Figure 20

[0333] Figure 20 shows exemplary representative experimental data evaluating the overall performance of predictions for native donor tissue.

Figure 21A

[0334] Figure 21A shows exemplary representative data demonstrating that an integrated presentation model predicts the cellular HLA class II ligandome. This represents the PPV at a hit-to-decoy ratio of 1:499 in the pan-DR dataset (also analyzed in Figures 30B and 32E). The predictors use binding predictions (NetMHCIIpan or neonmhc2) and optionally employ gene expression, gene bias (according to Figure 32A), and overlap with HLA-DQ peptides observed so far. For each candidate peptide, the binding score was calculated as the maximum value across HLA-DR alleles in the genotype of the sample.

Figure 21B

[0335] Figure 21B shows exemplary representative data showing the predictive performance of tumor-derived peptides as identified using SILAC, presented by dendritic cells (analyzed from cell lysates), using the same hit:decoy ratio and performance metrics as in Figure 21A with and without the use of processing features.

Figure 21C

[0336] Figure 21C shows exemplary expression and gene bias scores of heavily labeled peptides (red dots, plotted according to K562 expression) observed in UV treatment experiments compared to lightly labeled peptides (gray dots, plotted according to DC expression).

Figure 21D

[0337] Figure 21D shows an exemplary illustration representing the overlap of genes of the source of heavily labeled peptides by lysate and UV treatment experiments. The names of the genes are color-coded by functional class.

Figure 22A

[0338] Figure 22A shows an exemplary flowchart representing the assay protocol disclosed herein for validating CD4+ T cells and T cell responses driven by HLA class II.

Figure 22B

[0339] Figure 22B shows the design of an exemplary HLA protein dimer construct (upper panel) and a graphical representation of an exemplary assay workflow (lower panel) for a peptide exchange assay. The figure discloses "10×His" as SEQ ID NO: 20.

Figure 23

[0340] Figure 23 shows an exemplary vector design for MHC-II expression for screening of new binding peptides, and an exemplary illustration of the expressed protein product. The figure discloses SEQ ID NO: 39 and discloses "10×His" as SEQ ID NO: 20.

Figure 24

[0341] Figure 24 shows an exemplary flowchart of transfection, purification and cleavage of placeholder peptides from the beta chain.

Figure 25A

[0342] Figure 25A shows an exemplary illustration of a vector encoding a CLIP peptide related to the increased secretion of expressed MHC-II peptide. The figure discloses SEQ ID NO: 21.

Figure 25B

[0343] Figure 25B shows exemplary graphical representations in shorter and longer forms of nucleic acids encoding CLIP0 and CLIP1, respectively. The figure discloses SEQ ID NO: 1 and SEQ ID NO: 21, respectively, in the order in which they appear.

Figure 25C

[0344] Figure 25C shows exemplary representative results of Coomassie gel analysis of alpha and beta chains with or without longer CLIP.

Figure 26A

[0345] Figure 26A shows an exemplary illustration of a TR-FRET assay.

Figure 26B

[0346] Figure 26B shows exemplary representative polarization data from an HLA class II peptide binding assay using a fluorescence resonance energy transfer (FRET) assay with specific peptides.

Figure 26C

[0347] Figure 26C shows exemplary representative polarization data from an HLA class II peptide binding assay using a fluorescence resonance energy transfer (FRET) assay with specific peptides.

Figure 26D

[0348] Figure 26D shows an exemplary percentage of peptide replacement bound by an MHC construct calculated from the increase in fluorescence.

Figure 26E

[0349] Figure 26E shows an exemplary percentage of peptide replacement bound by an MHC construct calculated from the increase in fluorescence.

Figure 26F

[0350] Figure 26F shows exemplary peptide exchange using an assay employing differential scanning fluorimetry (DSF). A graphical representation is shown of an exemplary mechanism for detecting thermal dissociation of peptides from MHC class II, which also dissociates the MHC class II heterodimer, resulting in fluorophore binding and high fluorescence. An exemplary schematic of the removal of a placeholder peptide by an epitope peptide is also shown. An exemplary melting curve plotted against temperature is also shown.

Figure 26G

[0351] Figure 26G shows an exemplary soluble HLA-DM construct for performing MHC class II peptide exchange and its use. The construct shown contains a CMV promoter, a coding sequence for the HLA-DM beta chain downstream of a secretion sequence (leader) and a coding sequence for the HLA-DM alpha chain, as well as a BAP sequence at the 3’ end of the beta chain coding sequence; a His tag at the 3’ end of the alpha chain coding sequence. The two chains are separated by an intervening ribosome skipping sequence. The construct was expressed in Expi-CHO cells, the protein was secreted into the medium, and the culture medium was purified. The figure discloses "10×His" as SEQ ID NO: 20.

Figure 26H

[0352] Figure 26H shows exemplary size exclusion chromatography data using HLA-sDM to perform peptide exchange.

Figure 27A

[0353] Figure 27A shows an exemplary illustration of the construction of an exemplary DRB tetramer repertoire.

Figure 27B

[0354] Figure 27B shows an exemplary illustration of the construction of an exemplary class II tetramer repertoire.

Figure 27C

[0355] Figure 27C shows an exemplary illustration of a summary of DRB tetramer repertoire coverage of DRB1 alleles for peptide exchange.

Figure 27D

[0356] Figure 27D shows exemplary coverage of human MHC class II allele production.

Figure 27E

[0357] Figure 27E shows exemplary results from tetramer staining of samples induced with Flu epitopes (memory response) or HIV epitopes (naïve response).

Figure 28A

[0358] Figure 28A shows an exemplary graphical representation of a method for evaluating peptides regarding HLA class II restriction by a fluorescence polarization assay that enables a screening method for rapidly identifying allelic restriction of epitope peptides. The assay principle shown in Figure 28A enables affinity measurements and measurements free of suspicion of peptide exchange.

Figure 28B

[0359] Figure 28B shows an exemplary summary (upper panel) of a plurality of assay conditions investigated in a fluorescence polarization assay using DRB1*01:01. Also shown is an illustration of a soluble MHC class II allele and a full-length MHC class II allele having a transmembrane domain in a surfactant micelle (lower panel), both of which were constructed using a placeholder peptide with a cleavable linker for use in the assay.

Figure 28C

[0360] Figure 28C shows an exemplary graphical representation of an assay for investigating the full-length and soluble alleles shown in the lower panel of Figure 28B. In short, both the full-length and soluble alleles are expressed in cells. The form of the membrane-bound full-length allele is recovered by permeabilizing the membrane, while the secreted form is recovered from the cell supernatant. The recovered class II HLA allele protein is purified by passing it through a nickel (Ni2+) column.

Figure 28D

[0361] Figure 28D shows exemplary data indicating that the purification method does not affect peptide potency. The average IC50 values from experiments using full-length HLA-DR1 purified with L243 and full-length HLA-DR1 purified with Ni2+ are shown on the left.

Figure 28E

[0362] Figure 28E shows exemplary data indicating that the choice of soluble form (sDR1) or full-length form (fDR1) does not affect peptide potency. The average IC50 values from experiments using the sDR1 form or fDR1 are shown on the left. FP, fluorescence polarization.

Figure 28F

[0363] Figure 28F shows an exemplary graphical representation of an exemplary evaluation of neonmhc2 and NetMHCIIpan predicted peptides in a binding assay and identification of discrepant peptides.

Figure 28G

[0364] Figure 28G shows exemplary fluorescence polarization binding screening data regarding the evaluation of neonmhc2 predicted peptides; also shown as a heat map of the percent inhibition of probe binding shown for each concentration of peptide used. Green indicates excellent binding proportional to the intensity of the color. Yellow indicates intermediate binding and red indicates poor binding, which are also indicated by the corresponding percent inhibition values.

Figure 28H

[0365] Figure 28H shows a summary of the evaluation of neonmhc2 predicted peptides in an exemplary binding assay.

Figure 29

[0366] Figure 29 shows exemplary average numbers of peptides from average MAPTAC™ experimental replicates (50 million cells) for each HLA allele.

Figure 30A

[0367] Figures 30A - 30C show exemplary binding core analyses regarding HLA class II MAPTAC™ allele + / - HLA - DM and the fidelity of multi - allele deconvolution. Figure 30A shows exemplary sequence logos of one representative HLA - DR, - DQ, and - DP allele with HLA - DM cotransfection (expi293 cell line) and with and without IEDB of MAPTAC™. The height of each amino acid is proportional to its frequency. Amino acids with a frequency greater than 10% are shown in color according to their chemical properties; all others are shown in gray. Peptides were aligned according to the GibbsCluster tool (supplementary method), and the logo represents all peptides that did not exactly match the entire motif (e.g., there are no peptides sequestered in the "trash" cluster).

Figure 30B

Figure 30C

Figure 31A

[0368] Figures 31A - 31F show exemplary architectures and benchmarking of the neonmhc2 binding prediction algorithm. Figure 31A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish monoallelic HLA class II peptides from length - matched scrambled decoys. The schematic shows an embedding layer of amino acid features, two convolutional layers of width 6, the presence of skip - to - end connections, and the use of a combination of average and max - pooling operations as inputs to the final logistic output node.

Figure 31B

Figure 31C

Figure 31D

Figure 31E

Figure 31F

Figure 32A

[0369] Figures 32A - 32E show the expression of exemplary genes and protein processing in the HLA class II tumor peptidome. Figure 32A shows exemplary results of the observed number of HLA class II peptides per gene and the predicted number determined by the co-analysis of datasets of colorectal cancer, melanoma, and ovarian cancer. The predicted number is derived by multiplying the length of the gene by the expression level. The predicted and observed numbers were summed across the relevant samples. Genes known to be present in plasma are marked according to their concentrations.

Figure 32B

Figure 32C

Figure 32D

Figure 32E

Figure 33A

[0370] Figures 33A - 33G show exemplary results of the identification and prediction of tumor antigens presented by dendritic cells. Figure 33A shows an exemplary graphical representation of an experimental workflow for identifying HLA-II ligands presented by DCs originating from cancer cells (K562). The cancer cells were grown in SILAC medium until completely incorporated, either lysed or irradiated, and then plated with monocyte-derived dendritic cells. The presented peptides were isolated with a pan-DR antibody and sequenced by LC-MS / MS.

Figure 33B

Figure 33C

Figure 33D

Figure 33E

Figure 33F

Figure 33G

Figure 34A

[0371] Figures 34A - 34B show exemplary characterization of MAPTAC™ data with respect to FIG. 29. FIG. 34A shows an exemplary HLA cell surface analysis by FACS of an Expi293 cell line transfected with a MAPTAC™ construct encoding affinity-tagged HLA-A*02:01 - BAP.

Figure 34B

Figure 35-1

[0372] Figure 35 shows an exemplary comparison of MAPTAC™ and IEDB logos with FIG. 30A. Measured, and NetMHCIIpan-predicted affinities of peptides observed by MS that did not display an excellent NetMHCIIpan score but were well supported by MS (scored peak intensity > 70 and nested set size ≧ 1).

Figure 35-2

Figure 35-3

Figure 36A

[0373] Figures 36A - 36C show exemplary analyses of the fidelity of HLA - DR1 MAPTAC™ data with respect to Figures 30A - 30C. Figure 36A shows exemplary NetMHCIIpan3.1 scores of HLA - DR1 MAPTAC™ peptides (green) (lengths 12 - 23) compared to 50,000 length - matched decoy peptides (blue) randomly sampled from the proteome for common allergens.

Figure 36B

Figure 36C

Figure 37A

[0374] Figures 37A - 37C and 37D (continuation of Figure 37C) show additional exemplary analyses of the MAPTAC™ motif with respect to Figures 30A - 30C. Figure 37A shows sequence logos derived from MAPTAC™ for experiments with and without HLA - DM cotransfection (expi293 cell line).

Figure 37B

Figure 37C

Figure 37D

Figure 38X-1

[0375] Figures 38X, 38Y, 38B - 38D show exemplary neonmhc2 performance statistics and flow staining of T cells with respect to Figures 31A - 31D. Figure 38X shows exemplary performance of neonmhc2 as a function of training dataset size. In the same manner as Figure 31B, PPV was evaluated using the same set of evaluation peptides; however, the training data was randomly downsampled to simulate smaller training datasets.

Figure 38X-2

Figure 38Y

Figure 38B

Figure 38C

Figure 38D

Figure 39A

[0376] Figures 39A - 39C show additional exemplary origin cell analyses of HLA class II with respect to Figures 32A - 32E. Figure 39A shows the exemplary percent rank neonmhc2 scores of HLA class II peptides observed in four PBMC samples profiled by pan-DR antibodies (RG1248, RG1104, RG1095, and HDSC from Figure 30B), depending on whether the gene of the peptide source is present in human plasma. For each peptide, the best (lowest) percent rank over the alleles present in the donor was used. For comparison, scores of proteome decoys with matching random lengths are shown. The box-and-whisker plots mark the 5, 25, 50, 75, and 95 percentiles.

Figure 39B

Figure 39C

Figure 40A-1

[0377] Figures 40A-40B show exemplary additional analysis of the processing motifs related to Figures 32A-32E. Figure 40A shows exemplary amino acid frequencies near the N-terminal and C-terminal peptide cleavage sites, compared to the average proteome frequency (applied to upstream positions U3-U1 and downstream positions D1-D3) or compared to the average peptide frequency (applied to internal positions N1-C1), observed in donor PBMCs, monocyte-derived dendritic cells, colorectal cancer, melanoma, ovarian cancer, and the expi293 cell line (used for most MAPTAC™ data generation).

Figure 40A-2

Figure 40B

Figure 41

[0378] Figure 41 shows an exemplary naming system used to refer to positions upstream of the peptide, within the peptide, and downstream of the peptide.

Figure 42A

[0379] Figure 42A shows a diagram representing an exemplary workflow for the analysis of peptides that are endogenously processed by nLC-MS / MS and presented by HLA-1 and HLA class II.

Figure 42B

[0380] Figure 42B shows a graph presenting the results of an exemplary experiment from the nLC-MS / MS analysis of tryptic peptides with or without FAIMS. Representative overlaps in the detection of HLA-1 and HLA class II peptides by nLC-MS / MS analysis with or without FAIMS at the analysis scale shown are also shown.

Figure 43A

[0381] Figure 43A shows exemplary HLA class I acidic and basic reversed-phase fractionated peptide detections with or without FAIMS.

Figure 43B

[0382] Figure 43B shows exemplary experimental results showing the detection of HLA class I binding signature peptides plotted over retention time.

Figure 44A

[0383] Figure 44A shows exemplary HLA class II acidic and basic reversed-phase fractionated peptide detections with or without FAIMS.

Figure 44B

[0384] Figure 44B shows exemplary experimental results showing the detection of HLA class II binding signature peptides plotted over retention time.

Figure 45A

[0385] Figures 45A and 45B show an exemplary graph of the cross-sizes of HLA class I binding peptides detected using the indicated method (left), and a Venn diagram of exemplary standard and optimized workflows for LC-MS / MS detection of HLA class I binding peptides (right).

Figure 45B

Figure 46A

[0386] Figures 46A and 46B show an exemplary graph of the cross-sizes of HLA class II binding peptides detected using the indicated method (left), and a Venn diagram of exemplary standard and optimized workflows for LC-MS / MS detection of HLA class II binding peptides (right).

Figure 46B

Figure 47A

[0387] Figure 47A shows a study in which MHC class II alleles covering a wide range of human populations are produced as soluble heterodimers with cleavable peptide holders from transiently transfected human cells. Upper panel, design of the soluble MHC class II construct (see Methods and Table 19). Lower left, schematic highlights or protein expression and purification strategies for generating MHCII proteins ready for epitope loading, multimerization, and flow cytometry staining. Example of protein purification of HLA-DRB4*01:03 / DRA*01:01 heterodimer biotinylated (via BirA) and digested with thrombin bound to the CLIP0 holder (PVSKMRMATPLLMQA). Lower right, gel filtration chromatogram and SDS-PAGE gel shown for pooled purified fractions for epitope loading and flow cytometry staining. Lower panel, gel filtration chromatogram and SDS-PAGE gel shown using the subsequently pooled purified fractions for epitope loading and flow cytometry staining indicated by the highlighted points marked "pooled".

Figure 47B

[0388] Figure 47B shows the European allele frequencies of MHC II alleles for which protein purification was demonstrated (Table 19).

Figure 48A

[0389] Figure 48A shows data demonstrating that soluble HLA-DM catalyzes rapid and demand-responsive universal MHC class II peptide exchange. The schematic (upper left) shows a probe-binding assay in which MHC II alleles loaded with placeholder-peptides are exchanged for high-affinity FITC-labeled peptide probes via soluble HLA-DM (catalyst). The graph on the right shows the percentage of peptide binding; binding of the FITC probe was measured via fluorescence polarization at four time points across three (non-catalyzed) catalyzed conditions (see methods). The percentage of peptide binding was normalized to the 24-hour soluble HLA-DM catalyzed condition. The FITC conjugation site is indicated by underlined boldface letters (Table 20). The lower left shows the peptide-binding characteristics of the murine MHC, H2-I-A(b).

Figure 48B

[0390] Figure 48B (left) is a schematic showing a fluorescence polarization competition assay for quantifying the IC50 and allele restriction of epitope peptides. The graph on the right shows the dose-response IC50 curves of epitopes derived from 31 SARS-CoV-2 spike (S) predicted by neonmhc2. For each allele, predicted binders (P1 - P4) and non-binders (P5 - P6) were competed with the FITC probe and peptide binding was measured via fluorescence polarization (Table 21).

Figure 49A-1

[0391] Figure 49A shows results demonstrating that detailed characterization of neoantigen-specific CD4+ T cells from individualized peptide vaccine clinical trials elucidates clonal populations with memory and activation phenotypes. Ex vivo staining of PBMCs from cancer patients by MHCII multimer flow cytometry in individualized peptide vaccine clinical trials for non-small cell lung cancer (NSCLC), melanoma, and bladder cancer (Alspach, E. et al., MHC-II neoantigens shape tumour immunity and response to immunotherapy. Nature 574, 696-701 (2019)). Where possible, multimers were combinatorially encoded (Tarke, A. et al., Impact of SARS-CoV-2 variants on the total CD4+ and CD8+ T cell reactivity in infected or vaccinated individuals. Cell Rep. Med. 2, 100355 (2021)).

Figure 49A-2

Figure 49B

[0392] Figure 49B shows results from the same study as in the study in FIG. 50A, showing the durability of the multimer-positive population over the course of treatment. PBMC samples from patients before vaccination (week 10) and after vaccination (weeks 20, 52, and 76 where applicable) were stained ex vivo.

Figure 49C

[0393] Figure 49C shows results from the same study as in the studies in FIGS. 50A and 50B. The left shows UMAP clustering of bulk CD4 T cells and three tetramer+ populations from NSCLC patient L7 based on CITE antibodies. The right shows the CD4 T cell phenotypes of the bulk and multimer-positive populations of NSCLC patient L7 based on the expression levels of CITE antibodies.

Figure 49D

[0394] Figure 49D shows the clonal distribution and abundance of TCRs sorted from NSCLC patient L7.

Figure 50A-1

[0395] Figure 50A shows flow cytometry data demonstrating SARS-CoV-2 antigen-specific CD4 T cells identified using MHC class II multimers. Ex vivo identification of SARS-CoV-2 antigen-specific CD4+ T cells in PBMCs from five convalescent COVID-19 donors using MHCII multimers. Epitopes derived from the SARS-CoV-2 spike (S), membrane (M), and nucleocapsid (N).

Figure 50A-2

Figure 50B

[0396] Figure 50B shows data characterizing CD4 T cells from the same study as in Figure 49A. The left panel shows that antigen-specific T cells are predominantly effector (EM) and central memory (CM). Naïve, effector, and memory subsets were based on the expression of CD45RA and CD62L. The right panel shows the expression of activation and inhibitory markers among SARS-CoV-2 antigen-specific CD4+ T cells, e.g., Lag3, TIM3, PD1, CD69, CD137, and ICOS. Expression is shown as fold change in mean fluorescence intensity (MFI) of SARS-CoV-2 antigen-specific CD4+ T cells relative to bulk CD4+ T cells for each donor.

Figure 51A

[0397] Figure 51A shows a schematic of a strategy for investigating any possible CD4+ responses via the pMHCII technology platform. Step 1. MHCII alleles can be purified in parallel with epitope identification (via computer prediction and / or immunogenicity screening). Step 2. Candidate epitope / allele pairs are verified using the FP assay. Step 3. The epitope peptide of interest is loaded onto MHCII via HLA-sDM to create the pMHCII antigen for staining. Step 4. pMHCII is multimerized via conjugation to fluorescent streptavidin and then combinatorially encoded to stain CD4+ T cells (where three separate antigen-specific CD4+ T cell populations are combinatorially encoded, each having a unique combination of two colors). Step 5. The stained antigen-specific CD4+ T cells may be further analyzed for expression markers via flow cytometry and / or sorted for single cell analysis.

Figure 51B

[0398] Figure 51B shows the purification of soluble HLA-DM from transiently transfected ExpiCHO cultures. Top left, design of the soluble HLA-DM construct (see methods). Bottom left is a schematic (along with a timeline) representing the purification workflow for the protein expression and purification strategy to secrete HLA-sDM from ExpiCHO cultures. The polyhistidine-tagged protein is purified directly from the culture medium using IMAC resin and used for downstream epitope loading. The construct shown in the figure was transiently transfected into ExpiCHO suspension cells and cultured for a total of 14 days, during which the soluble HLA-DM protein was secreted directly into the culture supernatant. Right is an SDS-PAGE gel of HLA-sDM purified by IMAC.

Figure 52-1

[0399] Figure 52 shows the results demonstrating that soluble HLA-DM catalyzes rapid MHC class II peptide exchange across many MHCII alleles. Binding of the FITC probe was measured via fluorescence polarization at four time points across three (non-catalyzed) catalytic conditions (see the method in Example 17). The percent peptide binding was normalized to the 24-hour soluble HLA-DM catalytic condition. The FITC conjugation sites for allele-specific probes are indicated by red underlines (see Table 20 for all validated FITC-probes).

Figure 52-2

Figure 53A

[0400] Figure 53A shows the results demonstrating that peptide-loaded MHCII tetramers are sensitive to rare antigen-specific CD4+ T cell populations and can be multiplexed to detect multiple antigens in a single sample. pMHCII tetramer staining of PBMCs from healthy donors stimulated with pp65116-129 using epitope-loaded DRB1*01:01 monomers conjugated to A. Klickmer (at a defined molar ratio of streptavidin:pMHCII) or streptavidin tetramers. CLIP / DRB1*01:01 conjugated to either multimer scaffold was used as a negative control. The inset table summarizes the antigen-specific CD4+ frequency and staining index.

Figure 53B

[0401] Figure 53B shows pMHCII tetramer staining of three epitopes (and a CLIP negative control) demonstrating sensitive detection of antigen-specific CD4+ T cells. PBMCs from healthy donors stimulated with influenza (HA1306-318), CMV (pp65116-129), and HIV (Gag262-276) were serially diluted with unstimulated PBMCs from the same donor. Linear regression plot between the observed multimer positive frequency and dilution for each antigen-specific CD4+ T cell; the dotted horizontal line represents the limit of detection based on the observed tetramer frequency of irrelevant (CLIP) pMHCII staining.

Figure 53C

[0402] Figure 53C shows the combinatorial encoding strategy, tetramer staining flow plots, and observed / predicted tetramer frequencies for three pMHCII antigens (and a CLIP negative control). PBMCs from healthy donors stimulated with influenza, CMV, and HIV epitopes were mixed at equal ratios and stained with the corresponding loaded DRB1*01:01 tetramers. The inset table summarizes the tetramer positive frequencies between single epitope and multi-epitope staining approaches encoded in combination. The flow cytometry plots demonstrate the pMHCII tetramer positive populations for all three antigens using combinatorial encoding (lower panel). The right panel shows pMHCII tetramer staining gated on either CD8+ (upper) or CD4+ (lower) T cells from healthy donor PBMCs.

Figure 54A

[0403] Figure 54A shows the gating scheme used to characterize pMHCII tetramer positive CD4+ T cells from convalescent COVID-19 donors for the data shown in Figures 54B - 54E.

Figure 54B

[0404] Figure 54B shows the staining of irrelevant peptides (CLIP, IGRP, and / or proinsulin) of PBMCs from COVID-19 convalescent donors M, Q, N, O, and P. Naive, effector, central memory (CM), and effector memory (EM) subsets were defined by CD45RA and CD62L surface expression. Populations were gated on CD4+ T cells. Quantification of activation and exhaustion markers for each tetramer positive population is shown. Histogram values are based on the mean fluorescence intensity (MFI) of each surface marker.

Figure 54C

[0405] Figure 54C shows data regarding the phenotypic characterization of bulk (gray, back) and multimer-positive (red, front) CD4 T cell populations from convalescent COVID-19 donor number N. Naive, effector, central memory (CM), and effector memory (EM) subsets were defined by CD45RA and CD62L surface expression. Populations were gated on CD4+ T cells. Quantification of activation and exhaustion markers for each tetramer-positive population is shown. Histogram values are based on the mean fluorescence intensity (MFI) of each surface marker.

Figure 54D

[0406] Figure 54D shows the phenotypic characterization of bulk (gray, back) and multimer-positive (red, front) CD4 T cell populations from convalescent COVID-19 donor number P. Naive, effector, central memory (CM), and effector memory (EM) subsets were defined by CD45RA and CD62L surface expression. Populations were gated on CD4+ T cells. Quantification of activation and exhaustion markers for each tetramer-positive population is shown. Histogram values are based on the mean fluorescence intensity (MFI) of each surface marker.

Figure 54E

[0407] Figure 54E shows the phenotypic characterization of bulk (gray, back) and multimer-positive (red, front) CD4 T cell populations from convalescent COVID-19 donor number O. Naive, effector, central memory (CM), and effector memory (EM) subsets were defined by CD45RA and CD62L surface expression. Populations were gated on CD4+ T cells. Quantification of activation and exhaustion markers for each tetramer-positive population is shown. Histogram values are based on the mean fluorescence intensity (MFI) of each surface marker.

Figure 55A

[0408] Figure 55A shows MHCII multimer analysis and sorting of antigen-specific CD4+ T cells from patients enrolled in an individualized peptide cancer vaccine trial. The upper panel shows the gating scheme for sorting multimer-positive cells for CITEseq and TCRseq analysis. The middle panel shows irrelevant peptide (CLIP) staining of PBMCs from NSCLC patient L7, bladder cancer patients B9 and B10, and melanoma patient M23. The lower panel is a UMAP analysis of CITE marker expression levels for bulk CD4+ T cells from NSCLC patient L7 and CD4+ T cells sorted with three pMHCII tetramers presented separately.

Figure 55B-1

[0409] Figure 55B (left) shows the expression levels and clustering of specific CITE markers overlaid on the total UMAP of CD4+ T cells sorted with bulk + all multimers. The right shows the expression levels and clustering of specific CITE markers overlaid on the total UMAP of CD4+ T cells sorted with bulk + all multimers.

Figure 55B-2

Figure 55C

[0410] Figure 55C (left) shows the UMAP distribution for the top TCR clones from CD4+ T cell populations sorted with each tetramer from NSCLC L7; (right) CD4+ T cell phenotype distribution of the top 5 TCR clones of each population sorted with each tetramer from NSCLC L7.

Figure 56

[0411] Figure 56 shows data demonstrating that high post-translational modification (PTM) of MHC class II proteins affects staining with labeled epitopes that can bind to MHC class II proteins (here, the exemplary MHC class II protein, DRB1*01:01, is shown). Comparison of the second column from the top with the first column shows that low PTM DRB1*01:01 confers superior staining performance compared to high PTM MHCII proteins. Similarly, low PTM DRB1*01:01 shows high fluorescence staining with exchanged epitopes comparable to the data in the central column.

DETAILED DESCRIPTION OF THE INVENTION

[0196]

[0412] All terms are intended to be understood as would be understood by one of ordinary skill in the art. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0197]

[0413] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0414] The various features of the present disclosure may be described in the context of a single embodiment, but the features may be provided separately or in any suitable combination. Conversely, the present disclosure may be described herein in the context of separate embodiments for clarity, but the disclosure may be implemented in a single embodiment.

[0198]

[0415] The present disclosure is based on the important discovery that the presentation of antigens, specifically cancer antigens, by specific HLA class II alpha and beta chain pairs can be predicted with high confidence using a novel computer-based machine learning HLA peptide presentation prediction model that enables the use of HLA class II-specific peptides to improve immunotherapy.

[0199]

[0416] In one aspect, the present disclosure provides a method for predicting peptides that can accurately dock or bind to specific HLA class II alpha and beta chain heterodimers such that high-fidelity binding of the peptides to HLA class II proteins (including alpha and beta chain heterodimers) ensures the presentation of specific peptides to T lymphocytes, thereby eliciting a specific immune response and avoiding all cross-reactivity and immune dysregulation states. Some recent studies have shown that CD4+ T cells also recognize HLA class II presentation ligands and contribute to tumor management. Cancer vaccines and other immunotherapies ideally utilize directing CD4+ T cell responses, but current efforts are stalled in HLA class II antigen prediction due to the insufficient accuracy of current prediction tools.

[0200]

[0417] In one aspect, the present disclosure is further maintained by means of the ability of an HLA class II protein to stimulate CD4+ T cell activation and immune memory when a peptide is administered therapeutically to a subject expressing a specific cognate HLA class II protein, and provides a method for predicting a peptide that can bind precisely to a specific HLA class II protein such that a robust immune response can be activated using the peptide. In some embodiments, the methods provided herein demonstrate an improvement in specific HLA class II protein prediction beyond currently available predictors. In some embodiments, the methods provided herein demonstrate an improvement of at least about 1.1-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 2-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 3-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 4-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 5-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 6-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 7-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 8-fold beyond currently available predictors in specific HLA class II protein prediction. In some embodiments, the methods provided herein demonstrate an improvement of at least about 9-fold beyond currently available predictors in specific HLA class II protein prediction.In some embodiments, the methods provided herein exhibit at least about a 10-fold improvement over currently available predictors in predicting specific HLA class II proteins. In some embodiments, the methods provided herein exhibit at least about a 15-fold improvement over currently available predictors in predicting specific HLA class II proteins. In some embodiments, the methods provided herein exhibit at least about a 20-fold improvement over currently available predictors in predicting specific HLA class II proteins. In some embodiments, the methods provided herein exhibit at least about a 30-fold improvement over currently available predictors in predicting specific HLA class II proteins. In some embodiments, the methods provided herein exhibit at least about a 40-fold improvement over currently available predictors in predicting specific HLA class II proteins. In some embodiments, the methods provided herein exhibit at least about a 50-fold improvement over currently available predictors in predicting specific HLA class II proteins. In some embodiments, the methods provided herein exhibit at least about a 60-fold improvement over currently available predictors in predicting specific HLA class II proteins.

[0201]

[0418] In one aspect, provided herein are methods of immunotherapy tailored or individualized for a particular patient. All subjects or patients express a particular array of HLA class I and HLA class II proteins. HLA typing is a well-known technique that enables determination of the specific repertoire of HLA proteins expressed by a subject. Once the HLA heterodimers expressed by a particular subject are understood, improved, refined, and reliable methods for predicting peptides that can bind with high fidelity to specific HLA class II alpha and beta chain heterodimers, as described herein, can ensure that a specific immune response can be specifically generated for a particular purpose in a subject.

[0202]

[0419] In this application, the use of the singular form includes the plural unless specifically stated otherwise. It should be noted that when used in this specification, the singular forms "a", "an", and "the" include plural referents unless the context clearly indicates otherwise. In this application, the use of "or" means "and / or" unless stated otherwise. Further, the use of the terms "including" and other forms, such as "include", "includes", and "included", is not limiting. The terms "one or more" or "at least one", for example, one or more members of a group of members, or at least one member(s), are self-evident using further exemplification, and in particular the terms refer to any one of the said members, or any two or more of the said members, for example, any ≧3, ≧4, ≧5, ≧6, or ≧7 etc. of the said members, and up to and including all of the said members.

[0203]

[0420] Reference herein to "some embodiments", "an embodiment", "one embodiment", or "other embodiments" means that a characteristic, structure, or feature described in connection with the embodiments is included in at least some embodiments of the present disclosure but not necessarily in all embodiments.

[0204]

[0421] As used in this specification and the claims (if any), the word "comprising" (and any form of comprising such as "comprise" and "comprises"), "having" (and any form of having such as "have" and "has"), "including" (and any form of including such as "includes" and "include") or "containing" (and any form of containing such as "contains" and "contain") is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed herein can be practiced in relation to any method or composition of the present disclosure, and vice versa. Further, the compositions of the present disclosure can be used to achieve the methods of the present disclosure.

[0205]

[0422] As used herein, the terms "about" or "approximately" when referring to a measurable value such as a parameter, amount, period of time, etc. mean a variation of less than or equal to + / - 20%, less than or equal to + / - 10%, less than or equal to + / - 5% or less than or equal to + / - 1% of the specified value, as long as such a variation is appropriate for practicing the present disclosure. It is understood that the value itself to which the modifier "about" or "approximately" refers is also expressly disclosed.

[0206]

[0423] The term "immune response" includes T cell-mediated and / or B cell-mediated immune responses that are affected by the modulation of T cell costimulation. Exemplary immune responses include T cell responses such as cytokine production and cytotoxic activity. In addition, the term immune response includes immune responses that are indirectly affected by T cell activation such as antibody production (humoral response) and activation of cytokine-responsive cells such as macrophages.

[0207]

[0424] "Receptor" is understood to mean a biological molecule or group of molecules that can bind to a ligand. Receptors serve to transmit information in cells, cell formations or organisms. Receptors include at least one receptor unit and can contain two or more receptor units, where each receptor unit can consist of a protein molecule, e.g., a glycoprotein molecule. Receptors have a structure that complements the structure of a ligand and can complex the ligand as a binding partner. Signaling information can be transmitted by conformational changes of the receptor following binding with the ligand on the cell surface. According to the present disclosure, receptors can refer to proteins of MHC class I and II that can form a receptor / ligand complex with a ligand, e.g., a peptide or peptide fragment of suitable length. Class I and class II MHC peptides encoded by HLA class I and class II alleles are often referred to herein as HLA class I and HLA class II peptides, or HLA class I and HLA class II peptides, or HLA class I class II proteins, or HLA class I and HLA class II proteins, or HLA class I and class II molecules, or common variants thereof, as will be appreciated in the context of discussion by those of skill in the art.

[0208]

[0425] A "ligand" is a molecule capable of forming a complex with a receptor. According to the present disclosure, a ligand is understood to mean, for example, a peptide or peptide fragment having a suitable length and a suitable binding motif in its amino acid sequence, so that the peptide or peptide fragment can bind to and form a complex with a protein of MHC class I or MHC class II (i.e., HLA class I and HLA class II protein).

[0209]

[0426] An "antigen" is a molecule that can stimulate an immune response and may be produced by cancer cells, infectious pathogens, or autoimmune diseases. Antigens recognized by either T cells, helper T lymphocytes (helper T (TH) cells), or cytotoxic T lymphocytes (CTLs) are not recognized as intact proteins but rather as small peptides that associate with HLA class I or class II proteins on the cell surface. In the course of a naturally occurring immune response, antigens that associate with and are recognized by HLA class II molecules on antigen-presenting cells (APCs) are obtained extracellularly, internalized, and processed into small peptides that associate with HLA class II molecules. APCs can also cross-present peptide antigens by processing exogenous antigens and presenting the processed antigens on HLA class I molecules. Antigens that give rise to peptides recognized in association with HLA class I MHC molecules are generally peptides produced intracellularly, and these antigens are processed and associate with class I MHC molecules. It is now understood that peptides that associate with a given HLA class I or class II molecule are characterized as having a common binding motif, and the binding motifs for a large number of different HLA class I and II molecules have been determined. Synthetic peptides corresponding to the amino acid sequence of a given antigen and containing the binding motif for a given HLA class I or II molecule can also be synthesized. These peptides can then be added to appropriate APCs and the APCs can be used in vitro or in vivo to stimulate helper T cell or CTL responses. The binding motifs, methods for synthesizing the peptides, and methods for stimulating helper T cell or CTL responses are all well known and readily available to those of skill in the art.

[0210]

[0427] The term "peptide" is used interchangeably herein with "mutated peptide" and "neoantigen peptide". Similarly, the term "polypeptide" is used interchangeably herein with "mutated polypeptide" and "neoantigen polypeptide". By "neoantigen" or "neoepitope" is meant a class of tumor antigens or tumor epitopes that arise from tumor-specific mutations in expressed proteins. The present disclosure further includes peptides containing tumor-specific mutations, peptides containing well-known tumor-specific mutations, and mutated polypeptides or fragments thereof identified by the methods of the present disclosure. These peptides and polypeptides are referred to herein as "neoantigen peptides" or "neoantigen polypeptides". The polypeptides or peptides can be of various lengths, in their neutral (uncharged) form, or in the form of salts, and can have no modifications such as glycosylation, side-chain oxidation, phosphorylation or post-translational modifications, or can have these modifications, provided that the modifications are under conditions that do not destroy the biological activity of the polypeptides described herein. In some embodiments, the neoantigen peptides of the present disclosure are: for HLA class I, 22 residues or less in length, such as about 8 to about 22 residues, about 8 to about 15 residues, or 9 or 10 residues; for HLA class II, 40 residues or less in length, such as about 8 to about 40 residues in length, about 8 to about 24 residues in length, about 12 to about 19 residues, or about 14 to about 18 residues. In some embodiments, the neoantigen peptide or neoantigen polypeptide contains a neoepitope.

[0211]

[0428] The term "epitope" includes any protein determinant that can specifically bind to an antibody, antibody peptide, and / or antibody-like molecule (including, without limitation, T cell receptors) as defined herein. Epitope determinants typically consist of chemically active surface groups of molecules such as amino acids or sugar side chains, and generally have specific three-dimensional structural features and specific charge features.

[0212]

[0429] A "T cell epitope" is a peptide sequence that can be bound by class I or II MHC molecules in the form of peptide-presenting MHC molecules or MHC complexes, and in this form is recognized and bound by cytotoxic T lymphocytes or helper T cells, respectively.

[0213]

[0430] As used herein, the term "antibody" includes IgG (including IgG1, IgG2, IgG3 and IgG4), IgA (including IgA1 and IgA2), IgD, IgE, IgM as well as IgY, and includes whole antibodies including single-chain whole antibodies and their antigen-binding (Fab) fragments. Antigen-binding antibody fragments include, but are not limited to, Fab, Fab' and F(ab')2, Fd (consisting of VH and CH1), single-chain variable fragments (scFv), single-chain antibodies, disulfide-bonded variable fragments (dsFv) and fragments containing either a VL or VH domain. Antibodies can be derived from any animal. Antigen-binding antibody fragments including single-chain antibodies can include the variable region(s) alone or in combination with all or part of the following: hinge region, CH1, CH2 and CH3 domains. Also included are any combinations of the variable region(s) with the hinge region, CH1, CH2 and CH3 domains. Antibodies can be, for example, monoclonal, polyclonal, chimeric, humanized as well as human monoclonal and polyclonal antibodies that specifically bind to an HLA-related polypeptide or an HLA-HLA-binding peptide (HLA peptide) complex. Those skilled in the art will recognize that various immunoaffinity techniques are suitable for concentrating soluble proteins, such as soluble HLA peptide complexes, or membrane-bound HLA-related polypeptides, such as those proteolytically cleaved from the membrane. These include techniques in which (1) one or more antibodies capable of specifically binding to a soluble protein are immobilized on a stationary or mobile substrate (e.g., plastic well or resin, latex or paramagnetic beads), and (2) a solution containing soluble proteins from a biological sample is passed over the antibody-coated substrate such that the soluble proteins can bind to the antibodies. The substrate containing the antibody and the bound soluble protein is separated from the solution and, optionally, the antibody and soluble protein are dissociated, for example, by varying the pH and / or ionic strength and / or ionic composition of the solution in which the antibody is soaked. Alternatively, immunoprecipitation techniques can be used in which the antibody and soluble protein are combined to form a macromolecular aggregate. The macromolecular aggregate can be separated from the solution by size exclusion techniques or by centrifugation.

[0214]

[0431] The term "immunopurification (IP)" (or immunoaffinity purification or immunoprecipitation) is a process known in the art and is widely used for the isolation of a desired antigen from a sample. Generally, the process involves contacting a sample containing the desired antigen with an affinity matrix to which an antibody against the antigen is covalently bound to a solid phase. The antigen in the sample will bind to the affinity matrix through immunochemical binding. The affinity matrix is then washed to remove all unbound molecular species. The antigen is removed from the affinity matrix by changing the chemical composition of the solution that contacts the affinity matrix. Immunopurification may be performed on a column containing the affinity matrix, in which case the solution is the eluate. Alternatively, when the affinity matrix is maintained as a suspension in a solution, immunopurification may be in a batch process. An important step in the process is the removal of the antigen from the matrix. This is generally achieved, for example, by increasing the ionic strength of the solution that contacts the affinity matrix by the addition of inorganic salts. A change in pH may also be effective in dissociating the immunochemical bond between the antigen and the affinity matrix.

[0215]

[0432] An "agent" is any small molecule chemical, antibody, nucleic acid molecule or polypeptide or fragment thereof.

[0433] A "change" or "variation" is an increase or a decrease. The change can be of the order of as little as 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30% or up to 40%, 50%, 60% or even, 70%, 75%, 80%, 90% or 100%.

[0216]

[0434] "Biological sample" is any tissue, cell, body fluid or other substance derived from an organism. As used herein, the term "sample" includes biological samples, e.g., any tissue, cell, body fluid or other substance derived from an organism. "Specifically bind" refers to a compound (e.g., a peptide) that recognizes and binds to a molecule (e.g., a polypeptide), but does not substantially recognize and bind to other molecules in a sample, e.g., a biological sample.

[0217]

[0435] "Capture reagent" refers to a reagent that specifically binds to a molecule (e.g., a nucleic acid molecule or a polypeptide) to select or isolate the molecule (e.g., a nucleic acid molecule or a polypeptide).

[0218]

[0436] As used herein, the terms "determine", "evaluate", "assay", "measure", "detect" and their grammatical equivalents refer to both quantitative and qualitative determinations, and for this reason the term "determine" is used interchangeably herein with "assay", "measure", etc. When a quantitative determination is intended, the phrase "determine the amount of" an analyte, etc. is used. When a qualitative and / or quantitative determination is intended, the phrases "determine the level of" an analyte or "determine" an analyte are used.

[0219]

[0437] "Fragment" is a portion of a protein or nucleic acid that is substantially identical to a reference protein or nucleic acid. In some embodiments, the portion retains at least 50%, 75%, 80%, 90%, 95% or even 99% of the biological activity of the reference protein or nucleic acid described herein.

[0220]

[0438] The terms "isolated", "purified", "biologically pure" and their grammatical equivalents refer to a substance that contains to varying degrees no components that are normally associated with it as found in nature. "Isolated" describes the degree of separation from its original source or environment. "Purified" describes a higher degree of separation than isolated. A "purified" or "biologically pure" protein substantially contains no other substances such that all impurities do not substantially affect the biological properties of the protein or cause other adverse results. That is, a nucleic acid or peptide of the present disclosure is purified when it is produced by recombinant DNA technology and substantially contains no cellular material, viral material or culture medium, or when chemically synthesized and substantially contains no chemical precursors or other chemicals. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" may describe that a nucleic acid or protein yields essentially one band in electrophoresis. For proteins that can be subjected to modifications such as phosphorylation or glycosylation, different modifications can result in different isolated proteins that can be purified separately.

[0221]

[0439] An "isolated" polypeptide (e.g., a peptide derived from an HLA peptide complex) or polypeptide complex (e.g., an HLA peptide complex) is a polypeptide or polypeptide complex of the present disclosure that has been separated from its natural, accompanying components. Typically, a polypeptide or polypeptide complex is isolated when it contains less than 60% by weight of the proteins and naturally occurring organic molecules that naturally accompany it. The preparation can be at least 75%, at least 90% or at least 99% by weight the polypeptide or polypeptide complex of the present disclosure. The isolated polypeptide or polypeptide complex of the present disclosure can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding one or more components of such a polypeptide or polypeptide complex, or by chemically synthesizing one or more components of the polypeptide or polypeptide complex. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis. In some cases, an HLA-encoded MHC class II protein (i.e., an MHC class II peptide) is referred to interchangeably in this document as an HLA class II protein (or an HLA class II peptide).

[0222]

[0440] The term "vector" refers to a nucleic acid molecule capable of mediating the transport or expression of heterologous nucleic acids. A plasmid is a molecular species of the type encompassed by the term "vector". Typically, a vector refers to a nucleic acid sequence containing an origin of replication and other entities necessary for replication and / or maintenance in a host cell. A vector capable of directing the expression of an operably linked gene and / or nucleic acid sequence is referred to herein as an "expression vector". Generally, useful expression vectors are often in the form of "plasmids", which refer to circular double-stranded DNA molecules, and in their vector form, they do not bind to chromosomes and typically contain entities for stable or transient expression of the encoded DNA. Other expression vectors that can be used in the methods disclosed herein include, but are not limited to, plasmids, episomes, bacterial artificial chromosomes, yeast artificial chromosomes, bacteriophages or viral vectors, and such vectors can be integrated into the host genome or replicate autonomously in the cell. A vector can be a DNA or RNA vector. Other forms of expression vectors well known to those skilled in the art that perform equivalent functions, such as self-replicating extrachromosomal vectors or vectors that can be integrated into the host genome, can also be used. Exemplary vectors are those capable of autonomous replication and / or expression of the nucleic acids to which they are linked.

[0223]

[0441] When used with reference to a fusion protein, the term "spacer" or "linker" refers to a peptide that links proteins including the fusion protein. Generally, a spacer has no specific biological activity other than linking between proteins or RNA sequences or maintaining a certain minimum distance or other spatial relationship. However, in some embodiments, the constituent amino acids of the spacer can be selected to affect some properties such as the folding of the molecule, net charge or hydrophobicity. Linkers suitable for use in the embodiments of the present disclosure are known to those skilled in the art and include, but are not limited to, linear or branched carbon linkers, heterocyclic carbon linkers or peptide linkers. The linker is used in some embodiments to separate two antigenic peptides at a distance sufficient to ensure that each antigenic peptide folds properly. Exemplary peptide linker sequences adopt a flexible and extended conformation and show no tendency to generate an indicated secondary structure. Typical amino acids in flexible protein regions include Gly, Asn and Ser. Substantially any permutation of amino acid sequences containing Gly, Asn and Ser is predicted to meet the above criteria for linker sequences. Other near-neutral amino acids such as Thr and Ala can also be used in linker sequences. Still other amino acid sequences that can be used as linkers are disclosed in Maratea et al., (1985), Gene 40:39-46; Murphy et al., (1986) Proc. Nat’l. Acad. Sci. USA 83:8258-62; U.S. Patent No. 4,935,233; and U.S. Patent No. 4,751,180.

[0224]

[0442] The term "neoplasm" refers to any disease that is caused by, or gives rise to, unduly high levels of cell division, unduly low levels of apoptosis, or both. Glioblastoma is a non-limiting example of a neoplasm or cancer. The terms "cancer" or "tumor" or "hyperproliferative disorder" refer to the presence of cells having characteristics typical of cells that give rise to cancer, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells may exist alone in an animal, or may be non-tumorigenic cancer cells, such as leukemia cells. Cancer includes, but is not limited to, B-cell cancers (e.g., multiple myeloma, Waldenström's macroglobulinemia), heavy chain diseases (e.g., alpha chain disease, gamma chain disease, and mu chain disease, etc.), benign M proteinemia and immunocyte amyloidosis, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer (e.g., metastatic, hormone-refractory prostate cancer), pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or appendix cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, cancers of the blood tissues, etc.Other non-limiting examples of cancer types applicable to the methods encompassed by the present disclosure include human sarcomas and carcinomas, such as fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, cystadenocarcinoma, medullary carcinoma, bronchogenic lung cancer, renal cell carcinoma, hepatoma, cholangiocarcinoma, liver cancer, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, bone cancer, brain tumor, testicular cancer, lung cancer, small cell lung cancer, bladder cancer, epithelial cancer, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, retinoblastoma; leukemias, such as acute lymphocytic leukemia and acute myeloid leukemia (myeloblastic, promyelocytic, myelomonocytic, monocytic and erythroleukemia); chronic leukemia (chronic myeloid (granulocytic) leukemia and chronic lymphocytic leukemia); and polycythemia vera, lymphomas (Hodgkin's disease and non-Hodgkin's disease), multiple myeloma, Waldenström's macroglobulinemia and heavy chain disease. In some embodiments, the cancer is an epithelial cancer, such as, but not limited to, bladder cancer, breast cancer, cervical cancer, colon cancer, gynecological cancer, kidney cancer, laryngeal cancer, lung cancer, oral cancer, head and neck cancer, ovarian cancer, pancreatic cancer, prostate cancer or skin cancer. In other embodiments, the cancer is breast cancer, prostate cancer, lung cancer or colon cancer. In yet other embodiments, the epithelial cancer is non-small cell lung cancer, non-papillary renal cell carcinoma, cervical cancer, ovarian cancer (e.g., serous ovarian cancer) or breast cancer. The epithelial cancer can be characterized in various other ways including, but not limited to, serous, endometrioid, mucinous, clear cell, Brenner or undifferentiated. In some embodiments, the present disclosure is used in the treatment, diagnosis and / or prognosis of lymphomas or subtypes thereof including, but not limited to, mantle cell lymphoma. Lymphoproliferative disorders are also considered to be proliferative diseases.

[0225]

[0443] The term "vaccine" is understood to mean a composition for generating immunity for the prevention and / or treatment of a disease (e.g., neoplasm / tumor / infectious pathogen / autoimmune disease). Thus, a vaccine is a pharmaceutical that contains an antigen and is intended to be used in humans or animals to generate specific defensive and protective substances by vaccination. A "vaccine composition" may contain a pharmaceutically acceptable excipient, carrier or diluent. Aspects of the present disclosure relate to the use of techniques in the preparation of antigen-based vaccines. In these embodiments, a vaccine is meant to refer to one or more disease-specific antigenic peptides (or the corresponding nucleic acids encoding them). In some embodiments, the antigen-based vaccine contains at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30 or more antigenic peptides.In some embodiments, the antigen-based vaccine contains 2 to 100, 2 to 75, 2 to 50, 2 to 25, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 3 to 100, 3 to 75, 3 to 50, 3 to 25, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 4 to 100, 4 to 75, 4 to 50, 4 to 25, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 5 to 100, 5 to 75, 5 to 50, 5 to 25, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 10, 5 to 9, 5 to 8, or 5 to 7 antigenic peptides. In some embodiments, the antigen-based vaccine contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 antigenic peptides. In some cases, the antigenic peptide is a neoantigenic peptide. In some cases, the antigenic peptide contains one or more neoepitopes.

[0226]

[0444] The term "pharmaceutically acceptable" refers to being approved or approvable by a regulatory agency of the federal or state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeias for use in animals including humans. "Pharmaceutically acceptable excipients, carriers or diluents" refer to excipients, carriers or diluents that can be administered with an agent to a subject, which, when administered in a dosage sufficient to deliver a therapeutically effective amount of the agent, do not destroy the pharmacological activity and are non-toxic. "Pharmaceutically acceptable salts" of the pooled disease-specific antigens listed herein may be acidic or basic salts that are generally considered in the art to be suitable for use in contact with human or animal tissues without undue toxicity, irritation, allergic response or other problems or complications. Such salts include salts of inorganic and organic acids with basic residues such as amines, and salts of alkalis or organic bases with acidic residues such as carboxylic acids. Specific pharmaceutical salts include, but are not limited to, salts of hydrochloric acid, phosphoric acid, hydrobromic acid, malic acid, glycolic acid, fumaric acid, sulfuric acid, sulfamic acid, sulfanilic acid, formic acid, toluenesulfonic acid, methanesulfonic acid, benzenesulfonic acid, ethanedisulfonic acid, 2-hydroxyethylsulfonic acid, nitric acid, benzoic acid, 2-acetoxybenzoic acid, citric acid, tartaric acid, lactic acid, stearic acid, salicylic acid, glutamic acid, ascorbic acid, pamoic acid, succinic acid, fumaric acid, maleic acid, propionic acid, hydroxymaleic acid, hydroiodic acid, phenylacetic acid, alkanoic acids such as acetic acid, HOOC-(CH2)n-COOH, where n is from 0 to 4, etc. Similarly, pharmaceutically acceptable cations include, but are not limited to, sodium, potassium, calcium, aluminum, lithium and ammonium. Those skilled in the art will understand from the present disclosure and knowledge in the art further pharmaceutically acceptable salts for the pooled disease-specific antigens provided herein, including those listed by Remington’s Pharmaceutical Sciences, 17th ed., Mack Publishing Company, Easton, PA, page 1418 (1985).Generally, pharmaceutically acceptable acid or base salts can be synthesized by any conventional chemical method from parent compounds containing basic or acidic moieties. Briefly, such salts can be prepared by reacting the free acid or base form of these compounds with a stoichiometric amount of the appropriate base or acid in a suitable solvent.

[0227]

[0445] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. Polynucleotides having substantial identity to the endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. "Hybridize" refers to the situation where a pair of nucleic acid molecules forms a double-stranded molecule with a complementary polynucleotide sequence or a part thereof under various stringency conditions (see, for example, Wahl, G.M. and S.L.Berger (1987) Methods Enzymol. 152:399; Kimmel, A.R. (1987) Methods Enzymol. 152:507). For example, stringent salt concentrations can typically be less than about 750 mM NaCl and 75 mM trisodium citrate, less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of an organic solvent, such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, or at least about 50% formamide. Stringent temperature conditions typically include a temperature of at least about 30°C, at least about 37°C, or at least about 42°C. It is known to those skilled in the art to vary additional parameters such as hybridization time, the concentration of a surfactant, such as sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA. Various levels of stringency are achieved by combining these various conditions as needed. In an exemplary embodiment, hybridization can occur at 30°C, 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another exemplary embodiment, hybridization can occur at 37°C, 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA).In another exemplary embodiment, hybridization can occur at 42°C, 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide and 200 μg / ml ssDNA. Useful variations under these conditions will be readily apparent to those skilled in the art. For most applications, the wash step after hybridization can also vary in stringency. Wash stringency conditions can be defined by salt concentration and by temperature. As noted above, wash stringency can be increased by decreasing the salt concentration or by increasing the temperature. For example, a stringent salt concentration for the wash step can be less than about 30 mM NaCl and 3 mM trisodium citrate or less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash step can include a temperature of at least about 25°C, at least about 42°C or at least about 68°C. In an exemplary embodiment, the wash step can be performed at 25°C, 30 mM NaCl, 3 mM trisodium citrate and 0.1% SDS. In another exemplary embodiment, the wash step can be performed at 42°C, 15 mM NaCl, 1.5 mM trisodium citrate and 0.1% SDS. In another exemplary embodiment, the wash step can be performed at 68°C, 15 mM NaCl, 1.5 mM trisodium citrate and 0.1% SDS. Additional variations under these conditions will be readily apparent to those skilled in the art.Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0228]

[0446] "Substantially identical" refers to a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). Such sequences can be at least 60%, 80% or 85%, 90%, 95%, 96%, 97%, 98% or even 99% or more identical at the amino acid level or in nucleic acids to the sequences used for comparison. Sequence identity is typically measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by specifying the degree of homology for various substitutions, deletions and / or other modifications. Typically conservative substitutions include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach for determining the degree of identity, the BLAST program can be used with probability scores between e-3 and e-m 〇 to indicate closely related sequences. "Reference" is the standard for comparison.

[0229]

[0447] The term "subject" or "patient" refers to an animal that is the subject of treatment, observation or experiment. By way of example only, subjects include, but are not limited to, mammals, including, but not limited to, humans or non-human mammals such as non-human primates, mice, cows, horses, dogs, sheep or cats.

[0230]

[0448] The terms "treat", "treated", "treating", "treatment", etc. mean to reduce, prevent or ameliorate a disorder and / or a symptom associated therewith (e.g., a neoplasm or tumor or infectious agent or autoimmune disease). "Treatment" can refer to administering therapy to a subject after the onset or suspected onset of a disease (e.g., cancer or infection by an infectious agent or autoimmune disease). "Treatment" includes the concept of "mitigating" which refers to reducing the frequency of onset or recurrence, or severity of any symptom or other adverse effect associated with the disease, and / or side effects associated with the treatment. The term "treatment" also encompasses the concept of "managing" which refers to reducing the severity of a disease or disorder in a patient, e.g., extending the life or increasing the likelihood of survival of a patient having the disease, or delaying its recurrence, e.g., extending the period of remission in a patient suffering from the disease. It is understood that treating a disorder or condition does not require that the disorder, condition or symptom associated therewith be completely eliminated, although this is not excluded.

[0231]

[0449] As used herein, the terms "prevent", "preventing", "prevention" and their grammatical equivalents mean avoiding or delaying the onset of symptoms associated with a disease or condition in a subject that is not symptomatic at the time of initiation of administration of an agent or compound.

[0232]

[0450] The term "therapeutic effect" refers to reducing to some extent one or more symptoms of a disorder (e.g., a neoplasm, tumor or infection by an infectious agent or autoimmune disease) or its attendant pathological conditions. As used herein, a "therapeutically effective amount" refers to the amount of an agent in a single or multiple dosages administered to a cell or subject that is effective to extend the survival likelihood of a patient having such a disorder, reduce, prevent or delay one or more signs or symptoms of such a disorder, beyond what would be predicted in the absence of such treatment. A "therapeutically effective amount" is intended to qualify the amount necessary to achieve a therapeutic effect. A physician or veterinarian having ordinary skill in the art can readily determine and prescribe the "therapeutically effective amount" (e.g., ED50) of the required pharmaceutical composition. For example, a physician or veterinarian can start with a level of the dosage of the compounds of the disclosure used in the pharmaceutical composition that is lower than that required to achieve the desired therapeutic effect and gradually increase the dosage until the desired effect is achieved. Diseases, conditions and disorders are used interchangeably herein.

[0233]

[0451] One skilled in the art will recognize that the terms "peptide tag", "affinity tag", "epitope tag", or "affinity acceptor tag" are used interchangeably herein. As used herein, the term "affinity acceptor tag" refers to an amino acid sequence that allows a tagged protein to be readily detected or purified, e.g., by affinity purification. Affinity acceptor tags are generally (but not necessarily) positioned at or near the N- or C-terminus of an HLA allele. A variety of peptide tags are known in the art. Non-limiting examples include polyhistidine tags (e.g., 4 to 15 contiguous His residues (SEQ ID NO: 4), e.g., 8 contiguous His residues (SEQ ID NO: 5)); poly-histidine-glycine tags; HA tags (e.g., Field et al., Mol. Cell. Biol., 8:2159, 1988); c-myc tags (e.g., Evans et al., Mol. Cell. Biol. 5:3610, 1985); herpes simplex virus glycoprotein D (gD) tags (e.g., Paborsky et al., Protein Engineering, 3:547, 1990); FLAG tags (e.g., Hopp et al., BioTechnology, 6:1204, 1988; U.S. Pat. Nos. 4,703,004 and 4,851,341); KT3 epitope tags (e.g., Martine et al., Science, 255:192, 1992); tubulin epitope tags (e.g., Skinner, Biol. Chem., 266:15173, 1991); T7 gene 10 protein peptide tags (e.g., Lutz-Freyemuth et al., Proc. Natl. Acad. Sci. USA, 87:6393, 1990); streptavidin tags (StrepTag™ or StrepTagII™; see, e.g., Schmidt et al., J. Mol. Biol., 255(5):753-766, 1996 or U.S. Pat. No. 5,506,121; commercially available from Sigma-Genosys); or vesicular stomatitis virus glycoprotein-derived VSV-G epitope tags; or V5 tags derived from a small epitope (Pk) found in the P and V proteins of the paramyxovirus of simian virus 5 (SV5).In some embodiments, an affinity acceptor tag is an "epitope tag", a type of peptide tag that adds a recognizable epitope (antibody binding site) to an HLA protein to provide binding of a corresponding antibody, thereby enabling identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are Protein A or Protein G that bind to IgG. In some embodiments, the matrix of IgG Sepharose 6 Fast Flow chromatography resin is covalently bound to human IgG. This resin allows for high flow rates for rapid and convenient purification of proteins having a Protein A tag. Numerous other tag components are well known to those skilled in the art, are contemplated, and are intended herein. Any peptide tag can be used as long as it can be expressed as an element of an HLA peptide complex with an affinity acceptor tag.

[0234]

[0452] As used herein, the term "affinity molecule" refers to a molecule or ligand that binds to an affinity acceptor peptide with chemical specificity. Chemical specificity is the ability of the binding site of a protein to bind to a particular ligand. The fewer ligands a protein can bind to, the higher its specificity. Specificity describes the strength of the binding between a given protein and ligand. This relationship can be described by the dissociation constant (KD) that characterizes the equilibrium between bound and unbound states of a protein-ligand system.

[0235]

[0453] The term "HLA peptide complex with an affinity acceptor tag" refers to a complex that includes an affinity acceptor peptide and includes a portion thereof that specifically binds to an HLA class I or class II-related peptide or a single allele recombinant HLA class I or class II peptide.

[0236]

[0454] As used herein, the terms "specific binding" or "specifically binds," when used with reference to the interaction between an affinity molecule and an affinity acceptor tag, or between an epitope and an HLA peptide, means that the interaction depends on the presence of a specific structure on the protein (e.g., an antigenic determinant or epitope); in other words, the affinity molecule generally recognizes and binds to a specific affinity acceptor peptide structure rather than to the protein in general.

[0237]

[0455] As used herein, the term "affinity" refers to a measure of the strength of binding between two members of a binding pair, e.g., between an "affinity acceptor tag" and an "affinity molecule" and between an HLA-binding peptide and an HLA class I or II molecule. KD is the dissociation constant and has units of moles. The affinity constant is the reciprocal of the dissociation constant. The affinity constant is sometimes used as a generic term to describe this chemical substance. This is a direct measure of the energy of binding. Affinity can be determined experimentally, e.g., by surface plasmon resonance (SPR) using a commercially available Biacore SPR apparatus. Affinity can also be expressed as the concentration at which 50% of the peptide is displaced, the inhibitory concentration 50 (IC50). Similarly, lnIC50 refers to the natural logarithm of the IC50. Koff refers to the off-rate constant, e.g., for the dissociation of an affinity molecule from an HLA peptide complex with an affinity acceptor tag.

[0238]

[0456] In some embodiments, the HLA peptide complex with an affinity acceptor tag includes a biotin acceptor peptide (BAP) and is immunopurified from complex cell mixtures using streptavidin / NeutrAvidin beads. The biotin-avidin / streptavidin binding is one of the strongest non-covalent interactions known in nature. This property is utilized as a biological tool for a wide range of applications such as the immunopurification of proteins to which biotin is covalently attached. In an exemplary embodiment, the nucleic acid sequence encoding the HLA allele incorporates a biotin acceptor peptide (BAP) as an affinity acceptor tag for immunopurification. BAP can be specifically biotinylated in vivo or in vitro at a single lysine residue within the tag (e.g., U.S. Pat. Nos. 5,723,584; 5,874,239; and 5,932,433; and UK Patent GB2370039). BAP is typically 15 amino acids in length and contains a single lysine as the biotin acceptor residue. In some embodiments, BAP is positioned at or near the N or C terminus of a single allele HLA peptide. In some embodiments, BAP is positioned between the heavy chain domain and the β2-microglobulin domain of an HLA class I peptide. In some embodiments, BAP is positioned between the β-chain domain and the α-chain domain of an HLA class II peptide. In some embodiments, BAP is positioned in the loop regions between the α1, α2, and α3 domains of the heavy chain of HLA class I, or between the α1 and α2 and β1 and β2 domains of the α-chain and β-chain of HLA class II, respectively. Exemplary constructs designed for HLA class I and II expression incorporating BAP for biotinylation and immunopurification are described in FIG. 2.

[0239]

[0457] As used herein, the term "biotin" refers to the compound biotin itself and its analogs, derivatives, and variants. Thus, the term "biotin" includes biotin (cis-hexahydro-2-oxo-1H-thieno[3,4]imidazole-4-pentanoic acid) and any of its derivatives and analogs, including biotin-like compounds. Such compounds include, for example, biotin-e-N-lysine, biocytin hydrazide, 2-iminobiotin, and amino or sulfhydryl derivatives of biotinyl-E-aminocaproic acid-N-hydroxysuccinimide ester, sulfosuccinimidyl iminobiotin, biotin bromoacetyl hydrazide, p-diazobenzoyl biocytin, 3-(N-maleimidopropionyl)biocytin, desthiobiotin, and the like. The term "biotin" also includes biotin variants that can specifically bind to one or more of rhizavidin, avidin, streptavidin, the tamavidin moiety, or other avidin-like peptides.

[0240]

[0458] As used herein, "PPV determination method" may refer to a presented PPV determination method. For example, the "PPV determination method" is a step of processing amino acid information of a plurality of test peptide sequences using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, to generate a plurality of test presentation predictions, where each test presentation prediction indicates the possibility that one or more proteins encoded by the class II HLA allele of a cell, e.g., the class II HLA allele of a subject's cell, can present a given test peptide sequence of the plurality of test peptide sequences. The plurality of test peptide sequences includes at least 500 test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained within a protein encoded by the genome of an organism, such organism being of the same species as the subject. The plurality of test peptide sequences has a ratio of less than 1 of the number of decoy peptide sequences to the number of hit peptide sequences, e.g., a ratio of 1:499 of at least one hit peptide sequence to at least 499 decoy peptide sequences; (b) identifying or calling the top percentage, e.g., the top 0.2%, of the plurality of test peptide sequences as being presented by the class II HLA allele of the cell; and (c) calculating the PPV of the HLA peptide presentation prediction model, where the PPV is the ratio of the plurality of test peptide sequences identified or called as being presented by the class II HLA allele of the cell that are peptides observed by mass spectrometry to be presented by the class II HLA allele of the cell. In some embodiments, the decoy peptides are of the same length, i.e., contain the same number of amino acids as the hit peptides. In some embodiments, the decoy peptides may contain one more or one less amino acid compared to the hit peptides. In some embodiments, the decoy peptides are peptides that are endogenous peptides. In some embodiments, the decoy peptides are synthetic peptides. In some embodiments, the decoy peptides areAn endogenous peptide identified by mass spectrometry as binding to a first MHC class I or class II protein, where the first MHC class I or class II protein is different from a second MHC class I or class II protein that binds to the hit peptide. In some embodiments, the decoy peptide may be a scrambled peptide. For example, the decoy peptide may include an amino acid sequence in which the amino acid positions are reconfigured within the length of the peptide compared to those of the hit peptide. In some embodiments, the PPV determination method may be a presentation PPV determination method. In some embodiments, the ratio of the number of hit peptide sequences to the number of decoy peptide sequences is about 1:10, 1:20, 1:50, 1:100, 1:250, 1:500, 1:1000, 1:1500, 1:2000, 1:2500, 1:5000, 1:7500, 1:10000, 1:25000, 1:50000 or 1:100000. In some embodiments, at least one hit peptide sequence includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences. In some embodiments, at least 499 decoy peptide sequences include at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700,It contains 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000 decoy peptide sequences. In some embodiments, at least 500 test peptide sequences are at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000,It contains 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000 test peptide sequences. In some embodiments, identifying or calling the top percentage of a plurality of test peptide sequences as being presented by the class II HLA alleles of the cells is the top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20% as being presented by the class II HLA alleles of the cells,Identifying or calling for 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%, 9.90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20%. In some embodiments, the cell is a single allele cell.,

[0241]

[0459] As used herein, "PPV determination method" may refer to a combined PPV determination method. For example, the "PPV determination method" is a step of processing amino acid information of a plurality of test peptide sequences using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, to generate a plurality of test presentation predictions, wherein each test binding presentation prediction indicates the possibility that one or more proteins encoded by a class II HLA allele of a cell, e.g., a class II HLA allele of a subject's cell, can bind to a given test peptide sequence of the plurality of test peptide sequences, and the plurality of test peptide sequences includes at least 20 test peptide sequences including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 19 decoy peptide sequences contained within a protein including at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and the plurality of test peptide sequences has a ratio of the number of hit peptide sequences to the number of decoy peptide sequences of less than 1, e.g., a ratio of 1:19 of at least one hit peptide sequence to at least 19 decoy peptide sequences; (b) identifying or calling the top percentage, e.g., the top 5%, of the plurality of test peptide sequences as binding to the HLA protein; and (c) calculating the PPV of the HLA peptide binding prediction model, where the PPV is the percentage of the plurality of test peptide sequences identified or called as binding to the class II HLA allele of the cell that are peptides observed by mass spectrometry as being presented by the class II HLA allele of the cell. In some embodiments, the ratio of the number of hit peptide sequences to the number of decoy peptide sequences is about 1:2, 1:3, 1:4, 1:5, 1:10, 1:20, 1:25, 1:30, 1:40, 1:50, 1:75, 1:100, 1:200, 1:250, 1:500 or 1:1000. In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20,It includes 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 hit peptide sequences. In some embodiments, at least 19 decoy peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000,It contains 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 31,000, 32,000, 33,000, 34,000, 35,000, 36,000, 37,000, 38,000, 39,000, 40,000, 41,000, 42,000, 43,000, 44,000, 45,000, 46,000, 47,000, 48,000, 49,000, 50,000, 52,500, 55,000, 57,500, 60,000, 62,500, 65,000, 67,500, 70,000, 72,500, 75,000, 77,500, 80,000, 82,500, 85,000, 87,500, 90,000, 92,500, 95,000, 97,500, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 325,000, 350,000, 375,000, 400,000, 425,000, 450,000, 475,000, 500,000, 600,000, 700,000, 800,000, 900,000 or 1,000,000 decoy peptide sequences. In some embodiments, at least 20 test peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600,Comprising 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000 test peptide sequences. In some embodiments, identifying or calling the top percentage of a plurality of test peptide sequences as presented by a class II HLA allele of a cell comprises identifying or calling the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39% or 40% as presented by a class II HLA allele of a cell. In some embodiments, the cell is a single allele cell.,

[0242] Human leukocyte antigen (HLA) system

[0460] The immune system can be classified into two functional subsystems: innate and adaptive immune systems. The innate immune system is the front line of defense against infection, and most potential pathogens are rapidly neutralized by this system before they can cause, for example, a significant infection. The adaptive immune system reacts to molecular structures called antigens of the invading organisms. Unlike the innate immune system, the adaptive immune system is highly specific to pathogens. Acquired immunity can also provide long-term protection; for example, individuals who have recovered from measles are protected from measles for the rest of their lives. There are two types of adaptive immune responses, including humoral and cell-mediated immune responses. In the humoral immune response, antibodies secreted into the body fluids by B cells bind to pathogen-derived antigens, leading to the elimination of pathogens through various mechanisms, such as complement-mediated lysis. In the cell-mediated immune response, T cells that can destroy other cells are activated. For example, when disease-related proteins are present in cells, they are proteolytically fragmented into peptides intracellularly. Certain cell proteins then attach themselves to the antigens or antigens or peptides formed in this manner and transport them to the surface of the cells where they are presented to the molecular defense mechanisms in the body's T cells. Cytotoxic T cells recognize these antigens and kill the cells bearing the antigens.

[0243]

[0461] The terms "major histocompatibility complex (MHC)", "MHC molecule", or "MHC protein" refer to proteins that bind to peptides resulting from proteolytic cleavage of protein antigens, present potential T cell epitopes, transport them to the cell surface, and can present peptides to specific cells, such as cytotoxic T lymphocytes or helper T cells. Human MHC is also called the HLA complex. Thus, the terms "human leukocyte antigen (HLA) system", "HLA molecule", or "HLA protein" refer to the gene complex that encodes MHC proteins in humans. The term MHC is referred to as the "H-2" complex in the mouse species. Those skilled in the art will recognize that the terms "major histocompatibility complex (MHC)", "MHC molecule", "MHC protein" and "human leukocyte antigen (HLA) system", "HLA molecule", "HLA protein" are used interchangeably herein.

[0244]

[0462] HLA proteins are classified into two types called HLA class I and HLA class II. The structures of the proteins of the two HLA classes are very similar, but they have very different functions. HLA class I proteins are present on the surface of most cells of the body, including most tumor cells. HLA class I proteins are loaded with antigens that normally arise from endogenous proteins or pathogens present inside the cell and are then presented to naive or cytotoxic T lymphocytes (CTLs). HLA class II proteins are present in antigen-presenting cells (APCs), including but not limited to dendritic cells, B cells, and macrophages. They mainly present peptides processed from external, e.g., extracellular antigen sources, to helper T cells. Most peptides that bind to HLA class I proteins are derived from cytoplasmic proteins produced in healthy host cells of the organism itself and usually do not stimulate an immune response.

[0245]

[0463] The HLA class I molecule (Figure 1) consists of two non-covalently linked polypeptide chains, the HLA-encoded α chain (heavy chain, 44 to 47 kD) and a non-HLA-encoded subunit called β2-microglobulin (or β2m), (12 kD). The α chain has three extracellular domains, α1, α2, and α3, as well as a transmembrane region, and its α1 and α2 regions can bind to peptides of about 7 to 13 amino acids (e.g., about 8 to 11 amino acids or 9 or 10 amino acids). HLA class I molecules bind to peptides having a suitable binding motif and present them to cytotoxic T lymphocytes. The HLA class I heavy chain may be the protein product of the HLA-A allele, also referred to as the HLA-A monomer, or the protein product of the HLA-B allele (similarly, the HLA-B monomer), or the protein product of the HLA-C allele (the HLA-C monomer), each of which complexes with β-2-microglobulin. α1 is present on the non-HLA protein β2m; β2m is encoded by the beta-2-microglobulin gene located on human chromosome 15. The α3 domain is connected to the transmembrane region and anchors the HLA class I molecule to the cell membrane. The presented peptide is held in the floor of the peptide-binding groove in the central region of the α1 / α2 heterodimer (a molecule composed of two non-identical subunits). HLA class I-A, HLA class I-B, or HLA class I-C are highly polymorphic. Each of the HLA class 1-A gene (named the HLA-A gene), the HLA class 1-B gene (named the HLA-B gene), and the HLA class 1-C gene (named the HLA-C gene) contains eight exons. Exon 1 encodes the leader peptide, exons 2 and 3 encode the α1 and α2 domains, exon 5 encodes the transmembrane region, and exons 6 and 7 encode the cytoplasmic tail. The polymorphisms in exons 2 and 3 are responsible for the peptide-binding specificity of each class 1 molecule. The HLA class I-B gene (HLA-B) has a large number of possible diversities, expression patterns, and presented antigens.This group is subdivided into the group encoded within the HLA locus, for example, HLA-E, HLA-F, HLA-G and those not, such as stress ligands like ULBP, Rae1 and H60. The antigens / ligands for many of these molecules remain unknown, but they can interact with CD8+ T cells, NKT cells and NK cells respectively.

[0246]

[0464] In some embodiments, the present disclosure utilizes non-classical HLA class I-E alleles. HLA-E molecules are recognized by natural killer (NK) cells and CD8+ T cells. HLA-E is expressed in almost all tissues including lung, liver, skin and placental cells. HLA-E expression is also detected in solid tumors (e.g., osteosarcoma and melanoma). HLA-E molecules bind to TCRs expressed on CD8+ T cells, resulting in T cell activation. It is also well-known that HLA-E binds to CD94 / NKG2 receptors expressed on NK cells and CD8+ T cells. CD94 can pair with several different isoforms of NKG2 to form receptors that can either inhibit (NKG2A, NKG2B) or promote (NKG2C) cell activation. HLA-E can bind to peptides derived from amino acid residues 3-11 of the leader sequences of most HLA-A, -B, -C, and -G molecules, but not to its own leader peptide. HLA-E has also been shown to present peptides derived from endogenous proteins similar to HLA-A, -B or -C alleles. Association of CD94 / NKG2A with HLA-E loaded with peptides derived from HLA class I leader sequences under physiological conditions usually induces inhibitory signals. Cytomegalovirus (CMV) utilizes a mechanism for evasion from NK cell immunosurveillance through the expression of the UL40 glycoprotein, mimicking the HLA-A leader. However, it has also been reported that CD8+ T cells can recognize HLA-E loaded with UL40 peptides derived from the CMV Toledo strain and play a role in defense against CMV. A number of studies have revealed some important functions of HLA-E in infectious diseases and cancer.

[0247]

[0465] Peptide antigens attach themselves to HLA class I molecules by competitive affinity binding in the endoplasmic reticulum before they are presented on the cell surface. In the present specification, the affinity of individual peptide antigens is directly related to their amino acid sequence and the presence of specific binding motifs at defined positions within the amino acid sequence. When the sequence of such a peptide is well known, for example, it is possible to manipulate the immune system against diseased cells using peptide vaccines.

[0248]

[0466] MHC molecules are highly polymorphic, i.e., there are a large number of MHC variants. Each variant is encoded by the diversity of genes encoding proteins, and each such variant gene is called an allele. In humans, MHC is well known as human leukocyte antigen (HLA) including three types of HLA class II molecules: DP, DQ, and DR. HLA class II peptides (Figure 1) have two chains, α and β, each having two domains - α1 and α2 and β1 and β2 - each chain having transmembrane domains, α2 and β2 respectively, which anchor the HLA class II molecule to the cell membrane. The peptide binding groove is formed from the heterodimer of α1 and β1. The most widely studied HLA-DR molecules have DRA and DRB, corresponding to the α and β domains respectively. DRB is diverse and DRA is almost identical. Thereby, the binding specificity of the DRB allele indicates that of the corresponding HLA-DR. Each MHC protein has its own binding specificity, meaning that the set of peptide bindings to MHC molecules may be different from those to another MHC molecule. Classical molecules present peptides to CD4+ lymphocytes. Non-classical molecules with intracellular functions, accessory molecules are not exposed on the cell membrane but are in the internal membranes in lysosomes and usually load antigen peptides onto classical HLA class II molecules.

[0249]

[0467] In the HLA class II system, phagocytic cells such as macrophages and immature dendritic cells take up entities into phagosomes by phagocytosis - B cells exhibit more general endocytosis into endosomes - and fuse with lysosomes that cleave the ingested proteins into numerous diverse peptides. Autophagy is another source of HLA class II peptides. Through the physicochemical dynamics in the intermolecular interaction with host-derived HLA class II variants encoded by the host genome, certain peptides exhibit immunodominance and are loaded onto HLA class II molecules. These are transported to the cell surface and externalized. The most studied subclasses of the HLA class II gene are: HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRA and HLA-DRB1.

[0250]

[0468] The presentation of peptides to CD4+ helper T cells by HLA class II molecules is required for the immune response to foreign antigens (Roche and Furuta, 2015). When activated, CD4+ T cells promote B cell differentiation and antibody production as well as CD8+ T cell (CTL) responses. CD4+ T cells also secrete cytokines and chemokines that activate other immune cells and induce differentiation. HLA class II molecules are heterodimers of α- and β-chains that interact to form a peptide-binding groove that is more open than the HLA class I peptide-binding groove (Unanue et al., 2016). Peptides that bind to HLA class II molecules are thought to have a 9-amino acid binding core that protrudes from the groove and includes residues adjacent to either the N- or C-terminal side (Jardetzky et al., 1996; Stern et al., 1994). These peptides are usually 12 - 16 amino acids in length and often contain 3 - 4 anchor residues at positions P1, P4, P6 / 7 and P9 of the binding register (Rossjohn et al., 2015).

[0251]

[0469] HLA alleles are expressed in a codominant fashion, meaning that alleles (variants) inherited from both parents are equally expressed. For example, if each individual carries two alleles of each of the three class I genes (HLA-A, HLA-B, and HLA-C), there are six possible different types of HLA class II that can be expressed. At the HLA class II locus, each individual inherits a pair of HLA-DP genes (DPA1 and DPB1, encoding the α and β chains), HLA-DQ (DQA1 and DQB1, for the α and β chains), one gene HLA-DRα (DRA1), and one or more genes HLA-DRβ (DRB1 and DRB3, -4, or -5). HLA-DRB1, for example, has more than approximately 400 known alleles. This means that a single heterozygous individual can inherit six or eight functional HLA class II alleles: more than three from each parent. Thus, the HLA genes are highly polymorphic; a large number of different alleles exist in different individuals within a population. The genes encoding HLA proteins have a vast potential for diversity, enabling each individual's immune system to respond to a wide range of foreign invaders. Some HLA genes have hundreds of identified versions (alleles), each given a specific number. In some embodiments, the HLA class I alleles are HLA-A*02:01, HLA-B*14:02, HLA-A*23:01, HLA-E*01:01 (non-classical). In some embodiments, the HLA class II alleles are HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*11:01, HLA-DRB*15:01, and HLA-DRB*07:01.

[0252]

[0470] The subject-specific HLA allele or the HLA genotype of the subject can be determined by any method well known in the art. In an exemplary embodiment, the HLA genotype is determined by any method described in International Patent Application No. PCT / US2014 / 068746, filed June 11, 2015, as WO2015085147, which is hereby incorporated by reference in its entirety. Briefly, the method can include determining a polymorphic genotype that can include generating an alignment of reads extracted from a sequencing dataset to a set of genetic criteria that includes allelic variants of polymorphic genes, determining a first posterior probability or a posterior probability-derived score for each allelic variant in the alignment, identifying the allelic variant having the maximum first posterior probability or posterior probability-derived score as the first allelic variant, identifying one or more overlapping reads aligned using the first allelic variant and one or more other allelic variants, determining a second posterior probability or a posterior probability-derived score for one or more other allelic variants using a weighting factor, identifying a second allelic variant by selecting the allelic variant having the maximum second posterior probability or posterior probability-derived score, wherein the first and second allelic variants define a genotype for the polymorphic gene, and steps that result in the output of the first and second allelic variants.

[0253]

[0471] In some embodiments, the MHC class II peptide: antigen peptide binding and presentation prediction methods described herein have the ability to predict binders from a broad repertoire of MHC class II peptides encoded by an individual's HLA alleles. In some embodiments, the MAPTAC technology is trained using a large database of HLA-matched peptides verified by mass spectrometry. In some embodiments, the large database of HLA-matched peptides verified by mass spectrometry contains over 1.2 x 106 such HLA-matched peptides. In some embodiments, the large database of HLA-matched peptides verified by mass spectrometry covers over 150 HLA alleles including both MHC class I and class II allele subtypes. In some embodiments, the database covers at least 95% of the US population for HLA-I and HLA-II (DR subtypes).

[0254]

[0472] As described herein, there is extensive evidence in both animals and humans that mutant epitopes are effective in inducing an immune response, that cases of spontaneous tumor regression or long-term survival correlate with CD8+ T cell responses to mutant epitopes, and that "immune editing" can be traced in mice and humans to changes in the expression of antigens with dominant mutations.

[0255]

[0473] Sequencing techniques have revealed that each tumor contains multiple patient-specific mutations that alter the protein-coding content of genes. Such mutations range from single amino acid changes (resulting from missense mutations) to the addition of long regions of novel amino acid sequences by frameshifts, read-through of stop codons, or translation of intron regions (novel open reading frame mutations; neoORFs), creating proteins that vary. These mutant proteins are valuable targets for the host immune response to tumors, which are distinct from the native proteins and are not subject to the immune dampening effects of self-tolerance. Thus, mutant proteins are likely to be immunogenic and are more specific to tumor cells compared to the patient's normal cells. In essence, short peptides (8-24 amino acids in length) containing cancer-related mutations are candidates for cancer immunotherapy.

[0256]

[0474] In some embodiments, the algorithms driving the prediction methods can be further utilized to call mutations with peptides. In some embodiments, the prediction methods can be used for determination of driver mutation status, and / or RNA expression status, and / or cleavage prediction within the peptides.

[0257]

[0475] The term "T cell" includes CD4+ T cells and CD8+ T cells. The term T cell also includes both type 1 helper T cells and type 2 helper T cells. As used herein, T cells are generally classified into two main classes: helper T (TH) cells and cytotoxic T lymphocytes (CTLs) by function and by cell surface antigens (cluster of differentiation antigens or CD) that also facilitate binding of the T cell receptor to an antigen.

[0258]

[0476] Mature helper T (TH) cells express the surface protein CD4 and are referred to as CD4+ T cells. Following T cell development, mature, naive T cells leave the thymus and begin to spread throughout the body, including the lymph nodes. Naive T cells are T cells that have not been exposed to an antigen they are programmed to respond to. Like all T cells, they express the T cell receptor-CD3 complex. The T cell receptor (TCR) consists of both constant and variable regions. The variable region determines which antigen the T cell can respond to. CD4+ T cells have a TCR with an affinity for MHC class II, proteins, and CD4 is involved in determining MHC affinity during maturation in the thymus. MHC class II proteins are found only on the surface of specialized antigen-presenting cells (APCs). Specialized antigen-presenting cells (APCs) are mainly dendritic cells, macrophages, and B cells, but dendritic cells are the only cell population that continuously (at all times) expresses MHC class II. Some APCs also bind undenatured (or unprocessed) antigens to their surface, such as follicular dendritic cells, but unprocessed antigens do not interact with T cells and are not involved in their activation. Peptide antigens that bind to HLA class I proteins are typically shorter than peptide antigens that bind to HLA class II proteins.

[0259]

[0477] Cytotoxic T lymphocytes (CTLs), also known as cytotoxic T cells, cytolytic T cells, CD8+ T cells or killer T cells, refer to lymphocytes that induce apoptosis in target cells. CTLs form antigen-specific conjugates with target cells through the interaction of the TCR with processed antigen (Ag) on the target cell surface, resulting in apoptosis of the target cell. Apoptotic bodies are removed by macrophages. The term "CTL response" is used to refer to the primary immune response mediated by CTL cells. Cytotoxic T lymphocytes have both a T cell receptor (TCR) and a CD8 molecule on their surface. The T cell receptor can recognize and bind to peptides complexed with HLA class I molecules. Each cytotoxic T-lymphocyte expresses a unique T cell receptor that can bind to a specific MHC / peptide complex. Most cytotoxic T cells express a T cell receptor (TCR) that can recognize a specific antigen. For the TCR to bind to the HLA class I molecule, the former must be accompanied by a glycoprotein called CD8 that binds to the constant portion of the HLA class I molecule. Therefore, these T cells are called CD8+ T cells. The affinity between CD8 and the MHC molecule keeps the T cell and the target cell closely bound during antigen-specific activation. CD8+ T cells, when activated, are recognized as T cells and are generally classified as having a pre-determined cytotoxic role within the immune system. However, CD8+ T cells also have the ability to produce some cytokines.

[0260]

[0478] The "T cell receptor (TCR)" is a cell surface receptor that participates in the activation of T cells in response to antigen presentation. TCRs are generally made from two chains, alpha and beta, which assemble to form a heterodimer and associate with CD3-transduced subunits to form the T cell receptor complex present on the cell surface. Each of the alpha and beta chains of the TCR consists of immunoglobulin-like N-terminal variable (V) and constant (C) regions, a hydrophobic transmembrane domain, and a short cytoplasmic region. The variable regions of the alpha and beta chains, with respect to immunoglobulin molecules, are generated by V(D)J recombination, creating a broad diversity of antigen specificities within the population of T cells. However, in contrast to immunoglobulins that recognize intact antigens, T cells are activated by processed peptide fragments that are associated with MHC molecules, introducing another aspect, known as MHC restriction, to antigen recognition by T cells. Recognition of MHC differences between donor and recipient through the T cell receptor leads to T cell proliferation and the potential for the development of GVHD. It has been shown that normal surface expression of the TCR depends on the coordinated synthesis and assembly of all seven components of the complex (Ashwell and Klusner 1990). Inactivation of TCRα or TCRβ can result in the removal of the TCR from the surface of T cells, preventing the recognition of allogeneic antigens and thereby GVHD. However, TCR disruption generally results in the removal of CD3 signaling components, altering additional means of T cell expansion.

[0261]

[0479] The term "HLA peptidome" refers to a pool of peptides that specifically interact with a particular HLA class and can encompass thousands of different sequences. The HLA peptidome includes peptide diversity and is derived from both normal and abnormal proteins expressed in the cell. Thereby, the HLA peptidome can be studied as a source of information about protein synthesis and degradation schemes within cancer cells, as well as for the identification of cancer-specific peptides for cancer immunotherapy. In some embodiments, the HLA peptidome is a pool of soluble HLA peptides (sHLA). In some embodiments, the HLA peptidome is a pool of membrane-associated HLA (mHLA).

[0262]

[0480] "Antigen-presenting cell" or "APC" includes professional antigen-presenting cells (e.g., B lymphocytes, macrophages, monocytes, dendritic cells, Langerhans cells), and other antigen-presenting cells (e.g., keratinocytes, endothelial cells, astrocytes, fibroblasts, oligodendrocytes, thymic epithelial cells, thyroid epithelial cells, glial cells (brain), pancreatic beta cells, and vascular endothelial cells). "Antigen-presenting cell" or "APC" is a cell that expresses major histocompatibility complex (MHC) molecules and can display exogenous antigens complexed with MHC on its surface.

[0263] Monoallelic HLA cell line

[0481] A single allelic cell line expressing either a single HLA class I allele, a single pair of HLA class II alleles, or a single pair of a single HLA class I allele and an HLA class II allele can be generated by transducing or transfecting a suitable cell population with a polynucleotide, e.g., a vector encoding a single HLA allele (Figure 2). Suitable cell populations include, for example, HLA class I-deficient cell lines in which a single HLA class I allele is exogenously expressed, HLA class II-deficient cell lines in which a single exogenous pair of HLA class II alleles is expressed, or class I and class II-deficient cells in which a single pair of a single HLA class I and / or class II allele is exogenously expressed. As an exemplary embodiment, the HLA class I-deficient B cell line is B721.221. However, it will be apparent to those skilled in the art that other cell populations deficient in HLA class I and / or HLA class II can be generated. An exemplary method for deleting / inactivating the endogenous HLA class I or HLA class II gene includes, for example, CRISPR-Cas9-mediated genome editing in THP-1 cells. In some embodiments, the population of cells is professional antigen-presenting cells, such as macrophages, B cells, and dendritic cells. The cells can be B cells or dendritic cells. In some embodiments, the cells are tumor cells or cells derived from a tumor cell line. In some embodiments, the cells are isolated from a patient. In some embodiments, the cells contain an infectious pathogen or a part thereof. In some embodiments, the population of cells is at least 10 cells 7including. In some embodiments, the cell population is further modified by increasing and / or decreasing the expression and / or activity of at least one gene. In some embodiments, the gene encodes a member of the immunoproteasome. The immunoproteasome is well known to be involved in the processing of HLA class I-binding peptides and includes the LMP2 (β1i), MECL-1 (β2i), and LMP7 (β5i) subunits. The immunoproteasome can also be induced by interferon gamma. Thus, in some embodiments, the cell population can be contacted with one or more cytokines, growth factors, or other proteins. The cells can be stimulated with inflammatory cytokines such as interferon gamma, IL-10, IL-6, and / or TNF-α. The cell population can also be subjected to various environmental conditions such as stress (heat stress, oxygen deprivation, glucose starvation, DNA-damaging agents, etc.). In some embodiments, the cells are contacted with one or more of a chemotherapeutic agent, radiation, targeted therapy, or immunotherapy. Thus, the methods disclosed herein can be used to study the effects of various genes or conditions on HLA peptide processing and presentation. In some embodiments, the conditions used are selected to match the condition of the patient in which the HLA peptide population is to be identified.

[0264]

[0482] The single HLA alleles of the present disclosure can be encoded and expressed using virus-based systems (e.g., adenovirus systems, adeno-associated virus (AAV) vectors, poxviruses or lentiviruses). Plasmids that can be used for adeno-associated virus, adenovirus and lentivirus delivery have been previously described (see, e.g., U.S. Pat. Nos. 6,955,808 and 6,943,019 and U.S. Patent Application No. 20080254008, which are incorporated herein by reference). Among the vectors that can be used in the practice of the present disclosure, integration into the host genome of the cell is possible using retroviral gene transfer methods, often resulting in long-term expression of the inserted transgene. In an exemplary embodiment, the retrovirus is a lentivirus. Additionally, high transduction efficiencies have been observed in a number of different T cell types and target tissues. The tropism of the retrovirus can be altered by incorporating foreign envelope proteins, expanding the potential target population of the target cells. The retrovirus can be engineered to allow conditional expression of the inserted transgene such that only a particular cell type is infected by the lentivirus. Cell type-specific promoters can be used to target expression in a particular cell type. The lentivirus vector is a retrovirus vector (and thus both lentivirus and retroviral vectors can be used in the practice of the present disclosure). Further, the lentivirus vector can transduce or infect non-dividing cells and typically yields high viral titers.

[0265]

[0483] The selection of the retroviral gene transfer system can depend on the target tissue. The retroviral vector is composed of cis-acting long terminal repeats having the packing ability for exogenous sequences up to 6 - 10 kb. The minimal cis-acting LTR is sufficient for vector replication and packing and is then used to incorporate the desired nucleic acid into the target cells to effect persistent expression. Widely used retroviral vectors that can be used in the practice of the present disclosure include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV) and combinations thereof (see, for example, Buchscher et al., (1992) J. Virol. 66:2731 - 2739; Johann et al., (1992) J. Virol. 66:1635 - 1640; Sommerfelt et al., (1990) Virol. 176:58 - 59; Wilson et al., (1998) J. Virol. 63:2374 - 2378; Miller et al., (1991) J. Virol. 65:2220 - 2224; PCT / US94 / 05700). Similarly, useful in the practice of the present disclosure are minimal non-primate lentiviral vectors such as lentiviral vectors based on equine infectious anemia virus (EIAV) (see, for example, Balagaan, (2006) J Gene Med; 8:275 - 285, Published online 21 November 2005 in Wiley InterScience DOI:10.1002 / jgm.845). The vector can have a cytomegalovirus (CMV) promoter driving the expression of the target gene. Thus, the present disclosure contemplates viral vectors, including retroviral vectors and lentiviral vectors, among the vectors (s) useful in the practice of the present disclosure.

[0266]

[0484] Any HLA allele can be expressed in a cell population. In an exemplary embodiment, the HLA allele is an HLA class I allele. In some embodiments, the HLA class I allele is an HLA-A allele or an HLA-B allele. In some embodiments, the HLA allele is an HLA class II allele. The sequences of HLA class I and class II alleles can be found in the IPD-IMGT / HLA database. Exemplary HLA alleles include, but are not limited to, HLA-A*02:01, HLA-B*14:02, HLA-A*23:01, HLA-E*01:01, HLA-DRB*01:01, HLA-DRB*01:02, HLA-DRB*11:01, HLA-DRB*15:01, and HLA-DRB*07:01.

[0267]

[0485] In some embodiments, the HLA allele is selected to correspond to a desired genotype. In some embodiments, the HLA allele is a mutant HLA allele and can be an allele that does not occur naturally or an allele that occurs naturally in an affected patient. The methods disclosed herein further have the advantage of identifying HLA-binding peptides for HLA alleles associated with various disorders and alleles that occur at low frequency. Thus, in some embodiments, the methods provided herein can identify HLA alleles even when they occur at a frequency of less than 1% within a population such as the Caucasian population.

[0268]

[0486] In some embodiments, the nucleic acid sequence encoding the HLA allele further comprises an affinity acceptor tag that can be used to immunopurify the HLA protein. Suitable tags are known in the art.In some embodiments, the affinity acceptor tag is a polyhistidine tag, a polyhistidine-glycine tag, a polyarginine tag, a polyaspartic acid tag, a polycysteine tag, a polyphenylalanine, a c-myc tag, a herpes simplex virus glycoprotein D (gD) tag, a FLAG tag, a KT3 epitope tag, a tubulin epitope tag, a T7 gene 10 protein peptide tag, a streptavidin tag, a streptavidin-binding peptide (SPB) tag, a Strep tag, a Strep tag II, an albumin-binding protein (ABP) tag, an alkaline phosphatase (AP) tag, a blue tongue virus tag (B tag), a calmodulin-binding peptide (CBP) tag, a chloramphenicol acetyltransferase (CAT) tag, a choline-binding domain (CBD) tag, a chitin-binding domain (CBD) tag, a cellulose-binding domain (CBP) tag, a dihydrofolate reductase (DHFR) tag, a galactose-binding protein (GBP) tag, a maltose-binding protein (MBP), a glutathione-S-transferase (GST), a Glu-Glu (EE) tag, a human influenza hemagglutinin (HA) tag, a horseradish peroxidase (HRP) tag, a NE tag, an HSV tag, a ketosteroid isomerase (KSI) tag, a KT3 tag, a LacZ tag, a luciferase tag, a NusA tag, a PDZ domain tag, an AviTag, a calmodulin tag, an E tag, an S tag, an SBP tag, a Softag 1, a Softag 3, a TC tag, a VSV tag, an Xpress tag, an Isopeptag, a SpyTag, a SnoopTag, a Profinity eXact tag, a protein C tag, an S1 tag, an S tag, a biotin carboxyl carrier protein (BCCP) tag, a green fluorescent protein (GFP) tag, a small ubiquitin-like modifier (SUMO) tag, a tandem affinity purification (TAP) tag, a HaloTag, a Nus tag, a thioredoxin tag, an Fc tag, a CYD tag, an HPC tag, a TrpE tag, a ubiquitin tag, a vesicular stomatitis virus (VSV) glycoprotein-derived VSV-G epitope tag or a V5 tag derived from a small epitope (Pk) found in the paramyxovirus P and V proteins of simian virus 5 (SV5).In some embodiments, the affinity acceptor tag is an "epitope tag", which is a type of peptide tag that adds an epitope (antibody binding site) recognizable to cause binding of the corresponding antibody to the HLA protein, thereby enabling identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are Protein A or Protein G that bind to IgG. In some embodiments, the affinity acceptor tag comprises a biotin acceptor peptide (BAP) or a human influenza hemagglutinin (HA) peptide sequence. Numerous other tag moieties are well known to those skilled in the art and can be envisioned and are contemplated herein. Any peptide tag can be used as long as it can be expressed as an element of the HLA peptide complex with an affinity acceptor tag.

[0269]

[0487] The methods provided herein include the step of isolating HLA peptide complexes from transferase or transduced cells by affinity pulldown of HLA constructs (Figure 3). In some embodiments, the complexes can be isolated using standard immunoprecipitation techniques well known in the art using commercially available antibodies. The cells can first be lysed. HLA class I-peptide complexes are isolated using an HLA class I-specific antibody, such as the W6 / 32 antibody, while HLA class II-peptide complexes can be isolated using an HLA class II-specific antibody, such as the M5 / 114.15.2 monoclonal antibody. In some embodiments, a single (or paired) HLA allele is expressed as a fusion protein with a peptide tag, and the HLA peptide complex is isolated using a binding molecule that recognizes the peptide tag.

[0270]

[0488] The method further includes the steps of isolating a peptide from the HLA peptide complex and sequencing the peptide. The peptide is isolated from the complex by any method well known to those skilled in the art, such as acid elution. On the other hand, any sequencing method may be used, and in some embodiments, methods using mass spectrometry such as liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or alternatively HPLC-MS or HPLC-MS / MS) are utilized. These sequencing methods are known to those skilled in the art and are outlined in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb;34(1):43~63.

[0271]

[0489] In some embodiments, the population of cells expresses one or more endogenous HLA alleles. In some embodiments, the population of cells is a population of engineered cells lacking one or more endogenous HLA class I alleles. In some embodiments, the population of cells is a population of engineered cells lacking endogenous HLA class I alleles. In some embodiments, the population of cells is a population of engineered cells lacking one or more endogenous HLA class II alleles. In some embodiments, the population of cells is a population of engineered cells lacking endogenous HLA class II alleles or a population of engineered cells lacking both endogenous HLA class I alleles and endogenous HLA class II alleles. In some embodiments, the population of cells includes cells enriched or sorted by, for example, fluorescence-activated cell sorting (FACS). In some embodiments, fluorescence-activated cell sorting (FACS) is used to sort the population of cells. In some embodiments, the population of cells has been pre-FACS sorted for cell surface expression of either HLA class I or class II, or both HLA class I and class II. For example, FACS can be used to sort a population of cells for cell surface expression of HLA class I alleles, HLA class II alleles, or combinations thereof.

[0272] Method for preparing an individualized cancer vaccine

[0490] When mutations are identified that are cancer-specific mutations present in the DNA of cancer cells but not in the normal cells of the same human subject, the mutations result in changes in one or more amino acids in the protein encoded by the DNA, and the mutations can be targets of the host immune response. The innate immune response is directed against the mutated protein, resulting in the destruction of the cancer cells expressing the protein. Due to the natural tolerance response and the immune-deficient environment in cancerous tissue, immunotherapy is a clinical path that attempts to enhance such immune responses so as to nullify the body's tolerance and immunosuppressive effects. Accordingly, proteins or peptides containing the mutations described above are suitable candidates for immunotherapy.

[0273]

[0491] The mutant protein is eaten and digested by specialized phagocytic cells that act as antigen-presenting cells (APCs) and is presented on the cell surface as an antigen in an antigen-presenting complex that includes major histocompatibility complex (MHC) proteins for T cell activation. Human MHC proteins are called human leukocyte antigens, HLA. MHC proteins may be MHC class I or class II proteins, and while there are some functional differences that result from the presentation of peptides by either class I or class II MHC proteins (HLA class I and HLA class II proteins), one notable difference is that the HLA class I-peptide complex presents antigens to cytotoxic CD8+ T cells, while the HLA class II peptide complex can also activate CD4+ T cells and result in a long-term immune response. CD8+ T cells are essential in the task of cell-by-cell elimination of diseased cells such as infected or tumor cells. CD4+ T cells have a more sustained effect on activation, and the most important of these is the generation of immune memory. CD4 subsets are recruited differently depending on the type of immunological threat, and multiple subsets with overlapping or non-overlapping functions may be recruited simultaneously. This helps to balance the immunological response to the threat of pathogens. In this regard, HLA class II peptide-mediated antigen presentation is persistent and affects the appropriate immune response. On the other hand, HLA class II binding to peptides can be promiscuous, thereby resulting in non-specific peptide binding and presentation to the immune system, which leads to abnormal immune responses such as autoimmunity.

[0274]

[0492] In one aspect, the present disclosure provides a method for predicting peptides that can accurately match or bind to specific HLA class II alpha and beta chain heterodimers such that high-fidelity binding of the peptides to HLA class II proteins (including alpha and beta chain heterodimers) ensures presentation of the specific peptides to T lymphocytes, thereby inducing a specific immune response and avoiding any cross-reactivity or immunological chaos.

[0275]

[0493] In one aspect, the present disclosure provides a method for predicting a peptide that can precisely bind to a specific HLA class II protein so that a more persistent and robust immune response can be activated with the peptide, by the ability of the HLA class II protein to activate CD4+ T cells and stimulate immune memory when the peptide is therapeutically administered to a subject expressing the specific cognate HLA class II protein. In some embodiments, a given peptide predicted to bind to an HLA class II protein with high specificity is a peptide containing a mutation, where the mutation is prevalent in the cancer or tumor cells of the subject; while the same HLA class II protein predicted to bind to the mutant peptide either (a) does not bind to the corresponding non-mutated wild-type peptide or (b) binds with a significantly lower affinity compared to its affinity for the mutant peptide of the subject. The preferential binding of HLA to the mutant peptide is advantageous in the development of immunotherapy because cells expressing the wild-type peptide are spared from immune attack by T cells reactive to the HLA-presented peptide. In some embodiments, the predicted peptide that specifically binds to an HLA class II protein is a peptide having a post-translational modification. Exemplary post-translational modifications include, but are not limited to: phosphorylation, ubiquitination, dephosphorylation, glycosylation, methylation or acetylation. In some embodiments, the predicted peptide is subjected to post-translational modification prior to use in immunotherapy.

[0276]

[0494] In some embodiments, the immunotherapies and strategies disclosed herein may also be applicable to suppressing unwanted immune activation, such as in autoimmune reactions. Specifically, a peptide identified as a potential binder to a specific HLA subtype may be engineered to bind to the specific HLA molecule, inducing tolerance rather than an immunogenic response.

[0277]

[0495] In one aspect, methods of immunotherapy tailored or individualized for a particular subject are provided herein. All subjects or patients express a particular array of HLA class I and HLA class II proteins. HLA typing is a known technique that enables determination of the specific repertoire of HLA proteins expressed by a subject. Once the HLA heterodimers expressed by a particular subject are understood, the improved, refined, and reliable methods described herein for predicting peptides that can bind with high fidelity to specific HLA class II alpha and beta chain heterodimers ensure that specific immune responses can be specifically generated for a subject as desired.

[0278]

[0496] Genes encoding HLA heterodimers are highly polymorphic with over 4,000 HLA class II alleles identified in the human population. From maternal and paternal HLA haplotypes, an individual may inherit different alleles for each of the HLA class II gene loci, and each HLA class II heterodimer is made from an α- and a β-chain. Due to the large number of α- and β-chain pairing combinations, the population of possible HLA heterodimers is very complex, especially for HLA-DP and HLA-DQ alleles. HLA class II heterodimers are translated in the endoplasmic reticulum (ER) and assembled into a stable complex with the invariant chain (Ii) derived from the protein CD74. Ii stabilizes the class II complex by allowing correct protein folding and enables transport of the HLA class II heterodimer to the endosome / lysosome compartment. Within these HLA class II loading compartments, Ii is proteolytically cleaved by cathepsin into a placeholder peptide called CLIP. CLIP is then exchanged for a higher affinity peptide by the chaperone HLA-DM, a non-classical HLA class II heterodimer, in a low pH environment. The HLA class II complex loaded with the high affinity peptide then resides in the trans-Golgi and finally on the cell surface for presentation to CD4+ T cells.

[0279]

[0497] Each HLA heterodimer is presumed to bind to thousands of peptides with allele-specific binding selectivity. In fact, each HLA allele is presumed to bind to approximately 1,000 to 10,000 unique peptides and present them to T cells. Considering such diversity in HLA binding, it is very difficult to accurately predict whether a peptide is likely to bind to a specific HLA allele. Little is known about the allele-specific peptide-binding characteristics of HLA class II molecules due to the heterogeneity of α- and β-chain pairing, the complexity of data limiting the ability to confidently assign core binding epitopes, and the lack of immunoprecipitation-grade allele-specific antibodies required for high-resolution biochemical analysis. Furthermore, analyzing peptide epitopes derived from a given HLA allele can be ambiguous when multiple HLA alleles are presented on the cell surface.

[0280]

[0498] The prediction of candidate neoantigens is mainly made for HLA class I epitopes (taking into account the availability of experimental data on class I prediction algorithms compared to class II), and furthermore, CD4+ T cell responses are often observed in both preclinical and clinical individualized neoantigen vaccination studies. These observations demonstrate that HLA class II epitope processing and presentation can also play an important role in cancer treatment. Although HLA class II prediction algorithms exist, they are inaccurate because the open-ended peptide-binding groove in the HLA class II heterodimer allows longer peptides (generally 15 - 40 amino acids) to bind, increasing the heterogeneity and complexity of epitope presentation. Therefore, further research is needed for a better understanding of the characteristics of the HLA class II peptide-binding core and the intracellular processes involved in class II epitope processing and presentation. The proteomics field is currently limited by the complexity of HLA class II heterodimer formation and the availability of immunoprecipitation-grade antibodies for HLA class II-peptide complex isolation. To overcome these challenges, a single-allele HLA profiling workflow has been developed that relies on LC-MS / MS for the characterization of allele-specific HLA class II-ligandomes for class II epitope presentation methods. The following definitions supplement those in the art and are applicable to this application and any related or unrelated matters, e.g., not attributable to any common-owned patents or applications. Any methods and materials similar or equivalent to those described herein may be used in the practice for the tests of this disclosure, and exemplary materials and methods are described herein. Accordingly, the terms used herein are for the purpose of describing specific embodiments and are not intended to be limiting.

[0281]

[0499] Disclosed herein is a method for preparing an individualized cancer vaccine. A method for preparing an individualized cancer vaccine includes identifying a peptide sequence comprising a mutation expressed in a subject's cancer cells; inputting amino acid position information of the identified peptide sequence into a machine learning HLA peptide presentation prediction model using a computer processor to generate a series of presentation predictions for the identified peptide sequence, wherein each presentation prediction indicates the probability that one or more proteins encoded by the class II MHC alleles of the subject's cancer cells present a given sequence of the identified peptide sequence; and selecting a subset of the identified peptide sequences based on the series of presentation predictions to prepare an individualized cancer vaccine.

[0282]

[0500] In some embodiments, one or more results obtained from the methods described herein can provide one or more quantitative values indicating one or more of the following: the likelihood of diagnostic accuracy, the likelihood of the presence of a condition in a subject, the likelihood that a subject will develop a condition, the likelihood of success of a particular treatment, or any combination thereof. In some embodiments, the methods described herein can predict the risk or likelihood of developing a condition. In some embodiments, the methods described herein can be an indicator of early diagnosis of developing a condition. In some embodiments, the methods described herein can confirm the diagnosis or presence of a condition. In some embodiments, the methods described herein can monitor the progression of a condition. In some embodiments, the methods described herein can monitor the effectiveness of treatment for a condition in a subject.

[0283] Method for identification of MHC-II peptides

[0501] In one aspect, a method is provided for identifying one or more peptides presented by MHC-II proteins for immune activation. In some embodiments, the one or more peptides comprise epitopes. In some embodiments, the method includes a computer prediction of the likelihood that a specific epitope will be presented by an MHC-II protein. In some embodiments, the method includes a computer prediction of the specificity of an epitope for MHC-II presentation. In some embodiments, the computer prediction method includes an evaluation of peptide-MHC interactions. In some embodiments, the computer prediction method includes a prediction of the allelic specificity of a peptide for antigen presentation. In some embodiments, the computer prediction method includes incorporation of bioinformatics information, such as nucleotide sequences, structural motifs of biomolecules, protein-protein interaction characteristics, and functional strengths such as immunogenicity. In some embodiments, the computer prediction method includes machine learning. A number of immunoinformatics methods for the prediction of peptide-MHC interactions have been developed for both MHC class I and II based on machine learning approaches, such as simple pattern motifs, support vector machines (SVM), hidden Markov models (HMM), neural network (NN) models, quantitative structure-activity relationship (QSAR) analysis, structure-based methods, and biophysical methods. These methods can be classified into two categories, namely, intra-allelic (allele-specific) and trans-allelic (pan-specific) methods. Intra-allelic methods are trained on a limited set of experimental peptide-binding data for a specific MHC molecule and are applied to the prediction of peptides that bind to that molecule. Due to the extreme polymorphism of MHC molecules, the presence of thousands of allelic variants, combined with the lack of sufficient experimental binding data, makes it impossible to construct a prediction model for each allele.Accordingly, trans-allelic and general-purpose methods, such as NetMHCIIpan (Karosiene E et al., NetMHCIIpan-3.0, a common pan-specific MHC class II prediction method including all three human MHC class II isotypes, HLA-DR, HLA-DP and HLA-DQ. Immunogenetics (2013) 65(10):711-24), and TEPITOPEpan (Zhang L, et al., TEPITOPEpan: extending TEPITOPE for peptide binding prediction covering over 700 HLA-DR molecules. PLoS One (2012) 7(2):e30483), were developed using binding data peptide binding data extended over a number of alleles or over species. Similar methods for MHC-I are also available, such as NetMHCpan and KISS.

[0284]

[0502] In some embodiments, the peptide sequence may not be expressed in the normal cells of the subject. In some embodiments, each and all of the cells of the subject may not be cancer cells. Cancer cells include, but are not limited to, thyroid cancer, adrenocortical cancer, anal cancer, aplastic anemia, cholangiocarcinoma, bladder cancer, bone cancer, bone metastasis, central nervous system (CNS) cancer, peripheral nervous system (PNS) cancer, breast cancer, Castleman disease, cervical cancer, pediatric non-Hodgkin lymphoma, lymphoma, colorectal cancer, endometrial cancer, esophageal cancer, Ewing family tumors (e.g., Ewing sarcoma), eye cancer, gallbladder cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors, gestational trophoblastic disease, hairy cell leukemia, Hodgkin disease, Kaposi sarcoma, kidney cancer, laryngeal and hypopharyngeal cancer, acute lymphoblastic leukemia, acute myeloid leukemia, pediatric leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, liver cancer, lung cancer, lung carcinoid tumors, non-Hodgkin lymphoma, male breast cancer, malignant mesothelioma, multiple myeloma, myelodysplastic syndromes, myeloproliferative disorders, nasal and paranasal cancer, nasopharyngeal cancer, neuroblastoma, oral and oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, pituitary tumor, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma (adult soft tissue cancer), melanoma skin cancer, non-melanoma skin cancer, stomach cancer, testicular cancer, thymic cancer, uterine cancer (e.g., uterine sarcoma), vaginal cancer, vulvar cancer or Waldenström's macroglobulinemia and can be produced through various cancers.

[0285]

[0503] Identifying involves comparing a DNA, RNA or protein sequence from the cancer cells of the subject to a DNA, RNA or protein sequence from the normal cells of the subject. The DNA, RNA or protein sequence from the cancer cells of the subject may be different from the DNA, RNA or protein sequence from the normal cells of the subject. Identifying can identify nucleic acid variants with high sensitivity.

[0286]

[0504] The machine learning HLA peptide presentation prediction model may include a plurality of prediction variables identified based at least on training data. The training data includes sequence information of the sequences of peptides presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information, which is related to HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the prediction variables.

[0287]

[0505] In some embodiments, the training data includes structured data, time series data, unstructured data, and relational data. The unstructured data may include voice data, image data, video, mechanical data, electrical data, chemical data, and combinations thereof for use in accurately simulating or training robotics or simulations. The time series data may include data from one or more of smart meters, smart home appliances, smart devices, monitoring systems, telemetry devices, or sensors. The relational data includes data from customer systems, enterprise systems, operational systems, websites, application program interfaces (APIs) accessible from the web, or any combination thereof. This may be done by the user through any method of inputting files or other data formats into the software or system.

[0288]

[0506] In some embodiments, the training data can be stored in a database. The database can be stored in a computer-readable format. A computer processor can be configured to access the data stored in a computer-readable memory. In some embodiments, a computer system can be used to analyze the data to obtain results. The results can be stored internally in a storage medium or remotely and communicated to a staff member such as a medical professional. In some embodiments, the computer system can be operably coupled to components for communicating the results. The components for communication can include wired and wireless components. Examples of wired communication components can include Universal Serial Bus (USB) connections, coaxial cable connections, Ethernet cables such as Cat5 or Cat6 cables, fiber optic cables or telephone lines. Examples of wireless communication components can include Wi-Fi receivers, components for accessing mobile data standards such as 3G or 4G LTE data signals or Bluetooth receivers. In some embodiments, all of these data in the storage medium are retrieved to build a data warehouse and stored in an archive.

[0289]

[0507] In some embodiments, the database includes an external database. The external database is a medical database, for example, but not limited to only this, Adverse Drug Effects Database, AHFS Supplemental File, Allergen Picklist File, Average WAC Pricing File, Brand Probability File, Canadian Drug File v2, Comprehensive Price History, Controlled Substances File, Drug Allergy Cross-Reference File, Drug Application File, Drug Dosing & Administration Database, Drug Image Database v2.0 / Drug Imprint Database v2.0, Drug Inactive Date File, Drug Indications Database, Drug Lab Conflict Database, Drug Therapy Monitoring System (DTMS) v2.It may be the 2 / DTMS Consumer Monographs, Duplicate Therapy Database, Federal Government Pricing File, Healthcare Common Procedure Coding System Codes (HCPCS) Database, ICD-10 Mapping Files, Immunization Cross-Reference File, Integrated A to Z Drug Facts Module, Integrated Patient Education, Master Parameters Database, Medi-Span Electronic Drug File (MED-File) v2, Medicaid Rebate File, Medicare Plans File, Medical Condition Picklist File, Medical Conditions Master Database, Medication Order Management Database (MOMD), Parameters to Monitor Database, Patient Safety Programs File, Payment Allowance Limit-Part B (PAL-B) v2.0, Precautions Database, RxNorm Cross-Reference File, Standard Drug Identifiers Database, Substitution Groups File, Supplemental Names File, Uniform System of Classification Cross-Reference File or Warning Label Database.

[0290]

[0508] In some embodiments, training data can also be obtained through other data sources. The data sources can include sensors or smart devices, such as appliances, smart meters, wearables, monitoring systems, data stores, customer systems, billing systems, financial systems, cloud source data, weather data, social networks or any other sensors, enterprise systems or data stores. Examples of smart meters or sensors can include meters or sensors located at customer facilities, or meters or centers located between the customer and the generation or source location. By incorporating data from a wide range of sources, the system can perform complex and detailed analysis. In some embodiments, the data sources can include sensors or databases for other medical platforms without limitation.

[0291]

[0509] HLA typing is conventionally performed either by serological methods using antibodies or by PCR-based methods, such as sequence-specific oligonucleotide probes (SSOP) or sequence-based typing (SBT). The former is hampered by potentially high levels of cross-reactivity and limited resolution performance, while the latter is plagued by difficulties related to the efficiency of PCR due to the very limited possibilities for primer placement caused by the position of the polymorphisms.

[0292]

[0510] In some embodiments, the array information is identified by either a sequencing method or mass spectrometry, such as a method using liquid chromatography - mass spectrometry (LC - MS or LC - MS / MS, or alternatively HPLC - MS or HPLC - MS / MS). These sequencing methods are known to those skilled in the art and are outlined in Medzihradszky KF and Chalkley RJ. Mass Spectrom Rev. 2015 Jan - Feb;34(1):43~63. In some embodiments, the mass spectrometry is single - allele mass spectrometry. In some embodiments, the mass spectrometry can be MS analysis, MS / MS analysis, LC - MS / MS analysis or a combination thereof. In some embodiments, MS analysis can be used to determine the mass of intact peptides. For example, determining can include determining the mass of intact peptides (e.g., MS analysis). In some embodiments, MS / MS analysis can be used to determine the mass of peptide fragments. For example, determining can include determining the mass of peptide fragments, which can be used to determine the amino acid sequence of a peptide or a portion thereof (e.g., MS / MS analysis). In some embodiments, the mass of peptide fragments can be used to determine the sequence of amino acids within a peptide. In some embodiments, LC - MS / MS analysis can be used to separate a complex peptide mixture. For example, determining can include separating a complex peptide mixture, e.g., by liquid chromatography, and determining the mass of intact peptides, the mass of peptide fragments or a combination thereof (e.g., LC - MS / MS analysis). This data can be used, for example, for peptide sequencing.

[0293]

[0511] In some embodiments, the training peptide sequence information includes amino acid position information of the training peptide. In some embodiments, the training peptide sequence information includes at most about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% or less of the sequence information of the peptides presented by the HLA proteins expressed in cells and identified by mass spectrometry. In some embodiments, the training peptide sequence information may include at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the sequence information of the peptides presented by the HLA proteins expressed in cells and identified by mass spectrometry.

[0294]

[0512] Any information and data may be paired with a subject who is the source of the information and data. The subject or a healthcare provider may search for information and data from storage or a server using the identity of the subject. The identity of the subject may include a patient's photograph, name, address, social security number, date of birth, telephone number, postal code, or any combination thereof. The identity of the subject may be encrypted and encoded into a visual graphical code. The visual graphical code may be a one-time barcode uniquely associated with the identity of the subject. The barcode may be a UPC barcode, EAN barcode, Code39 barcode, Code128 barcode, ITF barcode, CodaBar barcode, GS1 DataBar barcode, MSI Plessey barcode, QR barcode, Datamatrix code, PDF417 code, or Aztec barcode. The visual graphical code may be configured to be shown on a display screen. The barcode may include a QR that is optically captured and machine-readable. The barcode may define elements such as the version, format, position, alignment, or timing of the barcode to enable reading and decoding of the barcode. The barcode may encode various types of information in any suitable format, such as binary, or information combining alphabet and numbers. The QR Code (Registered Trademark) may have various symbol sizes as long as the QR Code (Registered Trademark) can be scanned by an image device from a reasonable distance. The QR Code (Registered Trademark) may be in any image format (e.g., EPS or SVG vector graph, PNG, TIF, GIF, or JPEG raster graphic format).

[0295]

[0513] In some embodiments, the function representing the relationship between the amino acid position information received as input and the presentability generated as output based on the amino acid position information and the predictor variables, where the predictor variables include linear or non-linear functions. The function can be, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLu activation function or other functions, such as saturated hyperbolic tangent, identity mapping, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sinc, Gaussian or sigmoid function, or any combination thereof.

[0296]

[0514] In some embodiments, the linear function can be obtained through linear regression. In some embodiments, linear regression is a method of predicting a target variable by fitting the best linear relationship between a dependent variable and an independent variable. The best fit can mean that the sum of all distances between the shape at each point and the actual observations is minimized. Linear regression can include simple linear regression or multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting the best linear relationship. The non-linear function can be obtained using non-linear regression. Non-linear regression can be a form of regression analysis where the observed data is a non-linear combination of model parameters and is modeled by a function that depends on one or more independent variables. Non-linear regression can include step functions, piecewise functions, splines and generalized additive models.

[0297]

[0515] In some embodiments, the presentability is presented by a one-dimensional value (e.g., probability). In some embodiments, the probability is configured to measure the likelihood that an event ca...

Claims

1. A method for producing an HLA class II tetramer or polymer containing an epitope, comprising the step of contacting purified soluble HLA-DM loaded with a peptide epitope with an HLA class II tetramer or polymer to form the HLA class II tetramer or polymer loaded with the peptide epitope.

2. The method according to claim 1, wherein the HLA class II tetramer or polymer comprises one of HLA-DR, HLA-DP, or HLA-DQ heterodimers, and each heterodimer comprises alpha and beta chains.

3. The method according to claim 1, comprising the step of expressing the alpha and beta chains of an HLA class II tetramer or polymer in a cell.

4. The method according to claim 3, wherein the expression step comprises expressing the alpha and beta chains of an HLA class II tetramer or polymer from a polynucleic acid molecule containing an alpha-coding sequence and a beta-coding sequence, wherein the alpha-coding sequence and the beta-coding sequence are separated by a ribosome skipping sequence or a sequence encoding a protease cleavage site.

5. The method according to claim 3, wherein the expression step includes expressing the alpha and beta chains of an HLA class II tetramer or polymer in a eukaryotic cell.

6. The method according to claim 1, comprising the step of purifying an HLA class II tetramer or polymer.

7. The method according to claim 6, wherein the purification step includes gel filtration chromatography.

8. A method for purifying an epitope-loaded HLA class II tetramer or polymer, comprising: (a) loading a peptide epitope onto an HLA class II tetramer or polymer by contacting a purified soluble HLA-DM loaded with a peptide epitope onto the HLA class II tetramer or polymer; and (b) purifying the epitope-loaded HLA class II tetramer or polymer using a suitable protein purification method, wherein the purification step comprises purifying a peptide epitope-loaded HLA class II tetramer or polymer having a low level of post-translational modification from an HLA class II tetramer or polymer having a high level of post-translational modification.

9. A vector comprising a sequence encoding the alpha chain of an HLA class II tetramer or polymer and a sequence encoding the beta chain of an HLA class II tetramer or polymer, wherein the sequence encoding the alpha chain and the sequence encoding the beta chain are separated by a ribosome skipping sequence or a sequence encoding a protease cleavage site.

10. A method for identifying epitopes of HLA class II tetramers or polymers, (a) Incubating an HLA class II tetramer or polymer in the presence of (i) soluble HLA-DM, (ii) a peptide probe containing a detectable label, and (iii) a candidate peptide epitope, thereby forming a first complex comprising the HLA class II tetramer or polymer and the peptide probe, and a second complex comprising the HLA class II tetramer or polymer and the candidate peptide epitope; (b) A step of measuring the labeling of the peptide probe; (c) Step of identifying candidate peptide epitopes as HLA class II tetramer or polymer epitopes based on (b) Methods that include...

11. The method according to claim 10, wherein the detectable label is a fluorescent label, and the step of measuring the label of the peptide probe includes the step of measuring the fluorescence polarization.

12. The method according to claim 10, wherein an HLA class II tetramer or polymer is loaded with a placeholder peptide before incubation.

13. The method according to claim 12, wherein the placeholder peptide is a peptide selected from the peptides in Table 19.

14. The method according to claim 10, wherein the candidate peptide epitope is encoded by the genome or exome of the subject, or by a pathogen or virus in the subject.

15. The method according to claim 10, wherein the peptide probe is a validated epitope of an HLA class II tetramer or polymer.

16. The method according to claim 10, wherein the peptide probe comprises the probe sequence of Table 20, and the HLA class II tetramer or polymer comprises a protein encoded by the corresponding HLA allele of Table 20.

17. A library of HLA class II tetramers or polymers and matched and validated epitopes of HLA class II tetramers or polymers, wherein the validated epitopes include the probe sequences of Table 20, and the HLA class II tetramers or polymers include proteins encoded by the corresponding HLA alleles of Table 20.

18. The library according to claim 17, wherein the verified epitopes are fluorescently labeled.

19. A composition comprising purified soluble HLA-DM, wherein the purified soluble HLA-DM is present at a concentration higher than 1 mg / L.

20. A method for producing purified soluble HLA-DM, comprising the steps of: expressing a polynucleic acid sequence encoding an HLA-DM protein in a eukaryotic cell, wherein the HLA-DM protein is secreted from the eukaryotic cell; and purifying the HLA-DM protein from the supernatant of the eukaryotic cell.

21. An HLA class II tetramer or polymer comprising one of HLA-DR, HLA-DP, or HLA-DQ heterodimers, wherein each heterodimer contains alpha and beta chains, the heterodimers are purified, and are present at a concentration higher than 16 mg / L.

22. (a) A step of processing amino acid information of multiple candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate multiple presentation predictions, wherein each candidate peptide sequence of the multiple candidate peptide sequences is encoded by the genome or exome of the subject or by a pathogen or virus in the subject, the multiple presentation predictions include an HLA presentation prediction for each of the multiple candidate peptide sequences, and each HLA presentation prediction indicates the possibility that one or more proteins encoded by a class II HLA allele of the cell of the subject can present a given candidate peptide sequence of the multiple candidate peptide sequences. A machine learning HLA peptide presentation prediction model: (A) Incubating an HLA class II tetramer or polymer in the presence of a training peptide epitope loaded with an epitope comprising (i) a soluble HLA-DM, (ii) a fluorescently labeled peptide probe, and (iii) an HLA class II tetramer or polymer, thereby forming a first complex comprising the HLA class II tetramer or polymer and the fluorescently labeled peptide probe, and a second complex comprising the HLA class II tetramer or polymer and the training peptide epitope; (B) Measuring polarization; and (C) Determining the affinity of the training peptide epitope to the HLA class II tetramer or polymer based on (B). The training is performed using training data that includes sequence information of the training peptide sequence identified by a method including the steps; and (b) Step of identifying peptide sequences of multiple peptide sequences presented by at least one of one or more proteins encoded by a class II HLA allele of the target cell, based on at least multiple presentation predictions. A method that includes this.

23. The HLA proteins are HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04: 01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA-D QB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01: 02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-D RB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10: 01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-D The method according to any one of claims 1 to 8, 10 to 16, and 20 to 22, wherein the HLA class II protein is selected from the group consisting of RB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, and HLA-DRB5*01:

01.

24. The HLA proteins are HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04 :01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA -DQB1*02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1* 01:02, HLA-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, H LA-DRB1*04:07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB 1*10:01, HLA-DRB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02 A composition comprising the library according to claim 17 or 18, wherein the HLA class II protein is selected from the group consisting of HLA-DRB1*13:03, HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, and HLA-DRB5*01:

01.

25. HLA proteins are HLA-DPB1*01:01 / HLA-DPA1*01:03, HLA-DPB1*02:01 / HLA-DPA1*01:03, HLA-DPB1*03:01 / HLA-DPA1*01:03, HLA-DPB1*04:01 / HLA-DPA1*01:03, HLA-DPB1*04:02 / HLA-DPA1*01:03, HLA-DPB1*06:01 / HLA-DPA1*01:03, HLA-DQB1*02:01 / HLA-DQA1*05:01, HLA-DQB1 *02:02 / HLA-DQA1*02:01, HLA-DQB1*06:02 / HLA-DQA1*01:02, HLA-DQB1*06:04 / HLA-DQA1*01:02, HLA-DRB1*01:01, HLA-DRB1*01:02, HL A-DRB1*03:01, HLA-DRB1*03:02, HLA-DRB1*04:01, HLA-DRB1*04:02, HLA-DRB1*04:03, HLA-DRB1*04:04, HLA-DRB1*04:05, HLA-DRB1*04 :07, HLA-DRB1*07:01, HLA-DRB1*08:01, HLA-DRB1*08:02, HLA-DRB1*08:03, HLA-DRB1*08:04, HLA-DRB1*09:01, HLA-DRB1*10:01, HLA-D RB1*11:01, HLA-DRB1*11:02, HLA-DRB1*11:04, HLA-DRB1*12:01, HLA-DRB1*12:02, HLA-DRB1*13:01, HLA-DRB1*13:02, HLA-DRB1*13:03 A composition comprising a vector comprising a sequence encoding an HLA class II tetramer or polymer according to claim 9, which is an HLA class II protein selected from the group consisting of HLA-DRB1*14:01, HLA-DRB1*15:01, HLA-DRB1*15:02, HLA-DRB1*15:03, HLA-DRB1*16:01, HLA-DRB3*01:01, HLA-DRB3*02:02, HLA-DRB3*03:01, HLA-DRB4*01:01, and HLA-DRB5*01:

01.

26. The method according to claim 10 or 11, wherein the detection sensitivity limit is 0.001% by a standard fluorescence measurement method.

27. Use of HLA class II tetramers or multimers loaded with peptide epitopes prepared according to the method of claim 1 for identifying antigen-specific CD4+ T cells.

28. Use of HLA class II tetramers or polymers loaded with a peptide epitope purified according to the method of claim 8 for identifying antigen-specific CD4+ T cells.

29. Use of the library of HLA class II tetramers or multimers and matched and validated epitopes of HLA class II tetramers or multimers according to claim 17 for identifying antigen-specific CD4+ T cells.

30. Use of the composition according to claim 19 for identifying antigen-specific CD4+ T cells.

31. Use of purified soluble HLA-DM prepared according to the method of claim 20 for identifying antigen-specific CD4+ T cells.

32. Use of the HLA class II tetramer or multimer according to claim 21 for identifying antigen-specific CD4+ T cells.