Methods and systems for predicting HLA class II specific epitopes and characterizing CD4+ T cells
A machine learning-based method using mass spectrometry data and quality metrics enhances the prediction of HLA class II epitopes, achieving high positive predictive value and improving therapeutic efficacy by accurately identifying cancer-specific antigens.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BIONTECH US INC
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-29
AI Technical Summary
Current methods for predicting HLA class II epitopes are not accurate and do not effectively translate high CD4+ T cell responses into therapeutic efficacy due to unclear presentation pathways and lack of robust quantification of gene expression, enzymatic cleavage, and localization bias, leading to inefficiencies in identifying truly presented HLA class II-binding epitopes.
A machine learning-based approach for predicting HLA class II epitopes using mass spectrometry data and quality metrics to improve the identification of allele-specific binding rules, incorporating gene expression, cleavage, and localization bias, and validating through T cell induction to enhance predictive models.
The method achieves a positive predictive value (PPV) of at least 0.07 to 0.99 in identifying HLA class II-binding epitopes, significantly improving the accuracy of predicting cancer-specific antigens and translating CD4+ T cell responses into therapeutic efficacy.
Smart Images

Figure 2026123277000026 
Figure 2026123277000027 
Figure 2026123277000028
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefits of U.S. Provisional Application No. 62 / 783,914 filed December 21, 2018; U.S. Provisional Application No. 62 / 826,827 filed March 29, 2019; U.S. Provisional Application No. 62 / 855,379 filed May 31, 2019; and U.S. Provisional Application No. 62 / 891,101 filed August 23, 2019, each of which is incorporated herein by reference in its entirety. [Background technology]
[0002] background The major histocompatibility complex (MHC) is a gene complex that encodes human leukocyte antigen (HLA) genes. HLA genes are expressed as protein heterodimers that are displayed on the surface of human cells in relation to circulating T cells. HLA genes are highly polymorphic, which allows for fine-tuning of the adaptive immune system. The adaptive immune response relies, in part, on the ability of T cells to identify and eliminate cells that display disease-related peptide antigens bound to human leukocyte antigen (HLA) heterodimers.
[0003] In humans, endogenous and exogenous proteins are processed by the proteasome, as well as by proteases and peptidases in the cytoplasm and endosomes / lysosomes, to form peptides, which are then presented by two classes of cell surface proteins encoded by MHC genes. These cell surface proteins are called human leukocyte antigens (HLA class I and class II), and the group of peptides that bind to them and elicit an immune response are called HLA epitopes. HLA epitopes are crucial components that enable the immune system to detect danger signals such as pathogen infection and transformation of self. CD4+ T cells recognize class II MHC (HLA-DR, HLA-DQ, and HLA-DP) epitopes displayed on antigen-presenting cells (APCs) such as dendritic cells and macrophages. Endogenous processing and presentation of HLA class II ligands is a complex process requiring various chaperones and subsets of enzymes, not all of which are well-characterized. Presentation of HLA class II peptides activates helper T cells, which subsequently promotes B cell differentiation, antibody production, and CTL response. Activated helper T cells also secrete cytokines and chemokines that activate other T cells and induce their differentiation.
[0004] Understanding the peptide bond preferences for each HLA class II heterodimer is key to successfully predicting which cancer or tumor-specific antigens are most likely to elicit a cancer or tumor-specific T cell response. Methods for identifying and isolating specific HLA class II-related peptides (e.g., novel antigen peptides) are needed. Such methodologies and isolated molecules would be useful in the development of therapeutics, including, but not limited to, immunotherapy. [Overview of the project] [Means for solving the problem]
[0005] overview The methods and compositions described herein can be used in a wide range of applications. For example, the methods and compositions described herein can be used to identify immunogenic antigen peptides, and can also be used for the development of drugs such as personalized medicines, as well as for the isolation and characterization of antigen-specific T cells.
[0006] CD4+ T cell responses may possess antitumor activity. Without class II prediction, CD4+ T cell responses may be shown at a high rate (e.g., 60% of SLP epitopes in a NeoVax study (49% in NT-001, see Ott et al., Nature, 2017 Jul 13; 547 (7662): 217-221), and Biont In tests conducted by ech, 48% of mRNA epitopes (see Sahin et al., Nature, 2017 Jul 13; 547 (7662): 222-226). These epitopes are generally It may not always be clear whether the HLA class II-binding epitopes are presented natively (by the tumor or by phagocytic DCs). It may be desirable to translate high CD4+ T response rates into therapeutic efficacy by improving the identification of truly presented HLA class II-binding epitopes.
[0007] The roles of gene expression, enzymatic cleavage, and pathway / localization bias may not be robustly quantified. It may be unclear whether the more relevant pathway is autophagy (HLA class II presentation by tumor cells) or phagocytosis (HLA class II presentation of tumor epitopes by APCs), but the majority of existing MS data can be presumed to originate from autophagy. NetMHCIIpan may be the current predictive standard, but it may not always be accurate. Of the three HLA class II loci (DR, DP, and DQ), data may only exist for a specific common allele of HLA-DR.
[0008] Regarding the learning of HLA Class II presentation rules, various data generation methods may exist, including the field standard and the proposed method. The field standard may include affinity measurements that could form the basis of a NetMHCIIpan predictor, resulting in low throughput and the need for radioactive reagents, and the processing role is lacking in the field standard. The proposed method may include mass spectrometry, in which case data from cell lines / tissues / tumors may be useful in determining processing rules related to autophagy, and single-allele MS may enable the determination of allele-specific binding rules (multiple-allele MS data are presumed to be too complex for efficient learning (Bassani-Sternberg. MCP. 2018)).
[0009] There are various ways to validate a new HLA class II predictor, including: validation against submitted MS data, which may be the default setting; retrospective vaccine trials (e.g., NT-001) where immunomonitoring data can be used to assess vaccine peptide loading against APC rather than tumor presentation, and the data can be spread thinly across many different alleles; biochemical affinity measurements that can be configured to obtain measurements for peptides that are mispredicted (for only 2-3 alleles); and ex by Neon-preferential epitopes and NetMHCIIpan-preferential epitopes. T cell induction can be configured to test the rate at which a T cell response is induced in vivo.
[0010] Regarding validation by T cell induction, the default approach may involve evaluating the neoORF of the incongruously predicted TCGA, with induction material including healthy donor APCs and T cells, and induction and reading may be performed by SLP (a peptide of approximately 15 mers). Random peptides may result in high response rates, and SLP may not adequately address processing. Possible solutions may include induction by mRNA.
[0011] Methods disclosed herein may include generating LC-MS / MS single-allele data for training allele-specific machine learning methods for epitope prediction. Such methods may include improving LC-MS / MS data quality by utilizing a set of quality metrics to strictly eliminate false positives that improve the performance of predictive models; identifying allele-specific HLA class II binding cores from HLA-ligandome LC-MS / MS datasets; improving HLA class II-ligand and epitope predictions by utilizing machine learning algorithms; and / or identifying biological variables, such as gene expression, cleavage, gene bias, cellular localization, and secondary structure, that influence HLA class II-ligand presentation and improve HLA class II epitope prediction.
[0012] A method is provided herein that includes the steps of (a) processing amino acid information of a plurality of candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate a plurality of presentation predictions, wherein each of the plurality of candidate peptide sequences is encoded by the genome or exome of the plurality of candidate peptide sequences, the plurality of presentation predictions include an HLA presentation prediction for each of the plurality of candidate peptide sequences, each HLA presentation prediction indicates the possibility that a given candidate peptide sequence among the plurality of candidate peptide sequences may be presented by one or more proteins encoded by a class II HLA allele of the cell of the subject, and the machine learning HLA peptide presentation prediction model is trained using training data which includes sequence information of training peptide sequences identified by mass spectrometry as being presented by HLA proteins expressed in training cells; and (b) identifying, based on at least a plurality of presentation predictions, a peptide sequence among the plurality of peptide sequences which is presented by at least one of one or more proteins encoded by a class II HLA allele of the cell of the subject, wherein the positive predictive value (PPV) of the machine learning HLA peptide presentation prediction model is at least 0.07 according to the presentation PPV determination method.
[0013] A method is provided herein that includes (a) processing amino acid information of a plurality of peptide sequences encoded by a genome or exome of interest using a machine learning HLA peptide binding prediction model to generate a plurality of binding predictions, wherein the plurality of binding predictions include an HLA binding prediction for each of a plurality of candidate peptide sequences, each binding prediction indicating the likelihood that one or more proteins encoded by a class II HLA allele of the cell of interest will bind to a given candidate peptide sequence among the plurality of candidate peptide sequences, and the machine learning HLA peptide binding prediction model is trained using training data which includes sequence information of peptide sequences identified to bind to HLA class II proteins or HLA class II protein analogs; and (b) identifying peptide sequences among the plurality of peptide sequences that have a probability greater than a threshold binding prediction probability value for binding to at least one of the one or more proteins encoded by a class II HLA allele of the cell of interest, based on at least a plurality of binding predictions, wherein the positive predictive value (PPV) of the machine learning HLA peptide binding prediction model is at least 0.1 according to the binding PPV determination method.
[0014] In some embodiments, a machine learning HLA peptide presentation prediction model is trained using training data that includes sequence information of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in training cells.
[0015] In some embodiments, the method includes ranking at least two peptides that have been identified as being presented by at least one of the one or more proteins encoded by a class II HLA allele of the cell in question, based on presentation prediction.
[0016] In some embodiments, the method involves selecting one or more peptides from two or more ranked peptides.
[0017] In some embodiments, the method involves selecting one or more peptides identified as being presented by at least one of the one or more proteins encoded by a class II HLA allele of the cell of interest.
[0018] In some embodiments, the method includes selecting one or more peptides from two or more peptides ranked based on presentation predictions.
[0019] In some embodiments, the positive predictive value (PPV) of the machine learning HLA peptide presentation prediction model is at least 0.07 if the amino acid information of multiple test peptide sequences is processed to generate multiple test presentation predictions, and each test presentation prediction indicates that a given test peptide sequence among the multiple test peptide sequences may be presented by one or more proteins encoded by class II HLA alleles of the target cell, and the multiple test peptide sequences comprise at least 500 test peptide sequences, including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 499 decoy peptide sequences contained in a protein encoded by the organism's genome, and the organism and the target are of the same species, and the multiple test peptide sequences comprise at least one hit peptide sequence and at least 499 decoy peptide sequences in a 1:499 ratio, and the machine learning HLA peptide presentation prediction model predicts that the top percentage of the multiple test peptide sequences will be presented by an HLA protein expressed in the cell.
[0020] In some embodiments, the positive predictive value (PPV) of the machine learning HLA peptide presentation prediction model is at least 0.1 if the amino acid information of multiple test peptide sequences is processed to generate multiple test binding predictions, and each test binding prediction indicates that one or more proteins encoded by class II HLA alleles of the target cell are likely to bind to a given test peptide sequence among the multiple test peptide sequences, and the multiple test peptide sequences include at least 20 test peptide sequences, each including (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, and (ii) at least 19 decoy peptide sequences contained within a protein containing at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in the cell, e.g., a single HLA protein expressed in the cell (e.g., a single-allele cell), and the multiple test peptide sequences include at least one hit peptide sequence and at least 19 decoy peptide sequences in a 1:19 ratio, and the machine learning HLA peptide presentation prediction model predicts that the top percentage of the multiple test peptide sequences will bind to an HLA protein expressed in the cell.
[0021] In some embodiments, there is no overlap in amino acid sequences between at least one hit peptide sequence and a decoy peptide sequence.
[0022] In some embodiments, the positive predictive value (PPV) of the machine learning HLA peptide presentation prediction model is at least 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49 , 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0. 75, 0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99.
[0023] In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 Contains hit peptide sequences of 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100.
[0024] In some embodiments, at least 499 decoy peptide sequences are at least 500 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 540 0, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 1 2000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 5 5000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000,It contains a decoy peptide sequence of 800,000, 900,000, or 1,000,000. Those skilled in the art will understand that the PPV changes with variations in the hit:decoy ratio.
[0025] In some embodiments, at least 500 test peptide sequences are at least 600, 70 0, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 320 0, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 57 00, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 1500 0, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 3 6000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000 Includes test peptide sequences of 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000, or 1000000.
[0026] In some embodiments, the top percentages are: top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%, 1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%. %, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00%, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5 0.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.30%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00 These are %, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%, 9.90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%.
[0027] In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 Contains hit peptide sequences of 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100.
[0028] In some embodiments, at least 19 decoy peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 310 0, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 560 0, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 81 00, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15 000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000 , containing decoy peptide sequences of 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000.
[0029] In some embodiments, at least 20 test peptide sequences include at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 44 00, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8 200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 260 00, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500,Includes test peptide sequences of 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000 or 1000000.
[0030] In some embodiments, the top percentages are the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40%.
[0031] In some embodiments, the PPV is greater than the respective PPV in the second column of Table 11 for the protein encoded by the corresponding HLA allele in Table 11. In some embodiments, the PPV is at least equal to the respective PPV in the third column of Table 11 for the protein encoded by the corresponding HLA allele in Table 11.
[0032] In some embodiments, the PPV is equal to or greater than the respective PPV in the second column of Table 12 for the protein encoded by the HLA class II allele.
[0033] In some embodiments, the PPV is greater than the respective PPV in the second column of Table 16 for the proteins encoded by the HLA class II allele.
[0034] In some embodiments, the object is a single object.
[0035] In some embodiments, the subject is a mammal.
[0036] In some embodiments, the subject is a human.
[0037] In some embodiments, the training cells are cells that express a single protein encoded by the class II HLA allele of the target cell.
[0038] In some embodiments, the training cells are single-allele HLA cells or cells expressing an HLA allele with an affinity tag.
[0039] In some embodiments, the target cells include cancer cells.
[0040] In some embodiments, the method is for identifying peptide sequences.
[0041] In some embodiments, the method is for selecting a peptide sequence.
[0042] In some embodiments, the method is for preparing cancer treatments.
[0043] In some embodiments, the method is for preparing a target-specific cancer treatment.
[0044] In some embodiments, the method is for preparing cancer cell-specific cancer therapies.
[0045] In some embodiments, each peptide sequence in a plurality of peptide sequences is associated with cancer.
[0046] In some embodiments, at least one of a plurality of peptide sequences is overexpressed by the target cancer cells.
[0047] In some embodiments, each peptide sequence of a plurality of peptide sequences is overexpressed by the target cancer cells.
[0048] In some embodiments, at least one of the multiple peptide sequences is a cancer cell-specific peptide.
[0049] In some embodiments, each peptide sequence in a plurality of peptide sequences is a cancer cell-specific peptide.
[0050] In some embodiments, each peptide sequence of a plurality of peptide sequences is expressed by the target cancer cells.
[0051] In some embodiments, at least one of the multiple peptide sequences is not encoded by the non-cancer cells in question.
[0052] In some embodiments, each peptide sequence in the plurality of peptide sequences is not encoded by the target non-cancer cells.
[0053] In some embodiments, at least one of the multiple peptide sequences is not expressed in the target non-cancer cells.
[0054] In some embodiments, each peptide sequence of the multiple peptide sequences is not expressed in the target non-cancer cells.
[0055] In some embodiments, the method includes obtaining a plurality of target peptide sequences.
[0056] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences of interest.
[0057] In some embodiments, the method involves obtaining a plurality of polynucleotide sequences of a subject that encode a plurality of peptide sequences encoded by the genome or exome of the subject or by a pathogen or virus in the subject.
[0058] In some embodiments, the method includes obtaining a target polynucleotide sequence encoding a target multiple peptide sequence encoded by the target genome or exome using a computer processor.
[0059] The method according to any one claim, wherein in some embodiments the method comprises obtaining a plurality of polynucleotide sequences of interest by genome sequencing or exome sequencing.
[0060] In some embodiments, the method includes obtaining a plurality of polynucleotide sequences of interest by whole-genome sequencing or whole-exome sequencing.
[0061] In some embodiments, the processing step includes being processed by a computer processor.
[0062] In some embodiments, the processing step includes generating a plurality of predictor variables based on amino acid information of at least a plurality of peptide sequences.
[0063] In some embodiments, a machine learning HLA-peptide presentation prediction model is used to process multiple predictor variables.
[0064] In some embodiments, one or more proteins encoded by the class II HLA alleles of the target cell are one or more proteins encoded by the class II HLA alleles expressed by the target.
[0065] In some embodiments, one or more proteins encoded by the class II HLA alleles of the target cell are one or more proteins encoded by the class II HLA alleles expressed by the target cancer cell.
[0066] In some embodiments, one or more proteins encoded by the class II HLA alleles of the cell in question are a single protein encoded by the class II HLA alleles of the cell in question.
[0067] In some embodiments, one or more proteins encoded by the class II HLA alleles of the cell in question are two, three, four, five, or six or more proteins encoded by the class II HLA alleles of the cell in question.
[0068] In some embodiments, one or more proteins encoded by the class II HLA alleles of the target cell are each protein encoded by the class II HLA alleles of the target cell.
[0069] In some embodiments, the method further includes the step of administering to a subject a composition comprising one or more selected subsets of peptide sequences.
[0070] In some embodiments, the step of identifying a plurality of peptide sequences includes comparing DNA, RNA, or protein sequences derived from the target cancer cells with DNA, RNA, or protein sequences derived from the target normal cells, where each of the plurality of peptides contains at least one mutation present in the target cancer cells but not in the target normal cells.
[0071] In some embodiments, a machine learning HLA-peptide presentation prediction model includes at least several predictor variables identified based on training data, wherein the training data includes training peptide sequence information including amino acid position information, and the training peptide sequence information includes several predictor variables associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information and the presentation likelihood generated as an output based on the amino acid position information and the several predictor variables.
[0072] In some embodiments, the identifying step comprises identifying a peptide sequence among a plurality of peptide sequences, wherein the probability of being presented by at least one of one or more proteins encoded by the class II HLA alleles of the target cell is greater than a threshold value of the presentation prediction probability value.
[0073] In some embodiments, for one or more of 0.2% of a plurality of test peptide sequences predicted to be presented by a machine learning HLA peptide presentation prediction model, the probability of being presented by at least one of one or more proteins encoded by the class II HLA alleles of the target cell is higher than the threshold value of the presentation prediction probability value.
[0074] In some embodiments, for each of 0.2% of a plurality of test peptide sequences predicted to be presented by a machine learning HLA peptide presentation prediction model, the probability of being presented by at least one of one or more proteins encoded by the class II HLA alleles of the target cell is higher than the threshold value of the presentation prediction probability value.
[0075] In some embodiments, the number of positives is constrained to be equal to the number of hits.
[0076] In some embodiments, the mass spectrometry is single-allele mass spectrometry.
[0077] In some embodiments, the peptide is presented by an HLA protein expressed in the cell through autophagy.
[0078] In some embodiments, the peptide is presented by an HLA protein expressed in the cell through phagocytosis.
[0079] In some embodiments, the plurality of predictor variables includes an expression level predictor of a source protein containing the peptide.
[0080] In some embodiments, the multiple predictor variables include a stability predictor for source proteins, including peptides.
[0081] In some embodiments, the multiple predictor variables include predictors of the degradation rate of source proteins, including peptides.
[0082] In some embodiments, the multiple predictor variables include a protein cleavage predictor for a source protein containing a peptide.
[0083] In some embodiments, the multiple predictor variables include cellular or tissue localization predictors for source proteins, including peptides.
[0084] In some embodiments, the multiple predictor variables include a predictor of the intracellular processing scheme of the source protein, which includes a predictor of whether the source protein is subjected to autophagy, phagocytosis, and intracellular transport, among other things.
[0085] In some embodiments, the quality of the training data is improved using multiple quality metrics.
[0086] In some embodiments, multiple quality metrics include general contaminating peptide removal, high scoring peak intensity, high score, and high mass accuracy.
[0087] In some embodiments, the scored peak intensity is at least 50%.
[0088] In some embodiments, the scored peak intensity is at least 60%.
[0089] In some embodiments, the score is at least 7.
[0090] In some embodiments, the mass accuracy is up to 5 ppm.
[0091] In some embodiments, the peptides presented by HLA proteins expressed in a cell are peptides presented by a single immunoprecipitated HLA protein expressed in the cell.
[0092] In some embodiments, the peptides presented by HLA proteins expressed in a cell are peptides presented by a single exogenous HLA protein expressed in the cell.
[0093] In some embodiments, the peptides presented by HLA proteins expressed in a cell are peptides presented by a single recombinant HLA protein expressed in the cell.
[0094] In some embodiments, the plurality of predictor variables includes peptide-HLA affinity predictor variables.
[0095] In some embodiments, the peptides presented by HLA proteins include peptides identified by searching an enzyme-specificity-free and modification-free peptide database.
[0096] In some embodiments, the peptides presented by HLA proteins include peptides identified by searching a peptide database using an inverse database search strategy.
[0097] In some embodiments, the HLA proteins include HLA-DR proteins, HLA-DQ proteins, or HLA-DP proteins.
[0098] In some embodiments, the HLA protein is HLA-DPB1 * 01:01 / HLA-DPA1 * HLA-DPB1 01:03 * 02:01 / HLA-DPA1 * HLA-DPB1 01:03 * 03:01 / HLA-DPA1 * HLA-DPB1 01:03 *04:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 04:02 / HLA-DPA1 * 01:03、HLA-DPB1 * 06:01 / HLA-DPA1 * 01:03、HLA-DQB1 * 02:01 / HLA-DQA1 * 05:01、HLA-DQB1 * 02:02 / HLA-DQA1 * 02:01、HLA-DQB1 * 06:02 / HLA-DQA1 * 01:02、HLA-DQB1 * 06:04 / HLA-DQA1 * 01:02、HLA-DRB1 * 01:01、HLA-DRB1 * 01:02、HLA-DRB1 * 03:01、HLA-DRB1 * 03:02、HLA-DRB1 * 04:01、HLA-DRB1 * 04:02、HLA-DRB1 * 04:03、HLA-DRB1 * 04:04、HLA-DRB1 * 04:05、HLA-DRB1 * 04:07、HLA-DRB1 * 07:01、HLA-DRB1 * 08:01、HLA-DRB1 * 08:02、HLA-DRB1 * 08:03、HLA-DRB1 * 08:04、HLA-DRB1 * 09:01、HLA-DRB1 * 10:01、HLA-DRB1 * 11:01、HLA-DRB1 * 11:02、HLA-DRB1 * 11:04、HLA-DRB1 * 12:01、HLA-DRB1 * 12:02、HLA-DRB1 * 13:01、HLA-DRB1 *13:02, HLA-DRB1 * 13:03, HLA-DRB1 * 14:01, HLA-DRB1 * 15:01, HLA-DRB1 * 15:02, HLA-DRB1 * 15:03, HLA-DRB1 * 16:01, HLA-DRB3 * 01:01, HLA-DRB3 * 02:02, HLA-DRB3 * 03:01, HLA-DRB4 * 01:01, HLA-DRB5 * It contains HLA class II proteins selected from the group consisting of 01:01.
[0099] In some embodiments, HLA-DR is DRA * It forms a pair with 01:01.
[0100] In some embodiments, the HLA protein is DPA * 01:03 / DPB * 04:01, DRB1 * 01:01, DRB1 * 01:02, DRB1 * 03:01, DRB1 * 04:01, DRB1 * 04:02, DRB1 * 04:04, DRB1 * 04:05, DRB1 * 07:01, DRB1 * 08:01, DRB1 * 08:02, DRB1 * 08:03, DRB1 * 09:01, DRB1 * 11:01, DRB1 * 11:02, DRB1 * 11:04, DRB1 * 12:01, DRB1 * 13:01, DRB1 * 13:02, DRB1 * 13:03, DRB1 * 14:01, DRB1 *15:01, DRB1 * 15:02, DRB1 * 15:03, DRB1 * 16:02, DRB3 * 01:01, DRB3 * 02:01, DRB3 * 02:02, DRB3 * 03:01, DRB4 * 01:01, DRB4 * 01:03 and DRB5 * It is an HLA class II protein selected from the group consisting of 01:01.
[0101] In some embodiments, the HLA-DR protein comprises DRA * 01:01 in the dimer.
[0102] In some embodiments, the HLA protein comprises DPB1 * 01:01, DPB1 * 02:01, DPB1 * 02:02, DPB1 * 03:01, DPB1 * 04:01, DPB1 * 04:02, DPB1 * 05:01, DPB1 * 06:01, DPB1 * 11:01, DPB1 * 13:01, DPB1 * It comprises an HLA-DP protein selected from the group consisting of 17:01.
[0103] * 06:03, A1 * 02:01+B1 * 02:02, A1 * 02:01+B1 * 03:03, A1 * 03:01+B1 * 03:02, A1 * 03:03+B1 * 03:01, A1 * 05:01+B1 * 02:01 and A1 * 05:05+B1 * It contains an HLA-DQ protein complex selected from the group consisting of 03:01.
[0105] In some embodiments, the peptides presented by the HLA protein include peptides identified by comparing the MS / MS spectrum of the HLA-peptide with the MS / MS spectra of one or more peptides or proteins in a peptide or protein database.
[0106] In some embodiments, the mutation is selected from the group consisting of point mutations, splice site mutations, frameshift mutations, readthrough mutations, and gene fusion mutations.
[0107] In some embodiments, the length of the peptide presented by the HLA protein is 15 to 40 amino acids.
[0108] In some embodiments, the peptide presented by the HLA protein includes peptides identified by comparing the MS / MS spectrum of the HLA-peptide with the MS / MS spectrum of one or more peptides or proteins in a peptide or protein database.
[0109] In some embodiments, personalized cancer therapy further includes an adjuvant.
[0110] In some embodiments, personalized cancer therapy further includes immune checkpoint inhibitors.
[0111] In some embodiments, the training data includes structured data, time-series data, unstructured data, relational data, or any combination thereof.
[0112] In some embodiments, unstructured data includes image data.
[0113] In some embodiments, relational data includes data from customer systems, enterprise systems, operational systems, websites, web-accessible application programming interfaces (APIs), or any combination thereof.
[0114] In some embodiments, training data is uploaded to a cloud-based database.
[0115] In some embodiments, training is performed using a convolutional neural network.
[0116] In some embodiments, the convolutional neural network includes at least two convolutional layers.
[0117] In some embodiments, the convolutional neural network includes at least one batch normalization step.
[0118] In some embodiments, the convolutional neural network includes at least one spatial dropout step.
[0119] In some embodiments, the convolutional neural network includes at least one global max pooling step.
[0120] In some embodiments, the convolutional neural network includes at least one high-density layer.
[0121] In some embodiments, the step of identifying the peptide sequence includes identifying the peptide sequence using mutations expressed in the target cancer cells.
[0122] In some embodiments, the step of identifying peptide sequences includes identifying peptide sequences that are not expressed in the normal cells of interest.
[0123] In some embodiments, the step of identifying a peptide sequence includes identifying a viral peptide sequence.
[0124] In some embodiments, the step of identifying the peptide sequence includes identifying the overexpressed peptide sequence.
[0125] A method for identifying HLA class II specific peptides for immunotherapy against a target, comprising the steps of: obtaining candidate peptides containing epitopes and multiple peptide sequences, each containing an epitope, using a computer processor; and processing the amino acid information of the multiple peptide sequences using a machine learning HLA-peptide presentation prediction model with a computer processor to generate presentation predictions for each of the multiple peptide sequences to immune cells, wherein each presentation prediction indicates the possibility that a given peptide sequence among the multiple peptide sequences may be presented by one or more proteins encoded by an HLA class II allele, and the machine learning HLA-peptide presentation prediction model is used to determine the presentation of peptides by HLA proteins expressed in cells and identified by mass spectrometry. A method is provided herein that includes the steps of: training using training data containing sequence information of columns; selecting a protein from one or more proteins encoded by the HLA class II alleles of a target cell that is predicted by a machine learning HLA-peptide presentation prediction model to bind to a candidate peptide, wherein the protein has a probability greater than a threshold of presentation prediction probability value for presentation of the candidate peptide to immune cells; contacting the selected protein with the selected protein, thereby causing the candidate peptide to compete with a placeholder peptide associated with the selected protein; and identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein, based on whether the candidate peptide replaces the placeholder.
[0126] In some embodiments, the obtaining step includes identifying candidate peptides, which includes comparing DNA, RNA, or protein sequences derived from target cancer cells with DNA, RNA, or protein sequences derived from target normal cells.
[0127] In some embodiments, the processing steps include identifying multiple predictor variables based on amino acid information of at least multiple peptide sequences, and processing the multiple predictor variables using a machine learning HLA-peptide presentation prediction model.
[0128] In some embodiments, a machine learning HLA-peptide presentation prediction model includes at least several predictor variables identified based on training data, wherein the training data is training peptide sequence information including amino acid position information, associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information and the presentation likelihood generated as an output based on the amino acid position information and the several predictor variables.
[0129] In some embodiments, the number of positives is constrained to be equal to the number of hits.
[0130] In some embodiments, the mass spectrometry is single-allele mass spectrometry.
[0131] In some embodiments, the multiple predictor variables include any one or more of the following: expression level predictors, stability predictors, degradation rate predictors, cleavage predictors, cell or tissue localization predictors, and intracellular processing method predictors, including autophagy, phagocytosis, and intracellular transport.
[0132] In some embodiments, the quality of the training data is improved using multiple quality metrics.
[0133] In some embodiments, multiple quality metrics include general contaminating peptide removal, high scoring peak intensity, high score, and high mass accuracy.
[0134] In some embodiments, the scored peak intensity is at least 50%.
[0135] In some embodiments, the scored peak intensity is at least 60%.
[0136] In some embodiments, the placeholder peptide is a CLIP peptide.
[0137] In some embodiments, the placeholder peptide is a CMV peptide.
[0138] In some embodiments, the 3 method further includes the step of measuring the IC50 of the substitution of a placeholder peptide with a target peptide.
[0139] In some embodiments, the IC50 of substitution of a placeholder peptide with a target peptide is less than 500 nM.
[0140] In some embodiments, at least one protein from one or more proteins encoded by the HLA class II alleles of the cells in question is an HLA class II tetramer or multimer.
[0141] In some embodiments, the target peptide is further identified by mass spectrometry.
[0142] In some embodiments, at least one protein encoded by the HLA class II allele of the cell in question is a recombinant protein.
[0143] In some embodiments, at least one protein encoded by the HLA class II allele of the target cell is expressed in eukaryotic cells.
[0144] In some embodiments, the peptide is presented by HLA proteins expressed in cells via autophagy.
[0145] In some embodiments, the peptide is presented by HLA proteins expressed in cells via phagocytosis.
[0146] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single immunoprecipitated HLA protein expressed in the cell.
[0147] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single exogenous HLA protein expressed in the cell.
[0148] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single recombinant HLA protein expressed in the cell.
[0149] In some embodiments, the multiple predictor variables include peptide-HLA affinity predictor variables.
[0150] In some embodiments, the peptides presented by the HLA protein include peptides identified by searching an enzyme-unspecific, unmodified peptide database.
[0151] In some embodiments, the peptides presented by the HLA protein include peptides identified by searching a peptide database using a reverse database search strategy.
[0152] In some embodiments, the HLA protein includes HLA-DR protein, HLA-DQ protein, or HLA-DP protein.
[0153] In some embodiments, the immunotherapy is cancer immunotherapy.
[0154] In some embodiments, the epitope is a cancer-specific epitope.
[0155] In some embodiments, at least one protein encoded by an HLA class II allele comprises at least an alpha-1 subunit and a beta-1 subunit of an HLA protein that exists in a dimeric form.
[0156] In some embodiments, the type of peptide is known.
[0157] In some embodiments, the type of peptide is unknown.
[0158] In some embodiments, the type of peptide is determined by mass spectrometry.
[0159] In some embodiments, the peptide exchange assay includes the detection of a peptide fluorescent probe or tag.
[0160] In some embodiments, the placeholder peptide is a CLIP peptide. In some embodiments, the placeholder peptide has the amino acid sequence PVSKMRMATPLLMQA.
[0161] In some embodiments, the polynucleic acid construct comprises an expression vector further comprising one or more of a promoter, a secretory signal, a dimerizing factor, a ribosome skipping sequence, and one or more tags for purification and / or detection.
[0162] In some embodiments, the placeholder peptide sequence is encoded by a nucleic acid sequence in the vector.
[0163] In some embodiments, the sequence encoding the cleavable domain is placed between the sequence encoding the placeholder peptide and the sequence encoding the HLA beta-1 peptide.
[0164] A method for assaying the immunogenicity of an MHC class II binding peptide is provided herein, comprising the steps of: selecting a protein encoded by an HLA class II allele, which is predicted by a machine learning HLA-peptide presentation prediction model to bind to the MHC class II binding peptide, wherein the machine learning HLA-peptide presentation prediction model is configured to generate presentation predictions for a given peptide sequence, the presentation predictions indicating the possibility that the given peptide sequence may be presented by one or more proteins encoded by an HLA class II allele, and the proteins have a probability greater than a threshold of the presentation prediction probability value for presentation of the MHC class II binding peptide; contacting the peptide with the selected protein, thereby causing the peptide to compete with a placeholder peptide associated with the selected protein, replacing the placeholder peptide, and thereby forming a complex comprising the HLA class II protein and the MHC class II binding peptide; contacting the complex with CD4+ T cells; and assaying for one or more CD4+ T cell activation parameters selected from the group consisting of cytokine induction, chemokine induction, and cell surface marker expression.
[0165] In some embodiments, the HLA class II allele is a tetramer or a multimer.
[0166] In some embodiments, the cytokine is IL-2.
[0167] A method for inducing CD4+ T cell activation in a subject for cancer immunotherapy, comprising the steps of: identifying a peptide sequence associated with cancer and containing cancer mutations, wherein the peptide sequence identification comprises comparing a DNA, RNA, or protein sequence derived from the cancer cells of the subject with a DNA, RNA, or protein sequence derived from normal cells of the subject; and selecting a protein encoded by an HLA class II allele that is normally expressed by the cells of the subject and is predicted by a machine learning HLA-peptide presentation prediction model to bind to a peptide, wherein the positive predictive value of the prediction model is at least 0.1%, 0.1% to 50%, or a maximum recall of 50%, and the protein is the identified peptide. A method is provided herein that includes the steps of: having a probability greater than a threshold of the predicted presentation probability value for the presentation of a cydoid sequence; contacting the identified peptide with a protein encoded by a selected HLA class II allele to verify whether the identified peptide competes for and replaces the placeholder peptide associated with the protein encoded by the selected HLA class II allele, with an IC50 value of less than 500 nM; optionally purifying the identified peptide; and administering an effective amount to a polypeptide or polynucleotide encoding a polypeptide containing the sequence of the identified peptide.
[0168] A method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject is provided herein, comprising the steps of: obtaining a plurality of peptide sequences of the polypeptide sequence by a computer processor; processing the amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model by a computer processor to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject, and the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information related to HLA proteins expressed in the cell; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject; and administering a composition comprising the drug to the subject.
[0169] A method for producing an HLA class II tetramer or multimer by conjugating four individual HLA protein alpha-1 and beta-1 heterodimers is provided herein, comprising the steps of: expressing a vector in eukaryotic cells containing nucleic acid sequences encoding the alpha and beta chains of the HLA protein, a secretory signal, a biotinylated motif, and at least one tag for identification or purification, thereby causing each HLA protein alpha-1 and beta-1 heterodimer to be secreted in a dimerized state, wherein the heterodimer is associated with a placeholder peptide; purifying the secreted heterodimer from cell culture medium; verifying peptide binding activity using a peptide exchange assay; adding streptavidin to thereby conjugate the heterodimer into a tetramer; and purifying the tetramer to obtain a yield greater than 1 mg / L. Multimers, such as pentamers, hexamers, or octamers, can also be produced and are similarly intended herein.
[0170] In some embodiments, the vector includes a CMV promoter.
[0171] In some embodiments, the vector includes a sequence encoding a placeholder peptide linked to a beta-1 chain via a cleavable site.
[0172] In some embodiments, the peptide exchange assay involves pre-cleaving the placeholder peptide from the beta chain.
[0173] In some embodiments, the severable portion is the thrombin severing portion.
[0174] In some embodiments, the peptide exchange assay is a FRET assay.
[0175] In some embodiments, purification is performed by one of the following methods: column chromatography, ion exchange chromatography, molecular sieve chromatography, affinity chromatography, or LC-MS.
[0176] This specification provides an HLA class II tetramer or polymer comprising either an HLA-DR heterodimer, an HLA-DP heterodimer, or an HLA-DQ heterodimer, wherein each heterodimer contains an alpha chain and a beta chain, the heterodimer is purified, and is present at a concentration greater than 1 mg / L.
[0177] In some embodiments, the HLA class II tetramer is selected from Tables 8A to 8C.
[0178] In some embodiments, the HLA class II tetramer comprises a heterodimer pair selected from the group consisting of HLA-DR protein, HLA-DP protein, and HLA-DQ protein.
[0179] In some embodiments, the HLA protein is HLA-DPB1* 01:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 02:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 03:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 04:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 04:02 / HLA-DPA1 * 01:03、HLA-DPB1 * 06:01 / HLA-DPA1 * 01:03、HLA-DQB1 * 02:01 / HLA-DQA1 * 05:01、HLA-DQB1 * 02:02 / HLA-DQA1 * 02:01、HLA-DQB1 * 06:02 / HLA-DQA1 * 01:02、HLA-DQB1 * 06:04 / HLA-DQA1 * 01:02、HLA-DRB1 * 01:01、HLA-DRB1 * 01:02、HLA-DRB1 * 03:01、HLA-DRB1 * 03:02、HLA-DRB1 * 04:01、HLA-DRB1 * 04:02、HLA-DRB1 * 04:03、HLA-DRB1 * 04:04、HLA-DRB1 * 04:05、HLA-DRB1 * 04:07、HLA-DRB1 * 07:01、HLA-DRB1 * 08:01、HLA-DRB1 * 08:02、HLA-DRB1 * 08:03、HLA-DRB1 * 08:04、HLA-DRB1 * 09:01、HLA-DRB1 * 10:01、HLA-DRB1* 11:01, HLA-DRB1 * 11:02, HLA-DRB1 * 11:04, HLA-DRB1 * 12:01, HLA-DRB1 * 12:02, HLA-DRB1 * 13:01, HLA-DRB1 * 13:02, HLA-DRB1 * 13:03, HLA-DRB1 * 14:01, HLA-DRB1 * 15:01, HLA-DRB1 * 15:02, HLA-DRB1 * 15:03, HLA-DRB1 * 16:01, HLA-DRB3 * 01:01, HLA-DRB3 * 02:02, HLA-DRB3 * 03:01, HLA-DRB4 * 01:01 and HLA-DRB5 * It is an HLA class II protein selected from the group consisting of 01:01.
[0180] In some embodiments, the heterodimer pair is expressed in eukaryotic cells.
[0181] In some embodiments, the heterodimer pair is encoded by a vector.
[0182] Provided herein is a vector comprising nucleic acid sequences encoding the alpha and beta chains of the HLA proteins described herein, a secretion signal, a biotinylated motif, and at least one tag for identification or purification, thereby causing each HLA protein alpha-1 and beta-1 heterodimer to be secreted in a dimerized state, wherein the secreted heterodimer is associated with a placeholder peptide as appropriate.
[0183] Cells containing the vectors described herein are provided herein.
[0184] In some embodiments, HLA class II heterodimers are secreted from eukaryotic cells into cell culture medium and further purified by one of the following methods: column chromatography, ion exchange chromatography, molecular sieve chromatography, affinity chromatography, or LC-MS.
[0185] A method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject is provided herein, comprising the steps of: obtaining a plurality of peptide sequences of the polypeptide sequence by a computer processor; processing the amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model by a computer processor to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject, the machine learning HLA-peptide presentation prediction model is trained using training data including sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0186] A method for screening a drug containing a polypeptide sequence for immunogenicity in a subject is provided herein, comprising the steps of: inputting amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction indicates the probability that the epitope sequence of a given peptide sequence will be presented by one or more proteins encoded by class I or class II MHC alleles in the cell of the subject, and the machine learning HLA-peptide presentation prediction model includes at least a plurality of predictor variables identified based on training data, the training data being sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and predictor variables; determining or predicting that each of the peptide sequences of the polypeptide sequence is not immunogenic to the subject based on the set of presentation predictions; and administering a composition containing the drug to the subject.
[0187] A method for screening a drug containing a polypeptide sequence for immunogenicity in a subject is provided herein, comprising the steps of: inputting amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction indicates the probability that the epitope sequence of a given peptide sequence will be presented by one or more proteins encoded by class I or class II MHC alleles in the cell of the subject, and the machine learning HLA-peptide presentation prediction model includes at least a plurality of predictor variables identified based on training data, the training data being sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and predictor variables; and determining or predicting, based on the set of presentation predictions, that at least one of the peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0188] A method for screening a drug containing a polypeptide sequence for immunogenicity in a subject is provided herein, comprising the steps of: obtaining a plurality of peptide sequences of the polypeptide sequence by a computer processor; processing the amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model by a computer processor to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject, and the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information related to HLA proteins expressed in the cell; determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject; and administering a composition containing the drug to the subject.
[0189] In some embodiments, the method further includes deciding not to administer the drug to the target.
[0190] In some embodiments, the drug comprises an antibody or a binding fragment thereof.
[0191] In some embodiments, the peptide sequence of the polypeptide sequence has a length of 8, 9, 10, 11, or 12 amino acids, and the protein encoded by the class I or class II MHC allele of the cell in question is the protein encoded by the class I MHC allele of the cell in question.
[0192] In some embodiments, the peptide sequence of the polypeptide sequence has a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids, and the protein encoded by the class I or class II MHC allele of the cell in question is the protein encoded by the class II MHC allele of the cell in question.
[0193] A method for treating a subject having an autoimmune disease or condition is provided herein, comprising the steps of (a) identifying or predicting an expressed protein epitope presented by class I or class II MHC on cells of the subject, wherein a complex comprising the identified or predicted epitope and class I or class II MHC is targeted by the subject's CD8 T cells or CD4 T cells; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells derived from the subject or allogeneic regulatory T cells; and (d) administering the regulatory T cells expressing the TCR to the subject.
[0194] In some embodiments, the autoimmune disease or condition is diabetes.
[0195] In some embodiments, the cells are pancreatic islet cells.
[0196] A method for treating a subject having an autoimmune disease or condition is provided herein, comprising the step of administering to the subject regulatory T cells expressing a T cell receptor (TCR) that binds to the complex, which comprises (i) an epitope of an expressed protein identified or predicted to be presented by class I or class II MHC on the subject's cells and (ii) class I or class II MHC, and which is targeted by the subject's CD8 T cells or CD4 T cells.
[0197] A computer system for identifying peptide sequences for targeted personalized cancer therapy is provided herein, comprising: a database configured to store a plurality of target peptide sequences; one or more computer processors operably coupled to the database, the computer system comprising the steps of: processing amino acid information of a plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class II MHC alleles of a cell of interest; the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; and one or more computer processors individually and collectively programmed to perform the step of selecting a subset of the plurality of peptide sequences for targeted personalized cancer therapy based on at least a plurality of presentation predictions.
[0198] A computer system for identifying HLA class II specific peptides for immunotherapy against a target, comprising: a database configured to store candidate peptides, each comprising epitopes and a plurality of peptide sequences, each comprising an epitope; and one or more computer processors operably coupled to the database, wherein the system processes amino acid information of a plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences to immune cells, wherein each presentation prediction indicates the possibility that a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by an HLA class II allele, and the machine learning HLA-peptide presentation prediction model includes sequence information of the peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry. A computer system is provided herein that includes one or more computer processors individually and collectively programmed to perform the following steps: a step of being trained using refinement data; a step of selecting from one or more proteins encoded by the HLA class II alleles of a target cell a protein is predicted by a machine learning HLA-peptide presentation prediction model to bind to a candidate peptide, wherein the protein has a probability greater than a threshold of the predicted presentation probability value for the presentation of the candidate peptide to immune cells; and a step of contacting the candidate peptide with the selected protein and thus identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein, based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide competes with the placeholder peptide associated with the selected protein.
[0199] A computer system for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, comprising: a database configured to store multiple peptide sequences of the polypeptide sequence; one or more computer processors operably coupled to the database, the computer system comprising: processing amino acid information of the multiple peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the multiple peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the multiple peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject, and the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information related to HLA proteins expressed in the cell; and one or more computer processors individually and collectively programmed to perform the steps of determining or predicting, based on the multiple presentation predictions, that each of the multiple peptide sequences of the polypeptide sequence is not immunogenic to the subject, and administering a composition comprising the drug to the subject.
[0200] A computer system for screening a drug comprising a polypeptide sequence for immunogenicity in a subject is provided herein, comprising: a database configured to store multiple peptide sequences of the polypeptide sequence; one or more computer processors operably coupled to the database, the computer system comprising the steps of: processing amino acid information of the multiple peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the multiple peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the multiple peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject; the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information of peptide sequences presented by HLA proteins expressed in the cell and identified by mass spectrometry; and one or more computer processors individually and collectively programmed to perform the steps of determining or predicting, based on the multiple presentation predictions, that at least one of the multiple peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0201] The following non-temporary computer-readable medium containing machine-executable code is provided herein to implement a method for identifying peptide sequences for a targeted personalized cancer treatment, which, when executed by one or more computer processors, comprises: obtaining a plurality of peptide sequences of interest; processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class II MHC alleles of the cell of interest, and the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; and selecting a subset of the plurality of peptide sequences for a targeted personalized cancer treatment based on at least a plurality of presentation predictions.
[0202] A method for identifying HLA class II specific peptides for immunotherapy against a target, which is performed by one or more computer processors, comprising the steps of: obtaining candidate peptides comprising an epitope and a plurality of peptide sequences, each comprising an epitope; and processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences against immune cells, wherein each presentation prediction indicates the possibility that a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by an HLA class II allele, and the machine learning HLA-peptide presentation prediction model uses training data comprising sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry. Provided herein is a non-temporary, computer-readable medium containing machine-executable code that implements a method comprising: a step of being trained to; a step of selecting from one or more proteins encoded by the HLA class II alleles of a target cell that are predicted by a machine learning HLA-peptide presentation prediction model to bind to a candidate peptide, wherein the protein has a probability greater than a threshold of presentation prediction probability value for presentation of the candidate peptide to immune cells; and a step of contacting the selected protein with the selected protein and thus identifying the candidate peptide as a peptide for immunotherapy specific to the selected protein, based on whether the candidate peptide replaces the placeholder peptide when the candidate peptide is made to compete with the placeholder peptide.
[0203] A non-transient, computer-readable medium containing machine-executable code is provided herein that implements a method comprising the steps of: a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, which, when executed by one or more computer processors, comprises: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject, and the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information related to HLA proteins expressed in the cell; and determining or predicting, based on the plurality of presentation predictions, that each of the plurality of peptide sequences of the polypeptide sequence is not immunogenic to the subject, and administering a composition containing the drug to the subject.
[0204] The following non-temporary computer-readable medium containing machine-executable code is provided herein that implements a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject, which, when executed by one or more computer processors, comprises the steps of: obtaining a plurality of peptide sequences of the polypeptide sequence; processing amino acid information of the plurality of peptide sequences using a machine learning HLA-peptide presentation prediction model to generate presentation predictions for each of the plurality of peptide sequences, wherein each presentation prediction indicates the possibility that the epitope sequence of a given peptide sequence among the plurality of peptide sequences may be presented by one or more proteins encoded by class I or class II MHC alleles of the cell of the subject, and the machine learning HLA-peptide presentation prediction model is trained using training data containing sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; and determining or predicting, based on the plurality of presentation predictions, that at least one of the plurality of peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0205] The method comprises the steps of: processing amino acid information of multiple candidate peptide sequences using a machine learning HLA peptide presentation prediction model to generate multiple presentation predictions, wherein each of the multiple candidate peptide sequences is encoded by the genome or exome of the target cell, the multiple presentation predictions include an HLA presentation prediction for each of the multiple candidate peptide sequences, each presentation prediction indicates the possibility that a given candidate peptide sequence may be presented by one or more proteins encoded by the class II HLA alleles of the target cell, and the machine learning HLA peptide presentation prediction model is trained using training data which includes sequence information relating to the sequences of training peptides identified by mass spectrometry as being presented by HLA proteins expressed in training cells; and, based on at least multiple presentation predictions, identifying peptide sequences among the multiple peptide sequences for which the probability of being presented by at least one of the one or more proteins encoded by the class II HLA alleles of the target cell is greater than a threshold of the presentation prediction probability value, processing amino acid information of multiple test peptide sequences to generate multiple test presentation predictions, each test presentation prediction indicates the possibility that the class II HLA allele of the target cell is The present invention provides a method in which, given a given test peptide sequence among a plurality of test peptide sequences may be presented by one or more proteins encoded by an HLA allele, the positive predictive value (PPV) of a machine learning HLA peptide presentation prediction model is at least 0.07, and the plurality of test peptide sequences comprises at least 500 test peptide sequences, each comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 499 decoy peptide sequences contained within a protein encoded by the genome of the organism, wherein the organism and the subject are of the same species, and the plurality of test peptide sequences comprise at least one hit peptide sequence and at least 499 decoy peptide sequences in a 1:499 ratio, and the machine learning HLA peptide presentation prediction model predicts that 0.2% of the plurality of test peptide sequences will be presented by an HLA protein expressed in a cell.
[0206] The method comprises the steps of: processing amino acid information of multiple peptide sequences encoded by the target genome or exome using a machine learning HLA-peptide binding prediction model to generate multiple binding predictions, wherein the multiple binding predictions include an HLA binding prediction for each of the multiple candidate peptide sequences, each binding prediction indicates the possibility that one or more proteins encoded by the class II HLA alleles of the target cell will bind to a given candidate peptide sequence among the multiple candidate peptide sequences, and the machine learning HLA peptide binding prediction model is trained using training data including sequence information of peptide sequences identified to bind to HLA class II proteins or HLA class II protein analogs; and identifying peptide sequences among the multiple peptide sequences that have a probability greater than a threshold binding prediction probability value for binding to at least one of the one or more proteins encoded by the class II HLA alleles of the target cell, based on at least multiple binding predictions, processing amino acid information of multiple test peptide sequences to generate multiple test binding predictions, each test binding prediction indicates the possibility that one or more proteins encoded by the class II HLA alleles of the target cell will bind to a given candidate peptide sequence among the multiple candidate peptide sequences The present invention provides a method in which, when it is shown that one or more proteins encoded by an HLA allele may bind to a given test peptide sequence among a plurality of test peptide sequences, the positive predictive value (PPV) of a machine learning HLA peptide binding prediction model is at least 0.1, the plurality of test peptide sequences comprises at least 50 test peptide sequences, each comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 19 decoy peptide sequences contained in a protein containing the peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, the organism and the subject are of the same species, the plurality of test peptide sequences comprise at least one hit peptide sequence and at least 19 decoy peptide sequences in a 1:19 ratio, and the machine learning HLA peptide presentation prediction model predicts that 5% of the plurality of test peptide sequences will bind to an HLA protein expressed in a cell.
[0207] In some embodiments, a machine learning HLA peptide presentation prediction model is trained using training data that includes sequence information relating to the sequence of a training peptide, which has been identified by mass spectrometry as being presented by an HLA protein expressed in training cells.
[0208] In some embodiments, the probability that one or more of the 0.2% of test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model will be presented by at least one of the one or more proteins encoded by the class II HLA alleles of the cells in question is higher than the threshold of the presentation prediction probability value.
[0209] In some embodiments, the probability that each of 0.2% of the multiple test peptide sequences predicted to be presented by the machine learning HLA peptide presentation prediction model will be presented by at least one of the one or more proteins encoded by the class II HLA alleles of the cells in question is higher than the threshold of the presentation prediction probability value.
[0210] In some embodiments, the PPV is greater than the respective PPV in the second column of Table 11 for the protein encoded by the corresponding HLA allele in Table 13. In some embodiments, the PPV is at least equal to the respective PPV in the third column of Table 11 for the protein encoded by the corresponding HLA allele in Table 11.
[0211] In some embodiments, the PPV is greater than the respective PPV in the second column of Table 12 for the proteins encoded by the HLA class II allele.
[0212] In some embodiments, the PPV is at least equal to the respective PPV in the second column of Table 16 for the protein encoded by the corresponding HLA allele in Table 16.
[0213] A method for preparing a personalized cancer therapy, comprising the steps of: identifying a peptide sequence associated with cancer, comprising comparing a DNA, RNA, or protein sequence derived from a target cancer cell with a DNA, RNA, or protein sequence derived from a target normal cell; and inputting amino acid position information of the identified peptide sequence into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of presentation predictions for the identified peptide sequence, wherein each presentation prediction represents the probability that one or more proteins encoded by the HLA class II alleles of the target cell will present a given sequence from among the identified peptide sequences, and the machine learning HLA-peptide presentation prediction model generates a set of predictions identified based on at least the training data. A method is provided herein that includes the steps of: a parameter including sequence information of peptide sequences identified by mass spectrometry, which are presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information, which is associated with HLA proteins expressed in cells; and a function representing the relationship between amino acid position information received as input and the presentability generated as output based on the amino acid position information and predictor variables; and selecting a subset of peptide sequences for preparing personalized cancer treatments, identified based on a set of presentation predictions, wherein the positive predictive value of the predictive model is at least 0.1%, 0.1% to 50%, or at least 0.1 with a recall of up to 50%.
[0214] A method is provided herein for training a machine learning HLA-peptide presentation prediction model, comprising the step of using a computer processor to input amino acid position sequences of HLA peptides isolated from one or more HLA-peptide complexes derived from cells expressing HLA class II alleles into the HLA-peptide presentation prediction model, wherein the machine learning HLA-peptide presentation prediction model includes, at least, a plurality of predictor variables identified based on training data, which are presented by HLA proteins expressed in cells and whose sequences are identified by mass spectrometry; training peptide sequence information, which includes amino acid position information of the training peptides and is associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentation possibilities generated as output based on the amino acid position information and predictor variables.
[0215] In some embodiments, the positive predictive value of the presented model is at least 0.25, with a recall of at least 0.1%, 0.1% to 50%, or up to 50%.
[0216] In some embodiments, the positive predictive value of the presented model is at least 0.4, with a recall of at least 0.1%, 0.1% to 50%, or up to 50%.
[0217] In some embodiments, the positive predictive value of the presented model is at least 0.6, with a recall of at least 0.1%, 0.1% to 50%, or up to 50%.
[0218] In some embodiments, the mass spectrometry is single-allele mass spectrometry.
[0219] In some embodiments, the peptide is presented by HLA proteins expressed in cells via autophagy.
[0220] In some embodiments, the peptide is presented by HLA proteins expressed in cells via phagocytosis.
[0221] In some embodiments, the quality of the training data is improved using multiple quality metrics.
[0222] In some embodiments, multiple quality metrics include general contaminating peptide removal, high scoring peak intensity, high score, and high mass accuracy.
[0223] In some embodiments, the scored peak intensity is at least 50%.
[0224] In some embodiments, the scored peak intensity is at least 60%.
[0225] In some embodiments, the score is at least 7.
[0226] In some embodiments, the mass accuracy is up to 5 ppm.
[0227] In some embodiments, the mass accuracy is up to 2 ppm.
[0228] In some embodiments, the skeletal cross-section score is at least 5.
[0229] In some embodiments, the skeletal cross-section score is at least 8.
[0230] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single immunoprecipitated HLA protein expressed in the cell.
[0231] In some embodiments, the peptide presented by the HLA protein expressed in the cell is a peptide presented by a single exogenous HLA protein expressed in the cell.
[0232] In some embodiments, the peptide presented by the HLA protein expressed in the cell is the peptide presented by a single recombinant HLA protein expressed in the cell.
[0233] In some embodiments, the multiple predictor variables include peptide-HLA affinity predictor variables.
[0234] In some embodiments, the multiple predictor variables include a source protein expression level predictor variable.
[0235] In some embodiments, the multiple predictor variables include peptide cleavage predictor variables.
[0236] In some embodiments, the training peptide sequence information includes sequences derived from peptides presented by HLA proteins, including peptides identified by searching a peptide database for enzyme-specific and unmodified peptides. In some embodiments, the peptides presented by HLA proteins include peptides identified by searching a novel peptide sequencing tool.
[0237] In some embodiments, the peptides presented by the HLA protein include peptides identified by searching a peptide database using a reverse database search strategy.
[0238] In some embodiments, the HLA protein comprises an HLA-DR protein and an HLA-DP protein or an HLA-DQ protein. In some embodiments, the HLA protein comprises an HLA-DR protein selected from the group consisting of an HLA-DR protein and an HLA-DP protein or an HLA-DQ protein. In some embodiments, the HLA protein comprises an HLA-DPB1 * 01:01 / HLA-DPA1 * 01:03, HLA-DPB1 * 02:01 / HLA-DPA1 * 01:03, HLA-DPB1 *03:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 04:01 / HLA-DPA1 * 01:03、HLA-DPB1 * 04:02 / HLA-DPA1 * 01:03、HLA-DPB1 * 06:01 / HLA-DPA1 * 01:03、HLA-DQB1 * 02:01 / HLA-DQA1 * 05:01、HLA-DQB1 * 02:02 / HLA-DQA1 * 02:01、HLA-DQB1 * 06:02 / HLA-DQA1 * 01:02、HLA-DQB1 * 06:04 / HLA-DQA1 * 01:02、HLA-DRB1 * 01:01、HLA-DRB1 * 01:02、HLA-DRB1 * 03:01、HLA-DRB1 * 03:02、HLA-DRB1 * 04:01、HLA-DRB1 * 04:02、HLA-DRB1 * 04:03、HLA-DRB1 * 04:04、HLA-DRB1 * 04:05、HLA-DRB1 * 04:07、HLA-DRB1 * 07:01、HLA-DRB1 * 08:01、HLA-DRB1 * 08:02、HLA-DRB1 * 08:03、HLA-DRB1 * 08:04、HLA-DRB1 * 09:01、HLA-DRB1 * 10:01、HLA-DRB1 * 11:01、HLA-DRB1 * 11:02、HLA-DRB1 * 11:04、HLA-DRB1 * 12:01、HLA-DRB1 *12:02, HLA-DRB1 * 13:01, HLA-DRB1 * 13:02, HLA-DRB1 * 13:03, HLA-DRB1 * 14:01, HLA-DRB1 * 15:01, HLA-DRB1 * 15:02, HLA-DRB1 * 15:03, HLA-DRB1 * 16:01, HLA-DRB3 * 01:01, HLA-DRB3 * 02:02, HLA-DRB3 * 03:01, HLA-DRB4 * 01:01 and HLA-DRB5 * It contains HLA-DR proteins selected from the group consisting of 01:01.
[0239] In some embodiments, the peptide presented by the HLA protein includes peptides identified by comparing the MS / MS spectrum of the HLA peptide with the MS / MS spectra of one or more HLA peptides in a peptide database.
[0240] In some embodiments, the mutation is selected from the group consisting of point mutations, splice site mutations, frameshift mutations, readthrough mutations, and gene fusion mutations.
[0241] In some embodiments, the peptide presented by the HLA protein has a length of 15 to 40 amino acids.
[0242] In some embodiments, the peptide presented by the HLA protein includes the peptide identified by the steps of (a) isolating one or more HLA complexes from a cell line expressing a single HLA class II allele; (b) isolating one or more HLA-peptides from one or more isolated HLA complexes; (c) obtaining MS / MS spectra for one or more isolated HLA-peptides; and (d) obtaining peptide sequences from a peptide database corresponding to the MS / MS spectra of one or more isolated HLA-peptides, wherein the sequences of one or more isolated HLA-peptides are identified by the steps of (d).
[0243] In some embodiments, personalized cancer therapy further includes an adjuvant.
[0244] In some embodiments, personalized cancer therapy further includes immune checkpoint inhibitors.
[0245] In some embodiments, the training data includes structured data, time-series data, unstructured data, relational data, or any combination thereof.
[0246] In some embodiments, unstructured data includes image data.
[0247] In some embodiments, relational data includes data from customer systems, enterprise systems, operational systems, websites, web-accessible application programming interfaces (APIs), or any combination thereof.
[0248] In some embodiments, training data is uploaded to a cloud-based database.
[0249] In some embodiments, the training step is performed using a convolutional neural network.
[0250] In some embodiments, the convolutional neural network includes at least two convolutional layers.
[0251] In some embodiments, the convolutional neural network (CNN) includes at least one batch normalization step.
[0252] In some embodiments, the convolutional neural network includes at least one spatial dropout step.
[0253] In some embodiments, the convolutional neural network includes at least one global max pooling step.
[0254] In some embodiments, the convolutional neural network includes at least one high-density layer.
[0255] In some embodiments, the step of identifying the peptide sequence includes identifying the peptide sequence using mutations expressed in the target cancer cells.
[0256] In some embodiments, the step of identifying peptide sequences includes identifying peptide sequences that are not expressed in the normal cells of interest.
[0257] In some embodiments, the step of identifying the peptide sequence includes identifying the overexpressed peptide sequence.
[0258] In some embodiments, the step of identifying peptide sequences includes identifying viral peptide sequences. In one embodiment, a method for identifying HLA class II specific peptides for target-specific immunotherapy, comprising the steps of: identifying candidate peptides containing an epitope; and using a computer processor, inputting amino acid information of a plurality of peptide sequences, each containing an epitope, into a machine learning HLA-peptide presentation prediction model to generate a set of HLA presentation predictions of peptide sequences to immune cells, wherein each presentation prediction represents the probability that one or more proteins encoded by the HLA class II alleles of the target cell will present a given peptide sequence containing an epitope, and the positive predictive value of the prediction model is at least 0.1, with a recall of at least 0.1%, 0.1% to 50%, or up to 50%, and the HLA class II alleles of the target cell A method is provided herein that includes the steps of: selecting a protein from among one or more proteins that is predicted by a predictive model to bind to a candidate peptide, wherein the protein has a probability greater than a threshold of the predicted presentation probability value for the presentation of the candidate peptide to immune cells; contacting the candidate peptide with a protein encoded by an HLA class II allele, thereby causing the candidate peptide to compete with a placeholder peptide associated with the protein encoded by the HLA class II allele; and identifying the candidate peptide as a peptide for immunotherapy specific to the protein encoded by the HLA class II allele, based on whether the placeholder peptide is replaced by the candidate peptide.
[0259] In some embodiments, the immunotherapy is cancer immunotherapy.
[0260] In some embodiments, the identification step includes comparing a DNA, RNA, or protein sequence derived from the target cancer cell with a DNA, RNA, or protein sequence derived from the target normal cell. In some embodiments, the epitope is a cancer-specific epitope.
[0261] In some embodiments, the at least one protein encoded by the HLA class II allele comprises at least the alpha-1 subunit and beta-1 subunit or a fragment thereof of the HLA protein, which exists in a dimeric form. In some embodiments, the placeholder peptide is the CLIP peptide. In some embodiments, the placeholder peptide is the CMV peptide. In some embodiments, the method further includes the step of measuring the IC50 of the substitution of the placeholder peptide with the target peptide. In some embodiments, the IC50 of the substitution of the placeholder peptide with the target peptide is less than 500 nM. In some embodiments, the at least one protein from one or more proteins encoded by the HLA class II allele of the cell of interest is an HLA class II tetramer or multimer. In some embodiments, the target peptide is further identified by mass spectrometry. In some embodiments, the at least one protein encoded by the HLA class II allele of the cell of interest is a recombinant protein. In some embodiments, the at least one protein encoded by the HLA class II allele of the cell of interest is expressed in eukaryotic cells.
[0262] In one embodiment, an assay method for verifying the specificity of a candidate peptide regarding binding to an HLA class II protein is provided herein, comprising the steps of: expressing a polynucleic acid construct in eukaryotic cells, the construct comprising a nucleic acid sequence encoding an HLA class II protein, which includes an alpha chain and a beta chain or a portion thereof, capable of binding to a peptide containing an MHC-II binding epitope, wherein the expressed HLA class II protein or a portion thereof remains associated with a placeholder peptide; isolating the HLA class II protein or a portion thereof expressed in eukaryotic cells; and performing a peptide exchange assay by (a) adding an increasing amount of the candidate peptide to determine whether the placeholder peptide associated with the HLA class II protein or a portion thereof is replaced by the candidate peptide; and (b) calculating the IC50 of the substitution reaction to determine the affinity of the candidate peptide to the HLA class II protein or a portion thereof compared with the placeholder peptide, thereby verifying the specificity of the candidate peptide regarding binding to an HLA class II protein.
[0263] In some embodiments, the type of peptide is known. In some embodiments, the type of peptide is unknown. In some embodiments, the type of peptide is determined by mass spectrometry.
[0264] In some embodiments, the peptide exchange assay includes detection of a peptide fluorescent probe or tag. In some embodiments, the placeholder peptide is a CLIP peptide.
[0265] In some embodiments, the polynucleic acid construct comprises an expression vector further comprising one or more of a promoter, a linker, one or more protease cleavage sites, a secretory signal, a dimerizing factor, a ribosome skipping sequence, and one or more tags for purification and / or detection.
[0266] In one embodiment, a method for assaying the immunogenicity of an MHC class II binding peptide is provided herein, comprising the steps of: selecting a protein encoded by an HLA class II allele, which is predicted to bind to the peptide by a machine learning HLA-peptide presentation prediction model, wherein the positive predictive value of the prediction model is at least 0.1%, with a recall of 0.1% to 50%, or up to 50%, and the protein has a probability greater than a threshold of the predicted presentation probability value for the presentation of the identified peptide sequence; contacting the peptide with the selected protein encoded by an HLA class II allele, thereby competing the peptide with a placeholder peptide associated with the selected protein encoded by an HLA class II allele, replacing the placeholder peptide, and thereby forming a complex comprising the HLA class II protein and the identified peptide; contacting the complex of the HLA class II protein and the identified peptide with CD4+ T cells; and assaying for one or more CD4+ T cell activation parameters selected from cytokine induction, chemokine induction, and cell surface marker expression.
[0267] In some embodiments, the HLA class II allele is a tetramer or a multimer. In some embodiments, the cytokine is IL-2. In some embodiments, the cytokine is IFN-gamma.
[0268] In one embodiment, a method for inducing CD4+ T cell activation in a subject for cancer immunotherapy, comprising the steps of: identifying a peptide sequence associated with cancer and containing cancer mutations, comprising comparing a DNA, RNA, or protein sequence derived from a cancer cell of the subject with a DNA, RNA, or protein sequence derived from a normal cell of the subject; and selecting a protein encoded by an HLA class II allele that is normally expressed by the cell of the subject and is predicted by a machine learning HLA-peptide presentation prediction model to bind to a peptide, wherein the positive predictive value of the prediction model is at least 0.1%, 0.1% to 50%, or at least 0.1% with a recall of up to 50%. A method is provided herein that includes the steps of: determining that the protein has a probability greater than a threshold of the predicted presentation probability value for the presentation of the identified peptide sequence; contacting the identified peptide with a protein encoded by a selected HLA class II allele to verify whether the identified peptide competes with and replaces a placeholder peptide associated with the selected HLA class II allele-encoded protein at an IC50 value of less than 500 nM; purifying the identified peptide; and administering the identified peptide to the target in an effective amount.
[0269] In one embodiment, a method for producing HLA class II tetramers or multimers is provided herein, comprising the steps of: expressing a vector in eukaryotic cells containing nucleic acid sequences encoding the alpha and beta chains of HLA proteins, a linker, one or more protease cleavage sites, a secretory signal, a biotinylated motif, and at least one tag for identification or purification, thereby causing each HLA protein alpha-1 and beta-1 heterodimer to be secreted in a dimerized state, wherein the heterodimer is associated with a placeholder peptide; purifying the secreted heterodimer from cell culture medium; verifying peptide binding activity using a peptide exchange assay; adding streptavidin to conjugate the heterodimer into a tetramer; purifying the tetramer; and obtaining a yield greater than 1 mg / L.
[0270] In some embodiments, the vector includes a CMV promoter. In some embodiments, the vector includes a sequence encoding a placeholder peptide linked to the beta-1 chain via a cleavable site. In some embodiments, the peptide exchange assay involves pre-cleaving the placeholder peptide from the beta chain. In some embodiments, the cleavable site is a thrombin cleavage site. In some embodiments, the peptide exchange assay is a FRET assay. In some embodiments, purification is performed by one of the following: column chromatography, batch chromatography, ion exchange chromatography, molecular sieve chromatography, affinity chromatography, or LC-MS.
[0271] In one embodiment, the present invention provides a composition comprising an HLA class II tetramer containing either an HLA-DR heterodimer, an HLA-DP heterodimer, or an HLA-DQ heterodimer, wherein each heterodimer comprises an alpha chain and a beta chain, is purified, and is present at a concentration higher than 0.25 mg / L. In some embodiments, the HLA class II tetramer comprises a heterodimer pair selected from the group consisting of an HLA-DR protein and a protein that can be selected from the group consisting of an HLA-DP protein or an HLA-DQ protein. In some embodiments, the HLA protein is HLA-DPB1 * 01:01 / HLA-DPA1 * 01:03, HLA-DPB1 * 02:01 / HLA-DPA1 * 01:03, HLA-DPB1 * 03:01 / HLA-DPA1 * 01:03, HLA-DPB1 * 04:01 / HLA-DPA1 * 01:03, HLA-DPB1 * 04:02 / HLA-DPA1 * 01:03, HLA-DPB1 * 06:01 / HLA-DPA1 * 01:03, HLA-DQB1 * 02:01 / HLA-DQA1 * 05:01, HLA-DQB1 * 02:02 / HLA-DQA1 * 02:01, HLA-DQB1 * 06:02 / HLA-DQA1 * 01:02, HLA-DQB1 * 06:04 / HLA-DQA1 * 01:02, HLA-DRB1 * 01:01, HLA-DRB1 * 01:02, HLA-DRB1 * 03:01, HLA-DRB1 * 03:02, HLA-DRB1 * 04:01, HLA-DRB1 * 04:02, HLA-DRB1 *04:03、HLA-DRB1 * 04:04、HLA-DRB1 * 04:05、HLA-DRB1 * 04:07、HLA-DRB1 * 07:01、HLA-DRB1 * 08:01、HLA-DRB1 * 08:02、HLA-DRB1 * 08:03、HLA-DRB1 * 08:04、HLA-DRB1 * 09:01、HLA-DRB1 * 10:01、HLA-DRB1 * 11:01、HLA-DRB1 * 11:02、HLA-DRB1 * 11:04、HLA-DRB1 * 12:01、HLA-DRB1 * 12:02、HLA-DRB1 * 13:01、HLA-DRB1 * 13:02、HLA-DRB1 * 13:03、HLA-DRB1 * 14:01、HLA-DRB1 * 15:01、HLA-DRB1 * 15:02、HLA-DRB1 * 15:03、HLA-DRB1 * 16:01、HLA-DRB3 * 01:01、HLA-DRB3 * 02:02、HLA-DRB3 * 03:01、HLA-DRB4 * 01:01、HLA-DRB5 * 01:01) The snowflake is on the ground.
[0272] In some embodiments, heterodimer pairs are expressed in eukaryotic cells. In some embodiments, heterodimer pairs are encoded by a vector. In some embodiments, the vector comprises nucleic acid sequences encoding the alpha and beta chains of the HLA protein, a secretion signal, a biotinylated motif, and at least one tag for identification or purification, so that each HLA protein alpha-1 and beta-1 heterodimer is secreted in a dimerized state, and the secreted heterodimer is associated with a placeholder peptide. In some embodiments, the vector comprises nucleic acid sequences encoding the alpha and beta chains of the HLA protein, a secretion signal, a biotinylated motif, and at least one tag for identification or purification, so that each HLA protein alpha-1 and beta-1 heterodimer is secreted in a dimerized state, and the secreted heterodimer is associated with a placeholder peptide.
[0273] In some embodiments, HLA class II heterodimers are secreted from eukaryotic cells into cell culture medium and purified by one of the following methods: column or batch chromatography, ion exchange chromatography, molecular sieve chromatography, affinity chromatography, or LC-MS.
[0274] In one embodiment, a method for screening a drug containing a polypeptide sequence for immunogenicity in a target, comprising the steps of using a computer processor to input amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction represents the probability that the epitope sequence of a given peptide sequence will be presented by one or more proteins encoded by HLA class I or class II alleles of the target cell, and the machine learning HLA-peptide presentation prediction model includes at least a plurality of predictor variables identified based on training data, and the training data is generated in cells A method is provided herein that includes the steps of: (b) determining or predicting, based on a set of presentation predictions, that each peptide sequence of the polypeptide sequence is not immunogenic to a subject; and (c) administering a composition comprising a drug to the subject.
[0275] In one embodiment, a method for screening a drug comprising a polypeptide sequence for immunogenicity in a subject is provided herein, comprising the steps of: (a) using a computer processor to input amino acid information of the peptide sequence of the polypeptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the peptide sequence, wherein each presentation prediction represents the probability that the epitope sequence of a given peptide sequence is presented by one or more proteins encoded by HLA class I or class II alleles of the cell of the subject, and the machine learning HLA-peptide presentation prediction model includes at least a plurality of predictor variables identified based on training data, the training data including sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and predictor variables; and (b) determining or predicting, based on the set of presentation predictions, that at least one of the peptide sequences of the polypeptide sequence is immunogenic to the subject.
[0276] In one embodiment, the method further includes deciding not to administer the drug to the subject.
[0277] In one embodiment, the drug comprises an antibody or a binding fragment thereof.
[0278] In one embodiment, the polypeptide sequence comprises a series of peptide sequences having lengths of 8, 9, 10, 11, or 12 amino acids, and the protein encoded by the HLA class I or class II allele of the cell in question is the protein encoded by the HLA class I allele of the cell in question.
[0279] In one embodiment, the polypeptide sequence comprises a series of peptide sequences having a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids, and the protein encoded by the HLA class I or class II allele of the cell in question is the protein encoded by the class II MHC allele of the cell in question.
[0280] In one embodiment, a method for treating a subject having an autoimmune disease or condition is provided herein, comprising the steps of: (a) identifying or predicting an expressed protein epitope presented by HLA class I or class II of cells of the subject, wherein a complex comprising the identified or predicted epitope and HLA class I or class II is targeted by the subject's CD8 T cells or CD4 T cells; (b) identifying a T cell receptor (TCR) that binds to the complex; (c) expressing the TCR in regulatory T cells derived from the subject or allogeneic regulatory T cells; and (d) administering the regulatory T cells expressing the TCR to the subject.
[0281] In one embodiment, the autoimmune disease or condition is diabetes mellitus.
[0282] In one embodiment, the cells are pancreatic islet cells.
[0283] In one embodiment, a method for treating a subject having an autoimmune disease or condition is provided herein, comprising the step of administering to the subject regulatory T cells expressing a T cell receptor (TCR) that binds to a complex comprising (i) an expressed protein epitope identified or predicted to be presented by HLA class I or class II of the subject's cells and (ii) HLA class I or class II, wherein the complex is targeted by the subject's CD8 T cells or CD4 T cells.
[0284] While only exemplary embodiments of this disclosure are shown, further aspects and advantages of this disclosure will be readily apparent to those skilled in the art from the following detailed description. As will be understood, other and different embodiments of this disclosure are possible, and some of its details may be modified in various obvious ways without departing from this disclosure. Accordingly, the figures and descriptions should be considered as factually illustrative and not limiting.
[0285] MAPTAC® can be used for high-throughput peptide binding assays, in which LC-MS / MS is used to measure peptides bound to HLA class II after isolation using MAPTAC® constructs under different conditions at different time points, for example, under heating at 37°C, to obtain sequences of peptide populations with different stabilities.
[0286] In one embodiment, a method for treating cancer in a subject, comprising the steps of: identifying a peptide sequence associated with cancer, including comparing a DNA, RNA, or protein sequence derived from a cancer cell of the subject with a DNA, RNA, or protein sequence derived from a normal cell of the subject; and using a computer processor to input amino acid information of the identified peptide sequence into a machine learning HLA-peptide presentation prediction model to generate a set of presentation predictions for the identified peptide sequence, wherein each presentation prediction represents the probability that one or more proteins encoded by the HLA class II allele of the cell of the subject will present a given sequence from among the identified peptide sequences, and the machine learning HLA-peptide presentation prediction model includes at least a plurality of predictor variables identified based on training data, and the training data is A method is provided herein that includes the steps of: a step of including sequence information of a peptide sequence presented by an HLA protein expressed in a cell and identified by mass spectrometry; training peptide sequence information including amino acid position information, associated with an HLA protein expressed in a cell; and a function representing the relationship between amino acid position information received as input and the presentability generated as output based on the amino acid position information and predictor variables; a step of selecting a subset of peptide sequences for preparing a personalized cancer treatment, identified based on a set of presentation predictions; and a step of administering a composition comprising one or more peptides to a subject, wherein the positive predictive value of the predictive model is at least 0.1%, 0.1% to 50%, or at least 0.1 with a recall of up to 50%.
[0287] In some embodiments, the machine learning HLA-peptide presentation prediction model includes sequence information of peptide sequences that are presented by HLA proteins expressed in cells and identified by mass spectrometry after reverse-phase offline fractionation.
[0288] In some embodiments, the predictive model shows an improvement of 1.1 to 100 times compared to NetMHCIIpan. Shows improvements of 43x, 44x, 45x, 50x, 55x, 60x, 61x, 62x, 63x, 64x, 65x, 66x, 67x, 68x, 69x, 70x, 71x, 72x, 73x, 74x, 75x, 76x, 77x, 78x, 79x, 80x, 81x, 8x, 83x, 84x, 85x, 86x, 87x, 88x, 89x, 90x, 91x, 92x, 93x, 94x, 95x, 96x, 97x, 98x, 99x, 100x or greater.
[0289] Embedding by reference All publications, patents, and patent applications referenced herein are incorporated by reference in the same way that individual publications, patents, or patent applications are specifically and individually incorporated by reference. To the extent that any publications, patents, or patent applications incorporated by reference conflict with any disclosures contained herein, this specification shall supersede and / or take precedence over any such conflicting material.
[0290] Novel features of the present invention are described in detail in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description, which includes exemplary embodiments in which the principles of the present invention are utilized, and the following accompanying drawings (also referred to herein as "Figures"). [Brief explanation of the drawing]
[0291] [Figure 1A]Figure 1A shows a peptide docked with an MHC class I protein.
[0292] [Figure 1B] Figure 1B shows an exemplary diagram representing a peptide docked with an MHC class II protein.
[0293] [Figure 2] Figure 2 shows an exemplary experimental technique for generating single-allele HLA class II-binding peptide data. HLA class II peptides are introduced into any cells, including those that do not express HLA class II, resulting in the expression of a specific HLA class II allele(s). A population of genetically engineered HLA-expressing cells is harvested, lysed, and their HLA-peptide complexes are tagged (e.g., biotinylated) and immunopurified (e.g., using biotin-streptavidin interactions). Single-HLA-specific HLA-associated peptides can then be eluted from these tagged (e.g., biotinylated) complexes and evaluated (e.g., by sequencing using high-resolution LC-MS / MS).
[0294] [Figure 3] Figure 3 shows exemplary sequence logo representations of HLA class II-DRB1*11:01-related peptides across Neon BAP, Expi293 cell line; Neon BAP, A375 cell line; IEDB, affinity <50 nM; and pan-HLA class II Ab, homozygous LCL. Figure 3 shows examples of MS-derived motifs matching known patterns and demonstrating consistency across transfected cell lines.
[0295] [Figure 4]Figure 4 is an illustrative depiction of the performance of the HLA class II binding predictor. Figure 4 is a bar plot showing the performance of the binding predictor (neonmhc2) and NetMHCIIpan applied to a validation dataset consisting of observed mass spectrometry peptides and decoy peptides generated in a 1:19 (hit:decoy) ratio by randomly shuffling hit peptides. For the NEON binding predictor neonmhc2, separate models are built for each MHC II allele shown. The height of the bars represents the positive predictive value (PPV), defined as the fraction of predicted conjugates that were actually hit peptides in the validation set. Alleles are sorted by the model's performance when predicting for that allele.
[0296] [Figure 5] Figure 5 illustrates the exemplary effect of the scoring peak intensity (SPI) threshold on the validation of the binding predictor. Figure 5 shows the performance of the HLA class II binding predictor when trained / validated with sets of peptides having various scoring peak intensity (SPI) cutoffs. For each trained allele-specific model, the model's performance is shown in the following three settings: training and evaluation on the dataset using observed MS hit peptides with an SPI greater than or equal to 70; training with peptides greater than or equal to 50 SPI and validation with peptides greater than or equal to 70 SPI; and training and validation with peptides greater than or equal to 50 SPI.
[0297] [Figure 6]Figure 6 shows an exemplary bar plot representing representative data from the number of peptides greater than or equal to the 70-scored peak intensity (SPI) cutoff observed by allele profiling by LC-MS / MS. Each bar represents the total number of observed peptides for the allele. Data is available for 35 HLA-DR alleles. The data for these 35 HLA-DR alleles has >95% population coverage for HLA-DR (USA allele frequencies).
[0298] [Figure 7A] Figure 7A shows the model PPV when applied to the test categories of the data for the given HLA class II alleles. The decoy peptide used was a scrambled sequence of positive (hit) peptide sequences with a hit-to-decoy ratio of 1:19. PPV was determined by identifying the top 5% scoring peptides in the test category and determining their fraction that were positive for binding to the proteins encoded by each HLA class II allele.
[0299] [Figure 7B-1] Figure 7B shows an exemplary predictive performance as a function of the training set size (a curve obtained by artificially downsampling the training set). Figure 7B generally shows that for the 35 HLA-DR alleles collected, increasing the training set size increases the PPV value. [Figure 7B-2] Figure 7B shows an exemplary predictive performance as a function of the training set size (a curve obtained by artificially downsampling the training set). Figure 7B generally shows that for the 35 HLA-DR alleles collected, increasing the training set size increases the PPV value. [Figure 7B-3]Figure 7B shows an exemplary predictive performance as a function of the training set size (a curve obtained by artificially downsampling the training set). Figure 7B generally shows that for the 35 HLA-DR alleles collected, increasing the training set size increases the PPV value.
[0300] [Figure 8] Figure 8 shows an exemplary graph demonstrating that predictions can be further improved by processing-related variables. Random peptide sequences observed by MS selected from protein-coding exomes can be distinguished. HLA class II presentation can be predicted by fitting logistic regression to the training data segment using binding strength (NetMHCIIpan or Neon predictor) and processing features (RNA-Seq expression and derived gene level bias terms). For separate evaluation segments, the exon locations overlapping with MS-observed MHC II peptides ("hits") can be scored in parallel with the locations of random exons not observed by MS (ratio of 1:499). The top 0.2% (1 / 500) can be called positive, and the positive predictive value can be evaluated using this threshold.
[0301] [Figure 9]Figure 9 shows an exemplary neural network architecture. Input peptides are shown as 20mers, and peptides shorter than this are marked with the letter "missing". Each peptide has a 31-dimensional embedding, and therefore the input to the neural network is a 20x31 matrix. Before processing by the neural network, feature normalization of the 20x31 matrix is performed based on the mean and standard deviation of the feature values in the training set. The first convolutional layer has a kernel of 9 amino acids and 50 filters (also called channels) and uses the Rectified Linear Unit (ReLU) activation function. This is followed by batch normalization, then spatial dropout with a dropout rate of 20%. This is followed by another convolutional layer with a kernel of 3 amino acids and 20 filters and using the ReLU activation function, then again batch normalization and spatial dropout with a dropout rate of 20%. Next, the maximally activated neurons in each of the 20 filters are taken and global maximal pooling is applied, and then these 20 values are passed through a fully connected (high density) layer with a single neuron using the sigmoid activation function. The output of this layer is treated as a coupled / uncoupled prediction. L2 regularization is applied to the weights of the first convolutional layer, the second convolutional layer, and the high-density layer, with weights of 0.05, 0.1, and 0.01, respectively. In the additional models used, the number of convolutional layers and the kernel size of each layer were varied.
[0302] [Figure 10] Figure 10 shows an exemplary computer-controlled system programmed or otherwise configured to implement the methods provided herein.
[0303] [Figure 11A] Figure 11A shows an illustrative overview of the MAPTAC® experimental workflow.
[0304] [Figure 11B] Figure 11B shows exemplary peptide counts per allele merged across repeated experiments.
[0305] [Figure 11C] Figure 11C shows an exemplary peptide length distribution for HLA class I and HLA class II alleles profiled by MAPTAC®.
[0306] [Figure 11D] Figure 11D shows exemplary cysteine frequencies per residue observed for MAPTAC® and IEDB (alleles DRB1*01:01, DRB1*03:01, DRB1*09:01, and DRB1*11:01), human proteome, and multiple allele MS data from previous publications.
[0307] [Figure 12A] Figure 12A shows the frequencies of HLA-DR, HLA-DP, and HLA-DQ alleles present in >1% of individuals in Caucasians, and the counts of peptides from the indicated sources, measured as strong conjugates (<50 nM).
[0308] [Figure 12B] Figure 12B shows an exemplary length distribution of IEDB peptides along with relevant HLA class II affinity measurements.
[0309] [Figure 12C]Figure 12C shows exemplary Western blots of A375 cell lines individually transfected with (1) Expi293, (2) HeLa, and (3) two HLA class I alleles and two HLA class II alleles: HLA-A*02:01, HLA-B*45:01, HLA-DRB1*01:01, and HLA-DRB1*11:01. Anti-biotin ligase epitope tags were blotted to the membrane to visualize biotin acceptor peptide (BAP) and anti-beta-tubulin as a loading control. The lanes correspond to the following fractions collected during the MAPTAC® protocol: Lane 1, input; Lane 2, biotinylated input; and Lane 3, input after pull-down.
[0310] [Figure 12D] Figure 12D shows exemplary amino acid frequencies per residue observed for multiple allele MS data from MAPTAC® and IEDB (alleles DRB1*01:01, DRB1*03:01, DRB1*09:01, and DRB1*11:01), the human proteome, and previous publications.
[0311] [Figure 12E] Figure 12E shows the frequencies of HLA-DR, HLA-DP, and HLA-DQ alleles present in >1% of individuals in Caucasians, and the counts of peptides from the indicated sources, measured as strong conjugates (<50 nM). This figure includes additional data compared to Figure 12A. The additional data was obtained from tools.iedb.org / main / datasets / .
[0312] [Figure 12F] Figure 12F shows exemplary amino acid frequencies per residue observed for MAPTAC(trademark) (reduced and alkylated), MAPTAC(trademark) (untreated), and IEDB (alleles DRB1*01:01, DRB1*03:01, DRB1*09:01, and DRB1*11:01), human proteome, and multiple allele MS data from previous publications.
[0313] [Figure 13] Figure 13 shows an exemplary representation of core-binding sequence logos for each MHC II allele in MAPTAC® and IEDB. The sequence logos are graphic representations, where the height of each amino acid is proportional to the frequency of occurrence of that amino acid in the peptide that binds to the MHC protein encoded by the allele. The lowest entropy position is indicated by color, and the color corresponds to the amino acid characteristic. The peptides are derived from the datasets shown and aligned according to a CNN-based predictor (method). The logos represent all peptides, including those that do not closely match the overall motif (e.g., no peptides are isolated in the "trash" cluster).
[0314] [Figure 14A] Figure 14A shows exemplary sequence logos for HLA-A*02:01 binding peptides (ligands) analyzed in two different cell lines (A375 & expi293) using different HLA-ligand profiling techniques, including binding assays, stability assays, soluble HLA (sHLA) mass spectrometry, single-allele mass spectrometry, and MAPTAC®.
[0315] [Figure 14B] Figure 14B shows exemplary fractions of MAPTAC® peptides exhibiting 0, 1, 2, 3, and 4 heuristically defined anchors.
[0316] [Figure 14C] Figure 14C shows an exemplary distribution of binding affinity predicted by NetMHCIIpan for peptides observed by MAPTAC® (20 peptides per allele, each with SPI > 70 and nested sets of size 2 or larger) and length-balanced decoys sampled from the proteome.
[0317] [Figure 15A] Figure 15A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish between a single-allele MHC peptide and a scrambled, length-balanced decoy. The schematic diagram shows the use of amino acid feature embedding, the use of two convolutional layers with different filter sizes, and global maximal pooling as input to the final logistic output node.
[0318] [Figure 15B] Figure 15B shows an exemplary result illustrating Kendall's tau statistic for the correlation of IEDB affinity measured using binding predictions from either neonmhc2 or NetMHCIIpan. The peptides evaluated include only those posted to the IEDB in the year following the publication of NetMHCIIpan.
[0319] [Figure 16] Figure 16 shows an illustrative depiction of neonmhc2's performance depending on the size of the training dataset.
[0320] [Figure 17A] Figure 17A shows exemplary cluster assignments for MAPTAC® peptides (20 per allele) spiked into pan-DR and pan-class II MHC MS datasets. The datasets were deconvolved using GibbsCluster. Each square represents one MAPTAC® peptide. The color of the square indicates the cluster to which it is assigned, and the gray bars indicate the alleles from which the peptide actually originates. The total number of clusters in the Gibbs cluster solution (right side) was selected using the mutual information (MI) metric. The MI score also determines how the samples are sorted; samples with high MI solutions appear higher.
[0321] [Figure 17B]Figure 17B shows exemplary core-binding sequence logos for multiple allele MS data deconvolved by GibbsCluster. Each set of peptides corresponds to a cluster that best aligns to MAPTAC® spike-in.
[0322] [Figure 17C] Figure 17C shows representative performance of models predicting holdout MAPTAC® peptides using either MAPTAC® data or deconvolved multi-allele data. For each allele, the larger of the two data sources (usually MAPTAC®) is downsampled, and therefore the predictors are based on an equal number of training examples. NetMHCIIpan performance is shown for further comparison.
[0323] [Figure 17D] Figure 17D shows exemplary core-binding sequence logos derived from multi-allele MS data from the indicated sources.
[0324] [Figure 18A] Figure 18A shows an exemplary graph of peptide versus source gene expression fraction (TPM) for peptides and random proteome decoys observed by MS (data replotted from Schuster et al. 2017).
[0325] [Figure 18B] Figure 18B shows exemplary observed and predicted counts of class II peptides per gene, determined by a collaborative analysis of colorectal cancer, melanoma, and ovarian cancer datasets (Loeffler et al., 2018, and Schuster et al., 2017). Predicted counts are derived by multiplying the gene length by the expression level. Predicted and observed counts were summed across relevant samples. Genes known to be present in plasma are marked according to their concentrations (inset).
[0326] [Figure 18C] Figure 18C shows an exemplary distribution of enrichment scores (ratio of observed to predicted observations, similar to Figure 18B) for genes related to autophagy.
[0327] [Figure 18D] Figure 18D shows an exemplary distribution of enrichment scores according to the localization of each source gene. The localization of source genes was determined using Uniprot (uniprot_sprot.dat).
[0328] [Figure 18E] Figure 18E shows exemplary data representing a comparison of predicted and observed frequencies of the fraction of the total number of peptides with MHC-II binding affinity sequestered based on intracellular localization characteristics.
[0329] [Figure 18F] Figure 18F shows representative data illustrating the relative agreement of peptides in observations of two different gene expression profiles. For each sample, gene-level peptide counts were modeled as a linear combination of bulk tumor gene expression and professional APC (macrophage) gene expression profiles. The ratio of the coefficients determines the relative agreement between each expression profile and the peptide repertoire. Error bars correspond to 95% confidence intervals calculated by computer using bootstrap resampling.
[0330] [Figure 19A] Figure 19A shows representative data of HLA-DRB1 expression levels in five study examples. Each dot represents the expression in the individual cell type of an individual patient, averaged across cells.
[0331] [Figure 19B]Figure 19B shows exemplary representative data of tumor and stromal HLA-DRB1 expression input from RNA-Seq in TCGA patients. The horizontal bars correspond to individual patients and are grouped by tumor type. Patients were included if they had mutations in HLA class II pathway genes (CIITA, CD74, or CTSSS) determined by DNA-based mutation calling. For each patient, the fraction of tumor-derived HLA-DRB1 expression was estimated as min(1,2f), where f is the fraction of RNA-Seq reads in CIITA, CD74, or CTSS showing mutations.
[0332] [Figure 19C] Figure 19C shows exemplary representative data from additional single-cell RNA-Seq studies, including biopsies before and after checkpoint blockade immunotherapy.
[0333] [Figure 20] Figure 20 shows representative experimental data that exemplifies the overall performance of predictions for natural donor tissue.
[0334] [Figure 21A] Figure 21A shows exemplary representative data illustrating an integrated presentation model for predicting the cellular HLA class II ligandome. This represents PPV with a hit-to-decoy ratio of 1:499 for the pan-DR dataset (also analyzed in Figures 30B and 32E). The predictor uses binding prediction (NetMHCIIpan or neonmhc2) and, where necessary, gene expression, gene bias (as shown in Figure 32A), and overlap with previously observed HLA-DQ peptides. For each candidate peptide, the binding score was calculated as the maximum value across HLA-DR alleles in the sample genotype.
[0335] [Figure 21B]Figure 21B shows exemplary representative data for tumor-derived peptides identified using SILAC presented by dendritic cells (analyzed from cell lysates), using the same hit:decoy ratio and performance metrics as in Figure 21A, illustrating predictive performance with and without the use of processing features.
[0336] [Figure 21C] Figure 21C shows exemplary expression and gene bias scores (red dots, plotted according to K562 expression) for heavily labeled peptides observed in UV-treated experiments, compared to lightly labeled peptides (gray dots, plotted according to DC expression).
[0337] [Figure 21D] Figure 21D shows an illustrative diagram illustrating duplication of heavily labeled peptide source genes by lysate and UV treatment experiments. Gene names are color-coded according to their functional class.
[0338] [Figure 22A] Figure 22A shows an exemplary flow chart representing the assay protocol disclosed herein for verifying HLA class II-driven CD4+ T cells and T cell responses.
[0339] [Figure 22B] Figure 22B shows an exemplary HLA protein dimer construct design for a peptide exchange assay (upper panel) and an illustrative diagram of an exemplary assay workflow (lower panel).
[0340] [Figure 23] Figure 23 shows an exemplary illustration of an exemplary vector design for MHC-II expression to screen for novel binding peptides, and a display of the expressed protein product.
[0341] [Figure 24]Figure 24 shows an illustrative flowchart of the transfection, purification, and cleavage from the beta chain of the placeholder peptide.
[0342] [Figure 25A] Figure 25A shows an exemplary illustration of a vector encoding the CLIP peptide, which is associated with increased secretion of expressed MHC-II peptides.
[0343] [Figure 25B] Figure 25B shows illustrative diagrams of the shorter and longer forms of the nucleic acids encoding CLIP0 and CLIP1, respectively.
[0344] [Figure 25C] Figure 25C shows exemplary representative results of Coomasiegel analysis of alpha and beta chains with and without longer clips.
[0345] [Figure 26A] Figure 26A shows an illustrative diagram of the TR-FRET assay.
[0346] [Figure 26B] Figure 26B shows exemplary representative polarization data from an HLA class II peptide binding assay using a fluorescence resonance energy transfer (FRET) assay with a specific peptide.
[0347] [Figure 26C] Figure 26C shows exemplary representative polarization data from an HLA class II peptide binding assay using a fluorescence resonance energy transfer (FRET) assay with a specific peptide.
[0348] [Figure 26D] Figure 26D shows exemplary percentage substitutions of MHC-construction-bound peptides calculated from fluorescence increase.
[0349] [Figure 26E] Figure 26E shows exemplary percentage substitutions of MHC-construction-bound peptides calculated from fluorescence increase.
[0350] [Figure 26F] Figure 26F shows an exemplary peptide exchange using a differential scanning fluorescence (DSF) assay. An exemplary mechanism for detecting peptide dissociation from MHC class II using heat is shown, which also dissociates MHC class II heterodimers, resulting in fluorophore binding and high fluorescence. An exemplary schematic diagram of the transfer of a placeholder peptide by an epitope peptide is also shown. An exemplary melting curve plotted against temperature is also presented.
[0351] [Figure 26G] Figure 26G shows an exemplary soluble HLA-DM construct and its use for performing MHC class II peptide exchange. The construct shown contains a CMV promoter, a secretory sequence (leader), downstream coding sequences for the HLA-DM beta chain and HLA-DM alpha chain, as well as a BAP sequence at the 3' end of the beta chain coding sequence; and a His tag at the 3' end of the alpha chain coding sequence. The two chains are separated by an intervening ribosome skipping sequence. The construct was expressed in Expi-CHO cells, and the secreted protein into the culture medium was purified.
[0352] [Figure 26H] Figure 26H shows exemplary molecular sieve chromatography data using HLA-sDM for peptide exchange.
[0353] [Figure 27A] Figure 27A shows an exemplary diagram illustrating the construction of an exemplary DRB tetramer repertoire.
[0354] [Figure 27B] Figure 27B shows an exemplary diagram illustrating the construction of an exemplary class II tetramer repertoire.
[0355] [Figure 27C] Figure 27C shows an illustrative diagram of a summary of DRB tetramer repertory coverage for DRB1 alleles for peptide exchange.
[0356] [Figure 27D] Figure 27D shows exemplary coverage of human MHC class II allele production.
[0357] [Figure 27E] Figure 27E shows exemplary results from tetramer staining of samples induced using a Flu epitope (memory response) or an HIV epitope (naive response).
[0358] [Figure 28A] Figure 28A shows an illustrative diagram of a method for evaluating peptides using a fluorescence polarization assay, which enables a rapid screening method for identifying allele restriction for epitope peptides with respect to HLA class II restriction. The assay principle in Figure 28A allows for affinity measurement and explicit measurement of peptide exchange.
[0359] [Figure 28B] Figure 28B shows an exemplary summary of the numerous assay conditions explored in the fluorescence polarization assay using DRB1*01:01 (upper panel). Diagrams of soluble MHC class II alleles and full-length MHC class II alleles with transmembrane domains within surfactant micelles are also shown (lower panel), both constructed using placeholder peptides with cleavable linkers for use in the assay.
[0360] [Figure 28C]Figure 28C shows an exemplary diagram of an assay for investigating the full-length and soluble alleles previously shown in the lower panel of Figure 28B. Briefly, both the full-length and soluble alleles are expressed in cells. The membrane-bound full-length allele form is recovered by permeabilization of the membrane, while the secreted form is recovered from the cell supernatant. The recovered class II HLA allele proteins are purified by passing them through a nickel (Ni2+) column.
[0361] [Figure 28D] Figure 28D shows exemplary data demonstrating that the purification method does not affect peptide potency. The left side shows the mean IC50 values from experiments using full-length HLA-DR1 purified by L243 and full-length HLA-DR1 purified by Ni2+.
[0362] [Figure 28E] Figure 28E shows exemplary data demonstrating that the choice between soluble form (sDR1) and full-length form (fDR1) does not affect peptide potency. The left side shows the mean IC50 values from experiments using either the sDR1 or fDR1 form. FP, fluorescence polarization.
[0363] [Figure 28F] Figure 28F illustrates an exemplary evaluation of peptide binding assays and identification of incompatible peptides using neonmhc2 and NetMHCIIpan.
[0364] [Figure 28G] Figure 28G shows exemplary fluorescence polarization binding screening data for evaluating peptides predicted by neonmhc2; also shown as a heatmap and as percentage inhibition of probe binding for each peptide concentration used. Green indicates good binding in proportion to the intensity of the color. Yellow indicates intermediate binding, and red indicates poor binding, also shown as the corresponding percentage inhibition value.
[0365] [Figure 28H] Figure 28H shows a summary of an exemplary binding assay for the evaluation of peptides predicted by neonmhc2.
[0366] [Figure 29] Figure 29 shows exemplary average counts of peptides from average MAPTAC™ experimental replicates (50 million cells) for each HLA allele.
[0367] [Figure 30A-B]Figures 30A–30C show the fidelity of exemplary binding core analysis and multi-allele backconvolution for the HLA class II MAPTAC® allele + / -HLA-DM. Figure 30A shows one representative exemplary sequence logo for the HLA-DR, HLA-DQ, and HLA-DP alleles according to MAPTAC® and IEDB, with or without HLA-DM co-transfection (expi293 cell line), where the height of each amino acid is proportional to its frequency. Amino acids with frequencies higher than 10% are shown in color according to their chemical properties; all other amino acids are shown in gray. Peptides were aligned according to the GibbsCluster tool (supplementary method), and the logos represent all peptides, including those that do not closely match the entire motif (e.g., no peptides are isolated in the "trash" cluster). Figure 30B shows an exemplary description of cluster assignment for MAPTAC® peptides (20 per allele) spiked into a pan-DR MS dataset. Deconvolution was performed on the dataset using GibbsCluster. Each colored square represents one MAPTAC™ peptide. The color of the square indicates the cluster to which it is assigned, and the gray bars indicate the allele from which the peptide originates. Figure 30C shows an exemplary graph illustrating peptide sharing for the alleles shown in Figure 30B, showing 0, 1, 2, 3, or 4 predicted residues at the anchor position. The anchor position was defined as the four positions with the lowest entropy, and "predicted" residues were defined as residues with a frequency of 10% or more at those positions. [Figure 30C]Figures 30A–30C show the fidelity of exemplary binding core analysis and multi-allele backconvolution for the HLA class II MAPTAC® allele + / -HLA-DM. Figure 30A shows one representative exemplary sequence logo for the HLA-DR, HLA-DQ, and HLA-DP alleles according to MAPTAC® and IEDB, with or without HLA-DM co-transfection (expi293 cell line), where the height of each amino acid is proportional to its frequency. Amino acids with frequencies higher than 10% are shown in color according to their chemical properties; all other amino acids are shown in gray. Peptides were aligned according to the GibbsCluster tool (supplementary method), and the logos represent all peptides, including those that do not closely match the entire motif (e.g., no peptides are isolated in the "trash" cluster). Figure 30B shows an exemplary description of cluster assignment for MAPTAC® peptides (20 per allele) spiked into a pan-DR MS dataset. Deconvolution was performed on the dataset using GibbsCluster. Each colored square represents one MAPTAC™ peptide. The color of the square indicates the cluster to which it is assigned, and the gray bars indicate the allele from which the peptide originates. Figure 30C shows an exemplary graph illustrating peptide sharing for the alleles shown in Figure 30B, showing 0, 1, 2, 3, or 4 predicted residues at the anchor position. The anchor position was defined as the four positions with the lowest entropy, and "predicted" residues were defined as residues with a frequency of 10% or more at those positions.
[0368] [Figure 31A]Figures 31A–31F show exemplary architectures and benchmarking of neonmhc2 binding prediction algorithms. Figure 31A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish between a single allele HLA class II peptide and scrambled, length-balanced decoys. The schematic diagram shows an amino acid feature embedding layer, two 6-width convolutional layers, the presence of skip connections to the end, and the use of a combination of mean pooling and max pooling operations as input to the final logistic output node. Figure 31B shows exemplary positive predictive values (PPVs) for NetMHCIIpan and neonmhc2, evaluated against segments of MAPTAC® data not used for training or hyperparameter optimization. For each allele, n peptides observed by MS were scored together with 19n length-balanced decoys sampled from the same set of source genes, and the top-ranked peptides (e.g., top 5%) n for each predictor were called positive. According to this evaluation protocol, the number of false positives and false negatives will always be equal, so PPV is equivalent to recall. Figure 31C shows exemplary NetMHCIIpan and neonmhc2 PPVs for the TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of confirmed immunogenic epitopes in the evaluated set. Figure 31D shows exemplary ex vivo T cell induction results for novel antigen peptides. The peptides were selected based on high neonmhc2 scores and low NetMHCIIpan scores for HLA-DRB1*11:01. Figure 31E shows a comparison of a model trained on single-allele MAPTAC data with deconvolved multi-allele data evaluated against holdout single-allele data. The values are as shown for neonmhc2 when the training dataset was downsampled to match the size of the deconvolved training set.Figure 31F shows NetMHCIIpan-v3.1, a predictor trained by reverse convolution, and PPV (with and without downsampling) for the neonmhc2 TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of immunogenic epitopes identified in the evaluated set. [Figure 31B]Figures 31A–31F show exemplary architectures and benchmarking of neonmhc2 binding prediction algorithms. Figure 31A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish between a single allele HLA class II peptide and scrambled, length-balanced decoys. The schematic diagram shows an amino acid feature embedding layer, two 6-width convolutional layers, the presence of skip connections to the end, and the use of a combination of mean pooling and max pooling operations as input to the final logistic output node. Figure 31B shows exemplary positive predictive values (PPVs) for NetMHCIIpan and neonmhc2, evaluated against segments of MAPTAC® data not used for training or hyperparameter optimization. For each allele, n peptides observed by MS were scored together with 19n length-balanced decoys sampled from the same set of source genes, and the top-ranked peptides (e.g., top 5%) n for each predictor were called positive. According to this evaluation protocol, the number of false positives and false negatives will always be equal, so PPV is equivalent to recall. Figure 31C shows exemplary NetMHCIIpan and neonmhc2 PPVs for the TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of confirmed immunogenic epitopes in the evaluated set. Figure 31D shows exemplary ex vivo T cell induction results for novel antigen peptides. The peptides were selected based on high neonmhc2 scores and low NetMHCIIpan scores for HLA-DRB1*11:01. Figure 31E shows a comparison of a model trained on single-allele MAPTAC data with deconvolved multi-allele data evaluated against holdout single-allele data. The values are as shown for neonmhc2 when the training dataset was downsampled to match the size of the deconvolved training set.Figure 31F shows NetMHCIIpan-v3.1, a predictor trained by reverse convolution, and PPV (with and without downsampling) for the neonmhc2 TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of immunogenic epitopes identified in the evaluated set. [Figure 31C]Figures 31A–31F show exemplary architectures and benchmarking of neonmhc2 binding prediction algorithms. Figure 31A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish between a single allele HLA class II peptide and scrambled, length-balanced decoys. The schematic diagram shows an amino acid feature embedding layer, two 6-width convolutional layers, the presence of skip connections to the end, and the use of a combination of mean pooling and max pooling operations as input to the final logistic output node. Figure 31B shows exemplary positive predictive values (PPVs) for NetMHCIIpan and neonmhc2, evaluated against segments of MAPTAC® data not used for training or hyperparameter optimization. For each allele, n peptides observed by MS were scored together with 19n length-balanced decoys sampled from the same set of source genes, and the top-ranked peptides (e.g., top 5%) n for each predictor were called positive. According to this evaluation protocol, the number of false positives and false negatives will always be equal, so PPV is equivalent to recall. Figure 31C shows exemplary NetMHCIIpan and neonmhc2 PPVs for the TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of confirmed immunogenic epitopes in the evaluated set. Figure 31D shows exemplary ex vivo T cell induction results for novel antigen peptides. The peptides were selected based on high neonmhc2 scores and low NetMHCIIpan scores for HLA-DRB1*11:01. Figure 31E shows a comparison of a model trained on single-allele MAPTAC data with deconvolved multi-allele data evaluated against holdout single-allele data. The values are as shown for neonmhc2 when the training dataset was downsampled to match the size of the deconvolved training set.Figure 31F shows NetMHCIIpan-v3.1, a predictor trained by reverse convolution, and PPV (with and without downsampling) for the neonmhc2 TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of immunogenic epitopes identified in the evaluated set. [Figure 31D]Figures 31A–31F show exemplary architectures and benchmarking of neonmhc2 binding prediction algorithms. Figure 31A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish between a single allele HLA class II peptide and scrambled, length-balanced decoys. The schematic diagram shows an amino acid feature embedding layer, two 6-width convolutional layers, the presence of skip connections to the end, and the use of a combination of mean pooling and max pooling operations as input to the final logistic output node. Figure 31B shows exemplary positive predictive values (PPVs) for NetMHCIIpan and neonmhc2, evaluated against segments of MAPTAC® data not used for training or hyperparameter optimization. For each allele, n peptides observed by MS were scored together with 19n length-balanced decoys sampled from the same set of source genes, and the top-ranked peptides (e.g., top 5%) n for each predictor were called positive. According to this evaluation protocol, the number of false positives and false negatives will always be equal, so PPV is equivalent to recall. Figure 31C shows exemplary NetMHCIIpan and neonmhc2 PPVs for the TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of confirmed immunogenic epitopes in the evaluated set. Figure 31D shows exemplary ex vivo T cell induction results for novel antigen peptides. The peptides were selected based on high neonmhc2 scores and low NetMHCIIpan scores for HLA-DRB1*11:01. Figure 31E shows a comparison of a model trained on single-allele MAPTAC data with deconvolved multi-allele data evaluated against holdout single-allele data. The values are as shown for neonmhc2 when the training dataset was downsampled to match the size of the deconvolved training set.Figure 31F shows NetMHCIIpan-v3.1, a predictor trained by reverse convolution, and PPV (with and without downsampling) for the neonmhc2 TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of immunogenic epitopes identified in the evaluated set. [Figure 31E-F]Figures 31A–31F show exemplary architectures and benchmarking of neonmhc2 binding prediction algorithms. Figure 31A shows an exemplary architecture of a convolutional neural network (CNN) trained to distinguish between a single allele HLA class II peptide and scrambled, length-balanced decoys. The schematic diagram shows an amino acid feature embedding layer, two 6-width convolutional layers, the presence of skip connections to the end, and the use of a combination of mean pooling and max pooling operations as input to the final logistic output node. Figure 31B shows exemplary positive predictive values (PPVs) for NetMHCIIpan and neonmhc2, evaluated against segments of MAPTAC® data not used for training or hyperparameter optimization. For each allele, n peptides observed by MS were scored together with 19n length-balanced decoys sampled from the same set of source genes, and the top-ranked peptides (e.g., top 5%) n for each predictor were called positive. According to this evaluation protocol, the number of false positives and false negatives will always be equal, so PPV is equivalent to recall. Figure 31C shows exemplary NetMHCIIpan and neonmhc2 PPVs for the TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of confirmed immunogenic epitopes in the evaluated set. Figure 31D shows exemplary ex vivo T cell induction results for novel antigen peptides. The peptides were selected based on high neonmhc2 scores and low NetMHCIIpan scores for HLA-DRB1*11:01. Figure 31E shows a comparison of a model trained on single-allele MAPTAC data with deconvolved multi-allele data evaluated against holdout single-allele data. The values are as shown for neonmhc2 when the training dataset was downsampled to match the size of the deconvolved training set.Figure 31F shows NetMHCIIpan-v3.1, a predictor trained by reverse convolution, and PPV (with and without downsampling) for the neonmhc2 TGEM dataset. For each allele, the top-ranked n peptides were called positive, where n is the number of immunogenic epitopes identified in the evaluated set.
[0369] [Figure 32A-E]Figures 32A–32E show exemplary gene representation and protein processing in the HLA class II tumor peptideome. Figure 32A shows exemplary results for observed and predicted counts of HLA class II peptides per gene, determined by a collaborative analysis of colorectal cancer, melanoma, and ovarian cancer datasets. Predicted counts are derived by multiplying the gene length by the expression level. Predicted and observed counts were summed across relevant samples. Genes known to be present in plasma were marked according to their concentration. Figure 32B shows exemplary results for predicted and observed frequencies of peptides per intracellular localization. Figure 32C shows exemplary results for the enrichment score distribution for genes regulated by the proteasome (ratio of observed to predicted observed, partly as in Figure 32B). Gene sets included those with known ubiquitination sites and those whose abundance increases upon application of proteasome inhibitors. Figure 32D shows three exemplary working models for how HLA class II peptides are processed: i) by destructive cleavage of the protein by cathepsin and other enzymes into peptide fragments to which HLA then binds; ii) by binding the protein or unfolded polypeptide to HLA, then by cleavage to the peptide length; or iii) by partial digestion of the protein followed by binding and further trimming. Each model corresponds to a different prediction method. Figure 32E shows the absolute increase in PPV observed for a logistic regression model including processing-related variables and neonmhc2 binding prediction, compared to a model using binding prediction only. The evaluation was performed on 11 samples profiled with HLA-DR antibodies (the same samples as analyzed in Figure 30B); each point corresponds to one sample. Asterisks indicate significant improvement according to a two-sided paired t-test (*: p<0.01, **: p<0.001, ***: p<0.0001). The same analysis is shown in Figure 40B, but instead, NetMHCIIpan is used as the base predictor. The methods for decoy selection and PPV calculation are the same as those used in Figure 31B.
[0370] [Figure 33A-C]Figures 33A–33G show exemplary results for the identification and prediction of tumor antigens presented by dendritic cells. Figure 33A shows an exemplary diagrammatic representation of an experimental workflow for identifying DC-presented HLA-II ligands derived from cancer cells (K562). Cancer cells were grown in SILAC medium for complete incorporation, lysed or irradiated, and then plated together with monocyte-derived dendritic cells. The presented peptides were isolated with a pan-DR antibody and sequenced by LC-MS / MS. Figure 33B shows exemplary data representing predictive performance for tumor-derived peptides presented by dendritic cells, using the same hit-to-decoy ratio and performance metrics as in Figure 21A. Performance is shown for NetMHCIIpan-based models and neonmhc2-based models, with and without the use of processing features. Figure 33C shows an exemplary gene expression distribution for the source genes of the heavily labeled peptide (red curve, plotted according to K562 expression) observed in UV-treated experiments, compared to the source genes of the lightly labeled peptide (gray curve, plotted according to DC expression). Figure 33D shows an exemplary graph of PPV at a hit-to-decoy ratio of 1:499 for predicting the presented tumor antigen with and without processing features, using a NetMHCIIpan-based model and a neonmhc2-based model. The data points, from left to right, represent the following: Sample: Donor 1 HOCl-treated cells: NetMHCIIpan consecutive expression; NetMHCIIpan consecutive expression + gene bias; NetMHCIIpan consecutive expression + gene bias + DQ duplication, full processing method; Donor 1, UV-treated: neonmhc2; neonmhc2 + threshold expression; neonmhc2 + consecutive expression; neonmhc2 + consecutive expression + gene bias; neonmhc2 + consecutive expression + gene bias + DQ duplication. Figure 33E shows the significance of the localization and functional class of various genes in the prediction of heavy (K562-derived) and light (DC-derived) peptides, respectively. The P-value is calculated according to logistic regression controlling for neonmhc2 binding score and source gene expression. The bar color indicates the sign related to the coefficient in the regression.Figure 33F shows an illustrative diagram of results showing duplication of tumor cell-derived peptide source genes in UV-treated and HOCl-treated experiments (colored by functional class). Figure 33G shows illustrative data demonstrating PPV for predicting tumor antigens presented in a second donor using a logistic model fitted to the heavily labeled peptides observed in a first donor. The model was fitted using neonmhc2 binding alone; binding and expression; or binding, expression, and a binary variable indicating whether the peptide originated from a mitochondrial gene. [Figure 33D-E]Figures 33A–33G show exemplary results for the identification and prediction of tumor antigens presented by dendritic cells. Figure 33A shows an exemplary diagrammatic representation of an experimental workflow for identifying DC-presented HLA-II ligands derived from cancer cells (K562). Cancer cells were grown in SILAC medium for complete incorporation, lysed or irradiated, and then plated together with monocyte-derived dendritic cells. The presented peptides were isolated with a pan-DR antibody and sequenced by LC-MS / MS. Figure 33B shows exemplary data representing predictive performance for tumor-derived peptides presented by dendritic cells, using the same hit-to-decoy ratio and performance metrics as in Figure 21A. Performance is shown for NetMHCIIpan-based models and neonmhc2-based models, with and without the use of processing features. Figure 33C shows an exemplary gene expression distribution for the source genes of the heavily labeled peptide (red curve, plotted according to K562 expression) observed in UV-treated experiments, compared to the source genes of the lightly labeled peptide (gray curve, plotted according to DC expression). Figure 33D shows an exemplary graph of PPV at a hit-to-decoy ratio of 1:499 for predicting the presented tumor antigen with and without processing features, using a NetMHCIIpan-based model and a neonmhc2-based model. The data points, from left to right, represent the following: Sample: Donor 1 HOCl-treated cells: NetMHCIIpan consecutive expression; NetMHCIIpan consecutive expression + gene bias; NetMHCIIpan consecutive expression + gene bias + DQ duplication, full processing method; Donor 1, UV-treated: neonmhc2; neonmhc2 + threshold expression; neonmhc2 + consecutive expression; neonmhc2 + consecutive expression + gene bias; neonmhc2 + consecutive expression + gene bias + DQ duplication. Figure 33E shows the significance of the localization and functional class of various genes in the prediction of heavy (K562-derived) and light (DC-derived) peptides, respectively. The P-value is calculated according to logistic regression controlling for neonmhc2 binding score and source gene expression. The bar color indicates the sign related to the coefficient in the regression.Figure 33F shows an illustrative diagram of results showing duplication of tumor cell-derived peptide source genes in UV-treated and HOCl-treated experiments (colored by functional class). Figure 33G shows illustrative data demonstrating PPV for predicting tumor antigens presented in a second donor using a logistic model fitted to the heavily labeled peptides observed in a first donor. The model was fitted using neonmhc2 binding alone; binding and expression; or binding, expression, and a binary variable indicating whether the peptide originated from a mitochondrial gene. [Figure 33F-G]Figures 33A–33G show exemplary results for the identification and prediction of tumor antigens presented by dendritic cells. Figure 33A shows an exemplary diagrammatic representation of an experimental workflow for identifying DC-presented HLA-II ligands derived from cancer cells (K562). Cancer cells were grown in SILAC medium for complete incorporation, lysed or irradiated, and then plated together with monocyte-derived dendritic cells. The presented peptides were isolated with a pan-DR antibody and sequenced by LC-MS / MS. Figure 33B shows exemplary data representing predictive performance for tumor-derived peptides presented by dendritic cells, using the same hit-to-decoy ratio and performance metrics as in Figure 21A. Performance is shown for NetMHCIIpan-based models and neonmhc2-based models, with and without the use of processing features. Figure 33C shows an exemplary gene expression distribution for the source genes of the heavily labeled peptide (red curve, plotted according to K562 expression) observed in UV-treated experiments, compared to the source genes of the lightly labeled peptide (gray curve, plotted according to DC expression). Figure 33D shows an exemplary graph of PPV at a hit-to-decoy ratio of 1:499 for predicting the presented tumor antigen with and without processing features, using a NetMHCIIpan-based model and a neonmhc2-based model. The data points, from left to right, represent the following: Sample: Donor 1 HOCl-treated cells: NetMHCIIpan consecutive expression; NetMHCIIpan consecutive expression + gene bias; NetMHCIIpan consecutive expression + gene bias + DQ duplication, full processing method; Donor 1, UV-treated: neonmhc2; neonmhc2 + threshold expression; neonmhc2 + consecutive expression; neonmhc2 + consecutive expression + gene bias; neonmhc2 + consecutive expression + gene bias + DQ duplication. Figure 33E shows the significance of the localization and functional class of various genes in the prediction of heavy (K562-derived) and light (DC-derived) peptides, respectively. The P-value is calculated according to logistic regression controlling for neonmhc2 binding score and source gene expression. The bar color indicates the sign related to the coefficient in the regression.Figure 33F shows an illustrative diagram of results showing duplication of tumor cell-derived peptide source genes in UV-treated and HOCl-treated experiments (colored by functional class). Figure 33G shows illustrative data demonstrating PPV for predicting tumor antigens presented in a second donor using a logistic model fitted to the heavily labeled peptides observed in a first donor. The model was fitted using neonmhc2 binding alone; binding and expression; or binding, expression, and a binary variable indicating whether the peptide originated from a mitochondrial gene.
[0371] [Figure 34A-B] Figures 34A–34B show exemplary characterization of MAPTAC® data related to Figure 29. Figure 34A shows exemplary HLA cell surface analysis by FACS of Expi293 cell lines transfected with a MAPTAC® construct encoding affinity-tagged HLA-A*02:01-BAP. Figure 34B shows exemplary HLA cell surface analysis by FACS of Expi293 cell lines transfected with a MAPTAC® construct encoding affinity-tagged HLA-DRB1*11:01-BAP (bottom). HLA cell surface expression of transfected Expi293 cells (orange) was compared with stained untransfected Expi293 (blue), unstained untransfected Expi293 (red), stained PBMCs (dark green), and unstained PBMCs (light green). W6 / 32 (pan-HLA class I) was used for all HLA class I staining, while REA332 (pan-HLA class II) was used for HLA class II staining.
[0372] [Figure 35]Figure 35 shows an exemplary comparison of the MAPTAC® and IEDB logos, related to Figure 30A. While it did not show a good NetMHCIIpan score, it was well supported by MS (scored peak intensity > 70 and nested set size ≥ 1), with measured affinity to the peptide observed by MS and affinity predicted by NetMHCIIpan.
[0373] [Figure 36A] Figures 36A–36C show exemplary analyses of HLA-DR1 MAPTAC® data fidelity related to Figures 30A–30C. Figure 36A shows exemplary NetMHCIIpan 3.1 scores for common alleles, comparing HLA-DR1 MAPTAC® peptides (green) (lengths 12–23) with 50,000 length-balanced decoy peptides (blue) randomly sampled from the proteome. Figure 36B shows exemplary measured and NetMHCIIpan-predicted affinities to peptides observed by exemplary MS that do not show good NetMHCIIpan scores but are well supported by MS (scored peak intensity > 70 and nested set size ≥ 1). Figure 36C shows exemplary HLA class II sequence logos for HLA-DRB1 alleles determined by MAPTAC® in different cell types. [Figure 36B]Figures 36A–36C show exemplary analyses of HLA-DR1 MAPTAC® data fidelity related to Figures 30A–30C. Figure 36A shows exemplary NetMHCIIpan 3.1 scores for common alleles, comparing HLA-DR1 MAPTAC® peptides (green) (lengths 12–23) with 50,000 length-balanced decoy peptides (blue) randomly sampled from the proteome. Figure 36B shows exemplary measured and NetMHCIIpan-predicted affinities to peptides observed by exemplary MS that do not show good NetMHCIIpan scores but are well supported by MS (scored peak intensity > 70 and nested set size ≥ 1). Figure 36C shows exemplary HLA class II sequence logos for HLA-DRB1 alleles determined by MAPTAC® in different cell types. [Figure 36C] Figures 36A–36C show exemplary analyses of HLA-DR1 MAPTAC® data fidelity related to Figures 30A–30C. Figure 36A shows exemplary NetMHCIIpan 3.1 scores for common alleles, comparing HLA-DR1 MAPTAC® peptides (green) (lengths 12–23) with 50,000 length-balanced decoy peptides (blue) randomly sampled from the proteome. Figure 36B shows exemplary measured and NetMHCIIpan-predicted affinities to peptides observed by exemplary MS that do not show good NetMHCIIpan scores but are well supported by MS (scored peak intensity > 70 and nested set size ≥ 1). Figure 36C shows exemplary HLA class II sequence logos for HLA-DRB1 alleles determined by MAPTAC® in different cell types.
[0374] [Figure 37A-B]Figures 37A–37C show additional exemplary analyses of MAPTAC® motifs related to Figures 30A–30C. Figure 37A shows MAPTAC®-derived sequence logos for experiments with and without HLA-DM co-transfection (expi293 cell line). Figure 37B shows sequence logos for several HLA class I alleles according to MAPTAC® and IEDB. Note that A*32:01 does not show high-frequency Q at P2 and C*03:03 does not show high-frequency Y at P9, which differs from previous studies using multiple allele reverse convolution; the logo for B*52:01 has not been previously published. Figure 37C shows exemplary alignments of peptides observed by MAPTAC® to the CD74 gene sequence. [Figure 37C] Figures 37A–37C show additional exemplary analyses of MAPTAC® motifs related to Figures 30A–30C. Figure 37A shows MAPTAC®-derived sequence logos for experiments with and without HLA-DM co-transfection (expi293 cell line). Figure 37B shows sequence logos for several HLA class I alleles according to MAPTAC® and IEDB. Note that A*32:01 does not show high-frequency Q at P2 and C*03:03 does not show high-frequency Y at P9, which differs from previous studies using multiple allele reverse convolution; the logo for B*52:01 has not been previously published. Figure 37C shows exemplary alignments of peptides observed by MAPTAC® to the CD74 gene sequence.
[0375] [Figure 38Ai]Figures 38Ai–38D show exemplary neonmhc2 performance statistics and T cell flow staining related to Figures 31A–31D. Figure 38Ai shows the performance of neonmhc2 according to the size of an exemplary training dataset. PPV was evaluated using the same evaluation peptide in the same manner as in Figure 31B; however, the training data was randomly downsampled to mimic a smaller training dataset. Figure 38Aii shows exemplary sequence logos for peptide clusters derived from the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash cluster allowed"). Figure 38B shows exemplary representative flow cytometry plots of IFN-γ expression by CD4+ cells from induced samples recall using neonmhc2 novel antigen peptide prediction. The delta value was calculated by subtracting the percentage of CD4+ cells expressing IFN-γ (+peptide) when recalled with the novel antigen from the percentage of CD4+ cells expressing IFN-γ (without peptide) when recalling in the absence of the novel antigen. The two flow plots on the left are CD4+ The figures show novel antigens that induced a T cell response (PEASLYGALSKGSGG) and novel antigens that did not induced a T cell response (PATYILILKEFCLVG). Figure 38C shows exemplary delta values from wells recalled with a single neonmhc2 novel antigen peptide. Peptides were considered induced hits if they had a positive response (delta response greater than 3%, highlighted). Figure 38D shows exemplary sequence logos for peptide clusters derived for the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash" cluster allowed). [Figure 38Aii]Figures 38Ai–38D show exemplary neonmhc2 performance statistics and T cell flow staining related to Figures 31A–31D. Figure 38Ai shows the performance of neonmhc2 according to the size of an exemplary training dataset. PPV was evaluated using the same evaluation peptide in the same manner as in Figure 31B; however, the training data was randomly downsampled to mimic a smaller training dataset. Figure 38Aii shows exemplary sequence logos for peptide clusters derived from the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash cluster allowed"). Figure 38B shows exemplary representative flow cytometry plots of IFN-γ expression by CD4+ cells from induced samples recall using neonmhc2 novel antigen peptide prediction. The delta value was calculated by subtracting the percentage of CD4+ cells expressing IFN-γ (+peptide) when recalled with the novel antigen from the percentage of CD4+ cells expressing IFN-γ (without peptide) when recalling in the absence of the novel antigen. The two flow plots on the left are CD4+ The figures show novel antigens that induced a T cell response (PEASLYGALSKGSGG) and novel antigens that did not induced a T cell response (PATYILILKEFCLVG). Figure 38C shows exemplary delta values from wells recalled with a single neonmhc2 novel antigen peptide. Peptides were considered induced hits if they had a positive response (delta response greater than 3%, highlighted). Figure 38D shows exemplary sequence logos for peptide clusters derived for the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash" cluster allowed). [Figure 38B-C]Figures 38Ai–38D show exemplary neonmhc2 performance statistics and T cell flow staining related to Figures 31A–31D. Figure 38Ai shows the performance of neonmhc2 according to the size of an exemplary training dataset. PPV was evaluated using the same evaluation peptide in the same manner as in Figure 31B; however, the training data was randomly downsampled to mimic a smaller training dataset. Figure 38Aii shows exemplary sequence logos for peptide clusters derived from the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash cluster allowed"). Figure 38B shows exemplary representative flow cytometry plots of IFN-γ expression by CD4+ cells from induced samples recall using neonmhc2 novel antigen peptide prediction. The delta value was calculated by subtracting the percentage of CD4+ cells expressing IFN-γ (+peptide) when recalled with the novel antigen from the percentage of CD4+ cells expressing IFN-γ (without peptide) when recalling in the absence of the novel antigen. The two flow plots on the left are CD4+ The figures show novel antigens that induced a T cell response (PEASLYGALSKGSGG) and novel antigens that did not induced a T cell response (PATYILILKEFCLVG). Figure 38C shows exemplary delta values from wells recalled with a single neonmhc2 novel antigen peptide. Peptides were considered induced hits if they had a positive response (delta response greater than 3%, highlighted). Figure 38D shows exemplary sequence logos for peptide clusters derived for the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash" cluster allowed). [Figure 38D]Figures 38Ai–38D show exemplary neonmhc2 performance statistics and T cell flow staining related to Figures 31A–31D. Figure 38Ai shows the performance of neonmhc2 according to the size of an exemplary training dataset. PPV was evaluated using the same evaluation peptide in the same manner as in Figure 31B; however, the training data was randomly downsampled to mimic a smaller training dataset. Figure 38Aii shows exemplary sequence logos for peptide clusters derived from the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash cluster allowed"). Figure 38B shows exemplary representative flow cytometry plots of IFN-γ expression by CD4+ cells from induced samples recall using neonmhc2 novel antigen peptide prediction. The delta value was calculated by subtracting the percentage of CD4+ cells expressing IFN-γ (+peptide) when recalled with the novel antigen from the percentage of CD4+ cells expressing IFN-γ (without peptide) when recalling in the absence of the novel antigen. The two flow plots on the left are CD4+ The figures show novel antigens that induced a T cell response (PEASLYGALSKGSGG) and novel antigens that did not induced a T cell response (PATYILILKEFCLVG). Figure 38C shows exemplary delta values from wells recalled with a single neonmhc2 novel antigen peptide. Peptides were considered induced hits if they had a positive response (delta response greater than 3%, highlighted). Figure 38D shows exemplary sequence logos for peptide clusters derived for the multi-allele HLA-DR ligandome using GibbsCluster (default settings; "trash" cluster allowed).
[0376] [Figure 39A-C]Figures 39A–39C show additional exemplary cellular origin analyses for HLA class II, related to Figures 32A–32E. Figure 39A shows exemplary percent rank neonmhc2 scores for HLA class II peptides observed in four PBMC samples (RG1248, RG1104, RG1095, and HDSC in Figure 30B) profiled with pan-DR antibodies, depending on whether the peptide source gene is present in human plasma. For each peptide, the best (lowest) percent rank across alleles present in the donor was used. Scores for randomly length-balanced proteome decoys are shown for comparison. Box plots show the 5th, 25th, 50th, 75th, and 95th percentiles. Figure 39B shows exemplary counts of observed and predicted peptides per gene for HLA class I, using the same methodological framework as in Figure 32A. The data corresponds to the same tumor type (colorectal, ovarian, and melanoma). Genes present in human plasma are highlighted in blue and sized according to their concentration. Figure 39C shows the relative agreement of exemplary peptide observations for two different gene expression profiles. For each sample, gene-level peptide counts were modeled as a linear combination of bulk tumor gene expression and professional APC gene expression profiles. The ratio of the coefficients determines the relative agreement of each expression profile with the peptide repertoire. Error bars correspond to 95% confidence intervals calculated by bootstrapping resampling.
[0377] [Figure 40A-B]Figures 40A–40B show additional exemplary analyses of processing motifs related to Figures 32A–32E. Figure 40A shows exemplary amino acid frequencies near the N-terminal and C-terminal peptide cut portions relative to mean proteome frequencies (applying upstream positions U3–U1 and downstream positions D1–D3) or mean peptide frequencies (applying internal positions N1–C1) observed in donor PBMCs, monocyte-derived dendritic cells, colorectal cancer, melanoma, ovarian cancer, and expi293 cell lines (used for generating the majority of MAPTAC® data). Figure 40B is the same as Figure 32E, but shows an analysis using NetMHCIIpan as the basic predictor. For eight samples profiled with HLA-DR antibodies (the same samples analyzed in Figure 31B), an absolute increase in PPV was observed for logistic regression models that included processing-related variables in addition to predictions by NetMHCIIpan (compared to the model using NetMHCIIpan alone). An asterisk indicates a significant improvement (*: p<0.01, **: p<0.001, ***: p<0.0001). A paired two-tailed t-test was performed.
[0378] [Figure 41] Figure 41 shows an exemplary naming system used to indicate the upstream, internal, and downstream locations of the peptide.
[0379] [Figure 42A] Figure 42A shows an exemplary workflow for the analysis of endogenously processed peptides presented by HLA-1 and HLA class II using nLC-MS / MS.
[0380] [Figure 42B] Figure 42B shows a graph illustrating exemplary experimental results from nLC-MS / MS analysis of trypsin peptides with and without FAIMS. Representative overlaps in the detection of HLA-1 and HLA class II peptides by nLC-MS / MS analysis with and without FAIMS at the indicated analytical scale are also shown.
[0381] [Figure 43A] Figure 43A shows exemplary HLA class I acidic and basic reversed-phase peptide detection with and without FAIMS.
[0382] [Figure 43B] Figure 43B shows exemplary experimental results illustrating the detection of a unique peptide bound to HLA class I, plotted against retention time.
[0383] [Figure 44A] Figure 44A shows exemplary HLA class II acid and basic reversed-phase peptide detection with and without FAIMS.
[0384] [Figure 44B] Figure 44B shows exemplary experimental results illustrating the detection of a unique peptide bound to HLA class II, plotted against retention time.
[0385] [Figure 45-1] Figure 45 shows an exemplary graph of the cross-size of HLA class I-binding peptides detected using the method described (left), as well as Venn diagrams of exemplary standard and optimized workflows for LC-MS / MS detection of HLA class I-binding peptides (right). [Figure 45-2] Figure 45 shows an exemplary graph of the cross-size of HLA class I-binding peptides detected using the method described (left), as well as Venn diagrams of exemplary standard and optimized workflows for LC-MS / MS detection of HLA class I-binding peptides (right).
[0386] [Figure 46-1]Figure 46 shows an exemplary graph of the cross-size of HLA class II-binding peptides detected using the method described (left), as well as Venn diagrams of exemplary standard and optimized workflows for LC-MS / MS detection of HLA class II-binding peptides (right). [Figure 46-2] Figure 46 shows an exemplary graph of the cross-size of HLA class II-binding peptides detected using the method described (left), as well as Venn diagrams of exemplary standard and optimized workflows for LC-MS / MS detection of HLA class II-binding peptides (right). [Modes for carrying out the invention]
[0387] Detailed explanation All terms are intended to be understood as they are understood by those skilled in the art. Unless otherwise defined, all scientific and technical terms used herein have the same meaning as they are generally understood by those skilled in the art to which this disclosure relates.
[0388] The section headings used in this specification are for structural purposes only and should not be construed as limiting the subject matter described herein.
[0389] While various features of this disclosure may be described in relation to a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, while this disclosure may be described in relation to separate embodiments for clarity in this specification, it may also be implemented in a single embodiment.
[0390] This disclosure is based on the important finding that a novel computer-based machine learning HLA-peptide presentation prediction model, which enables the use of HLA class II specific peptides for improved immunotherapy, can predict with high confidence the presentation of antigens, specifically cancer antigens, by specific HLA class II alpha and beta chain pairs.
[0391] In one embodiment, the disclosure provides a method for predicting peptides that can precisely pair or bind to specific HLA class II alpha-beta heterodimers such that high-fidelity binding of the peptide to HLA class II proteins (composed of alpha-beta heterodimers) ensures that the specific peptide is presented to T lymphocytes, thereby eliciting a specific immune response and avoiding any cross-reactivity or indiscriminate immunization. Several recent studies have also shown that CD4+ T cells can recognize ligands presented by HLA class II and contribute to tumor control. While cancer vaccines and other immunotherapies would ideally leverage the direction of CD4+ T cell responses, current attempts completely ignore HLA class II antigen prediction due to the insufficient accuracy of existing predictive tools.
[0392] In one embodiment, the Disclosure provides a method for predicting peptides that can precisely bind to a specific HLA class II protein, such that when the peptide is therapeutically administered to a subject expressing a specific analogue of the HLA class II protein, the peptide can activate a more sustained and robust immune response by leveraging the HLA class II protein's ability to activate CD4+ T cells and stimulate immunological memory. In some embodiments, the methods provided herein demonstrate an improvement over currently available predictors with respect to the prediction of a specific HLA class II protein. In some embodiments, the methods provided herein demonstrate an improvement of at least about 1.1 times over currently available predictors with respect to the prediction of a specific HLA class II protein. In some embodiments, the methods provided herein demonstrate an improvement of at least about 2 times over currently available predictors with respect to the prediction of a specific HLA class II protein. In some embodiments, the methods provided herein demonstrate an improvement of at least about 3 times over currently available predictors with respect to the prediction of a specific HLA class II protein. In some embodiments, the methods provided herein demonstrate an improvement of at least about 4 times over currently available predictors with respect to the prediction of a specific HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 5-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 6-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 7-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 8-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 9-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein.In some embodiments, the methods provided herein demonstrate at least about 10-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 15-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 20-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 30-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 40-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 50-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein. In some embodiments, the methods provided herein demonstrate at least about 60-fold improvement over currently available predictors with respect to the prediction of a particular HLA class II protein.
[0393] In one embodiment, a method of immunotherapy tailored or individualized for a specific target is presented herein. Any target or patient expresses a specific array of HLA class I and HLA class II proteins. HLA typing is a well-known technique that allows for the determination of a specific repertoire of HLA proteins expressed by a target. Once the HLA heterodimers expressed by a particular target are known, the improved, refined, and reliable method described herein for predicting peptides that can bind with high fidelity to specific HLA class II alpha and beta heterodimers ensures that a specific immune response tailored specifically for the target can be produced.
[0394] In this application, unless otherwise specified, the use of the singular form includes the plural form. It should be noted that, as used herein, the singular forms “a,” “an,” and “the” encompass multiple referents unless the context clearly indicates otherwise. In this application, the use of “or” means “and / or” unless otherwise specified. Furthermore, the use of the term “including,” as well as other forms such as “include,” “includes,” and “included,” is not limited to this application. The terms “one or more” or “at least one,” for example, one or more or at least one member of a group of members, are self-evident and, by further illustration, encompass, among other things, any one of the members, or any two or more of the members, for example, any of the members ≥ 3, ≥ 4, ≥ 5, ≥ 6, or ≥ 7, and at most all of the members.
[0395] References to “some embodiments,” “a certain embodiment,” “one embodiment,” or “other embodiments” in this specification mean that features, structures, or characteristics described in relation to an embodiment are included in at least some embodiments of this disclosure, but not necessarily in all embodiments.
[0396] Where used herein and in the claims, the words “comprising” (and any form of “comprising,” such as “comprise” and “comprises”), “having” (and any form of “having,” such as “have” and “has”), “including” (and any form of “including,” such as “includes” and “include”), or “containing” (and any form of “containing,” such as “contains” and “contain”) are inclusive or open-ended, and no additional, unenumerated elements or method steps are excluded. Any embodiment considered herein may be implemented with respect to any method or composition of the Disclosure, and vice versa. Furthermore, the methods of the Disclosure may be implemented using the compositions of the Disclosure.
[0397] When the terms “about” or “approximately” are used herein to refer to measurable values such as parameters, quantities, or durations, they shall encompass variations of less than + / -20%, less than + / -10%, less than + / -5%, or less than + / -1% from the specified value, to the extent that such variations are appropriate for the implementation of this disclosure. It should be understood that the values referred to by the modifiers “about” or “approximately” are themselves specifically disclosed.
[0398] The term "immune response" includes T cell-mediated and / or B cell-mediated immune responses influenced by the modulation of T cell costimulation. Exemplary immune responses include T cell responses, such as cytokine production and cytotoxicity. Furthermore, the term "immune response" includes immune responses indirectly influenced by T cell activation, such as antibody production (humoral response) and cytokine-responsive cell activation, such as macrophage activation.
[0399] The term "receptor" should be understood to mean a biomolecule or group of molecules that can bind to a ligand. Receptors can be useful in the transmission of information in cells, cell formation, or organisms. A receptor comprises at least one receptor unit, and may contain two or more receptor units, in which case each receptor unit may consist of a protein molecule, e.g., a glycoprotein molecule. A receptor has a structure that complements the structure of the ligand and can form a complex with its ligand as a binding partner. Signaling information can be transmitted by changes in the conformation of the receptor after it has bound to a ligand on the surface of a cell. According to this disclosure, a receptor may refer to MHC class I and class II proteins that can form a receptor / ligand complex with a ligand, e.g., a peptide or peptide fragment of appropriate length. The class I and class II MHC peptides encoded by the HLA class I and class II alleles are, as will be well understood in the context of the considerations of those skilled in the art, often referred to herein as HLA class I peptides and HLA class II peptides, or HLA class I and HLA class II peptides, or HLA class I and class II proteins, or HLA class I and HLA class II proteins, or HLA class I and class II molecules, or common variants thereof.
[0400] A "ligand" is a molecule that can form a complex with a receptor. According to this disclosure, a ligand means, for example, a peptide or peptide fragment that has an appropriate length and an appropriate binding motif within its amino acid sequence, and it should be understood that the peptide or peptide fragment can bind to and form a complex with MHC class I or MHC class II proteins (i.e., HLA class I and HLA class II proteins).
[0401] An "antigen" is a molecule that can stimulate an immune response and may be produced by cancer cells, infectious agents, or autoimmune diseases. Antigens recognized by T cells, whether helper T lymphocytes (T helper (TH) cells) or cytotoxic T lymphocytes (CTLs), are not recognized as intact proteins, but rather as small peptides associated with HLA class I or class II proteins on the cell surface. In the course of naturally occurring immune responses, antigens recognized by association with HLA class II molecules on antigen-presenting cells (APCs) are acquired from outside the cell, translocated internally, and processed into small peptides that associate with HLA class II molecules. APCs can also cross-present peptide antigens by processing exogenous antigens and presenting the processed antigens on HLA class I molecules. Antigens that produce peptides recognized by association with HLA class I MHC molecules are generally intracellularly produced peptides, and these antigens are processed and associate with class I MHC molecules. It is now understood that peptides that associate with a given HLA class I or class II molecule are characterized by having a common binding motif, and binding motifs for numerous different HLA class I and class II molecules have been determined. Synthetic peptides can also be synthesized that correspond to the amino acid sequence of a given antigen and contain the binding motif of a given HLA class I or class II molecule. These peptides can then be attached to a suitable APC, and the APC can be used to stimulate a T helper cell or CTL response in vitro or in vivo. The binding motifs, methods for synthesizing peptides, and methods for stimulating T helper cell or CTL responses are all known to those skilled in the art and readily available.
[0402] The term “peptide” is used herein interchangeably with “mutant peptide” and “novel antigenic peptide.” Similarly, the term “polypeptide” is used herein interchangeably with “mutant polypeptide” and “novel antigenic polypeptide.” “Novel antigen” or “neoepitope” means a class of tumor antigens or tumor epitopes arising from tumor-specific mutations in expressed proteins. This disclosure further includes peptides containing tumor-specific mutations, known tumor-specific mutations, and peptides containing mutant polypeptides or fragments thereof identified by the methods of this disclosure. These peptides and polypeptides are referred herein as “novel antigenic peptide” or “novel antigenic polypeptide.” Polypeptides or peptides may be of various lengths, may be in either a neutral (uncharged) form or a salt form, and may have or may not have modifications such as glycosylation, side-chain oxidation, phosphorylation, or any post-translational modifications, and may be subjected to conditions in which the modification does not destroy the biological activity of the polypeptide described herein. In some embodiments, the novel antigenic peptides of the Disclosure may, with respect to HLA class I, consist of 22 residues or less in length, e.g., from about 8 to about 22 residues, from about 8 to about 15 residues, or 9 or 10 residues; and with respect to HLA class II, may consist of 40 residues or less in length, e.g., from about 8 to about 40 residues, from about 8 to about 24 residues, from about 12 to about 19 residues, or from about 14 to about 18 residues. In some embodiments, the novel antigenic peptide or novel antigenic polypeptide contains a neoepitope.
[0403] The term "epitope" includes any protein determinants that can specifically bind to antibodies, antibody peptides, and / or antibody-like molecules (including, but not limited to, T cell receptors) as defined herein. Epitope determinants generally consist of chemically active surface groups of molecules, such as amino acids or sugar side chains, and generally possess specific three-dimensional structural features as well as specific charge features.
[0404] A "T cell epitope" is a peptide sequence to which a class I or class II MHC molecule can bind to a peptide-presenting MHC molecule or MHC complex, and which is then recognized and bound by cytotoxic T lymphocytes or helper T cells, respectively, in these forms.
[0405] The term “antibody,” as used herein, includes IgG (including IgG1, IgG2, IgG3, and IgG4), IgA (including IgA1 and IgA2), IgD, IgE, IgM, and IgY, and also encompasses all antibodies, including single-chain whole antibodies, and their antigen-binding (Fab) fragments. Antigen-binding antibody fragments include, but are not limited to, Fab, Fab', and F(ab')2, Fd (consisting of VH and CH1), single-chain variable fragments (scFv), single-chain antibodies, disulfide-linked variable fragments (dsFv), and fragments containing either a VL domain or a VH domain. Antibodies may originate from any animal of origin. Antigen-binding antibody fragments, including single-chain antibodies, may contain variable regions, either individually or in whole or in part, the following: hinge regions, CH1 domains, CH2 domains, and CH3 domains. This includes any combination of variable regions (multiple) and hinge regions, CH1 domains, CH2 domains, and CH3 domains. Antibodies may be monoclonal antibodies, polyclonal antibodies, chimeric antibodies, humanized antibodies, and human monoclonal and polyclonal antibodies that specifically bind to HLA-associated polypeptides or HLA-HLA-binding peptide (HLA-peptide) complexes. Those skilled in the art will understand that various immunoaffinity techniques are suitable for enriching soluble proteins such as soluble HLA-peptide complexes or membrane-bound HLA-associated polypeptides, for example, those cleaved from membranes by proteolysis. These include (1) immobilizing one or more antibodies capable of specifically binding to soluble proteins onto a fixed or mobile substrate (e.g., plastic wells or resins, latex, or paramagnetic beads), and (2) passing a solution containing soluble proteins derived from a biological sample through the antibody-coated substrate to conjugate the soluble proteins to the antibodies. The substrate containing the antibody and bound soluble protein is separated from the solution, and if necessary, the antibody and soluble protein are separated, for example, by changing the pH and / or ionic strength and / or ionic composition of the solution in which the antibody is immersed.Alternatively, an immunoprecipitation technique can be used, in which antibodies and soluble proteins are combined to form polymer aggregates. These polymer aggregates can then be separated from the solution by molecular sieving or centrifugation.
[0406] The term "immunopurification (IP)" (or immunoaffinity purification or immunoprecipitation) is a well-known and widely used process in the art for isolating a desired antigen from a sample. Generally, this process involves contacting a sample containing the desired antigen with an affinity matrix on which antibodies against the antigen are covalently attached to a solid phase. The antigen in the sample then binds to the affinity matrix through immunochemical binding. The affinity matrix is then washed to remove any unbound species. The antigen is removed from the affinity matrix by changing the chemical composition of the solution in contact with the affinity matrix. Immunopurification can be performed using a column containing the affinity matrix, in which case the solution is the eluent. Alternatively, immunopurification may be a batch process, in which case the affinity matrix is maintained as a suspension in solution. A crucial step in this process is the removal of the antigen from the matrix. This is generally achieved by increasing the ionic strength of the solution in contact with the affinity matrix, for example, by adding an inorganic salt. Changing the pH may also be effective in dissociating the immunochemical binding between the antigen and the affinity matrix.
[0407] The "active substance" is any small molecule chemical compound, antibody, nucleic acid molecule, polypeptide, or fragment thereof.
[0408] A "change" or "alteration" is an increase or decrease. A change can be as large as 1%, 2%, 3%, 4%, 5%, 10%, 20%, 30%, or 40%, 50%, 60%, or even 70%, 75%, 80%, 90%, or 100%.
[0409] A “biological sample” is any tissue, cell, fluid, or other material derived from a living organism. As used herein, the term “sample” includes any biological sample such as any tissue, cell, fluid, or other material derived from a living organism. “Specifically binding” means that a compound (e.g., a peptide) recognizes and binds to a certain molecule (e.g., a polypeptide) in a sample, e.g., a biological sample, but substantially does not recognize or bind to other molecules.
[0410] A "capture reagent" refers to a reagent that specifically binds to a molecule (e.g., a nucleic acid molecule or polypeptide) for the purpose of selecting or isolating that molecule.
[0411] As used herein, the terms “determine,” “evaluate,” “assay,” “measure,” and “detect,” and their grammatical equivalents, refer to both quantitative and qualitative determinations. Therefore, the term “determine” is used herein interchangeably with “assay,” “measure,” etc. When a quantitative determination is intended, phrases such as “determine the quantity” of the analyte are used. When a qualitative and / or quantitative determination is intended, phrases such as “determine the level” of the analyte or “detect” the analyte are used.
[0412] A “fragment” is a portion of a protein or nucleic acid that is substantially identical to a reference protein or nucleic acid. In some embodiments, the portion retains at least 50%, 75%, or 80%, or 90%, 95%, or even 99% of the biological activity of the reference protein or nucleic acid described herein.
[0413] The terms “isolated,” “purified,” and “biologically pure,” and their grammatical equivalents, refer to a material that is free to varying degrees of its native components, which are typically associated with it. “Isolating” indicates a degree of separation from its original source or surroundings. “Purifying” indicates a higher degree of separation than isolation. A “purified” or “biologically pure” protein is sufficiently free of other material, and therefore no impurities substantially affect the biological properties of the protein or cause other harmful consequences. That is, the nucleic acids or peptides of this disclosure are purified if they are substantially free of cellular material, viral material, or culture medium if prepared by recombinant DNA techniques, or chemical precursors or other chemicals if chemically synthesized. Purity and homogeneity are generally determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term “purified” may indicate that essentially one band is produced from the nucleic acid or protein on an electrophoretic gel. With respect to proteins that can be subjected to modification, such as phosphorylation or glycosylation, different modifications can produce different isolated proteins, which can then be purified separately.
[0414] An "isolated" polypeptide (e.g., a peptide derived from an HLA-peptide complex) or polypeptide complex (e.g., an HLA-peptide complex) is a polypeptide or polypeptide complex of the present disclosure isolated from naturally occurring components. Generally, a polypeptide or polypeptide complex is isolated if it does not contain at least 60% by weight of naturally occurring proteins and naturally occurring organic molecules. The preparation may contain at least 75% by weight, at least 90% by weight, or at least 99% by weight of the polypeptide or polypeptide complex of the present disclosure. An isolated polypeptide or polypeptide complex of the present disclosure can be obtained, for example, by extraction from a natural source, by expressing recombinant nucleic acids encoding one or more components of such polypeptide or polypeptide complex, or by chemically synthesizing one or more components of the polypeptide or polypeptide complex. Purity can be measured by any reasonable method, for example, by column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis. In some cases, MHC class II proteins (i.e., MHC class II peptides) encoded by HLA alleles are interchangeably referred to as HLA class II proteins (or HLA class II peptides) within this document.
[0415] The term “vector” refers to a nucleic acid molecule capable of transporting or mediating the expression of heterologous nucleic acids. Plasmids are one of the genera that fall under the term “vector.” A vector generally refers to a nucleic acid sequence containing a replication origin and other entities necessary for replication and / or maintenance in a host cell. A vector capable of directing the expression of a operably linked gene and / or nucleic acid sequence is referred to herein as an “expression vector.” Generally, useful expression vectors are often in the form of a “plasmid,” which refers to a circular double-stranded DNA molecule that does not bind to a chromosome in its vector form and generally contains entities necessary for the stable or transient expression of the encoded DNA. Other expression vectors that can be used in the methods disclosed herein include, but are not limited to, plasmids, episomes, bacterial artificial chromosomes, yeast artificial chromosomes, bacteriophages, or viral vectors, which may be integrated into the host genome or capable of autonomous replication in cells. A vector may be a DNA vector or an RNA vector. Other forms of expression vectors performing equivalent functions, known to those skilled in the art, such as self-replicating extrachromosomal vectors or vectors that can be incorporated into the host genome, may also be used. Exemplary vectors are those that can autonomously replicate and / or express nucleic acids to which they are ligated.
[0416] The terms “spacer” or “linker,” when used in relation to fusion proteins, refer to peptides that join the proteins constituting the fusion protein. Generally, spacers have no particular biological activity other than joining protein or RNA sequences or preserving some minimum distance or other spatial relationship between protein or RNA sequences. However, in some embodiments, the amino acids that make up the spacer can be selected to influence several molecular properties such as molecular folding, net charge, or hydrophobicity. Linkers suitable for use in certain embodiments of this disclosure are well known to those skilled in the art and include, but are not limited to, linear or branched carbon linkers, heterocyclic carbon linkers, or peptide linkers. Linkers are used to separate two antigenic peptides, in some embodiments, by a distance sufficient to ensure that each antigenic peptide folds appropriately. Exemplary peptide linker sequences adopt a flexible, extended conformation and do not tend to produce ordered secondary structures. Typical amino acids within the flexible protein region include Gly, Asn, and Ser. It is predicted that virtually any permutation of amino acid sequences containing Gly, Asn, and Ser will satisfy the above linker sequence criteria. Other nearly neutral amino acids, such as Thr and Ala, can also be used in linker sequences. Further amino acid sequences that can be used as linkers are discussed by Maratea et al. (1985), Gene 40: 39-46;Murphy et al. (1986) Proc. Nat'l. Acad. Sci. USA 83: 8258-62; U.S. Patent No. 4,935,233; and U.S. Patent No. 4,751 This is disclosed in publication number 180.
[0417] The term “neoplastic” refers to any disease caused by or resulting from an inappropriately high level of cell division, an inappropriately low level of apoptosis, or both. Glioblastoma is one non-definite example of neoplastic or cancer. The terms “cancer,” “tumor,” or “hyperproliferative disorder” refer to the presence of cells that have characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic ability, rapid growth and proliferation rates, as well as certain characteristic morphological features. Cancer cells are often in the form of tumors, but such cells can exist alone in animals or may be non-tumor-forming cancer cells, such as leukemia cells. Cancers include, but are not limited to, B-cell cancers (e.g., multiple myeloma, Waldenström macroglobulinemia), heavy chain diseases (e.g., alpha chain disease, gamma chain disease, and muon chain disease), benign monoclonal immunoglobulinemia, and immune cell amyloidosis, melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer (e.g., metastatic prostate cancer, hormone-refractory prostate cancer), pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or appendiceal cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, and hematological tissue cancers.Other non-limiting examples of cancer types applicable to the methods contained herein include human sarcomas and carcinomas, e.g., fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endosarcoma, lymphangiosarcoma, lymphangiosarcoma, synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, colorectal cancer, pancreatic cancer, breast cancer, ovarian cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, cystadenocarcinoma, medullary carcinoma, bronchogenic lung cancer, renal cell carcinoma, hepatoma, cholangiocarcinoma, liver cancer, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, bone cancer, brain tumor, testicular cancer, Lung cancer, small cell lung cancer, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal glandoma, hemangioblastoma, acoustic neuroma, oligobranchial glioma, meningioma, melanoma, neuroblastoma, retinoblastoma; leukemia, such as acute lymphoblastic leukemia and acute myeloid leukemia (myeloblastic leukemia, preosteoblastic leukemia, myelomonocytic leukemia, monocytic leukemia and erythroleukemia); chronic leukemia (chronic myeloid (granulocytic) leukemia and chronic lymphocytic leukemia); as well as polycythemia vera, lymphoma (Hodgkin's disease and non-Hodgkin's disease), multiple myeloma, Waldenström macroglobulinemia, and heavy chain disease. In some embodiments, cancer is an epithelial cancer, including but not limited to bladder cancer, breast cancer, cervical cancer, colon cancer, gynecological cancer, renal cancer, laryngeal cancer, lung cancer, oral cancer, head and neck cancer, ovarian cancer, pancreatic cancer, prostate cancer, or skin cancer. In other embodiments, cancer is breast cancer, prostate cancer, lung cancer, or colon cancer. In yet another embodiment, epithelial cancer is non-small cell lung cancer, non-papillary renal cell carcinoma, cervical cancer, ovarian cancer (e.g., serous ovarian cancer), or breast cancer. Epithelial cancer can be characterized in various other ways, including but not limited to serous, endometrioid, mucinous, clear cell, Brenner, or undifferentiated. In some embodiments, this disclosure is used for the treatment, diagnosis, and / or prognosis of lymphoma or its subtypes, including but not limited to mantle cell lymphoma. Lymphoproliferative disorders are also considered proliferative disorders.
[0418] The term “vaccine” should be understood to mean a composition for producing immunity for the prevention and / or treatment of a disease (e.g., neoplastic / tumor / infectious / autoimmune disease). Therefore, a vaccine is a pharmaceutical product containing an antigen intended for use in humans or animals to produce a specific protective and protective substance by vaccination. A “vaccine composition” may include pharmaceutically acceptable excipients, carriers, or diluents. Aspects of this disclosure relate to the use of such techniques in the preparation of antigen-based vaccines. In these embodiments, vaccine refers to one or more disease-specific antigenic peptides (or corresponding nucleic acids encoding them). In some applications, the antigen-based vaccine contains at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, at least seventeen, at least eighteen, at least nineteen, at least twenty, at least twenty-one, at least twenty-two, at least twenty-three, at least twenty-four, at least twenty-five, at least twenty-six, at least twenty-seven, at least twenty-eight, at least twenty-nine, at least thirty, or more antigenic peptides.In some embodiments, antigen-based vaccines include 2-100, 2-75, 2-50, 2-25, 2-20, 2-19, 2-18, 2-17, 2-16, 2-15, 2-14, 2-13, 2-12, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 2-4, 3-100, 3-75, 3-50, 3-25, 3-20, 3-19, 3-18, 3-17, 3-16, 3-15, 3-14, 3-13, 3-12, 3-10, 3-9, 3-8, and 3- It contains 7, 3-6, 3-5, 4-100, 4-75, 4-50, 4-25, 4-20, 4-19, 4-18, 4-17, 4-16, 4-15, 4-14, 4-13, 4-12, 4-10, 4-9, 4-8, 4-7, 4-6, 5-100, 5-75, 5-50, 5-25, 5-20, 5-19, 5-18, 5-17, 5-16, 5-15, 5-14, 5-13, 5-12, 5-10, 5-9, 5-8, or 5-7 antigenic peptides. In some embodiments, the antigen-based vaccine contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 antigenic peptides. In some cases, the antigenic peptides are novel antigenic peptides. In some cases, the antigenic peptides contain one or more neoepitopes.
[0419] The term “pharmaceutically acceptable” means that it is approved or eligible for approval by a federal or state regulatory authority or is listed in the United States Pharmacopeia or other pharmacopoeias that are generally recognized for use in animals, including humans. “pharmaceutically acceptable excipients, carriers or diluents” means excipients, carriers or diluents that can be administered to a subject together with the active substance, do not destroy the pharmacological activity of the active substance, and are nontoxic when administered in a dose sufficient to deliver a therapeutic amount of the active substance. “pharmaceutically acceptable salts” of pooled disease-specific antigens as described herein may be acid or base salts that are generally considered in the art to be suitable for use in contact with human or animal tissue without excessive toxicity, irritation, allergic reactions, or other problems or complications. Such salts include inorganic and organic acid salts of basic residues such as amines, and alkali or organic salts of acidic residues such as carboxylic acids. Specific pharmaceutical salts include, but are not limited to, hydrochloric acid, phosphoric acid, hydrobromic acid, malic acid, glycolic acid, fumaric acid, sulfuric acid, sulfamic acid, sulfanilic acid, formic acid, toluenesulfonic acid, methanesulfonic acid, benzenesulfonic acid, ethanedisulfonic acid, 2-hydroxyethylsulfonic acid, nitric acid, benzoic acid, 2-acetoxybenzoic acid, citric acid, tartaric acid, lactic acid, stearic acid, salicylic acid, glutamic acid, ascorbic acid, pamoic acid, succinic acid, fumaric acid, maleic acid, propionic acid, hydroxymaleic acid, hydroiodic acid, phenylacetic acid, alkanic acid, and, for example, salts of acids such as acetic acid and HOOC-(CH2)n-COOH (where n is 0 to 4). Similarly, pharmaceutically acceptable cations include, but are not limited to, sodium, potassium, calcium, aluminum, lithium, and ammonium. Those skilled in the art will understand from this disclosure and the knowledge in the art any further pharmaceutically acceptable salts for the pooled disease-specific antigens provided herein, including those enumerated by Remington's Pharmaceutical Sciences, 17th ed., Mack Publishing Company, Easton, PA, p.1418 (1985).Generally, pharmaceutically acceptable acid or base salts can be synthesized from parent compounds containing a basic or acidic moiety by any conventional chemical method. Briefly, such salts can be prepared by reacting these compounds in free acid or base form with a stoichiometric amount of a suitable base or acid in a suitable solvent.
[0420] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecules encoding the polypeptide or a fragment thereof. Such nucleic acid molecules do not need to be 100% identical to the endogenous nucleic acid sequence, but generally exhibit substantial identity. Polynucleotides having substantial identity with the endogenous sequence can generally hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridizing" refers to the case where nucleic acid molecules pair between complementary polynucleotide sequences or parts thereof under various stringency conditions to form a double-stranded molecule (see, for example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152: 399; Kimmel, AR (1987) Methods Enzymol. 152: 507). For example, the stringent salt concentration is usually about 750 mg. The mixture may consist of less than M NaCl and 75 mM trisodium citrate, less than approximately 500 mM NaCl and 50 mM trisodium citrate, or less than approximately 250 mM NaCl and 25 mM trisodium citrate. Low-stringency hybridization can be achieved without the use of organic solvents, such as formamide, while high-stringency hybridization can be achieved in the presence of at least approximately 35% formamide, or at least approximately 50% formamide. Stringent temperature conditions may typically include temperatures of at least approximately 30°C, at least approximately 37°C, or at least approximately 42°C. Variable additional parameters such as hybridization time, the concentration of surfactants, such as sodium dodecyl sulfate (SDS), and whether or not to include carrier DNA are well known to those skilled in the art. By combining these various conditions as needed, various levels of stringency can be achieved. In exemplary embodiments, hybridization can be performed at 30°C with 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another exemplary embodiment, hybridization can be performed at 37°C with 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In yet another exemplary embodiment, hybridization can be performed at 42°C with 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those skilled in the art. For the vast majority of applications, stringency can also be varied with respect to the washing step after hybridization. Wash stringency conditions can be defined by salt concentration and temperature. As described above, washing stringency can be increased by lowering the salt concentration or by raising the temperature. For example, a stringent salt concentration in the washing step may be less than approximately 30 mM NaCl and 3 mM trisodium citrate, or less than approximately 15 mM NaCl and 1.5 mM trisodium citrate.Stringent temperature conditions for the washing step may include temperatures of at least about 25°C, at least about 42°C, or at least about 68°C. In an exemplary embodiment, the washing step can be carried out at 25°C with 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another exemplary embodiment, the washing step can be carried out at 42°C with 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In yet another exemplary embodiment, the washing step can be carried out at 68°C with 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art, for example, Benton and Davis (Science 196: 180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72: 3961, 1975); Ausubel. This information is found in et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0421] "Substantially identical" means that a polypeptide or nucleic acid molecule exhibits at least 50% identity with respect to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or a reference nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). Such a sequence may be at least 60%, 80%, or 85%, 90%, 95%, 96%, 97%, 98%, or even 99% or greater identity with respect to the sequence used for comparison, at the amino acid or nucleic acid level. Sequence identity is generally determined by sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer). The degree of identity is measured using the following software (BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX program, Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions generally include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. An exemplary method for determining the degree of identity is to use the BLAST program with probability scores between e-3 and em° that indicate closely related sequences. "Reference" is the comparison standard.
[0422] The terms “subject” or “patient” refer to an animal being treated, observed, or experimented on. Subjects, by definition, include, but are not limited to, human or non-human mammals, such as non-human primates, mice, cattle, equids, canids, sheep, or felines.
[0423] The terms “treat,” “treated,” “treating,” and “treatment” refer to reducing, preventing, or improving a disorder and / or its associated symptoms (e.g., neoplasia or tumor or infectious agent or autoimmune disease). “Treatment” may refer to administering treatment to a subject after the onset or suspected onset of a disease (e.g., cancer or infectious agent or autoimmune disease). “Treatment” includes the concept of “alleviating,” which refers to reducing the frequency or severity of any symptoms or other pathological effects associated with the disease, and / or side effects associated with treatment. The term “treating” also includes the concept of “managing,” which refers to reducing the severity of a disease or disorder in a patient, e.g., extending the lifespan or survival of a patient with the disease, or delaying its relapse, e.g., lengthening the period of remission in a patient with the disease. While not exclusionary, it should be understood that treating a disorder or condition does not require the complete elimination of the disorder, condition, or associated symptoms.
[0424] The terms “prevent,” “preventing,” and “prevention,” and their grammatical equivalents, as used herein, mean to avoid or delay the onset of symptoms associated with a disease or condition in a subject in whom symptoms are not present at the time of administration of the active substance or compound.
[0425] The term “therapeutic effect” refers to the reduction of one or more symptoms of a disorder (e.g., neoplasia, tumor, or infection or autoimmune disease caused by an infectious agent) or an associated pathological deviation to some degree. “Therapeutic effective dose,” as used herein, refers to the amount of an active agent, administered to cells or subjects in single or multiple doses, that is effective in extending the survival of a patient with such disorder, reducing, preventing, or delaying one or more signs or symptoms of the disorder, beyond what would be expected in the absence of such treatment. “Therapeutic effective dose” is intended to limit the amount required to achieve a therapeutic effect. A physician or veterinarian with ordinary art in the art can readily determine and prescribe the “therapeutic effective dose” (e.g., ED50) of a required pharmaceutical composition. For example, a physician or veterinarian may start the dose of a compound of the disclosure used in a pharmaceutical composition at a level lower than the dose required to achieve the desired therapeutic effect and gradually increase the dose until the desired effect is achieved. Diseases, conditions, and disorders are used interchangeably herein.
[0426] Those skilled in the art will understand that the terms “peptide tag,” “affinity tag,” “epitope tag,” or “affinity acceptor tag” are used interchangeably herein. As used herein, the term “affinity acceptor tag” refers to an amino acid sequence that enables the tagged protein to be easily detected or purified, for example, by affinity purification. Affinity acceptor tags are generally (but not necessarily) located at or near the N-terminus or C-terminus of an HLA allele. Various peptide tags are well known in the art. Non-limiting examples include: polyhistidine tags (e.g., 4-15 consecutive His residues, e.g., 8 consecutive His residues); polyhistidine-glycine tags; HA tags (e.g., Field et al., Mol. Cell. Biol., 8: 2159, 1988); c-myc tags (e.g., Evans et al., Mol. Cell. Biol., 5: 3610, 1985); herpes simplex virus glycoprotein D (gD) tags (e.g., Paborsky et al., Protein Engineering, 3: 547, 1990); FLAG tags (examples) For example, Hopp et al., BioTechnology, 6: 1204, 1988; U.S. 4,703,0 Nos. 04 and 4,851,341); KT3 epitope tag (e.g., Martine et al., Science, 255: 192, 1992); tubulin epitope tag (e.g., Skinner, Biol. Chem., 266: 15173, 1991); T7 gene 10 protein peptide tag (example) For example, Lutz-Freyemuth et al., Proc. Natl. Acad. Sci. USA, 87: 6393, 1990); streptavidin tag (StrepTag® or StrepTagII) (Trademark); see, for example, Schmidt et al., J. Mol. Biol., 255 (5):753-766, 1996 or U.S. Patent No. 5,506,121; also commercially available from Sigma-Genosys); or the VSV-G epitope tag derived from the Vesicular Stomatis viral glycoprotein; or the V5 tag derived from a small epitope (Pk) found in the P and V proteins of the simian virus 5 (SV5) paramyxovirus. In some embodiments, affinity acceptor tags are “epitope tags,” which are a type of peptide tag that adds a recognizable epitope (antibody-binding site) to an HLA-protein, resulting in the binding of the corresponding antibody, thereby enabling identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags are protein A or protein G that bind to IgG. In some embodiments, the matrix of IgG Sepharose 6 Fast Flow chromatography resin is covalently coupled to human IgG. This resin enables high flow rates for the rapid and convenient purification of proteins tagged with protein A. Numerous other tag portions are known to those skilled in the art, can be conceived by those skilled in the art, and are intended herein. Any peptide tag can be used as long as it can be expressed as an element of an HLA-peptide complex tagged with an affinity acceptor.
[0427] As used herein, the term “affinity molecule” refers to a molecule or ligand that binds to an affinity acceptor peptide with chemical specificity. Chemical specificity is the ability of a protein’s binding site to a specific ligand. The fewer ligands a protein can bind to, the greater its specificity. Specificity describes the strength of binding between a given protein and ligand. This relationship can be described by the dissociation constant (KD), which characterizes the equilibrium between the bound and unbound states of a protein-ligand system.
[0428] The term "affinity acceptor-tagged HLA-peptide complex" refers to a complex comprising an HLA class I or class II related peptide or a portion thereof, and a single allele recombinant HLA class I or class II peptide specifically bound to it, which contains an affinity acceptor peptide.
[0429] The terms "specific binding" or "specifically binding," when used in relation to the interaction between affinity molecules and affinity acceptor tags or epitopes and HLA peptides, mean that the interaction depends on the presence of a specific structure on the protein (e.g., an antigenicity-determining factor or epitope); in other words, affinity molecules recognize and bind to a specific affinity acceptor peptide structure, rather than to the protein in general.
[0430] As used herein, the term “affinity” refers to a criterion for evaluating the strength of the binding between two members of a binding pair, e.g., “affinity acceptor tag” and “affinity molecule,” and between an HLA-binding peptide and an HLA class I or class II molecule. KD is the dissociation constant, with units of volume molar concentration. The affinity constant is the inverse of the dissociation constant. The affinity constant is sometimes used as a general term to describe this chemical entity. The affinity constant is a direct criterion for evaluating the energy of the binding. Affinity can be experimentally determined, for example, by surface plasmon resonance (SPR) using commercially available Biacore SPR units. Affinity can also be expressed as inhibitory concentration 50 (IC50), which is the concentration at which 50% of the peptide is replaced. Similarly, lnIC50 refers to the natural logarithm of IC50. off This refers, for example, to the dissociation rate constant for the dissociation of affinity molecules from an HLA-peptide complex tagged with an affinity acceptor.
[0431] In some embodiments, affinity acceptor-tagged HLA-peptide complexes contain biotin acceptor peptides (BAPs) and are immunopurified from the complex cell mixture using streptavidin / NeutrAvidin beads. The biotin-avidin / streptavidin bond is the strongest non-covalent interaction known in nature. This property is utilized as a biological tool for a wide range of applications, such as the immunopurification of proteins to which biotin is covalently attached. In exemplary embodiments, biotin acceptor peptides (BAPs) are incorporated as affinity acceptor tags for immunopurification into nucleic acid sequences encoding HLA alleles. BAPs can be used to specifically biotinize a single lysine residue within the tag in vivo or in vitro (e.g., U.S. Patents 5,723,584; 5,874,239; and 5,932,433; and UK Patent GB2370039). BAP is typically 15 amino acids long and contains a single lysine residue as a biotiny acceptor residue. In some embodiments, BAP is positioned at or near the N-terminus or C-terminus of a single allele HLA peptide. In some embodiments, BAP is positioned between the heavy chain domain and the β2 microglobulin domain of an HLA class I peptide. In some embodiments, BAP is positioned between the β-chain domain and the α-chain domain of an HLA class II peptide. In some embodiments, BAP is positioned within the loop region between the α1, α2, and α3 domains of the heavy chain of an HLA class I peptide, or between the α1 and α2 domains and between the β1 and β2 domains of the α-chain and β-chain, respectively, of an HLA class II peptide. Exemplary constructs and immunopurifications designed for HLA class I and class II expression incorporating BAP for biotinylation are shown in Figure 2.
[0432] As used herein, the term “biotin” refers to the compound biotin itself, as well as its analogs, derivatives, and variants. Thus, the term “biotin” includes biotin (cis-hexahydro-2-oxo-1H-thieno[3,4]imidazole-4-pentanoic acid) and any derivatives and analogs, including biotin-like compounds. Examples of such compounds include biotin-eN-lysine, biocitin hydrazide, amino or sulfhydryl derivatives of 2-iminobiotin, and biotinyl-E-aminocaproic acid-N-hydroxysuccinimide ester, sulfosuccinimidoiminobiotin, biotin bromoacetylhydrazide, p-diazobenzoylbiocitin, 3-(N-maleimidopropionyl)biocitin, desthiobiotin, and the like. The term "biotin" also includes biotin variants that can specifically bind to one or more of the following: rizavidin, avidin, streptavidin, tamavidin moieties, or other avidin-like peptides.
[0433] As used herein, “PPV determination method” may refer to a presentation PPV determination method. For example, a “PPV determination method” is a method that (a) processes amino acid information of multiple test peptide sequences using an HLA peptide presentation prediction model, such as a machine learning HLA peptide presentation prediction model, in order to generate multiple test presentation predictions, wherein each test presentation prediction determines the class II HLA allele of the cell in question, such as a class II HLA allele of the cell in question. The possibility is shown that a given test peptide sequence among a plurality of test peptide sequences may be presented by one or more proteins encoded by an HLA allele, and the plurality of test peptide sequences comprises (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 500 test peptide sequences including at least 499 decoy peptide sequences contained in a protein encoded by the genome of an organism such as an organism of the same species as the subject, and the plurality of test peptide sequences has a ratio of less than one hit peptide sequence to the number of decoy peptide sequences, for example, a ratio of 1:499 for at least one hit peptide sequence to at least 499 decoy peptide sequences; (b) a step of identifying or calling the top percentage of the plurality of test peptide sequences, for example, 0.2% of the top plurality of test peptide sequences, as being presented by the cell's class II HLA allele; and (c) a step of calculating the PPV of an HLA peptide presentation prediction model, wherein the PPV is identified by mass spectrometry as being presented by the cell's class II HLA allele, among a plurality of A method may include the steps of: a peptide observed to be presented by an HLA allele, a fraction of a test peptide sequence identified or called to be presented by a cell's class II HLA allele, and a method of the peptide observed to be presented by an HLA allele, a fraction of a test peptide sequence identified or called to be presented by a cell's class II HLA allele, and a hit peptide. In some embodiments, the decoy peptide is the same length, i.e., contains the same number of amino acids as the hit peptide. In some embodiments, the decoy peptide may include the steps of:The decoy peptide is an endogenous peptide identified by mass spectrometry to bind to a first MHC class I or class II protein, and the first MHC class I or class II protein is distinct from the second MHC class I or class II protein that binds to the hit peptide. In some embodiments, the decoy peptide may be a scrambled peptide, for example, the decoy peptide may contain an amino acid sequence in which amino acid positions are rearranged relative to the amino acid positions of the hit peptide within the length of the peptide. In some embodiments, the PPV determination method may be a presentation PPV determination method. In some embodiments, the ratio of the number of hit peptide sequences to the number of decoy peptide sequences is approximately 1:10, 1:20, 1:50, 1:100, 1:250, 1:500, 1:1000, 1:1500, 1:2000, 1:2500, 1:5000, 1:7500, 1:10000, 1:25000, 1:50000, or 1:100000. In some embodiments, at least one hit peptide sequence is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 Contains hit peptide sequences of 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100. In some embodiments, at least 499 decoy peptide sequences are at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700,4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 87 00, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 320 00, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87 Contains decoy peptide sequences of 500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000, or 1000000. In some embodiments, at least 500 test peptide sequences are at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100,4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200 , 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000 ,30000,31000,32000,33000,34000,35000,36000,37000,38000,39000,40000,41000,42000,43000,44000,45000,46000,47000,48000,49000,50000,52500,55000,57500,60000,62500,65000,67500,70000,72500,75000,77500,80000,82500, Includes test peptide sequences of 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000, or 1000000. In some embodiments, the step of identifying or calling the top percentages of multiple test peptide sequences as those presented by the cell's class II HLA alleles is performed as follows: top 0.20%, 0.30%, 0.40%, 0.50%, 0.60%, 0.70%, 0.80%, 0.90%, 1.00%, 1.10%, 1.20%, 1.30%, 1.40%, 1.50%, 1.60%, 1.70%,1.80%, 1.90%, 2.00%, 2.10%, 2.20%, 2.30%, 2.40%, 2.50%, 2.60%, 2.70%, 2.80%, 2.90%, 3.00%, 3.10%, 3.20%, 3.30%, 3.40%, 3.50%, 3.60%, 3.70%, 3.80%, 3.90%, 4.00 %, 4.10%, 4.20%, 4.30%, 4.40%, 4.50%, 4.60%, 4.70%, 4.80%, 4.90%, 5.00%, 5.10%, 5.20%, 5.30%, 5.40%, 5.50%, 5.60%, 5.70%, 5.80%, 5.90%, 6.00%, 6.10%, 6.20%, 6.3 0%, 6.40%, 6.50%, 6.60%, 6.70%, 6.80%, 6.90%, 7.00%, 7.10%, 7.20%, 7.30%, 7.40%, 7.50%, 7.60%, 7.70%, 7.80%, 7.90%, 8.00%, 8.10%, 8.20%, 8.30%, 8.40%, 8.50%, 8 This includes identifying or calling 0.60%, 8.70%, 8.80%, 8.90%, 9.00%, 9.10%, 9.20%, 9.30%, 9.40%, 9.50%, 9.60%, 9.70%, 9.80%, 9.90%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20% as those presented by the cell's class II HLA alleles. In some embodiments, the cells are single-allele cells.
[0434] As used herein, “PPV determination method” may refer to a binding PPV determination method. For example, a “PPV determination method” is a method that (a) processes amino acid information of multiple test peptide sequences using an HLA peptide binding prediction model, such as a machine learning HLA peptide binding prediction model, in order to generate multiple test binding predictions, wherein each test binding prediction determines the class II HLA allele of the cell in question, such as a class II HLA allele of the cell in question. The possibility has been shown that one or more proteins encoded by an HLA allele can bind to a given test peptide sequence among a plurality of test peptide sequences, wherein the plurality of test peptide sequences comprises at least 20 test peptide sequences, each comprising (i) at least one hit peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, and (ii) at least 19 decoy peptide sequences contained within a protein comprising at least one peptide sequence identified by mass spectrometry as being presented by an HLA protein expressed in a cell, wherein the plurality of test peptide sequences have a ratio of less than one hit peptide sequence to the number of decoy peptide sequences, for example, a 1:19 ratio of at least one hit peptide sequence to at least 19 decoy peptide sequences, and (b) identifying or calling the top percentage of the plurality of test peptide sequences, for example, 5% of the top plurality of test peptide sequences, as binding to an HLA protein, and (c) calculating the PPV of an HLA peptide binding prediction model, wherein the PPV is a peptide among the plurality that has been observed by mass spectrometry as being presented by a cell's class II HLA allele, and is a cell's class II This method may refer to a fraction of test peptide sequences identified or called as binding to an HLA allele. In some embodiments, the ratio of the number of hit peptide sequences to the number of decoy peptide sequences is approximately 1:2, 1:3, 1:4, 1:5, 1:10, 1:20, 1:25, 1:30, 1:40, 1:50, 1:75, 1:100, 1:200, 1:250, 1:500, or 1:1000. In some embodiments, at least one hit peptide sequence is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19,20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, Contains hit peptide sequences 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100. In some embodiments, at least 19 decoy peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200 ,3300,3400,3500,3600,3700,3800,3900,4000,4100,4200,4300,4400,4500,4600,4700,4800,4900,5000,5100,5200,5300,5400,5500,5600,5700,5800 , 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700, 6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400 , 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000,19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 40000, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 7000 Contains decoy peptide sequences of 0, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000, or 1000000. In some embodiments, at least 20 test peptide sequences are at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 37 00, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, 5100, 5200, 5300, 5400, 5500, 5600, 5700, 5800, 5900, 6000, 6100, 6200, 6300, 6400, 6500, 6600, 6700,6800, 6900, 7000, 7100, 7200, 7300, 7400, 7500, 7600, 7700, 7800, 7900, 8000, 8100, 8200, 8300, 8400, 8500, 8600, 8700, 8800, 8900, 9000, 9100, 9200, 9300, 9400, 9500, 9600, 9700, 9800, 9900, 10000, 11000, 12000, 13000, 14000, 15000, 16000, 17000, 18000, 19000, 20000, 21000, 22000, 23000, 24000, 25000, 26000, 27000, 28000, 29000, 30000, 31000, 32000, 33000, 34000, 35000, 36000, 37000, 38000, 39000, 4000 0, 41000, 42000, 43000, 44000, 45000, 46000, 47000, 48000, 49000, 50000, 52500, 55000, 57500, 60000, 62500, 65000, 67500, 70000, 72500, 75000, 77500, 80000, 82500, 85000, 87500, 90000, 92500, 95000, 97 Contains test peptide sequences of 500, 100000, 125000, 150000, 175000, 200000, 225000, 250000, 275000, 300000, 325000, 350000, 375000, 400000, 425000, 450000, 475000, 500000, 600000, 700000, 800000, 900000, or 1000000. In some embodiments, the step of identifying or calling the top percentage of multiple test peptide sequences as those presented by the cell's class II HLA alleles includes identifying or calling the top 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% as those presented by the cell's class II HLA alleles. In some embodiments, the cells are single-allele cells.
[0435] Human leukocyte antigen (HLA) system The immune system can be classified into two functional subsystems: the innate immune system and the adaptive immune system. The innate immune system is the first line of defense against infection, and the vast majority of potential pathogens are rapidly neutralized by the innate immune system before they can cause, for example, a notable infection. The adaptive immune system reacts to molecular structures called antigens of invading organisms. Unlike the innate immune system, the adaptive immune system is highly specific to pathogens. Adaptive immunity can also provide long-lasting protection; for example, a person who recovers from measles will be protected from measles for life. There are two types of adaptive immune responses, which include humoral immune responses and cell-mediated immune responses. In humoral immune responses, antibodies secreted into the body fluid by B cells bind to antigens derived from pathogens, thereby leading to the elimination of pathogens through various mechanisms, such as complement-mediated lysis. In cell-mediated immune responses, T cells that can destroy other cells are activated. For example, if disease-related proteins are present in cells, these proteins are fragmented into peptides by proteolysis within the cell. Next, specific intracellular proteins themselves attach to the antigen or the peptide thus formed, transporting them to the cell surface where they are presented to the molecular defense mechanisms within the body's T cells. Cytotoxic T cells recognize these antigens and kill the cells possessing them.
[0436] The terms “Major Histocompatibility Complex (MHC),” “MHC molecule,” or “MHC protein” refer to proteins that represent potential T cell epitopes capable of binding to peptides resulting from the cleavage of protein antigens by proteolysis, transport them to the cell surface, and present the peptides to specific cells, such as cytotoxic T lymphocytes or helper T cells. Human MHC is also referred to as the HLA complex. Therefore, the terms “Human Leukocyte Antigen (HLA) System,” “HLA molecule,” or “HLA protein” refer to the gene complex encoding MHC proteins in humans. In mouse species, the term MHC is referred to as the “H-2” complex. It will be understood by those skilled in the art that the terms “Major Histocompatibility Complex (MHC),” “MHC molecule,” “MHC protein,” and “Human Leukocyte Antigen (HLA) System,” “HLA molecule,” and “HLA protein” are used interchangeably herein.
[0437] HLA proteins are classified into two types, known as HLA class I and HLA class II. While the structures of the two HLA class proteins are very similar, they have entirely different functions. HLA class I proteins are present on the surface of almost all cells in the body, including the vast majority of tumor cells. HLA class I proteins are typically loaded with antigens originating from endogenous proteins or pathogens present inside the cell, and then presented to naive or cytotoxic T lymphocytes (CTLs). HLA class II proteins are present on antigen-presenting cells (APCs), including but not limited to dendritic cells, B cells, and macrophages. HLA class II proteins primarily present external antigen sources, such as peptides processed from outside the cell, to helper T cells. The majority of peptides to which HLA class I proteins bind originate from cytoplasmic proteins produced in the healthy host cells of the organism itself and do not usually stimulate an immune response.
[0438] An HLA class I molecule (Figure 1) consists of two non-covalently linked polypeptide chains: an HLA-encoded α-chain (heavy chain, 44-47 kD) and a non-HLA-encoded subunit called β2-microglobulin (or β2m) (12 kD). The α-chain has three extracellular domains, α1, α2, and α3, as well as a transmembrane region. The α1 and α2 regions can bind to peptides of approximately 7-13 amino acids (e.g., approximately 8-11 amino acids, or 9 or 10 amino acids). The HLA class I molecule binds to peptides with appropriate binding motifs and presents them to cytotoxic T lymphocytes. The HLA class I heavy chain can be the protein product of the HLA-A allele (also called the HLA-A monomer), or the protein product of the HLA-B allele (similarly, the HLA-B monomer), or the protein product of the HLA-C allele (HLA-C monomer), each of which forms a complex with β-2-microglobulin. α1 sits on top of the non-HLA protein β2m; β2m is encoded by the β-2-microglobulin gene located on human chromosome 15. The α3 domain connects to the transmembrane region, thereby tethering the HLA class I molecule to the cell membrane. The presented peptide is held at the bottom of the peptide-binding groove in the central region of the α1 / α2 heterodimer (a molecule composed of two non-identical subunits). HLA class IA, HLA class IB, or HLA class IC are highly polymorphic. Each of the HLA class 1-A gene (referred to as the HLA-A gene), HLA class 1-B gene (referred to as the HLA-B gene), and HLA class 1-C gene (referred to as the HLA-C gene) contains eight exons, with exon 1 encoding the leader peptide, exons 2 and 3 encoding the α1 and α2 domains, exon 5 encoding the transmembrane region, and exons 6 and 7 encoding the cytoplasmic tail. Polymorphisms in exons 2 and 3 are responsible for the peptide bond specificity of each class molecule. The HLA class IB gene (HLA-B) has many possible variations, expression patterns, and presented antigens.This group is further subdivided into those encoded within HLA loci, e.g., HLA-E, HLA-F, HLA-G, and is not a stress ligand such as ULBP, Rae1, and H60. Many antigens / ligands of these molecules remain unknown, but they may interact with CD8+ T cells, NKT cells, and NK cells, respectively.
[0439] In some embodiments, this disclosure utilizes non-classical HLA class IE alleles. HLA-E molecules are recognized by natural killer (NK) cells and CD8+ T cells. HLA-E is expressed in almost all tissues, including lung, liver, skin, and placental cells. HLA-E expression is also detected in solid tumors (e.g., osteosarcoma and melanoma). HLA-E molecules bind to the TCR expressed on CD8+ T cells, resulting in T cell activation. It is also known that HLA-E binds to the CD94 / NKG2 receptor expressed on NK cells and CD8+ T cells. CD94 can pair with several different isoforms of NKG2 to form receptors with the potential to inhibit cell activation (NKG2A, NKG2B) or receptors with the potential to promote cell activation (NKG2C). HLA-E can bind to peptides derived from amino acid residues 3-11 of the leader sequences of most HLA-A, HLA-B, HLA-C, and HLA-G molecules, but cannot bind to its own leader peptide. HLA-E has also been shown to present peptides derived from endogenous proteins, similar to the HLA-A, HLA-B, and HLA-C alleles. Under physiological conditions, the association of CD94 / NKG2A with HLA-E loaded with HLA class I leader sequence peptides typically induces an inhibitory signal. Cytomegalovirus (CMV) utilizes a mechanism to escape NK cell immune surveillance by expressing the UL40 glycoprotein, which mimics the HLA-A leader. However, it has also been reported that CD8+ T cells can recognize HLA-E loaded with the UL40 peptide derived from the CMV Toledo strain and play a role in defense against CMV. Several studies have revealed some important functions of HLA-E in infectious diseases and cancer.
[0440] Peptide antigens are presented on the cell surface after attaching themselves to HLA class I molecules within the endoplasmic reticulum by competitive affinity binding. Here, the affinity of individual peptide antigens is directly linked to their amino acid sequence and the presence of specific binding motifs at defined positions within that sequence. When the sequence of such peptides is known, it is possible to manipulate the immune system against diseased cells, for example, using peptide vaccines.
[0441] MHC molecules are highly polymorphic, meaning many MHC variants exist. Each variant is encoded by a variation in the protein-coding gene, and each such variant gene is called an allele. In humans, MHC is known as human leukocyte antigen (HLA) and is associated with three types of HLA class II molecules: DP, DQ, and DR. The HLA class II peptide (Figure 1) has two chains, α and β, each with two domains—α1 and α2 and β1 and β2, respectively, and each chain has transmembrane domains, α2 and β2, respectively, thereby linking the HLA class II molecule to the cell membrane. A peptide-binding groove is formed from the heterodimer of α1 and β1. The most widely studied HLA-DR molecules have DRA and DRB, corresponding to the α and β domains, respectively. DRB is diverse, while DRA is nearly identical. Therefore, the binding specificity of the DRB allele indicates the binding specificity of the corresponding HLA-DR. Each MHC protein possesses its own binding specificity; that is, the set of peptide binding properties for an MHC molecule may differ from that for another MHC molecule. Classical molecules present peptides to CD4+ lymphocytes. Non-classical molecules, which are accessory factors with intracellular functions, are not exposed on the cell membrane but reside within the lysosome membrane and typically load antigenic peptides onto classical HLA class II molecules.
[0442] In the HLA class II system, phagocytic cells such as macrophages and immature dendritic cells take up entities into phagosomes via phagocytosis (although B cells exhibit more common endocytosis into endosomes), where they fuse with lysosomes. The ingested proteins are then cleaved by lysosomal acidic enzymes into many different peptides. Autophagy is another source of HLA class II peptides. Due to physicochemical dynamics in intermolecular interactions with host-generated HLA class II variants encoded in the host genome, certain peptides exhibit immunodominance and are loaded into HLA class II molecules. These are then transported to the cell surface and extruded. The most studied subclasses of HLA class II genes are HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRA, and HLA-DRB1.
[0443] The presentation of peptides to CD4+ helper T cells by HLA class II molecules is necessary for the immune response to foreign antigens (Roche and Furuta, 2015). CD4+ T cells, When activated, CD4+ T cells promote B cell differentiation and antibody production, as well as the CD8+ T cell (CTL) response. CD4+ T cells also secrete cytokines and chemokines that activate other immune cells and induce their differentiation. HLA class II molecules are heterodimers of α- and β-chains that interact to form a peptide-binding groove that is wider than that of HLA class I peptide-binding grooves (Unanue et al., 2016). Peptide molecules bound to HLA class II are thought to have a 9-amino acid binding core with adjacent residues protruding from the groove at either the N-terminus or C-terminus (Jardetzky et al., 1996; Stern et al., 1994). These peptides are typically 12-16 amino acids long and often contain 3-4 anchor residues at the P1, P4, P6 / 7, and P9 positions of the binding register (Rossjohn et al., 2015).
[0444] HLA alleles are expressed in a codominant manner, meaning that alleles (variants) inherited from both parents are expressed equally. For example, each person has two alleles for each of the three class I genes (HLA-A, HLA-B, and HLA-C), and therefore can express six different types of HLA class II. At the HLA class II locus, each person inherits a pair of HLA-DP genes (DPA1 and DPB1, encoding the α and β chains), HLA-DQ (DQA1 and DQB1, for the α and β chains), one gene HLA-DRα (DRA1), and one or more genes HLA-DRβ (DRB1 and DRB3, DRB4, or DRB5). For example, HLA-DRB1 has nearly 400 or more known alleles. This means that a heterozygous individual can inherit six or eight functional HLA class II alleles: three or more from each parent. Therefore, HLA genes are highly polymorphic; many different alleles exist in individuals within different populations. The genes encoding HLA proteins have many possible variations, thereby enabling each person's immune system to react to a wide range of foreign invaders. Some HLA genes have hundreds of identified versions (alleles), each of which is assigned a specific number. In some embodiments, the HLA class I allele is HLA-A * 02:01, HLA-B * 14:02, HLA-A * 23:01, HLA-E * It is 01:01 (non-classical). In some embodiments, the HLA class II allele is HLA-DRB * 01:01, HLA-DRB * 01:02, HLA-DRB * 11:01, HLA-DRB * 15:01, and HLA-DRB * It is 07:01.
[0445] The target-specific HLA allele or HLA genotype of the subject can be determined by any method known in the art. In exemplary embodiments, the HLA genotype is determined by any method described in International Patent Application PCT / US2014 / 068746, published on June 11, 2015, as W02015085147, which is incorporated herein by reference in its entirety. Briefly, the method comprises determining a polymorphic genotype, which may include the steps of: generating an alignment of reads extracted from a sequencing dataset to a gene reference set containing allele variants of the polymorphic gene; determining a first posterior probability or posterior probability-derived score for each allele variant in the alignment; identifying the allele variant having the highest first posterior probability or posterior probability-derived score as the first allele variant; identifying one or more duplicate reads that align with the first allele variant and one or more other allele variants; determining a second posterior probability or posterior probability-derived score for one or more other allele variants using weighting factors; identifying a second allele variant by selecting the allele variant having the highest second posterior probability or posterior probability-derived score, wherein the genotype of the polymorphic gene is defined by the first and second allele variants; and obtaining the output of the first and second allele variants.
[0446] In some embodiments, the MHC class II peptide:antigenic peptide binding and presentation prediction methods described herein can predict conjugates from a large repertoire of MHC class II peptides encoded by individual HLA alleles. In some embodiments, the MAPTAC technology is trained using a large database of mass spectrometry-validated HLA-compatible peptides. In some embodiments, the large database of mass spectrometry-validated HLA-compatible peptides is 1.2 × 10⁻⁶ 6It contains more such HLA-compatible peptides than species. In some embodiments, a large database of HLA-compatible peptides validated by mass spectrometry includes more than 150 HLA alleles, including both MHC class I and class II allele subtypes. In some embodiments, the database includes at least 95% of the US population for HLA-I and HLA-II (DR subtype).
[0447] As described herein, there is ample evidence in both animals and humans that mutant epitopes are effective in inducing immune responses, that spontaneous tumor regression or long-term survival correlates with CD8+ T cell responses to mutant epitopes, and that "immunoediting" can be tracked to alter the expression of dominant mutant antigens in mice and humans.
[0448] Sequencing techniques have revealed that each tumor contains numerous patient-specific mutations that alter the protein-coding content of genes. Such mutations create altered proteins ranging from single-amino acid changes (caused by missense mutations) to the addition of long regions of novel amino acid sequences resulting from frameshifts, termination codon read-throughs, or translation of intron regions (novel open-reading frame mutations; neoORFs). These mutant proteins, unlike native proteins, do not weaken self-tolerant immunity and are therefore valuable targets for the host immune response to tumors. Consequently, mutant proteins are more likely to be immunogenic and more specific to tumor cells compared to the patient's normal cells. Essentially, short peptides (8–24 amino acid length) containing cancer-related mutations are candidates for cancer immunotherapy.
[0449] In some embodiments, algorithms driving the prediction method can be further utilized for mutation calling on the peptide. In some embodiments, the prediction method can be used to determine the state of driver mutations within the peptide, and / or the state of RNA expression, and / or cleavage prediction.
[0450] The term "T cell" includes CD4+ T cells and CD8+ T cells. The term "T cell" also includes both T helper type 1 T cells and T helper type 2 T cells. As used herein, T cells are generally classified by function and cell surface antigen (cluster classification antigen, or CD), and T cells are also classified into two main classes: helper T (TH) cells and cytotoxic T lymphocytes (CTLs), which facilitate the binding of T cell receptors to antigens.
[0451] Mature helper T (TH) cells express the surface protein CD4 and are called CD4+ T cells. After T cell development, mature naive T cells leave the thymus and begin to spread throughout the body, including the lymph nodes. Naive T cells are T cells that have never been exposed to an antigen to which they are programmed to respond. Like all T cells, naive T cells express the T cell receptor-CD3 complex. The T cell receptor (TCR) consists of both a constant region and a variable region. The variable region determines which antigens the T cell can respond to. CD4+ T cells have a TCR with affinity for MHC class II, and the protein and CD4 are involved in determining MHC affinity during maturation in the thymus. MHC class II proteins are generally found only on the surface of specialized antigen-presenting cells (APCs). Specialized antigen-presenting cells (APCs) are primarily dendritic cells, macrophages, and B cells, with dendritic cells being the only cell group that constitutively (at any given time) expresses MHC class II. While some APCs, such as follicular dendritic cells, also bind native (or unprocessed) antigens to their surface, unprocessed antigens do not interact with T cells and are not involved in T cell activation. Peptide antigens that bind to HLA class I proteins are generally shorter than peptide antigens that bind to HLA class II proteins.
[0452] Cytotoxic T lymphocytes (CTLs) are cytotoxic T cells, cytolytic T cells, CD8+ Cytotoxic T lymphocytes (CTLs), also known as T cells or killer T cells, are lymphocytes that induce apoptosis in targeted cells. CTLs form antigen-specific conjugates with target cells through the interaction of the TCR and processed antigen (Ag) on the surface of the target cell, resulting in apoptosis of the targeted cell. Apoptotic bodies are eliminated by macrophages. The term "CTL response" is used to refer to the primary immune response mediated by CTL cells. Cytotoxic T lymphocytes have both a T cell receptor (TCR) and a CD8 molecule on their surface. The T cell receptor can recognize and bind to peptides that have formed complexes with HLA class I molecules. Each cytotoxic T lymphocyte expresses a unique T cell receptor that can bind to a specific MHC / peptide complex. The vast majority of cytotoxic T cells express a T cell receptor (TCR) that can recognize a specific antigen. For the TCR to bind to an HLA class I molecule, a glycoprotein called CD8 must be associated with the TCR, which binds to the constant portion of the HLA class I molecule. Therefore, these T cells are called CD8+ T cells. The affinity between CD8 and MHC molecules maintains a close binding between the T cell and the target cell during antigen-specific activation. CD8+ T cells, once activated, are recognized as T cells and are generally classified as having a predetermined cytotoxic role within the immune system. However, CD8+ T cells also possess the ability to produce several cytokines.
[0453] The T cell receptor (TCR) is a cell surface receptor involved in the activation of T cells in response to antigen presentation. Generally, the TCR consists of two chains, alpha and beta, which assemble to form a heterodimer and associate with a CD3 transduction subunit to form the T cell receptor complex present on the cell surface. Each alpha and beta chain of the TCR consists of an immunoglobulin-like N-terminal variable (V) and constant (C) region, a hydrophobic transmembrane domain, and a short intracytoplasmic region. Regarding immunoglobulin molecules, the variable regions of the alpha and beta chains are generated by V(D)J recombination, creating a wide diversity of antigen specificity within the T cell population. However, in contrast to immunoglobulins that recognize intact antigens, T cells are activated by processed peptide fragments associated with MHC molecules, introducing an extra element known as MHC constraint to T cell antigen recognition. Recognition of MHC mismatch between donor and recipient via the T cell receptor leads to T cell proliferation and the potential development of GVHD. The normal surface expression of TCR has been shown to depend on the harmonious synthesis and assembly of all seven components of the complex (Ashwell and Klusner 1990). As a result of inactivation of TCRα or TCRβ, The disappearance of the TCR from the surface of T cells can result, thereby preventing alloantigen recognition and thus GVHD. However, TCR disruption generally leads to the disappearance of CD3 signaling components, altering the means by which further T cells proliferate.
[0454] The term "HLA peptidome" refers to a pool of peptides that specifically interact with a particular HLA class and can encompass thousands of different sequences. The HLA peptidome includes a diversity of peptides derived from both normal and abnormal proteins expressed in cells. Therefore, the HLA peptidome can be studied to identify cancer-specific peptides, for the development of tumor immunotherapies, and as a source of information about protein synthesis and degradation schemes within cancer cells. In some embodiments, the HLA peptidome is a pool of soluble HLA peptides (sHLA). In some embodiments, the HLA peptidome is a pool of membrane-bound HLA (mHLA).
[0455] "Antigen-presenting cells" or "APCs" include professional antigen-presenting cells (e.g., B lymphocytes, macrophages, monocytes, dendritic cells, Langerhans cells) as well as other antigen-presenting cells (e.g., keratinocytes, endothelial cells, astrocytes, fibroblasts, oligodendrocytes, thymic epithelial cells, thyroid epithelial cells, glial cells (brain), pancreatic beta cells, and vascular endothelial cells). "Antigen-presenting cells" or "APCs" are cells that express major histocompatibility complex (MHC) molecules and are capable of presenting foreign antigens on their surface that have formed complexes with MHC.
[0456] Single-allele HLA cell line Monoallelic cell lines expressing a single HLA class I allele, a single pair of HLA class II alleles, or a single HLA class I allele and a single pair of HLA class II alleles can be generated by transfecting or transfecting a suitable cell population with a polynucleic acid encoding a single HLA allele, such as a vector (Figure 2). Suitable cell populations include, for example, HLA class I-deficient cell lines exogenously expressing a single HLA class I allele, HLA class II-deficient cell lines expressing a single exogenous pair of HLA class II alleles, or class I and class II-deficient cell lines exogenously expressing a single pair of HLA class I and / or class II alleles. An exemplary embodiment is the HLA class I-deficient B cell line, B721.221. However, it will be apparent to those skilled in the art that other cell populations lacking HLA class I and / or HLA class II can be generated. Typical methods for deleting / inactivating endogenous HLA class I or HLA class II genes include, for example, CRISPR-Cas9-mediated genome editing in THP-1 cells. In some embodiments, the cell population is professional antigen-presenting cells such as macrophages, B cells, and dendritic cells. The cells may be B cells or dendritic cells. In some embodiments, the cells are tumor cells or cells derived from tumor cell lines. In some embodiments, the cells are isolated from a patient. In some embodiments, the cells contain an infectious agent or a portion thereof. In some embodiments, the cell population comprises at least 10⁷ cells. In some embodiments, the cell population is further modified, for example, by increasing or decreasing the expression and / or activity of at least one gene. In some embodiments, the gene encodes a member of the immunoproteasome. The immunoproteasome is known to be involved in the processing of HLA class I binding peptides and to include the LMP2(β1i), MECL-1(β2i), and LMP7(β5i) subunits. Immunopresomes can also be induced by interferon-gamma.Therefore, in some embodiments, a population of cells can be exposed to one or more cytokines, growth factors, or other proteins. Cells can be stimulated with inflammatory cytokines such as interferon-gamma, IL-10, IL-6, and / or TNF-α. A population of cells can also be subjected to various environmental conditions, such as stress (heat stress, oxygen deficiency, glucose starvation, DNA damaging agents, etc.). In some embodiments, cells can be exposed to one or more of the following: chemotherapeutic agents, radiation, targeted therapy, or immunotherapy. Therefore, the effects of various genes or conditions on the processing and presentation of HLA peptides can be tested using the methods disclosed herein. In some embodiments, the conditions used are selected to suit the patient's condition in which the population of HLA peptides is being identified.
[0457] The single HLA-allele of this disclosure can be encoded and expressed using a virus-based system (e.g., adenovirus system, adeno-associated virus (AAV) vector, poxvirus, or lentivirus). Plasmids that can be used for delivery by adeno-associated virus, adenovirus, and lentivirus have been previously described (see, for example, U.S. Patent Nos. 6,955,808 and 6,943,019, and U.S. Patent Application No. 20080254008, incorporated herein by reference). Among the vectors that can be used in the implementation of this disclosure, integration into the host genome of cells is possible using retroviral gene transfer methods, often resulting in long-term expression of the inserted transgene. In exemplary embodiments, the retrovirus is a lentivirus. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues. The tropism of retrovirals can be modified by incorporating exogenous envelope proteins to increase the potential target population of target cells. Retroviruses can be engineered to enable conditional expression of inserted transgenes so that lentiviruses infect only certain cell types. Cell type-specific promoters can be used to target expression in specific cell types. Lentiviral vectors are retroviral vectors (and therefore both lentiviral and retroviral vectors can be used in the implementation of this disclosure). Furthermore, lentiviral vectors can transduce or infect non-dividing cells and generally produce high viral titers.
[0458] The selection of a retroviral gene transfer system may depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats (LTRs) capable of packaging exogenous sequences up to 6–10 kb. The smallest cis-acting LTR is sufficient for replication and vector packaging and is therefore used to incorporate the desired nucleic acid into target cells for permanent expression. Widely used retroviral vectors that can be used in the implementation of this disclosure include those based on mouse leukemia virus (MuLV), gibbon leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (e.g., Buchscher et al., (1992) J. Virol. 66: 2731-2739; Johann et al., (1992) J. Virol. 66: 1635-1640; Sommnerfelt et al. al., (1990) Virol. 176: 58-59;Wilson et al., (1998) J. Virol. 63: 2374-2378;Miller et al., (1991) J. Virol. 65: 2220-2224;PCT / See US94 / 05700). Also, the smallest non-primate lentiviral vectors, such as lentiviral vectors based on equine infectious anemia virus (EIAV), are useful for carrying out this disclosure (see, for example, Balagaan, (2006) J Gene Med; 8: 275-285, Published online 21 November 2005 in Wiley InterScience DOI: 10.1002 / jgm.845). The vector drives the expression of the target gene with cytomegalovirus. (CMV) promoter may be present. Therefore, this disclosure intends viral vectors, including retroviral vectors and lentiviral vectors, among vectors useful for carrying out this disclosure.
[0459] Any HLA allele can be expressed in a cell population. In exemplary embodiments, the HLA allele is an HLA class I allele. In some embodiments, the HLA class I allele is an HLA-A allele or an HLA-B allele. In some embodiments, the HLA allele is an HLA class II allele. The sequences of HLA class I and class II alleles can be found in the IPD-IMGT / HLA database. Examples of exemplary HLA alleles include, but are not limited to, HLA-A * 02:01, HLA-B * 14:02, HLA-A * 23:01, HLA-E * 01:01, HLA-DRB * 01:01, HLA-DRB * 01:02, HLA-DRB * 11:01, HLA-DRB * 15:01, and HLA-DRB * 07:01 is one example.
[0460] In some embodiments, HLA alleles are selected to correspond to the desired genotype. In some embodiments, the HLA alleles may be mutant HLA alleles, which may be alleles that do not naturally exist in affected individuals or alleles that do exist naturally. The methods disclosed herein have the additional advantage of identifying HLA-binding peptides for HLA alleles associated with various disorders, as well as for low-frequency alleles. Thus, in some embodiments, the methods provided herein can identify HLA alleles even if they exist at a frequency of less than 1% within a population, for example, in the Caucasian population.
[0461] In some embodiments, the nucleic acid sequence encoding the HLA allele further includes affinity acceptor tags that can be used to immunopurify the HLA protein. Suitable tags are well known in the art. In some embodiments, affinity acceptor tags include polyhistidine tags, polyhistidine-glycine tags, polyarginine tags, polyaspartic acid tags, polycysteine tags, polyphenylalanine tags, c-myc tags, herpes simplex virus glycoprotein D (gD) tags, FLAG tags, KT3 epitope tags, tubulin epitope tags, T7 gene 10 protein peptide tags, streptavidin tags, streptavidin-binding peptide (SPB) tags, Strep-tags, Strep-tag II, albumin-binding protein (ABP) tags, alkaline phosphatase (AP) tags, blue tongue virus tags (B-tags), calmodulin-binding peptide (CBP) tags, and chloramphenicol. Acetyltransferase (CAT) tag, choline-binding domain (CBD) tag, chitin-binding domain (CBD) tag, cellulose-binding domain (CBP) tag, dihydrofolate reductase (DHFR) tag, galactose-binding protein (GBP) tag, maltose-binding protein (MBP), glutathione-S-transferase (GST), Glu-Glu (EE) tag, human influenza hemagglutinin (HA) tag, horseradish peroxidase (HRP) tag, NE- tag, HSV tag, ketosteroid isomerase (KSI) tag, KT3 tag, LacZ tag, luciferase tag, NusA tag, PDZ domain tag, AviTag, calmodulin- tag, E- tag, S- tag, SBP- tag, Softag1, Softag3, TC tag, VSV- tag, Xpress tag, Isopeptag, SpyTag, SnoopTag, ProfinityExamples include eXact tags, protein C tags, S1-tags, S-tags, biotin-carboxyl carrier protein (BCCP) tags, green fluorescent protein (GFP) tags, small ubiquitin-like modifier (SUMO) tags, tandem affinity purification (TAP) tags, HaloTag, Nus-tags, thioredoxin-tags, Fc-tags, CYD tags, HPC tags, TrpE tags, ubiquitin tags, VSV-G epitope tags derived from Vescular Stomatis virus glycoproteins, or V5 tags derived from small epitopes (Pk) found in the P and V proteins of the paramyxovirus 5 (SV5). In some embodiments, affinity acceptor tags are “epitope tags,” which are a type of peptide tag that adds a recognizable epitope (antibody-binding site) to an HLA-protein, resulting in binding to the corresponding antibody, thereby enabling identification or affinity purification of the tagged protein. Non-limiting examples of epitope tags include protein A or protein G that bind to IgG. In some embodiments, affinity acceptor tags include biotin acceptor peptide (BAP) or human influenza hemagglutinin (HA) peptide sequences. Numerous other tag portions are known to those skilled in the art, can be conceived by those skilled in the art, and are intended herein. Any peptide tag can be used as long as it can be expressed as an element of an HLA-peptide complex tagged with an affinity acceptor.
[0462] The methods provided herein include isolating HLA-peptide complexes from cells transfected or transduced with an affinity pulldown of an HLA construct (Figure 3). In some embodiments, the complexes can be isolated using commercially available antibodies and standard immunoprecipitation techniques known in the art. The cells can first be lysed. HLA class I-peptide complexes can be isolated using HLA class I-specific antibodies such as W6 / 32 antibodies, while HLA class II-peptide complexes can be isolated using HLA class II-specific antibodies such as M5 / 114.15.2 monoclonal antibodies. In some embodiments, a single (or pair of) HLA alleles is expressed as a fusion protein with a peptide tag, and the HLA-peptide complexes are isolated using a binding molecule that recognizes the peptide tag.
[0463] The method further comprises isolating the peptide from the HLA-peptide complex and sequencing the peptide. The peptide is isolated from the complex by any method known to those skilled in the art, such as acid elution. Any sequencing method can be used, but in some embodiments, a method using mass spectrometry is employed, for example, liquid chromatography-mass spectrometry (LC-MS or LC-MS / MS, or alternatively HPLC-MS or HPLC-MS / MS). These sequencing methods are well known to those skilled in the art, and Medzihradszky KF and This is outlined in Chalkley RJ. Mass Spectrom Rev. 2015 Jan-Feb; 34 (1): 43-63.
[0464] In some embodiments, the cell population expresses one or more endogenous HLA alleles. In some embodiments, the cell population is a population of cells lacking one or more engineered endogenous HLA class I alleles. In some embodiments, the cell population is a population of cells lacking one or more engineered endogenous HLA class I alleles. In some embodiments, the cell population is a population of cells lacking one or more engineered endogenous HLA class II alleles. In some embodiments, the cell population is a population of cells lacking an engineered endogenous HLA class II allele or a population of cells lacking both an engineered endogenous HLA class I allele and an endogenous HLA class II allele. In some embodiments, the cell population includes cells enriched or sorted, for example, by fluorescence-activated cell sorting (FACS). In some embodiments, fluorescence-activated cell sorting (FACS) is used to sort the cell population. In some embodiments, the cell population is pre-sorted by FACS for cell surface expression of either HLA class I or class II, or both HLA class I and class II. For example, FACS can be used to sort a population of cells based on the cell surface expression of HLA class I alleles, HLA class II alleles, or combinations thereof.
[0465] Method for preparing personalized cancer vaccines If cancer-specific mutations are identified—that is, mutations present in the DNA of cancer cells of the same human subject but not in normal cells—and if the mutation leads to a change in one or more amino acids in the DNA-encoded protein, then these mutations can be targeted to the host immune response. The innate immune response can be directed to the mutated protein, thereby leading to the destruction of cancer cells expressing that protein. Due to the innate tolerance response and immunodeficiency environment in cancerous tissue, immunotherapy is a clinical pathway that attempts to enhance such immune responses and neutralize the body's tolerance and immunosuppressive effects. Therefore, proteins or peptides containing the above-mentioned mutations are suitable candidates for immunotherapy.
[0466] Mutant proteins are taken up by professional phagocytic cells that act as antigen-presenting cells (APCs), cleaved, and displayed on the cell surface as antigen-presenting complexes containing major histocompatibility complex (MHC) proteins as antigens for T cell activation. Human MHC proteins are also known as human leukocyte antigens, or HLA. MHC proteins can be MHC-class I or class II proteins, and some functional differences stem from the presentation of peptides by either class I or class II MHC proteins (HLA class I and HLA class II proteins). One notable difference is that HLA class I peptide complexes present antigens to cytotoxic CD8+ T cells, while HLA class II peptide complexes can also activate CD4+ T cells and lead to a sustained immune response. CD8+ T cells are essential for the cell-to-cell elimination of disease cells such as infected cells or tumor cells. CD4+ T cells, once activated, have a more sustained effect, most importantly in the generation of immunological memory. CD4 subsets are differentially recruited depending on the type of immunological threat, and multiple subsets with overlapping or different functions may be co-mobilized. This helps maintain equilibrium in the immunological response to pathogenic threats. In this regard, HLA class II peptide-mediated antigen presentation results in a sustained, regulated immune response. On the other hand, the binding of HLA class II to peptides can be promiscuous, and therefore, nonspecific peptide binding and presentation to the immune system can lead to ectopic immune responses such as autoimmunity.
[0467] In one embodiment, the disclosure provides a method for predicting peptides that can precisely pair or bind to a specific HLA class II alpha-beta heterodimer such that the peptide's high-fidelity binding to the HLA class II protein (composed of an alpha-beta heterodimer) ensures that the specific peptide is presented to T lymphocytes, thereby eliciting a specific immune response and avoiding any cross-reactivity or indiscriminate immunization.
[0468] In one embodiment, the disclosure provides a method for predicting peptides that can precisely bind to a specific HLA class II protein, such that when the peptide is therapeutically administered to a target expressing a specific homologous HLA class II protein, the HLA class II protein's ability to activate CD4+ T cells and stimulate immunological memory can activate a more sustained and robust immune response using the peptide. In some embodiments, a given peptide predicted to bind to an HLA class II protein with high specificity is a peptide containing a mutation, the mutation being prevalent in the target cancer or tumor cells; on the other hand, the same HLA class II protein predicted to bind to the mutant peptide either (a) does not bind to the corresponding unmutated wild-type peptide, or (b) binds separately with low affinity compared to its affinity for binding to the target mutant peptide. Preferential binding of HLA to mutant peptides is advantageous for the development of immunotherapies, as cells expressing the wild-type peptide are spared from immune attack by T cells reactive to the HLA-presented peptide. In some embodiments, the peptide predicted to bind specifically to an HLA class II protein is a peptide with post-translational modifications. Exemplary post-translational modifications include, but are not limited to, phosphorylation, ubiquitination, dephosphorylation, glycosylation, methylation, or acetylation. In some embodiments, the predicted peptide is subjected to post-translational modification before use in immunotherapy.
[0469] In some embodiments, the immunotherapies and strategies disclosed herein may also be applicable to suppressing undesirable immune activation, such as in autoimmune reactions. Specifically, peptides identified as potential conjugates for specific HLA subtypes can be modified to bind to specific HLA molecules and induce tolerance rather than trigger an immunogenic response.
[0470] In one embodiment, a method of immunotherapy tailored or individualized for a specific target is presented herein. Any target or patient expresses a specific array of HLA class I and HLA class II proteins. HLA typing is a well-known technique that allows for the determination of a specific repertoire of HLA proteins expressed by a target. Once the HLA heterodimers expressed by a particular target are known, the improved, refined, and reliable method described herein for predicting peptides that can bind with high fidelity to specific HLA class II alpha and beta heterodimers ensures that a specific immune response tailored specifically for the target can be produced.
[0471] The genes encoding HLA heterodimers are highly polymorphic, with more than 4,000 HLA class II allele variants identified across the human population. From maternal and paternal HLA haplotypes, individuals can inherit different alleles for each HLA class II locus, and each HLA class II heterodimer consists of an α-chain and a β-chain. Particularly for the HLA-DP and HLA-DQ alleles, the population of possible HLA heterodimers is highly complex due to the numerous combinations of α and β chain pairing. HLA class II heterodimers are converted to the endoplasmic reticulum (ER) and assembled with an invariant chain (Ii) derived from the protein CD74 to form a stable complex. Ii stabilizes the class II complex by enabling proper protein folding and allows for the export of HLA class II heterodimers into endosomal / lysosomal compartments. Within these HLA class II loading compartments, Ii is cleaved by cathepsin proteolysis to form placeholder peptides called CLIPs. Next, in a low pH environment, CLIP is replaced with a higher-affinity peptide by the non-classical HLA class II heterodimer chaperone HLA-DM. The HLA class II complex, now loaded with the higher-affinity peptide, then moves towards the trans-Golgi and finally towards the cell surface for display to CD4+ T cells.
[0472] Each HLA heterodimer is estimated to bind to thousands of peptides with allele-specific binding preference. In fact, each HLA allele is estimated to bind to approximately 1,000–10,000 unique peptides and present them on T cells. Given such diversity in HLA binding, accurately predicting whether a peptide may bind to a particular HLA allele is extremely difficult. Little is known about the allele-specific peptide-binding characteristics of HLA class II molecules due to heterogeneity in α- and β-chain pairing, the complexity of the data limiting the ability to confidently assign core-binding epitopes, and the lack of immunoprecipitation-grade allele-specific antibodies required for high-resolution biochemical analysis. Furthermore, analysis of peptide epitopes derived from a given HLA allele presents ambiguity when numerous HLA alleles are presented on the cell surface.
[0473] While the prediction of candidate novel antigens is primarily performed for HLA class I epitopes (due to the availability of experimental data on class I prediction algorithms compared to class II), CD4+ T cell responses are frequently observed in both preclinical and clinically personalized novel antigen vaccination trials. These findings demonstrate that HLA class II epitope processing and presentation may also play a crucial role in cancer treatment. Although HLA class II prediction algorithms exist, they are inaccurate due to the fact that the open-ended peptide-binding grooves of HLA class II heterodimers allow for the binding of longer peptides (generally 15-40 amino acids), thereby increasing the heterogeneity and complexity of epitope presentation. Therefore, further research is needed to better understand the HLA class II peptide-binding core and the cellular process characteristics involved in class II epitope processing and presentation. The proteomics field is currently limited by the complexity of HLA class II heterodimer formation and the availability of immunoprecipitation-grade antibodies for isolating HLA class II-peptide complexes. To overcome these challenges, we developed a single-allele HLA profiling workflow that relies on LC-MS / MS for characterizing allele-specific HLA class II-ligandomes for class II epitope prediction methods. The following definitions are supplementary to those skilled in the art and apply to this application, not relating to any related or unrelated patents or applications, for example, any jointly owned patents or applications. Any methods and materials similar to or equivalent to those described herein may be used in carrying out the disclosure to test the present disclosure, but exemplary materials and methods are described herein. Accordingly, the terminology used herein is for illustrative purposes only and not limited to describing specific embodiments.
[0474] A method for preparing a personalized cancer vaccine is disclosed herein. The method for preparing a personalized cancer vaccine may include the steps of: identifying a peptide sequence using a mutation expressed in a target cancer cell; inputting amino acid position information of the identified peptide sequence into a machine learning HLA-peptide presentation prediction model using a computer processor to generate a set of presentation predictions for the identified peptide sequence, wherein each presentation prediction represents the probability that one or more proteins encoded by a class II MHC allele of the target cancer cell will present a given sequence from among the identified peptide sequences; and selecting a subset of the identified peptide sequences based on the set of presentation predictions in order to prepare a personalized cancer vaccine.
[0475] In some embodiments, one or more results obtained by the methods described herein result in a quantitative value or a value indicating one or more of the following: the likelihood of diagnostic accuracy, the likelihood of the presence of a condition in the subject, the likelihood of a condition occurring in the subject, the likelihood of success of a particular treatment, or any combination thereof. In some embodiments, the methods described herein can predict the risk or likelihood of a condition occurring. In some embodiments, the methods described herein can be an early diagnostic indicator for the occurrence of a condition. In some embodiments, the methods described herein can diagnose or confirm the presence of a condition. In some embodiments, the methods described herein can monitor the exacerbation of a condition. In some embodiments, the methods described herein can monitor the effectiveness of a treatment for a condition in the subject.
[0476] Method for identifying MHC-II peptides In one embodiment, a method for identifying one or more peptides presented by MHC-II proteins for immune activation is presented herein. In some embodiments, the one or more peptides include an epitope. In some embodiments, the method involves computer prediction of the likelihood that a specific epitope will be presented by an MHC-II protein. In some embodiments, the method involves computer prediction of the specificity of the epitope to MHC-II presentation. In some embodiments, the computer prediction method involves evaluation of peptide-MHC interactions. In some embodiments, the computer prediction method involves prediction of the allele specificity of the peptide to antigen presentation. In some embodiments, the computer prediction method involves the incorporation of bioinformatics information, such as functional titers including nucleotide sequences, structural motifs of biomolecules, protein-protein interaction features, and immunogenicity. In some embodiments, the computer prediction method involves machine learning. Numerous immunoinformatics methods have been developed to predict peptide-MHC interactions for both MHC class I and class II, based on machine learning techniques such as simple pattern motifs, support vector machines (SVMs), hidden Markov models (HMMs), neural network (NN) models, quantitative structure-activity relationship (QSAR) analysis, structure-based methods, and biophysical methods. These methods can be divided into two categories: intra-allele (allele-specific) methods and trans-allele (general-specific) methods. Intra-allele methods train a model for a specific MHC molecule on a limited set of experimental peptide-binding data and apply the peptide-binding prediction to that molecule. Because MHC molecules are highly polymorphic and thousands of allele variants exist, coupled with a lack of sufficient experimental binding data, it is impossible to build a predictive model for each allele.Therefore, using peptide-binding data that extends across many alleles or species, we can develop trans-allele and general-purpose methods, such as NetMHCIIpan(Karosiene E et al., NetMHCIIpan-3.0, a common pan-specific MHC class II prediction method including all three human MHC class II isotypes, HLA-DR, HLA-DP). and HLADQ. Immunogenetics (2013) 65(10):711-24), and TEPITOPEpan (Zhang L, et al., TEPITOPEpan: extending TEPITOPE for peptide binding prediction covering over 700 HLA-DR molecules. PLoS One (2012) 7(2):e30483) have been developed. MHC- A similar method can be used for I.
[0477] In some embodiments, the peptide sequence may not be expressed in the normal cells of the subject. In some embodiments, all cells of the subject may not be cancer cells. Cancer cells include, but are not limited to, thyroid cancer, adrenocortical cancer, anal cancer, aplastic anemia, cholangiocarcinoma, bladder cancer, bone cancer, bone metastases, central nervous system (CNS) cancer, peripheral nervous system (PNS) cancer, breast cancer, Castleman disease, cervical cancer, childhood non-Hodgkin lymphoma, lymphoma, colon cancer / rectal cancer, endometrial cancer, esophagus cancer, Ewing family tumors (e.g., Ewing sarcoma), eye cancer, gallbladder cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors, gestational trophoblastic disease, hairy cell leukemia, Hodgkin's disease, Kaposi's sarcoma, and kidney cancer. It may be produced by different cancers, including cancer, laryngeal and hypopharyngeal cancer, acute lymphoblastic leukemia, acute myeloid leukemia, childhood leukemia, chronic lymphoblastic leukemia, chronic myeloid leukemia, liver cancer, lung cancer, pulmonary carcinoid tumors, non-Hodgkin lymphoma, male breast cancer, malignant mesothelioma, multiple myeloma, myelodysplastic syndrome, myeloproliferative disorders, nasal cavity and paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oral and oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penile cancer, pituitary tumors, prostate cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma (adult soft tissue cancer), melanoma skin cancer, non-melanoma skin cancer, gastric cancer, testicular cancer, thymic cancer, uterine cancer (e.g., uterine sarcoma), vaginal cancer, vulvar cancer, or Waldenström macroglobulinemia.
[0478] The identification step may include comparing DNA, RNA, or protein sequences derived from target cancer cells with DNA, RNA, or protein sequences derived from target normal cells. The DNA, RNA, or protein sequences derived from target cancer cells may differ from those derived from target normal cells. The identification step can identify highly sensitive nucleic acid variants.
[0479] A machine learning HLA-peptide presentation prediction model may include at least several predictor variables identified based on training data. The training data may include sequence information of peptide sequences presented by HLA proteins expressed in cells and identified by mass spectrometry; training peptide sequence information including amino acid position information, associated with HLA proteins expressed in cells; and a function representing the relationship between the amino acid position information received as input and the presentation probability generated as output based on the amino acid position information and predictor variables.
[0480] In some embodiments, training data may further include structured data, time-series data, unstructured data, and relational data. Unstructured data may include audio data, image data, video, mechanical data, electrical data, chemical data, and any combination thereof for use in accurate simulation or training of robotics or simulations. Time-series data may include data from one or more of the following: smart meters, smart appliances, smart devices, monitoring systems, telemetry devices, or sensors. Relational data may include data from customer systems, enterprise systems, operational systems, websites, web-accessible application programming interfaces (APIs), or any combination thereof. This can be done by any method by which a user inputs files or other data formats into the software or system.
[0481] In some embodiments, training data can be stored in a database. The database can be stored in a computer-readable format. A computer processor can be configured to access the data stored in computer-readable memory. In some embodiments, the data can be analyzed using a computer system to obtain results. The results can be stored remotely or internally on a storage medium and communicated with personnel such as drug application specialists. In some embodiments, the computer system can be operably coupled with components for transmitting results. These transmission components may include wired and wireless components. Examples of wired transmission components include Universal Serial Bus (USB) connections, coaxial cable connections, Ethernet® cables such as Cat5 or Cat6 cables, fiber optic cables, or telephone lines. Examples of wireless transmission components include Wi-Fi receivers, components for accessing mobile data standards such as 3G or 4G LTE data signals, or Bluetooth® receivers. In some embodiments, all of this data in the storage medium is collected and archived to create a data warehouse.
[0482] In some embodiments, the database includes an external database. The external database may be a medical database, for example, an Adverse Drug database, but is not limited to this. Effects Database, AHFS Supplemental File, Allergen Picklist File, Average WAC Pricing File, Brand Probability File, Canadian Drug File v2、Comprehensive Price History、Controlled Substances File、Drug Allergy Cross-Reference File、Drug Application File、Drug Dosing & Administration Database、Drug Image Database v2.0 / Drug Imprint Database v2.0、Drug Inactive Date File、Drug Indications Database、Drug Lab Conflict Database、Drug Therapy Monitoring System(DTMS)v2.2 / DTMS Consumer Monographs、Duplicate Therapy Database、Federal Government Pricing File、Healthcare Common Procedure Coding System Codes(HCPCS)Database、ICD-10 Mapping Files、Immunization Cross-Reference File、Integrated A to Z Drug Facts Module、Integrated Patient Education、Master Parameters Database、Medi-Span Electronic Drug File(MED-File)v2、Medicaid Rebate File,Medicare Plans File、Medical Condition Picklist File,Medical Conditions Master Database, Medication Order Management Database(MOMD), Parameters to Monitor Database, Patient Safety Programs File, Payment Allowance Limit-Part B(PAL-B)v2.0, Precautions Database, RxNorm Cross-Reference File, Standard Drug Identifiers Database,Substitution Groups File, Supplemental Names File, Uniform System of Classification Cross-Reference File, or Warning Label Can be a Database.
[0483] In some embodiments, training data may also be obtained through other data sources. These data source...
Claims
[Claim 1] The invention described in the specification.