Neoantigen identification, manufacture, and use
Patent Information
- Application Number
- JP2025014472
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2016-11-23
- Filing Date
- 2025-01-31
- Publication Date
- 2025-11-10
AI Technical Summary
The prior art presents low positive predictive values (PPV) when predicting and selecting neoplasmic antigens, resulting in a low success rate of nascent antigen vaccine and failing to fully consider a variety of antigen sources, including the effects of mutations in splice factors and changes in protease cleavage sites.
Optimized tumor full-epipedic and transcriptomic analytical methods were used, combined with deep learning models, including statistical regression and peptide-allele map mapping, to identify and select antigens with high PPV. These models are able to deal with the independence of multiple MHC alleles and take into account the effects of the antigen and the protease cleavage site.
It improves the accuracy and success rate of antigen prediction, ensures that more patients can receive antigens that can trigger antitumor immune responses, and enhances the effectiveness of the nascent antigen vaccine.
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 62 / 268,333, filed December 16, 2015, U.S. Provisional Patent Application No. 62 / 317,823, filed April 4, 2016, U.S. Provisional Patent Application No. 62 / 379,986, filed August 26, 2016, U.S. Provisional Patent Application No. 62 / 394,074, filed September 13, 2016, and U.S. Provisional Patent Application No. 62 / 425,995, filed November 23, 2016, each of which is incorporated by reference in its entirety for all purposes. [Background technology]
[0002] background Therapeutic vaccines based on tumor-specific neoantigens hold great promise as the next generation of personalized cancer immunotherapy 1~3 Cancers with high mutational burden, such as non-small cell lung cancer (NSCLC) and melanoma, are attractive targets for such therapies, especially given their relatively large potential for neoantigen generation. 4、5 Early evidence suggests that neoantigen-based vaccination can elicit T cell responses. 6 and that neoantigen-targeted cell therapy can induce tumor regression under certain circumstances in selected patients. 7 is shown.
[0003] One question for neoantigen vaccine design is which of the many coding mutations present in the target tumor can generate the "best" therapeutic neoantigen, e.g., an antigen that can elicit anti-tumor immunity and induce tumor regression.
[0004] Initial methods are proposed to incorporate mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC-binding affinity of candidate neoantigen peptides. 8However, these proposed methods may not be able to model the entire epitope generation process, which involves many steps in addition to gene expression and MHC binding (e.g., TAP trafficking, proteasomal cleavage, and / or TCR recognition). 9 Therefore, existing methods likely suffer from a reduced positive predictive value (PPV) (Figure 1A).
[0005] Indeed, analyses of peptides presented by tumor cells performed by several groups have shown that less than 5% of peptides predicted to be presented using gene expression and MHC binding affinity can be found on tumor surface MHC. 10、11 (Figure 1B). This poor correlation between binding prediction and MHC presentation is further strengthened by the recent observation of a lack of improvement in the predictive accuracy of binding-restricted neoantigens for checkpoint inhibitor response beyond mutation count alone. 12 .
[0006] This low positive predictive value (PPV) of existing methods for predicting presentation presents a problem for neoantigen-based vaccine design. If a vaccine is designed with a prediction with a low PPV, the majority of patients are unlikely to receive therapeutic neoantigens (even assuming that all presented peptides are immunogenic), and even fewer patients are likely to receive more than one. Thus, neoantigen vaccination with current methods is unlikely to be successful in a substantial number of subjects with tumors (Figure 1C).
[0007] In addition, previous approaches have generated candidate neoantigens using only cis-acting mutations, which occur in multiple tumor types and result in aberrant splicing of many genes. 13 , and have largely neglected to consider additional sources of nascent ORFs, including mutations in splicing factors and mutations that create or remove protease cleavage sites.
[0008] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard tumor analysis approaches may unintentionally promote sequence artifacts or germline polymorphisms as neoantigens that result in inefficient use of vaccine potential or autoimmune risk, respectively. Summary of the Invention
[0009] overview Optimized approaches for identifying and selecting neoantigens for personalized cancer vaccines are disclosed herein. First, we address optimized tumor exome and transcriptome analysis approaches for identifying neoantigen candidates using next generation sequencing (NGS). These methods build on standard approaches to NGS tumor analysis to ensure that we advance neoantigen candidates with the highest sensitivity and specificity across all categories of genomic alterations. Second, we present novel approaches for high PPV neoantigen selection to overcome specificity issues and ensure that neoantigens advanced for inclusion in vaccines are more likely to elicit anti-tumor immunity. These approaches include, depending on the embodiment, trained statistical regression or peptide-allele mapping and nonlinear deep learning models that jointly model per-allele motifs for multiple peptide lengths that share statistical strength across peptides of various lengths. Nonlinear deep learning models can be specifically designed and trained to treat different MHC alleles in the same cell as independent, thereby addressing the problem with linear models where they would have interfered with each other. Finally, additional considerations for personalized vaccine design and manufacturing based on neoantigens are addressed. [The present invention 1001] A method for identifying one or more neoantigens derived from tumor cells of a subject that are likely to be presented on the tumor cell surface, comprising the steps of: obtaining at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from tumor cells of a subject, the tumor nucleotide sequencing data being used to obtain data representing a peptide sequence for each of a set of neoantigens, and wherein the peptide sequence of each neoantigen comprises at least one alteration that makes it different from a corresponding wild-type parent peptide sequence; inputting the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each of the neoantigens will be presented by one or more MHC alleles on the tumor cell surface of tumor cells of the subject, the set of numerical likelihoods having been determined based at least on the received mass spectrometry data; and Selecting a subset of the set of neoantigens based on said set of numerical likelihoods to generate a set of selected neoantigens. [The present invention 1002] The method of claim 10, wherein the number of sets of selected neoantigens is 20. [The present invention 1003] The presented model is the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in the peptide sequence; the likelihood of presentation on the tumor cell surface of such a peptide sequence containing a particular amino acid at a particular position by a particular one of the MHC alleles of the pair; Any of the methods 1001 to 1002 of the present invention, which expresses a dependency between. [The present invention 1004] The step of inputting a peptide sequence comprises: applying the one or more presentation models to the peptide sequences of the corresponding neoantigens to generate a dependency score for each of the one or more MHC alleles that indicates whether the MHC allele presents the corresponding neoantigen based on at least the position of an amino acid in the peptide sequence of the corresponding neoantigen. Any of the methods of 1001 to 1003 of the present invention. [The present invention 1005] transforming the dependency scores to generate a corresponding per-allele likelihood for each MHC allele, the likelihood being that the corresponding MHC allele will present the corresponding neoantigen; and Combining the allele likelihoods to generate a numerical likelihood The method of the present invention 1004 further comprises: [The present invention 1006] The method of claim 1005, wherein the step of transforming the dependency scores models the presentation of peptide sequences of corresponding neoantigens as mutually exclusive. [The present invention 1007] Transforming the combination of dependency scores to generate a numerical likelihood. The method according to any one of 1004 to 1006 of the present invention, further comprising: [The present invention 1008] 1007. The method of claim 1007, wherein the step of transforming the combination of dependency scores models the presentation of peptide sequences of corresponding neoantigens as interference between MHC alleles. [The present invention 1009] The set of numerical likelihoods is further specified by at least an allele non-interaction property, and applying an allele-non-interacting model of the one or more presentation models to the allele-non-interacting feature to generate a dependency score for the allele-non-interacting feature that indicates whether a peptide sequence of a corresponding neoantigen is presented based on the allele-non-interacting feature. The method of any one of 1004 to 1008 of the present invention, further comprising: [The present invention 1010] combining the dependency score for each MHC allele in the one or more MHC alleles with the dependency score for the allele non-interacting property; transforming the combined dependency scores for each MHC allele to generate a corresponding per-allele likelihood for the MHC allele, the likelihood being that the corresponding MHC allele will present the corresponding neoantigen; and Combining the allele likelihoods to generate a numerical likelihood The method of the present invention 1009 further comprises: [The present invention 1011] Transforming the combination of the dependency scores for each of the MHC alleles and the dependency scores for the allele-non-interacting feature to generate a numerical likelihood. Any of the methods 1009 to 1010 of the present invention further comprising: [The present invention 1012] Any of the methods of claims 1001 to 1011, wherein the set of numerical parameters for the representation model is trained based on a training dataset comprising at least a set of training peptide sequences identified as present in a multiplicity of samples, and one or more MHC alleles associated with each training peptide sequence, and the training peptide sequences are identified through mass spectrometry on isolated peptides eluted from MHC alleles derived from the multiplicity of samples. [The present invention 1013] The method of claim 1012, wherein the training dataset further comprises data on mRNA expression levels in tumor cells. [The present invention 1014] The method of any of claims 1012 to 1013, wherein said sample comprises a cell line engineered to express a single MHC class I or class II allele. [The present invention 1015] The method of any of claims 1012 to 1014, wherein said sample comprises a cell line engineered to express multiple MHC class I or class II alleles. [The present invention 1016] 16. The method of any of claims 1012 to 1015, wherein said samples are obtained from multiple patients or comprise human cell lines derived from said patients. [The present invention 1017] The method of any of claims 1012 to 1016, wherein said samples comprise fresh or frozen tumor samples obtained from multiple patients. [The present invention 1018] The method of any one of claims 1012 to 1017, wherein said samples comprise fresh or frozen tissue samples obtained from multiple patients. [The present invention 1019] The method of any of claims 1012 to 1018, wherein said sample comprises a peptide identified using a T cell assay. [The present invention 1020] The training data set is the peptide abundances of the set of training peptides present in the sample; Peptide lengths of the set of training peptides in the sample Any of the methods of 1012 to 1019, further comprising data relating to the above. [The present invention 1021] Any of the methods of claims 1012 to 1020, wherein the training dataset is generated by comparing a set of training peptide sequences via alignment to a database containing a set of known protein sequences, and the set of training protein sequences is longer than and includes the training peptide sequences. [The present invention 1022] Any of the methods of claims 1012 to 1021, wherein the training dataset is generated based on performing mass spectrometry on the cell line or having mass spectrometry completed to obtain at least one of exome, transcriptome, or whole genome peptide sequencing data from the cell line, the peptide sequencing data comprising at least one protein sequence that includes the alteration. [The present invention 1023] Any of the methods of claims 1012 to 1022, wherein the training dataset is generated based on obtaining at least one of normal nucleotide sequencing data of the exome, transcriptome, and whole genome from a normal tissue sample. [The present invention 1024] The method of any of claims 1012 to 1023, wherein the training dataset further comprises data relating to proteomic sequences associated with said sample. [The present invention 1025] The method of any of claims 1012 to 1024, wherein the training data set further comprises data relating to an MHC peptidome sequence associated with said sample. [The present invention 1026] 1026. The method of any of claims 1012 to 1025, wherein the training dataset further comprises data relating to peptide-MHC binding affinity measurements for at least one of the isolated peptides. [The present invention 1027] The method of any of claims 1012 to 1026, wherein the training dataset further comprises data relating to peptide-MHC binding stability measurements for at least one of the isolated peptides. [The present invention 1028] The method of any one of claims 1012 to 1027, wherein the training dataset further comprises data relating to a transcriptome associated with said sample. [The present invention 1029] The method of any of claims 1012 to 1028, wherein the training dataset further comprises data relating to a genome associated with said sample. [The present invention 1030] 1029. The method of any of claims 1012 to 1029, wherein the training peptide sequences are in the k-mer range, where k is between 8 and 15 in length. [The present invention 1031] The method of any of claims 1012 to 1030, further comprising encoding the peptide sequence using a one-hot encoding scheme. [The present invention 1032] The method of claim 1031, further comprising encoding the training peptide sequences using a left-padded one-hot encoding scheme. [The present invention 1033] The present invention includes carrying out any one of steps 1001 to 1032, and obtaining a tumor vaccine comprising the set of selected neoantigens; and Administering the tumor vaccine to the subject The method of treating a subject having a tumor further comprises: [The present invention 1034] The present invention includes carrying out any one of steps 1001 to 1033, and Producing or having produced a tumor vaccine comprising a set of selected neoantigens. The method of producing a tumor vaccine further comprises: [The present invention 1035] A set of neoantigens selected according to any one of the methods of the present inventions 1001 to 1032. A tumor vaccine comprising: [The present invention 1036] The vaccine of the present invention 1035, wherein said tumor vaccine comprises one or more of a nucleotide sequence, a polypeptide sequence, RNA, DNA, a cell, a plasmid, or a vector. [The present invention 1037] The vaccine of any one of claims 1035 to 1036, wherein the tumor vaccine comprises one or more neoantigens presented on the surface of a tumor cell. [The present invention 1038] The vaccine of any of claims 1035 to 1037, wherein the tumor vaccine comprises one or more neoantigens that are immunogenic in a subject. [The present invention 1039] The vaccine of any of claims 1035 to 1038, wherein said tumor vaccine does not comprise one or more neoantigens that induce an autoimmune response against normal tissue in a subject. [The present invention 1040] The vaccine of any one of claims 1035 to 1039, wherein the tumor vaccine further comprises an adjuvant. [The present invention 1041] The vaccine of any one of claims 1035 to 1040, wherein the tumor vaccine further comprises an excipient. [The present invention 1042] Selection of the set of selected neoantigens Selecting neoantigens that have an increased likelihood of being presented on the tumor cell surface compared to neoantigens not selected based on a presentation model. Any of the methods of 1001 to 1041 of the present invention. [The present invention 1043] Selection of the set of selected neoantigens Selecting neoantigens that have an increased likelihood of inducing a tumor-specific immune response in a subject compared to neoantigens that are not selected based on a presentation model. Any of the methods of the present invention 1001 to 1042, comprising: [The present invention 1044] Selection of the set of selected neoantigens Selecting neoantigens that have an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) compared to neoantigens that are not selected based on a presentation model and optionally, the APC is a dendritic cell (DC). [The present invention 1045] Selection of the set of selected neoantigens Selecting neoantigens that have a reduced likelihood of being subject to inhibition via central or peripheral tolerance compared to neoantigens not selected based on a presentation model Any of the methods of 1001 to 1044 of the present invention. [The present invention 1046] Selection of the set of selected neoantigens Selecting neoantigens that have a reduced likelihood of being able to induce an autoimmune response against normal tissue in a subject compared to neoantigens not selected based on a presentation model. Any of the methods of 1001 to 1045 of the present invention. [The present invention 1047] Any of the methods of claims 1001 to 1046, wherein the exome or transcriptome nucleotide sequencing data is obtained by performing sequencing on tumor tissue. [The present invention 1048] Any of the methods of claims 1001 to 1047, wherein the sequencing is next generation sequencing (NGS) or any massively parallel sequencing approach. [The present invention 1049] The method of any of claims 1001 to 1048, wherein the set of numerical likelihoods is further specified by at least an MHC allele interaction characteristic comprising at least one of the following: a. The predicted affinity of the MHC allele to the neoantigen-encoded peptide; b. predicted stability of neoantigen-encoded peptide-MHC complexes; c. The sequence and length of the neoantigen-encoded peptide; d. The probability of presentation of neoantigen-encoded peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means; e. The expression level of a particular MHC allele in the subject in question (e.g., as measured by RNA-seq or mass spectrometry); f. the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other distinct subjects expressing that particular MHC allele; g. The overall neoantigen-encoded peptide sequence-independent probability of presentation by MHC alleles in the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other separate subjects. [The present invention 1050] The method of any of claims 1001-1049, wherein the set of numerical likelihoods is further specified by at least an MHC allele non-interacting property comprising at least one of the following: a. C-terminal and N-terminal sequences adjacent to the neoantigen-encoded peptide within the source protein sequence; b. The presence of protease cleavage motifs in neoantigen-encoded peptides, optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry); c. The turnover rate of the source protein as measured in the appropriate cell type; d. The length of the source protein, optionally taking into account the specific splice variants ("isoforms") most highly expressed in tumor cells, as measured by RNA-seq or proteomic mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data; e. The level of expression of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which may be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry); f. Expression of the source gene of the neoantigen-encoding peptide (e.g., as measured by RNA-seq or mass spectrometry); g. Typical tissue-specific expression of the source genes of neoantigen-encoding peptides during various stages of the cell cycle; h. A comprehensive catalog of the properties of the source protein and / or its domains, such as can be found in, for example, uniProt or the PDB http: / / www.rcsb.org / pdb / home / home.do; i. Features that describe the nature of the domain of the source protein that contains the peptide, such as secondary or tertiary structure (e.g., alpha helices versus beta sheets); alternative splicing, j. the probability of presentation of the source protein-derived peptide of the neoantigen-encoded peptide in question in other separate subjects; k. the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias; l. Expression of different gene modules / pathways as measured by RNASeq that informs about the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs) (not necessarily including the source protein of the peptides); m. the copy number of the gene that provides the neoantigen-encoding peptide in the tumor cell; n. the probability that the peptide will bind to TAP or the measured or predicted binding affinity of the peptide to TAP; o. Expression levels of TAP in tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, immunohistochemistry); p. Presence or absence of tumor mutations, including but not limited to: Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3; ii. in genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome); peptides whose presentation relies on components of the antigen presentation machinery that are affected by loss-of-function mutations in the tumor have a reduced probability of presentation. q. Presence or absence of functional germline polymorphisms, including but not limited to: i. in genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome); r. Tumor type (e.g., NSCLC, melanoma), s. Clinical tumor subtype (e.g., squamous cell lung cancer vs. non-squamous), t. Smoking history, u. Typical expression of peptide source genes in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. [The present invention 1051] Any of the methods of 1001 to 1050, wherein at least one mutation is a frameshift or non-frameshift insertion deletion (indel), a missense or nonsense substitution, a splice site alteration, a genomic rearrangement or gene fusion, or any genomic or expression alteration that results in a de novo ORF. [The present invention 1052] Any of the methods of claims 1001 to 1051, wherein the tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer. [The present invention 1053] The method of any of claims 1001 to 1052, further comprising the step of obtaining a tumor vaccine comprising the selected set of neoantigens or a subset thereof, and optionally further comprising the step of administering the tumor vaccine to a subject. [The present invention 1054] The method of any one of claims 1001 to 1053, wherein at least one of the neoantigens in the set of selected neoantigens, when in polypeptide form, comprises at least one of the following: Binding affinity to MHC with IC50 value less than 1000 nM; for MHC class 1 polypeptides, 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids in length; the presence of a sequence motif within or near the polypeptide in the parent protein sequence that promotes proteasomal cleavage; and Presence of sequence motifs that facilitate TAP transport. [The present invention 1055] 1. A method for generating a model for identifying one or more neoantigens likely to be presented on a tumor cell surface of a tumor cell, comprising performing the steps of: receiving mass spectrometry data, the data relating to a plurality of isolated peptides eluted from major histocompatibility complex (MHC) peptides from a plurality of samples; obtaining a training data set by at least identifying a set of training peptide sequences present in the sample and one or more MHCs associated with each training peptide sequence; Training a set of numerical parameters of a presentation model using a training dataset comprising training peptide sequences, the presentation model providing a number of numerical likelihoods that a peptide sequence derived from a tumor cell will be presented by one or more MHC alleles on the surface of the tumor cell. [The present invention 1056] The presented model is the presence of a particular amino acid at a particular position in the peptide sequence; The likelihood of a peptide sequence containing a particular amino acid at a particular position being presented by one of the MHC alleles on a tumor cell. The method of the present invention 1055 represents the dependency between. [The present invention 1057] The method of any of claims 1055 to 1056, wherein said sample comprises a cell line engineered to express a single MHC class I or class II allele. [The present invention 1058] The method of any of claims 1055 to 1057, wherein said sample comprises a cell line engineered to express multiple MHC class I or class II alleles. [The present invention 1059] The method of any of claims 1055 to 1058, wherein said samples are obtained from multiple patients or comprise human cell lines derived from said patients. [The present invention 1060] The method of any of claims 1055 to 1059, wherein said samples comprise fresh or frozen tumor samples obtained from multiple patients. [The present invention 1061] The method of any of claims 1055 to 1060, wherein said sample comprises a peptide identified using a T cell assay. [The present invention 1062] The training data set is the peptide abundances of the set of training peptides present in the sample; Peptide lengths of the set of training peptides in the sample Any of the methods of claims 1055 to 1061, further comprising data relating to the method. [The present invention 1063] The step of obtaining a training data set comprises: Obtaining a set of training protein sequences based on the training peptide sequences, the set of training protein sequences being longer than and including the training peptide sequences, by comparing the set of training peptide sequences through alignment against a database comprising a set of known protein sequences. Any of the methods of the present invention 1055 to 1062, comprising: [The present invention 1064] The step of obtaining a training data set comprises: Mass spectrometry has been performed or completed on the cell line to obtain at least one of the following nucleotide sequencing data: exome, transcriptome, or whole genome from the cell line The method of any one of claims 1055 to 1063, wherein the nucleotide sequencing data comprises at least one protein sequence comprising a mutation. [The present invention 1065] Training a set of parameters of the representation model, Encode the training peptide sequences using a one-hot encoding scheme Any of the methods of claims 1055 to 1064 of the present invention. [The present invention 1066] Obtaining at least one of normal exome, transcriptome, and whole genome nucleotide sequencing data from a normal tissue sample; and training a set of parameters of the proposed model using the normal nucleotide sequencing data; Any of the methods of claims 1055 to 1065, further comprising: [The present invention 1067] The method of any one of claims 1055 to 1066, wherein the training data set further comprises data relating to proteomic sequences associated with said sample. [The present invention 1068] 1068. The method of any of claims 1055 to 1067, wherein the training data set further comprises data relating to an MHC peptidome sequence associated with said sample. [The present invention 1069] The method of any of claims 1055 to 1068, wherein the training dataset further comprises data relating to peptide-MHC binding affinity measurements for at least one of the isolated peptides. [The present invention 1070] 1069. The method of any of claims 1055 to 1069, wherein the training dataset further comprises data relating to peptide-MHC binding stability measurements for at least one of the isolated peptides. [The present invention 1071] The method of any one of claims 1055 to 1070, wherein the training dataset further comprises data relating to a transcriptome associated with said sample. [The present invention 1072] 1072. The method of any one of claims 1055 to 1071, wherein the training dataset further comprises data relating to a genome associated with said sample. [The present invention 1073] Training the set of numerical parameters comprises: To perform a logistic regression on a set of parameters Any of the methods of claims 1055 to 1072, further comprising: [The present invention 1074] The method of any of claims 1055 to 1073, wherein the training peptide sequences are in the k-mer range, where k is between 8 and 15 in length. [The present invention 1075] Training a set of numerical parameters of the representation model, Encode the training peptide sequences using a left-padded one-hot encoding scheme Any of the methods of claims 1055 to 1074. [The present invention 1076] Training the set of numerical parameters comprises: Determining values for a set of parameters using a deep learning algorithm Any of the methods of claims 1055 to 1075, further comprising: [The present invention 1077] 1. A method for generating a model for identifying one or more neoantigens likely to be presented on a tumor cell surface of a tumor cell, comprising performing the steps of: receiving mass spectrometry data, the data relating to a number of isolated peptides eluted from major histocompatibility complex (MHC) peptides derived from a number of fresh or frozen tumor samples; obtaining a training dataset by at least identifying a set of training peptide sequences that are present in the tumor sample and that are presented on one or more MHC alleles associated with each training peptide sequence; Obtaining a set of training protein sequences based on the training peptide sequences; and Training a set of numerical parameters of a presentation model using training protein sequences and training peptide sequences, the presentation model providing a number of numerical likelihoods that a peptide sequence derived from a tumor cell will be presented by one or more MHC alleles on the surface of the tumor cell. [The present invention 1078] The presented model is the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in the peptide sequence; the likelihood of presentation on the tumor cell surface of such a peptide sequence containing a particular amino acid at a particular position by a particular one of the MHC alleles of the pair; The method of the present invention 1077 represents the dependency between. [Brief description of the drawings]
[0010] These and other features, aspects, and advantages of the present invention will become better understood with regard to the following description and accompanying drawings.
[0011] [Figure 1A] Current clinical approaches to neoantigen identification are presented. [Figure 1B] It is shown that less than 5% of the predicted binding peptides are displayed on tumor cells. [Figure 1C] 1 illustrates the impact of the specificity problem on neoantigen prediction. [Figure 1D] We show that binding prediction is not sufficient for neoantigen identification. [Figure 1E] Probability of MHC-I presentation as a function of peptide length. [Figure 1F] 1 shows exemplary peptide spectra generated from Promega dynamic range standards. [Figure 1G] We show how adding features increases the positive predictive value of the model. [Figure 2A] 1 is a schematic of an environment for identifying the likelihood of peptide presentation in a patient according to one embodiment. [Figure 2B]According to one aspect, a method for obtaining presentation information is described. [Figure 2C] According to one aspect, a method for obtaining presentation information is described. [Diagram 3] FIG. 1 is a high-level block diagram illustrating computer logic components of a presentation specific system, according to one embodiment. [Figure 4] 1 illustrates an exemplary set of training data, according to one embodiment. [Diagram 5] 1 illustrates an exemplary network model relating to MHC alleles. [Figure 6A] 1 illustrates an exemplary network model shared by MHC alleles. [Figure 6B] 1 illustrates an exemplary network model shared by MHC alleles. [Figure 7] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 8] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 9] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 10] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 11] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 12] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 13A] Performance results for various exemplary proposed models are described. [Figure 13B] Performance results for various exemplary proposed models are described. [Figure 13C] Performance results for various exemplary proposed models are described. [Figure 13D]Performance results for various exemplary proposed models are described. [Figure 13E] Performance results for various exemplary proposed models are described. [Figure 13F] Performance results for various exemplary proposed models are described. [Figure 13G] Performance results for various exemplary proposed models are described. [Figure 13H] Performance results for various exemplary proposed models are described. [Figure 13I] Performance results for various exemplary proposed models are described. [Figure 13J] Performance results for various exemplary proposed models are described. [Figure 14] An exemplary computer for implementing the entities illustrated in FIGS. 1 and 3 is described. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Detailed Description I. Definition In general, the terms used in the claims and the specification are intended to be interpreted as having the plain meaning understood by a person skilled in the art. Certain terms are defined below to provide additional clarity. In the event of a conflict between the plain meaning and the definition provided, the definition provided shall prevail.
[0013] As used herein, the term "antigen" is a substance that induces an immune response.
[0014] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes it different from the corresponding wild-type parent antigen, for example, through a mutation in tumor cells or a tumor cell-specific post-translational modification. A neoantigen can include a polypeptide sequence or a nucleotide sequence. A mutation can include a frameshift or non-frameshift insertion deletion (indel), a missense or nonsense substitution, a splice site change, a genomic rearrangement or gene fusion, or any genomic or expression change that results in a neo-ORF. A mutation can also include a splice variant. A tumor cell-specific post-translational modification can include aberrant phosphorylation. A tumor cell-specific post-translational modification can also include a proteasome-generated spliced antigen. See Liepe et al., A large fraction of HLA class I ligands are proteasome-generated spliced peptides; Science. 2016 Oct 21; 354(6310):354-358.
[0015] As used herein, the term "tumor neoantigen" is a neoantigen that is present in tumor cells or tissues of a subject, but is not present in corresponding normal cells or tissues of the subject.
[0016] As used herein, the term "neoantigen-based vaccine" is a vaccine construct that is based on one or more neoantigens, for example multiple neoantigens.
[0017] As used herein, the term "candidate neoantigen" is a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.
[0018] As used herein, the term "coding region" is the portion of a gene that codes for a protein.
[0019] As used herein, the term "coding mutation" is a mutation that occurs in a coding region.
[0020] As used herein, the term "ORF" means open reading frame.
[0021] As used herein, the term "neo-ORF" is a tumor-specific ORF that has arisen from a mutation or other abnormality such as splicing.
[0022] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.
[0023] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon.
[0024] As used herein, the term "frameshift mutation" is a mutation that causes an alteration in the frame of a protein.
[0025] As used herein, the term "insertion / deletion" refers to the insertion or deletion of one or more nucleic acids.
[0026] As used herein, the term percent "identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences that have a specified percentage of nucleotides or amino acid residues that are the same when compared and aligned for maximum correspondence as determined using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN, or other algorithms available to those of skill in the art) or by visual inspection. Depending on the application, the percent "identity" can exist over a region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.
[0027] For sequence comparison, typically, one sequence serves as a reference sequence, with which test sequence is compared.When using sequence comparison algorithm, test sequence and reference sequence are input into computer, partial sequence coordinates are designated if necessary, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the percent sequence identity of test sequence to reference sequence based on designated program parameters.Alternatively, sequence similarity or difference can be established by the presence or absence of specific nucleotides in combination, or amino acids at selected sequence positions (e.g., sequence motifs) for translated sequences.
[0028] Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (see generally, Ausubel et al., infra).
[0029] One example of an algorithm that is suitable for determining percent sequence identity and sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.
[0030] As used herein, the term "non-stop or read-through" is a mutation that results in the removal of the natural stop codon.
[0031] As used herein, the term "epitope" is a specific portion of an antigen that is typically bound by an antibody or T-cell receptor.
[0032] As used herein, the term "immunogenic" is the ability to elicit an immune response, for example, via T cells, B cells, or both.
[0033] As used herein, the terms "HLA binding affinity," "MHC binding affinity," refer to the affinity of binding between a specific antigen and a specific MHC allele.
[0034] As used herein, the term "bait" is a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.
[0035] As used herein, the term "variant" refers to a difference between a nucleic acid of interest and a reference human genome used as a control.
[0036] As used herein, the term "variant calling" is the algorithmic determination, typically from sequencing, of the presence of a variant.
[0037] As used herein, the term "polymorphism" refers to a germline variant, ie, a variant that is found in all DNA-bearing cells of an individual.
[0038] As used herein, the term "somatic variant" is a variant that occurs in the non-germline cells of an individual.
[0039] As used herein, the term "allele" is a version of a gene or a version of a gene sequence or a version of a protein.
[0040] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.
[0041] As used herein, the term "nonsense-mediated decay" or "NMD" is the degradation of mRNA by a cell due to a premature stop codon.
[0042] As used herein, the term "truncal mutation" is a mutation that begins early in the development of a tumor and is present in a substantial portion of the cells of the tumor.
[0043] As used herein, the term "subclonal mutation" is a mutation that begins late in the development of a tumor and is present in only a subset of cells of the tumor.
[0044] As used herein, the term "exome" is a subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.
[0045] As used herein, the term "logistic regression" is a regression model for binary data of statistical origin in which the logit of the probability that the dependent variable is equal to 1 is modeled as a linear function of the dependent variable.
[0046] As used herein, the term "neural network" is a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinearities, typically trained via stochastic gradient descent and backpropagation.
[0047] As used herein, the term "proteome" is the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.
[0048] As used herein, the term "peptidome" is the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome can also refer to the properties of a cell or a collection of cells (e.g., a tumor peptidome refers to the union of the peptidomes of all cells that comprise a tumor).
[0049] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.
[0050] As used herein, the term "dextramer" is a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.
[0051] As used herein, the term "tolerance or immune tolerance" is a state of immune non-responsiveness to one or more antigens, e.g., self-antigens.
[0052] As used herein, the term "central tolerance" is tolerance that is affected in the thymus, either by deleting autoreactive T cell clones or by promoting their differentiation into immunosuppressive regulatory T cells (Tregs).
[0053] As used herein, the term "peripheral tolerance" is tolerance that is affected in the periphery by downregulating or anergizing autoreactive T cells that survived central tolerance, or by promoting these T cells to differentiate into Tregs.
[0054] The term "sample" can include a single cell, or multiple cells, or fragments of cells, or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.
[0055] The term "subject" includes cells, tissues, or organisms, either male or female, human or non-human, whether in vivo, ex vivo, or in vitro. The term subject includes mammals, including humans.
[0056] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.
[0057] The term "clinical factor" refers to a measurement of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values that can be obtained from the evaluation of a subject or a sample (or a population of samples) from a subject under a determined condition. A clinical factor can also be predicted by other parameters, such as markers and / or gene expression surrogates. A clinical factor can include tumor type, tumor subtype, and smoking history.
[0058] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigens; FFPE: formalin-fixed paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.
[0059] It should be noted that as used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0060] Any terms not directly defined herein will be understood to have the meaning generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide additional guidance to the practitioner in describing the compositions, devices, methods, etc. of the present invention aspects and how to make or use them. It will be recognized that the same thing may be said in more than one way. Thus, alternative language and synonyms may be used for any one or more of the terms discussed herein. No importance should be placed on whether a term is detailed or discussed herein. Some synonyms or interchangeable methods, materials, etc. are provided. The detailed description of one or a few synonyms or equivalents does not exclude the use of other synonyms or equivalents unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the present invention aspects herein.
[0061] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.
[0062] II. Methods for identifying neoantigens Disclosed herein is a method for identifying the neoantigens from a target tumor that are likely to be presented on the cell surface of the tumor and / or are likely to be immunogenic.For example, one such method may include the following steps: obtain at least one of tumor nucleotide sequencing data of exome, transcriptome or whole genome from the tumor cells of the target, the tumor nucleotide sequencing data is used to obtain data representing each peptide sequence of a set of neoantigens, and the peptide sequence of each neoantigen comprises at least one change that makes it different from its corresponding wild-type parent peptide sequence; input the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each neoantigen is presented by one or more MHC alleles on the tumor cell surface of the target tumor cells or by cells present in the tumor, the set of numerical likelihoods being determined at least based on the received mass spectrometry data; and select a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens.
[0063] The presented model can include a statistical regression or machine learning (e.g., deep learning) model trained against a set of reference data (also called training data set) that includes a set of corresponding labels, the set of reference data being obtained from each of a large number of separate subjects, some of which may have tumors, and includes at least one of the following: data representing the exome nucleotide sequence from tumor tissue, data representing the exome nucleotide sequence from normal tissue, data representing the transcriptome nucleotide sequence from tumor tissue, data representing the proteome sequence from tumor tissue, and data representing the MHC peptidome sequence from tumor tissue, and data representing the MHC peptidome sequence from normal tissue. The reference data can further include mass spectrometry data, sequencing data, RNA sequencing data, and proteomics data for synthetic proteins, normal and tumor human cell lines, and single allele cell lines that are engineered to express predetermined MHC alleles, which are then exposed to fresh and frozen primary samples, and T cell assays (e.g., ELISPOT). In certain aspects, the set of reference data includes each form of reference data.
[0064] The proposed model can include a set of features derived at least in part from a set of reference data, the set of features including at least one of an allele-dependent feature and an allele-independent feature. In certain aspects, each feature is included.
[0065] Characteristics of dendritic cell presentation to naive T cells can include at least one of the following: the characteristics described above. The dose and type of antigen in the vaccine (e.g., peptide, mRNA, virus, etc.): (1) the route by which dendritic cells (DCs) take up the antigen type (e.g., endocytosis, micropinocytosis); and / or (2) the efficacy with which the antigen is taken up by DCs. The dose and type of adjuvant in the vaccine. The length of the vaccine antigen sequence. The number and site of vaccine administrations. The baseline immune function of the patient (e.g., as measured by recent history of infection, blood counts, etc.). For RNA vaccines: (1) the turnover rate of mRNA protein products in dendritic cells; (2) the rate of translation of mRNA after ingestion by dendritic cells as measured in in vitro or in vivo experiments; and / or (3) the number or rounds of translation of mRNA after ingestion by dendritic cells as measured by in vivo or in vitro experiments. Optionally, the presence of protease cleavage motifs in the peptide (e.g., as measured by RNA-seq or mass spectrometry), which gives additional weight to proteases typically expressed in dendritic cells. The level of proteasome and immunoproteasome expression in typical activated dendritic cells (which may be measured by RNA-seq, mass spectrometry, immunohistochemistry, or other standard techniques). The level of expression of a particular MHC allele in the individual in question (e.g., as measured by RNA-seq or mass spectrometry), optionally specifically measured in activated dendritic cells or other immune cells. Optionally, the probability of peptide presentation by a particular MHC allele in other individuals expressing the particular MHC allele, optionally specifically measured in activated dendritic cells or other immune cells. Optionally, the probability of peptide presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals, optionally specifically measured in activated dendritic cells or other immune cells.
[0066] The immune tolerance escape trait can include at least one of the following: direct measurement of the self peptidome via protein mass spectrometry performed on one or several cell types; estimation of the self peptidome by taking the union of all k-mer (e.g., 5-25) subsequences of self proteins; estimation of the self peptidome using a model of presentation similar to the presentation model described above applied to all non-mutated self proteins, optionally accounting for germline variants.
[0067] Ranking can be performed using a number of neoantigens provided by at least one model based at least in part on numerical likelihood. After ranking, selection can be performed to select a subset of the ranked neoantigens according to a selection criterion. After selection, the subset of ranked peptides can be provided as an output.
[0068] The set of selected neoantigens may be twenty in number.
[0069] The presentation model may represent the dependency between the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in a peptide sequence, and the likelihood of presentation of such a peptide sequence containing the particular amino acid at the particular position on the tumor cell surface by that particular one of the MHC alleles.
[0070] The methods disclosed herein can also include applying one or more presentation models to a peptide sequence of a corresponding neoantigen to generate a dependency score for each of one or more MHC alleles indicating whether the MHC allele presents the corresponding neoantigen based on at least the position of an amino acid in the peptide sequence of the corresponding neoantigen.
[0071] The methods disclosed herein can also include transforming the dependency scores to generate a corresponding per-allele likelihood for each MHC allele, which indicates the likelihood that the corresponding MHC allele will present the corresponding neoantigen; and combining the per-allele likelihoods to generate a numerical likelihood.
[0072] The process of transforming the dependency scores can model the presentation of peptide sequences of corresponding neoantigens as mutually exclusive.
[0073] The methods disclosed herein can also include transforming the combination of dependency scores to generate a numerical likelihood.
[0074] Transforming the combination of dependency scores can model the presentation of the peptide sequences of the corresponding neoantigens as interference between MHC alleles.
[0075] The set of numerical likelihoods can be further specified by at least the allele non-interacting feature, and the methods disclosed herein can also apply an allele non-interacting model of one or more presentation models to the allele non-interacting feature to generate a dependency score for the allele non-interacting feature indicating whether the corresponding neoantigen peptide sequence is presented based on the allele non-interacting feature.
[0076] The methods disclosed herein can also include combining the dependency scores for each MHC allele with the dependency scores for the allele-non-interacting property in one or more MHC alleles; transforming the combined dependency scores for each MHC allele to generate a corresponding per-allele likelihood for the MHC allele, which indicates the likelihood that the corresponding MHC allele will present the corresponding neoantigen; and combining the per-allele likelihoods to generate a numerical likelihood.
[0077] The methods disclosed herein can also include a step of transforming the combination of the dependency scores for each of the MHC alleles and the dependency scores for the allele-non-interacting trait to generate a numerical likelihood.
[0078] The set of numerical parameters for the proposed model can be trained based on a training dataset that includes at least a set of training peptide sequences identified as present in multiple samples and one or more MHC alleles associated with each training peptide sequence, the training peptide sequences being identified through mass spectrometry on isolated peptides eluted from the MHC alleles derived from the multiple samples.
[0079] Samples can also include cell lines engineered to express a single MHC class I or class II allele.
[0080] The sample may also include cell lines engineered to express multiple MHC class I or class II alleles.
[0081] Samples can also include human cell lines obtained or derived from multiple patients.
[0082] Samples can also include fresh or frozen tumor samples obtained from multiple patients.
[0083] Samples can also include fresh or frozen tissue samples obtained from multiple patients.
[0084] The sample may also include peptides identified using a T cell assay.
[0085] The training dataset may further include data relating to the peptide abundance of the set of training peptides present in the sample; the peptide lengths of the set of training peptides in the sample.
[0086] The training data set may be generated by comparing a set of training peptide sequences via alignment to a database containing a set of known protein sequences, the set of training protein sequences being longer than and including the training peptide sequences.
[0087] The training dataset may be generated based on performing nucleotide sequencing on the cell line or having nucleotide sequencing completed to obtain at least one of exome, transcriptome, or whole genome sequencing data from the cell line, the sequencing data including at least one nucleotide sequence that includes the alteration.
[0088] The training dataset may be generated based on obtaining at least one of exome, transcriptome, or whole genome normal nucleotide sequencing data from a normal tissue sample.
[0089] The training dataset may further include data relating to proteomic sequences associated with the sample.
[0090] The training dataset may further comprise data relating to MHC peptidome sequences associated with the sample.
[0091] The training dataset may further include data relating to peptide-MHC binding affinity measurements for at least one of the isolated peptides.
[0092] The training dataset may further include data relating to peptide-MHC binding stability measurements for at least one of the isolated peptides.
[0093] The training dataset may further include data relating to the transcriptome associated with the sample.
[0094] The training dataset may further include data relating to the genome associated with the sample.
[0095] Training peptide sequences may be in the range of lengths k-mers, where k is between 8 and 15 for MHC class I, or between 9 and 30 for MHC class II.
[0096] The methods disclosed herein can also include encoding the peptide sequence using a one-hot encoding scheme.
[0097] The methods disclosed herein can also include encoding the training peptide sequences using a left-padded one-hot encoding scheme.
[0098] A method for treating a subject having a tumor, comprising carrying out the steps of claim 1 and further comprising obtaining a tumor vaccine comprising a set of selected neoantigens, and administering the tumor vaccine to the subject.
[0099] Also disclosed herein is a method for producing a tumor vaccine, comprising the steps of: obtaining at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from tumor cells of a subject, wherein the tumor nucleotide sequencing data is used to obtain data representing a peptide sequence of each of a set of neoantigens, and wherein the peptide sequence of each neoantigen comprises at least one mutation that makes it different from a corresponding wild-type parent peptide sequence; inputting the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each of the neoantigens will be presented by one or more MHC alleles on the tumor cell surface of the tumor cells of the subject, wherein the set of numerical likelihoods has been identified based at least on the received mass spectrometry data; and selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens; and producing, or production has been completed, a tumor vaccine comprising the set of selected neoantigens.
[0100] Also disclosed herein is a tumor vaccine comprising a set of selected neoantigens selected by performing a method comprising the steps of: obtaining at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from tumor cells of a subject, wherein the tumor nucleotide sequencing data is used to obtain data representing a peptide sequence of each of the set of neoantigens, and wherein the peptide sequence of each neoantigen comprises at least one mutation that makes it different from a corresponding wild-type parent peptide sequence; inputting the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each of the neoantigens will be presented by one or more MHC alleles on the tumor cell surface of the tumor cells of the subject, wherein the set of numerical likelihoods has been identified based at least on the received mass spectrometry data; and selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens; and producing or has completed production of a tumor vaccine comprising the set of selected neoantigens.
[0101] The tumor vaccine may comprise one or more of a nucleotide sequence, a polypeptide sequence, RNA, DNA, a cell, a plasmid, or a vector.
[0102] A tumor vaccine may comprise one or more neoantigens displayed on the surface of tumor cells.
[0103] A tumor vaccine may include one or more neoantigens that are immunogenic in the subject.
[0104] A tumor vaccine may not include one or more neoantigens that induce an autoimmune response in a subject against normal tissues.
[0105] The tumor vaccine may include an adjuvant.
[0106] The tumor vaccine may include an excipient.
[0107] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of being presented on the tumor cell surface relative to neoantigens that are not selected based on the presentation model.
[0108] The methods disclosed herein may also include selecting a neoantigen that has an increased likelihood of being able to induce a tumor-specific immune response in a subject compared to a neoantigen that is not selected based on the presentation model.
[0109] The methods disclosed herein may also include selecting a neoantigen that has an increased likelihood of being presented to naive T cells by a professional antigen-presenting cell (APC) compared to a neoantigen that is not selected based on a presentation model, optionally where the APC is a dendritic cell (DC).
[0110] The methods disclosed herein may also include the selection of neoantigens that have a reduced likelihood of being subject to inhibition via central or peripheral tolerance compared to neoantigens that are not selected based on a presentation model.
[0111] The methods disclosed herein may also include the selection of neoantigens that have a reduced likelihood of being able to induce an autoimmune response against normal tissue in a subject compared to neoantigens that are not selected based on the presentation model.
[0112] Exome or transcriptome nucleotide sequencing data may be obtained by performing sequencing on tumor tissue.
[0113] Sequencing may be next generation sequencing (NGS) or any massively parallel sequencing approach.
[0114] The set of numerical likelihoods may be further specified by at least MHC allele interaction properties including at least one of the following: the predicted affinity with which the MHC allele and the neoantigen-encoded peptide bind; the predicted stability of the neoantigen-encoded peptide-MHC complex; the sequence and length of the neoantigen-encoded peptide; the probability of presentation of a neoantigen-encoded peptide with a similar sequence in cells from other individuals expressing the particular MHC allele, as assessed by mass spectrometry proteomics or other means; the expression level of the particular MHC allele in the subject in question (e.g., as measured by RNA-seq or mass spectrometry); the overall neoantigen-encoded peptide sequence-independent probability of presentation by the particular MHC allele in other distinct individuals expressing the particular MHC allele; the overall neoantigen-encoded peptide sequence-independent probability of presentation by MHC alleles in the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other distinct subjects.
[0115] The set of numerical likelihoods is further specified by at least MHC allele non-interacting properties, including at least one of the following: C-terminal and N-terminal sequences flanking the neoantigen-encoded peptide within the source protein sequence; the presence of protease cleavage motifs in the neoantigen-encoded peptide, optionally weighted according to the expression of the corresponding protease in tumor cells (as measured by RNA-seq or mass spectrometry); the turnover rate of the source protein, as measured in the appropriate cell type; the length of the source protein, optionally taking into account the specific splice variants ("isoforms") most highly expressed in tumor cells, as measured by RNA-seq or proteomic mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data; the expression of the proteasome, immunoproteasome, thymoproteasome in tumor cells (which may be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry), or other protease expression levels; expression of the source gene of the neoantigen-encoded peptide (e.g., as measured by RNA-seq or mass spectrometry); typical tissue-specific expression of the source gene of the neoantigen-encoded peptide during different stages of the cell cycle; a comprehensive catalog of properties of the source protein and / or its domains, such as can be found in, for example, uniProt or the PDB http: / / www.rcsb.org / pdb / home / home.do; properties describing the nature of the domain of the source protein that contains the peptide, e.g., secondary or tertiary structure (e.g., alpha helix vs. beta sheet); alternative splicing; the probability of presentation of the peptide from the source protein of the neoantigen-encoded peptide in question in other distinct subjects; the probability that the peptide will be undetected or over-represented by mass spectrometry due to technical bias; expression of various gene modules / pathways (not necessarily containing the source protein of the peptide), as measured by RNASeq, that give information about the status of the tumor cells, stroma, or tumor infiltrating lymphocytes (TILs);The number of copies of the source gene of the neoantigen-encoding peptide in the tumor cells; the probability that the peptide binds to TAP or the binding affinity of the peptide to TAP, measured or predicted; the expression level of TAP in the tumor cells (which may be measured by RNA-seq, proteomic mass spectrometry, immunohistochemistry); the presence or absence of tumor mutations, including but not limited to: driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, NTRK3, and genes encoding proteins involved in antigen presentation. in genes that are involved in the differentiation of HLA-related proteins (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery that are affected by loss-of-function mutations in the tumor have a reduced probability of presentation; including but not limited to the presence or absence of functional germline polymorphisms in genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO , HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome); tumor type (e.g., NSCLC, melanoma);clinical tumor subtype (e.g., squamous cell lung cancer vs. non-squamous); smoking history; typical expression of the peptide source gene in the relevant tumor type or clinical subtype, optionally stratified by driver mutations;
[0116] The at least one mutation may be a frameshift or non-frameshift indel, a missense or nonsense substitution, a splice site alteration, a genomic rearrangement or gene fusion, or any genomic or expression alteration that results in a de novo ORF.
[0117] The tumor cells may be selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0118] The methods disclosed herein also include obtaining a tumor vaccine comprising the selected set of neoantigens or a subset thereof, and optionally further comprising administering the tumor vaccine to a subject.
[0119] At least one of the neoantigens in the set of selected neoantigens, when in polypeptide form, may comprise at least one of the following: binding affinity to MHC with an IC50 value of less than 1000 nM, a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class 1 polypeptides, the presence of a sequence motif within or near the polypeptide in the parent protein sequence that promotes proteasomal cleavage, and the presence of a sequence motif that promotes TAP transport.
[0120] Also disclosed herein is a method for generating a model for identifying one or more neoantigens likely to be presented on the tumor cell surface of a tumor cell, comprising the steps of: receiving mass spectrometry data comprising data relating to a number of isolated peptides eluted from major histocompatibility complexes (MHC) derived from a number of samples; obtaining a training dataset by at least identifying a set of training peptide sequences present in the samples and one or more MHCs associated with each training peptide sequence; training a set of numerical parameters of a presentation model using the training dataset comprising the training peptide sequences, wherein the presentation model provides a number of numerical likelihoods that a peptide sequence derived from a tumor cell will be presented by one or more MHC alleles on the tumor cell surface.
[0121] The presentation model may represent the dependency between the presence of a particular amino acid at a particular position in a peptide sequence and the likelihood of presentation of a peptide sequence containing a particular amino acid at a particular position by one of the MHC alleles on a tumor cell.
[0122] Samples can also include cell lines engineered to express a single MHC class I or class II allele.
[0123] The sample may also include cell lines engineered to express multiple MHC class I or class II alleles.
[0124] Samples can also include human cell lines obtained or derived from multiple patients.
[0125] Samples can also include fresh or frozen tumor samples obtained from multiple patients.
[0126] The sample may also include peptides identified using a T cell assay.
[0127] The training dataset may further include data relating to the peptide abundances of the set of training peptides present in the sample; the peptide lengths of the set of training peptides in the sample.
[0128] The methods disclosed herein can also include obtaining a set of training protein sequences based on the training peptide sequences, the training protein sequences being longer than and including the training peptide sequences, by comparing the set of training peptide sequences via alignment against a database comprising a set of known protein sequences.
[0129] The methods disclosed herein can also include performing or having mass spectrometry completed on the cell line to obtain at least one of exome, transcriptome, or whole genome nucleotide sequencing data from the cell line, wherein the nucleotide sequencing data comprises at least one protein sequence that includes the mutation.
[0130] The methods disclosed herein can also include encoding the training peptide sequences using a one-hot encoding scheme.
[0131] The methods disclosed herein may also include obtaining at least one of normal nucleotide sequencing data of the exome, transcriptome, and whole genome from a normal tissue sample; and using the normal nucleotide sequencing data to train a set of parameters of the proposed model.
[0132] The training dataset may further include data relating to proteomic sequences associated with the sample.
[0133] The training dataset may further comprise data relating to MHC peptidome sequences associated with the sample.
[0134] The training dataset may further include data relating to peptide-MHC binding affinity measurements for at least one of the isolated peptides.
[0135] The training dataset may further include data relating to peptide-MHC binding stability measurements for at least one of the isolated peptides.
[0136] The training dataset may further include data relating to the transcriptome associated with the sample.
[0137] The training dataset may further include data relating to the genome associated with the sample.
[0138] The methods disclosed herein may also include performing a logistic regression of the set of parameters.
[0139] Training peptide sequences may be in the range of lengths k-mers, where k is between 8 and 15 for MHC class I, or between 9 and 30 for MHC class II.
[0140] The methods disclosed herein may also include encoding the training peptide sequences using a left-padded one-hot encoding scheme.
[0141] The methods disclosed herein may also include determining values for the set of parameters using a deep learning algorithm.
[0142] Disclosed herein is a method for identifying one or more neoantigens likely to be presented on the tumor cell surface of tumor cells, comprising performing the following steps: receiving mass spectrometry data including data relating to a number of isolated peptides eluted from major histocompatibility complexes (MHC) derived from a number of fresh or frozen tumor samples; obtaining a training dataset by identifying at least a set of training peptide sequences that are present in the tumor sample and presented on one or more MHC alleles associated with each training peptide sequence; obtaining a set of training protein sequences based on the training peptide sequences; and training a set of numerical parameters of a presentation model using the training protein sequences and training peptide sequences, wherein the presentation model provides a number of numerical likelihoods that a peptide sequence derived from a tumor cell will be presented by one or more MHC alleles on the tumor cell surface.
[0143] The presentation model may represent the dependency between the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in a peptide sequence, and the likelihood of presentation of such a peptide sequence containing a particular amino acid at a particular position on the tumor cell surface by that particular one of the MHC alleles of the pair.
[0144] The methods disclosed herein can also include a step of selecting a subset of neoantigens, each of which has an increased likelihood of being presented on the cell surface of a tumor compared to one or more distinct tumor neoantigens.
[0145] The methods disclosed herein can also include a step of selecting a subset of neoantigens, each of which has an increased likelihood of inducing a tumor-specific immune response in a subject compared to one or more distinct tumor neoantigens.
[0146] The methods disclosed herein can also include a step of selecting a subset of neoantigens, each of which has an increased likelihood of being presented to naive T cells by a professional antigen-presenting cell (APC) relative to one or more distinct tumor neoantigens, and optionally the APC is a dendritic cell (DC).
[0147] The methods disclosed herein can also include a step of selecting a subset of neoantigens, each of which is selected because it has a reduced likelihood of being subject to inhibition via central or peripheral tolerance compared to one or more distinct tumor neoantigens.
[0148] The methods disclosed herein can also include a step of selecting a subset of neoantigens, each of which has a reduced likelihood of being able to induce an autoimmune response against normal tissue in a subject compared to one or more distinct tumor neoantigens.
[0149] The methods disclosed herein can also include a step of selecting a subset of neoantigens, each of which has a reduced likelihood of being differentially post-translationally modified in tumor cells versus APCs, and optionally, the APCs are dendritic cells (DCs).
[0150] The practice of the methods herein employs conventional methods of protein chemistry, biochemistry, recombinant DNA techniques, and pharmacology within the skill of the art, unless otherwise indicated. Such techniques are fully explained in the literature. See, for example, TE Creighton, Proteins: Structures and Molecular Properties (WH Freeman and Company, 1993);AL Lehninger, Biochemistry (Worth Publishers, Inc., current addition);Sambrook, et al., Molecular Cloning: A Laboratory Manual (2nd Edition, 1989);Methods In Enzymology (S. Colowick and N. Kaplan eds., Academic Press, Inc.);Remington's Pharmaceutical Sciences, 18th Edition (Easton, Pennsylvania: Mack Publishing Company, 1990);Carey and Sundberg Advanced Organic Chemistry 3rd Edition (East, Pennsylvania: Mack Publishing Company, 1990); rd Ed. (Plenum Press) Vols A and B (1992).
[0151] III. Identification of tumor-specific mutations in neoantigens Also disclosed herein is a method for identifying certain mutations (e.g., variants or alleles present in cancer cells).In particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject with cancer, but may not be present in normal tissues from the subject.
[0152] Genetic mutations in tumors can be considered useful for tumor immunological targeting if they cause changes in the amino acid sequence of a protein exclusively in tumors. Useful mutations include: (1) non-synonymous mutations that result in different amino acids in the protein; (2) read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; (3) splice site mutations that result in the inclusion of an intron in mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) frameshift mutations or deletions that result in new open reading frames with new tumor-specific protein sequences. Mutations can also include one or more of non-frameshift insertions and deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in new ORFs.
[0153] For example, mutated peptides or mutated polypeptides resulting from splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.
[0154] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0155] A variety of methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. Several techniques have been described, including, for example, dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize the amplification of target gene regions, typically by PCR. Still other methods are based on the generation of small signal molecules by invasive cleavage followed by mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.
[0156] PCR-based detection means can include multiplex amplification of multiple markers simultaneously. For example, it is well known in the art to select PCR primers to generate PCR products that do not overlap in size and can be analyzed simultaneously. Alternatively, it is possible to amplify different markers with primers that are differentially labeled and therefore can be differentially detected. Of course, hybridization-based detection means allows the differential detection of multiple PCR products in a sample. Other techniques that allow multiplex analysis of multiple markers are known in the art.
[0157] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA. For example, single nucleotide polymorphisms can be detected by using specialized exonuclease-resistant nucleotides, as disclosed, for example, in Mundy, CR (US Pat. No. 4,656,127). According to the method, a primer that is complementary to the allele sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a specific animal or human. If the polymorphic site on the target molecule contains a nucleotide that is complementary to the specific exonuclease-resistant nucleotide derivative present, the derivative is incorporated onto the end of the hybridized primer. Due to such incorporation, the primer becomes resistant to exonucleases, thereby allowing its detection. Since the identity of the exonuclease-resistant derivative of the sample is known, the knowledge that the primer has become resistant to exonucleases reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of exogenous sequence data.
[0158] Solution-based methods can be used to determine the identity of the nucleotide at a polymorphic site. Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087). As in the method of Mundy in U.S. Patent No. 4,656,127, a primer is used that is complementary to the allelic sequence immediately 3' to the polymorphic site. The method uses a labeled dideoxynucleotide derivative that becomes incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site to determine the identity of the nucleotide at that site.
[0159] An alternative method known as Genetic Bit Analysis or GBA has been described by Goelet, P. et al. (PCT Application No. 92 / 15712). The method of Goelet, P. et al. uses a mixture of labeled terminators and primers that are complementary to the sequence 3' of the polymorphic site. The labeled terminators that are incorporated are therefore determined by and complementary to the nucleotides present at the polymorphic site of the target molecule being evaluated. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which the primers or the target molecule are immobilized on a solid phase.
[0160] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, JS et al., Nucl. Acids. Res. 17:7779-7784 (1989);Sokolov, BP, Nucl. Acids Res. 18:3671 (1990);Syvanen, A.-C., et al., Genomics 8:684-692 (1990);Kuppuswamy, MN et al., Proc. Natl. Acad. Sci. (USA) 88:1143-1147 (1991);Prezant, TR et al., Hum. Mutat. 1:159-164 (1992);Ugozzoli, L. et al., GATA 9:107-112 (1992);Nyren, P. et al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, the signal is proportional to the number of deoxynucleotides incorporated, so that polymorphisms occurring in runs of the same nucleotides can result in a signal proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).
[0161] Numerous initiatives obtain sequence information directly from millions of individual molecules of DNA or RNA in parallel. Real-time single molecule sequencing by synthesis techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent strands of DNA that are complementary to the template to be sequenced. In one method, oligonucleotides 30-50 bases in length are covalently anchored at their 5' ends to a glass coverslip. These anchored strands serve two functions. First, they act as capture sites for the target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis of sequence reading. The capture primers serve as fixed location sites for multiple cycles of synthesis, detection, and sequencing using chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mix, rinsing, imaging, and cleavage of the dye. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. As the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing-by-synthesis techniques also exist.
[0162] Any suitable sequencing-by-synthesis platform can be used to identify mutations. As mentioned above, four major sequencing-by-synthesis platforms are currently available: Genome Sequencer from Roche / 454 Life Sciences, 1G Analyzer from Illumina / Solexa, SOLiD system from Applied BioSystems, and Heliscope system from Helicos Biosciences. Sequencing-by-synthesis platforms are also described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the multiple nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. The capture sequence (also called the universal capture sequence) is a nucleic acid sequence that is complementary to the sequence attached to the support, which can double as a universal primer.
[0163] As an alternative to capture sequences, members of a coupling pair (e.g., antibody / antigen, receptor / ligand, or an avidin-biotin pair, e.g., as described in U.S. Patent Application Publication No. 2006 / 0252077) can be linked to each fragment and captured onto a surface coated with the respective second member of the coupling pair.
[0164] Following capture, the sequence can be analyzed by single molecule detection / sequencing, e.g., as described in the Examples and in U.S. Pat. No. 7,283,337, including template-dependent sequencing by synthesis. In sequencing by synthesis, the surface-bound molecules are exposed to a multitude of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides that are incorporated into the 3' end of the growing strand. This can be done in real time and in a step-and-repeat mode. For real-time analysis, a different optical label can be incorporated for each nucleotide and multiple lasers can be utilized for stimulation of the incorporated nucleotides.
[0165] Sequencing can also include other massively parallel sequencing or next generation sequencing (NGS) techniques and platforms.Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, Thermo PGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION.Additional similar current massively parallel sequencing technologies and future generations of these technologies can be used.
[0166] Any cell type or tissue can be utilized to obtain nucleic acid samples for use in the methods described herein. For example, DNA or RNA samples can be obtained from tumors or body fluids, such as blood obtained by known techniques (e.g., venipuncture), or saliva. Alternatively, nucleic acid testing can be performed on dry samples (e.g., hair or skin). In addition, a sample can be obtained from tumors for sequencing, and another sample can be obtained from normal tissues for sequencing, if the normal tissues are of the same tissue type as the tumor. A sample can be obtained from tumors for sequencing, and another sample can be obtained from normal tissues for sequencing, if the normal tissues are of a different tissue type from the tumor.
[0167] The tumor can include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0168] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. Peptides can be acid eluted from tumor cells or from HLA molecules immunoprecipitated from tumors and then identified using mass spectrometry.
[0169] IV. Neoantigens Neoantigens can comprise nucleotides or polynucleotides. For example, neoantigens can be RNA sequences that code for a polypeptide sequence. Neoantigens useful in vaccines can thus comprise nucleotide sequences or polypeptide sequences.
[0170] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the neoantigens comprise nucleotide sequences (e.g., DNA or RNA) that encode the associated polypeptide sequences.
[0171] The one or more polypeptides encoded by the neoantigen nucleotide sequence can include at least one of the following: binding affinity to MHC with an IC50 value of less than 1000 nM, a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class 1 peptides, the presence of a sequence motif within or near the peptide that promotes proteasomal cleavage, and the presence of a sequence motif that promotes TAP transport.
[0172] One or more neoantigens can be present on the surface of the tumor.
[0173] The neoantigen or neoantigens can be immunogenic in a tumor-bearing subject, for example, capable of eliciting a T cell or B cell response in the subject.
[0174] One or more neoantigens that induce an autoimmune response in a subject can be eliminated from consideration in the context of generating a vaccine for a tumor-bearing subject.
[0175] The size of the at least one neoantigenic peptide molecule can include, but is not limited to, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120 or more amino acid molecule residues, and any range derivable therein. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.
[0176] Neoantigenic peptides and polypeptides can be 15 residues or less in length, typically between about 8 and about 11 residues, particularly 9 or 10 residues, for MHC class I; and can be 15 to 24 residues for MHC class II.
[0177] If desired, longer peptides can be designed in several ways. In one example, where the presentation likelihood of a peptide on an HLA allele is predicted or known, the longer peptide can consist of either (1) an individual presented peptide with a 2-5 amino acid extension toward the N-terminus and C-terminus of each corresponding gene product; (2) a concatenation of some or all of the presented peptides, each with an extended sequence. In another example, where sequencing reveals long (longer than 10 residues) neo-epitope sequences present in the tumor (e.g., due to frameshifts, read-throughs, or inclusion of introns resulting in novel peptide sequences), the longer peptide will (3) consist of the entire stretch of novel tumor-specific amino acids, thus avoiding the need for computational or in vitro test-based selection of the strongest HLA-presented shorter peptide. In both examples, the use of longer peptides may allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.
[0178] Neo-antigenic peptides and polypeptides can be presented on HLA proteins.In some aspects, neo-antigenic peptides and polypeptides are presented on HLA proteins with stronger affinity than wild-type peptides.In some aspects, neo-antigenic peptides or polypeptides can have IC50 of at least 5000nM or less, at least 1000nM or less, at least 500nM or less, at least 250nM or less, at least 200nM or less, at least 150nM or less, at least 100nM or less, at least 50nM or less.
[0179] In some aspects, the neoantigenic peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.
[0180] Also provided is a composition comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two distinct peptides. The at least two distinct peptides can be derived from the same polypeptide. By distinct polypeptide, it is meant that the peptides differ by length, amino acid sequence, or both. The peptides are derived from any polypeptide that is known or found to contain tumor-specific mutations. Suitable polypeptides from which neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC serves as a curator of comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some aspects, the tumor-specific mutations are driver mutations for a particular cancer type.
[0181] Neoantigenic peptides and polypeptides with desired activity or properties can be modified to provide improved pharmacological characteristics while increasing or at least retaining substantially all of the biological activity of the unmodified peptide, such as binding to a desired MHC molecule and activating appropriate T cells, certain desired attributes. By way of example, neoantigenic peptides and polypeptides can be subjected to various modifications, such as either conservative or non-conservative substitutions, which may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. By conservative substitution, it is meant replacing an amino acid residue with another that is biologically and / or chemically similar, for example, one hydrophobic residue with another hydrophobic residue, or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effect of single amino acid substitutions may also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2d Ed. (1984).
[0182] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful in increasing the stability of peptides and polypeptides in vivo. Stability can be assayed in a number of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, for example, Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol is generally as follows: Pooled human serum (type AB, non-heat inactivated) is delipidated by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small amounts of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction sample is cooled (4° C.) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.
[0183] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. As an example, the ability of a peptide to induce CTL activity can be enhanced by linkage to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is typically composed of relatively small neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is typically selected, for example, from Ala, Gly, or other neutral spacers of non-polar amino acids or neutral polar amino acids. It will be understood that the spacer, if present, need not be composed of the same residues and thus may be a hetero- or homo-oligomer. If present, the spacer will usually be at least 1 or 2 residues, more usually 3-6 residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.
[0184] The neoantigenic peptide can be linked to the T helper peptide either directly or via a spacer at either the amino or carboxy terminus of the peptide. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxoid 830-843, influenza 307-319, malaria circumsporozoite 382-398 and 378-389.
[0185] Proteins or peptides can be produced by any technique known to those skilled in the art, including expressing proteins, polypeptides, or peptides through standard molecular biology techniques, isolating proteins or peptides from natural sources, or chemically synthesizing proteins or peptides. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes can be found in computerized databases that have been previously disclosed and are known to those skilled in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located at the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those skilled in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those skilled in the art.
[0186] In a further aspect, the neoantigen comprises a nucleic acid (e.g., polynucleotide) encoding a neoantigenic peptide or a portion thereof. The polynucleotide can be, for example, a single-stranded and / or double-stranded polynucleotide, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), e.g., a polynucleotide having a phosphorothioate backbone, or in either a natural or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further aspect provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host, although such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.
[0187] IV. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., a tumor-specific immune response. Vaccine compositions typically include multiple neoantigens, e.g., selected using the methods described herein. Vaccine compositions can also be referred to as vaccines.
[0188] The vaccine can contain 1-30 peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine may contain from 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 109, 102, 104, 105, 106, 107, 108, 109, 109, 109 It may contain 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine contains 1-30 neoantigen sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 110 It can contain 6, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.
[0189] In one embodiment, the different peptides and / or polypeptides, or the nucleotide sequences encoding them, are selected so that the peptides and / or polypeptides can bind to different MHC molecules, such as different MHC class I molecules. In some aspects, one vaccine composition comprises coding sequences for peptides and / or polypeptides that can bind to the most frequently occurring MHC class I molecules. Thus, the vaccine composition can comprise different fragments that can bind to at least two preferred, at least three preferred, or at least four preferred MHC class I molecules.
[0190] The vaccine composition may generate a specific cytotoxic T cell response and / or a specific helper T cell response.
[0191] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are given herein below. The composition can be combined with a carrier, such as, for example, a protein, or an antigen-presenting cell, such as, for example, a dendritic cell (DC), which can present peptides to T cells.
[0192] An adjuvant is any substance whose incorporation into a vaccine composition increases or otherwise modifies the immune response to a neoantigen. The carrier can be a scaffold structure, such as a polypeptide or polysaccharide, to which the neoantigen can be bound. Optionally, the adjuvant is covalently or non-covalently conjugated.
[0193] The ability of adjuvant to increase immune response to antigen is typically manifested by a significant or substantial increase in immune-mediated reaction or a reduction in disease symptoms.For example, the increase in humoral immunity is typically manifested by a significant increase in the titer of antibody generated against antigen, and the increase in T cell activity is typically manifested in an increase in cell proliferation, or cellular cytotoxicity, or cytokine secretion.Adjuvant can also change immune response, for example, by changing a predominantly humoral or Th response to a predominantly cellular or Th response.
[0194] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, Imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA 206, Montanide ISA 50V, Montanide Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila's QS21 stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponin, mycobacterial extracts and synthetic bacterial cell wall mimics, and other proprietary adjuvants such as Ribi's Detox. Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are useful. Several immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparation have been previously described (Dupuis M, et al., Cell Immunol. 1998; 186(1):18-27; Allison AC; Dev Biol Stand. 1998; 92:3-11). Cytokines can also be used.Several cytokines have been directly linked to influencing dendritic cell migration to lymphoid tissues (e.g., TNF-α), accelerating dendritic cell maturation into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J Immunother Emphasis Tumor Immunol. 1996 (6):414-418).
[0195] CpG immunostimulatory oligonucleotides have also been reported to enhance the effect of adjuvants in a vaccine setting. Other TLR binding molecules, such as RNA that binds to TLR 7, TLR 8, and / or TLR 9, may also be used.
[0196] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies, such as cyclophosphamide, sunitinib, bevacizumab, celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony stimulating factors, such as granulocyte-macrophage colony stimulating factor (GM-CSF, sargramostim).
[0197] Vaccine compositions can include more than one different adjuvant. Additionally, therapeutic compositions can include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.
[0198] The carrier (or excipient) can be present independent of the adjuvant. The function of the carrier can be, for example, to increase the activity or immunogenicity, to provide stability, to increase biological activity, or to increase serum half-life, particularly to increase the molecular weight of the variant. In addition, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, for example, a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, a serum protein such as keyhole limpet hemocyanin, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is acceptable and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be a dextran, for example, sepharose.
[0199] Cytotoxic T cells (CTL) recognize antigen in the form of peptide bound to MHC molecules rather than intact foreign antigen itself. MHC molecules themselves are located on the cell surface of antigen-presenting cells. Therefore, activation of CTL is possible when the trimeric complex of peptide antigen, MHC molecules and APC exists. Correspondingly, it can enhance immune response not only when peptide is used for activating CTL, but also when APC with each MHC molecule is added in addition. Thus, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.
[0200] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including but not limited to second, third, or hybrid second / third generation lentiviruses, and recombinant lentiviruses of any generation designed to target a specific cell type or receptor (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human (See, for example, Ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880). Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequences may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guerin). BCG vectors are described in Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0201] IV.A. Additional Considerations for Vaccine Design and Manufacturing IV.A.1. Determination of a set of peptides covering all tumor subclones Truncal peptides, meaning those presented by all or most of the tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides that are predicted to be highly likely to be presented and immunogenic, or if the number of truncal peptides that are predicted to be highly likely to be presented and immunogenic is low enough that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .
[0202] IV.A.2. Neoantigen prioritization After applying all of the above neoantigen filters, more candidate neoantigens may still be available for vaccine inclusion than the vaccine technology can accommodate. Additionally, uncertainties may remain about various aspects of the neoantigen analysis, and compromises may exist between various properties of the candidate vaccine neoantigens. Therefore, instead of pre-determined filters at each stage of the selection process, an integral multidimensional model can be considered, in which the candidate neoantigens are placed in a space with at least the following axes, and the selection is optimized using an integral approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmune risk is typically favorable) 2. Probability of sequencing artifacts (lower artifact probability is typically preferred) 3. Probability of immunogenicity (a higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probabilities of presentation are typically preferred) 5. Gene Expression (higher expression is typically preferred) 6. HLA gene coverage (a higher number of HLA molecules involved in the presentation of a set of neoantigens may decrease the probability that a tumor will evade immune attack through downregulation or mutation of HLA molecules)
[0203] V. Methods of Treatment and Manufacturing Also provided are methods of inducing a tumor-specific immune response in a subject, vaccinating against a tumor, treating and / or alleviating symptoms of cancer in a subject by administering to the subject one or more neoantigens, such as multiple neoantigens identified using the methods disclosed herein.
[0204] In some aspects, the subject is diagnosed with cancer or is at risk of developing cancer.The subject can be a human, a dog, a cat, a horse, or any animal for which tumor-specific immune response is desired.The tumor can be any solid tumor, such as breast, ovarian, prostate, lung, kidney, stomach, colon, testicular, head and neck, pancreas, brain, melanoma, and other tissue organ tumors, and hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.
[0205] The neoantigen can be administered in an amount sufficient to induce a CTL response.
[0206] The neoantigens can be administered alone or in combination with other therapeutic agents, such as chemotherapeutic agents, radiation, or immunotherapy. Any suitable therapeutic treatment for the particular cancer can be administered.
[0207] In addition, the subject can be further administered with an anti-immunosuppressive / immunostimulatory substance, such as a checkpoint inhibitor.For example, the subject can be further administered with an anti-CTLA antibody or anti-PD-1 or anti-PD-L1.The blockade of CTLA-4 or PD-L1 with an antibody can enhance the immune response against cancerous cells in the patient.In particular, CTLA-4 blockade has been shown to be effective when vaccination protocols are adopted.
[0208] The optimal amount of each neoantigen to be included in the vaccine composition and the optimal dosing regimen can be determined. For example, the neoantigen or its variants can be prepared for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Methods of injection include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administration of the vaccine composition are known to those skilled in the art.
[0209] Vaccines can be edited so that the selection, number, and / or amount of neoantigens present in the composition are tissue, cancer, and / or patient specific. As an example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. Selection can depend on the specific type of cancer, the state of the disease, earlier treatment regimes, the immune status of the patient, and of course the HLA haplotype of the patient. Furthermore, vaccines can contain components that are personalized according to the personal needs of a particular patient. Examples include changing the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjusting for secondary treatments after the first round or scheme of treatment.
[0210] For compositions to be used as vaccines for cancer, neoantigens with similar normal self-peptides that are abundantly expressed in normal tissues can be avoided or present in low amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a high amount of a particular neoantigen, the respective pharmaceutical composition for the treatment of this cancer can be present in high amount and / or can include more than one neoantigen specific to this particular neoantigen or pathway of this neoantigen.
[0211] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In therapeutic applications, the compositions are administered to patients in an amount sufficient to elicit an effective CTL response against tumor antigens and to cure or at least partially halt symptoms and / or complications. The amount adequate to achieve this is defined as a "therapeutically effective dose". The amount effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the weight and general health of the patient, and the judgment of the prescribing physician. It should be kept in mind that the compositions can generally be used in severe disease states, i.e., life-threatening or potentially life-threatening situations, especially when cancer has metastasized. In such instances, it may be possible, and the treating physician may feel desirable, to administer substantial excesses of these compositions, taking into account the minimization of foreign material and the relatively non-toxic nature of neoantigens.
[0212] For therapeutic use, administration can begin upon detection or surgical removal of the tumor, followed by boosting doses until at least symptoms are substantially abated, and for a period thereafter.
[0213] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, e.g., intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against the tumor. Disclosed herein are compositions for parenteral administration that include a solution of a neoantigen, the vaccine composition being dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. A variety of aqueous carriers can be used, e.g., water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional, well-known sterilization techniques, or can be sterile filtered. The resulting aqueous solutions can be packaged for use as is, or lyophilized, the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may contain pharma- ceutically acceptable auxiliary substances required to approximate physiological conditions, such as pH adjusting and buffering agents, isotonicity agents, wetting agents, and the like, e.g., sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.
[0214] Neoantigens can also be administered via liposomes, which target them to specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigen to be delivered is incorporated as part of the liposome, either alone or with a molecule that binds to a receptor that is predominant among lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Thus, liposomes filled with the desired neoantigen can be directed to the site of lymphoid cells, where they then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid lability, and stability of the liposomes in the bloodstream. A variety of methods are available for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9; 467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.
[0215] For targeting to immune cells, the ligand to be incorporated into the liposomes can include, for example, an antibody or fragment thereof specific for a cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., in doses that vary according to, among other things, the mode of administration, the peptide being delivered, and the stage of the disease being treated.
[0216] For therapeutic or immunization purposes, the peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient. Numerous methods are conveniently used to deliver nucleic acids to a patient. For example, the nucleic acid can be delivered directly as "naked DNA". This approach is described, for example, in Wolff et al., Science 247: 1465-1468 (1990), and in U.S. Pat. Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery, as described, for example, in U.S. Pat. No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, DNA can be attached to particles, such as gold particles. Approaches for delivering nucleic acid sequences can include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.
[0217] Nucleic acid can also be delivered by complexing with cationic compounds such as cationic lipids.Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO 96 / 18372; 9324640 WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833; 9106309 WOAWO 91 / 06309; and Felgner et al., Proc. Natl. Acad. Sci. USA 84: 7413-7414 (1987).
[0218] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including but not limited to second, third, or hybrid second / third generation lentiviruses, and recombinant lentiviruses of any generation designed to target a specific cell type or receptor (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human (See, for example, Ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880). Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequences may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guerin). BCG vectors are described in Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors, such as Salmonella typhi vectors, useful for therapeutic administration or immunization of neoantigens will be apparent to those skilled in the art from the description herein.
[0219] The means of administering the nucleic acid uses a minigene construct that encodes one or more epitopes. To generate a DNA sequence (minigene) encoding the selected CTL epitope for expression in human cells, the amino acid sequence of the epitope is reverse translated. A human codon usage table is used to guide the codon selection for each amino acid. The DNA sequences encoding these epitopes are directly adjacent to generate a continuous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse translated and included in the minigene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of the CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides that encode the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.
[0220] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is the reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). A variety of methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) can also be complexed with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.
[0221] Also disclosed herein is a method of producing a tumor vaccine comprising performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising multiple neoantigens or a subset of multiple neoantigens.
[0222] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector disclosed herein (e.g., a vector comprising at least one sequence encoding one or more neoantigens) can include culturing a host cell under conditions suitable for expressing the neoantigen or vector, the host cell comprising at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatographic techniques, electrophoretic techniques, immunological techniques, precipitation techniques, dialysis techniques, filtration techniques, concentration techniques, and chromatofocusing techniques.
[0223] The host cell can include Chinese Hamster Ovary (CHO) cells, NS0 cells, yeast, or HEK293 cells. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to at least one nucleic acid sequence encoding the neoantigen or vector. In certain embodiments, the isolated polynucleotide can be a cDNA.
[0224] VI. Identification of neoantigens VI.A. Identification of candidate neoantigens A research methodology for NGS analysis of tumor and normal exomes and transcriptomes is described and applied in the specific space of neoantigens. 6、14、15The examples below consider certain optimizations for greater sensitivity and specificity for identifying neoantigens in a clinical setting. These optimizations can be grouped into two areas: those related to laboratory processes and those related to NGS data analysis.
[0225] VI.A.1. Laboratory process optimization The process improvements presented herein support the concepts developed for reliable assessment of cancer driver genes in targeted cancer panels. 16 This addresses the challenges in high-precision neoantigen discovery from low tumor content and small volume clinical specimens by expanding the current approach to the whole-exome and whole-transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) unique average coverage across the tumor exome to detect mutations present at low mutant allele frequency due to either low tumor content or subclonal status. 2. Less than 5% of bases are covered by less than 100x to minimize missing potential neoantigens, e.g. a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for areas that are not sufficiently covered by targeting uniform coverage across the tumor exome. 3. Targeting uniform coverage across the normal exome with less than 5% of bases covered under 20x to minimize the chance of potential neoantigens remaining unclassified for somatic / germline status (and therefore not usable as TSNAs). 4. To minimize the total amount of sequencing required, sequence capture probes are designed only for the coding regions of genes, since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplemental probes for HLA genes that are GC-rich and not well captured by standard exome sequencing 18 . b. Elimination of genes predicted to produce few or no candidate neoantigens due to factors such as poor expression, suboptimal digestion by the proteasome, or atypical sequence characteristics. 5. Tumor RNA is also sequenced at high depth (>100M reads) to allow for variant detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples is sequenced using probe-based enrichment with the same or similar probes used to capture the exome in DNA. 19 It is extracted using
[0226] VI.A.2. Optimizing NGS Data Analysis Analytical method improvements address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically allow for customization relevant for identifying neoantigens in the clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment, as it contains multiple MHC region assemblies that better reflect population polymorphism, as opposed to earlier genome releases. 2. Various programs 5 A single variant caller by merging results from 20 Overcoming the limitations. a. Single nucleotide variants and indels are detected in tumor DNA, tumor RNA, and normal DNA with a range of tools including: Strelka 21 and Mutect 22 and programs based on comparison of tumor and normal DNA, such as; 23 , and UNCeqR, programs that incorporate tumor DNA, tumor RNA, and normal DNA. b. Indels are mutated in Strelka and ABRA 24This is determined by a program that performs local reassembly, such as C. Structural rearrangements are 25 or Breakseq 26 It is determined using specialized tools such as 3. To detect and prevent sample swapping, variant calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. Extensive filtering of artificial calls is performed, for example by: a. Removal of variants found in normal DNA, potentially with relaxed detection parameters in the case of low coverage, and permissive proximity criteria in the case of indels. b. Removal of variants due to poor mapping quality or poor base quality 27 . c. Removal of variants arising from re-emerging sequencing artifacts, even if not observed in the corresponding normal 27 Examples include variants that are found primarily on one strand. d. Removal of variants detected in a set of unrelated controls 27 . 5.seq2HLA 28 , ATHLETES 29 or Optitype, and also combining exome and RNA sequencing data 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing, such as long-read DNA sequencing. 30 or adaptation of methods for linking RNA fragments to maintain continuity 31 Includes. 6. Robust detection of nascent ORFs arising from tumor-specific splice variants using CLASS 32 , Bayesembler 33 , StringTie 34This is done by assembling transcripts from RNA-seq data using Cufflinks, or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate the entire transcripts from each experiment). 35 Although commonly used for this purpose, it frequently produces an incredibly large number of splice variants, many of which are much shorter than the full-length gene, and may not be able to recover a simple positive control, SpliceR, which can mutate the coding sequence and potentially nonsense-mediated decay mechanisms to reintroduce mutant sequences. 36 and MAMBA 37 Gene expression is determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels are determined using ASE 38 or HTSeq 39 Potential filtering steps include: a. Removal of candidate nascent ORFs that are thought to be poorly expressed. b. Removal of candidate nascent ORFs predicted to undergo nonsense-mediated decay (NMD). 7. Candidate neoantigens observed only in RNA (e.g., neo-ORFs) that cannot be directly verified as tumor-specific are classified as likely to be tumor-specific according to additional parameters, for example, by considering the following: a. Presence of cis-acting frameshift or splice site mutation support in tumor DNA only. b. The presence of confirmed trans-acting mutations in splicing factors in tumor DNA only. As an example, the gene that exhibited the most differential splicing in three independently published experiments with R625 mutant SF3B1 was 40 The second experiment examined uveal melanoma cell lines. 41, and a third study looked at breast cancer patients. 42 Nevertheless, there was a match. c. For novel splicing isoforms, the presence of confirmatory “novel” splice-junction reads in the RNASeq data. d. For a de novo rearrangement, the presence of conclusive evidence of exon-proximal reads in tumor DNA that are not present in normal DNA. e.GTEx 43 etc. from the gene expression compendium (i.e., making germline origin less likely). 8. Complementing reference genome alignment-based analyses by comparing tumor and normal reads (or k-mers derived from such reads) of assembled DNA to directly avoid alignment and annotation-based errors and artifacts (e.g., for somatic variants occurring near germline variants or repeat context indels).
[0227] In samples with polyadenylated RNA, the presence of viral and microbial RNA in RNA-seq data will be used to identify additional factors that may predict patient response using RNA CoMPASS. 44 or a similar method is used.
[0228] VI.B. HLA Peptide Isolation and Detection Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) methods after lysis and solubilization of tissue samples (55-58). The clarified lysates were used for HLA-specific IP.
[0229] Immunoprecipitation was performed using antibodies coupled to beads, where the antibodies are specific for HLA molecules. For pan-class I HLA immunoprecipitation, a pan-class I CR antibody is used, and for class II HLA-DR, an HLA-DR antibody is used. The antibodies are covalently attached to NHS-Sepharose beads during overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP (59, 60).
[0230] The clarified tissue lysate is added to the antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is stored for additional experiments, including additional IPs. The IP beads are washed to remove non-specific binding and the HLA / peptide complexes are eluted from the beads using standard techniques. Protein components are removed from the peptides using molecular weight spin columns or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and in some cases stored at -20°C prior to MS analysis.
[0231] The dried peptides are reconstituted in HPLC buffer suitable for reversed-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution in a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of peptide mass / charge (m / z) are collected at high resolution in an Orbitrap detector, and then MS2 low-resolution scans are collected in an ion trap detector after HCD fragmentation of selected ions. Additionally, MS2 spectra can be acquired using either CID or ETD fragmentation methods, or any combination of the three techniques to obtain a larger amino acid coverage of the peptide. MS2 spectra can also be measured with high resolution mass accuracy in an Orbitrap detector.
[0232] MS2 spectra from each analysis are searched against protein databases using Comet (61, 62) and peptide identifications are scored using Percolator (63-65).
[0233] VI.B.1. Study of MS detection limits for comprehensive HLA peptide sequencing Using the peptide YVYVADVAAK, what the limit of detection was was determined using various amounts of peptide loaded onto the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol. (Table 1) The results are shown in Figure 1F. These results indicate that the lowest limit of detection (LoD) was in the attomolar range (10 -18 ), a dynamic range spanning five orders of magnitude, and a signal-to-noise ratio in the low femtomole range (10 -15 ) appears to be sufficient for sequencing. TIFF2025065197000002.tif74169
[0234] VII. Presented Model VII.A. System Overview 2A is a schematic of an environment 100 for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for introducing a presentation identification system 160, which itself includes a presentation information store 165.
[0235] The presentation identification system 160 is a computer model or a system embodied in a computational system such as that discussed below with respect to FIG. 14 that receives a peptide sequence associated with a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the set of associated MHC alleles. This is useful in a variety of contexts. An example of one specific application of the presentation identification system 160 is to receive a nucleotide sequence of a candidate neoantigen associated with a set of MHC alleles from tumor cells of a patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the associated MHC alleles of the tumor and / or induce an immunogenic response in the immune system of the patient 110. Those candidate neoantigens that have a high likelihood as determined by the system 160 can be selected for inclusion in the vaccine 118, such that an anti-tumor immune response can be elicited from the immune system of the patient 110 that provided the tumor cells.
[0236] The presentation identification system 160 determines the presentation likelihood through one or more presentation models. Specifically, the presentation models generate a likelihood that a given peptide sequence will be presented for a set of associated MHC alleles, the likelihood being generated based on the presentation information stored in the storage device 165. For example, the presentation models may be TIFF2025065197000003.tif3128 is the set of HLA-A alleles on the cell surface of the sample. * 02:01, HLA-B * 07:02, HLA-B * 08:03, HLA-C * 01:04, HLA-A * 06:03, HLA-B * 01:04 may generate a likelihood of being presented. Presentation information 165 contains information about whether these peptides bind to various types of MHC alleles so that the peptides are presented by the MHC alleles, which is determined in the model according to the position of the amino acid in the peptide sequence. Based on presentation information 165, the presentation model can predict whether an unrecognized peptide sequence will be presented in combination with a related set of MHC alleles.
[0237] VII.B. Presentation information 2 illustrates a method of obtaining presentation information according to one embodiment. Presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. Allele interaction information includes information that influences presentation of peptide sequences that is dependent on the type of MHC allele. Allele non-interaction information includes information that influences presentation of peptide sequences that is independent of the type of MHC allele.
[0238] VII.B.1. Allelic interaction information The allele interaction information primarily includes identified peptide sequences that are known to be presented by one or more identified MHC molecules from humans, mice, etc. Of note, this may or may not include data obtained from tumor samples. The presented peptide sequences may be identified from cells expressing a single MHC allele. In this example, the presented peptide sequences are generally collected from a single allele cell line engineered to express a predetermined MHC allele and then exposed to a synthetic protein. The peptides presented on the MHC alleles are isolated by techniques such as acid elution and identified through mass spectrometry. Figure 2B shows the peptide sequences of a cell expressing a predetermined MHC allele HLA-A. * 01:01 Exemplary peptides presented above An example of this is shown in TIFF2025065197000004.tif3128, in which a peptide was isolated and identified through mass spectrometry. In this situation, the direct association between the presented peptide and the MHC protein to which it bound is conclusively known, since the peptide is identified through cells engineered to express a single, predefined MHC protein.
[0239] The presented peptide sequences may also be collected from cells expressing multiple MHC alleles. Typically in humans, six different types of MHC molecules are expressed on cells. Such presented peptide sequences may be identified from multi-allelic cell lines that have been engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal tissue samples or tumor tissue samples. In this particular example, MHC molecules can be immunoprecipitated from normal or tumor tissue. Peptides presented on multiple MHC alleles can also be isolated by techniques such as acid elution and identified through mass spectrometry. Figure 2C shows six exemplary peptides. TIFF2025065197000005.tif14153 is the identified MHC allele HLA-A * 01:01, HLA-A* 02:01, HLA-B * 07:02, HLA-B * 08:01, HLA-C * 01:03, and HLA-C * An example of this is presented above at 01:04, where a peptide was isolated and identified through mass spectrometry. In contrast to monoallelic cell lines, the bound peptide is isolated from the MHC molecule before it is identified, so the direct association between the presented peptide and the MHC protein to which it is bound may be unknown.
[0240] Allele interaction information can also include mass spectrometry ion current, which depends on both the concentration of peptide-MHC molecule complex and the ionization efficiency of peptide.Ionization efficiency varies from peptide to peptide in a sequence-dependent manner.In general, ionization efficiency varies from peptide to peptide over about two orders of magnitude, while the concentration of peptide-MHC complex varies over a larger range.
[0241] The allele interaction information can also include measurements or predictions of binding affinity between a given MHC allele and a given peptide. One or more affinity models can generate such predictions. For example, returning to the example shown in FIG. 1D, the representation 165 can include a peptide TIFF2025065197000006.tif3128 and allele HLA-A * 01:01 may include a predicted binding affinity value of 1000 nM between 01:01. Few peptides with IC50s greater than 1000 nM are presented by the MHC, with lower IC50 values increasing the probability of presentation.
[0242] The allele interaction information can also include a measure or prediction of the stability of the MHC complex. One or more stability models can generate such a prediction. More stable peptide-MHC complexes (i.e., complexes with a longer half-life) are more likely to be presented in high copy number on tumor cells and on antigen-presenting cells that encounter the vaccine antigen. For example, returning to the example shown in FIG. 2C, the presentation information 165 may include a measure or prediction of the stability of the molecule HLA-A. * A stability prediction of 1 hour half-life for 01:01 may be included.
[0243] Allele interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at a faster rate are more likely to be presented on the cell surface in high concentration.
[0244] Allele interaction information can also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with lengths between 8 and 15 peptides. 60-80% of presented peptides have a length of 9. A histogram of presented peptide lengths from several cell lines is shown in FIG. 5.
[0245] Allele interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoded peptide, and the presence or absence of specific post-translational modifications on the neoantigen-encoded peptide. The presence of a kinase motif influences the probability of post-translational modifications that may enhance or interfere with MHC binding.
[0246] Allelic interaction information can also include expression or activity levels of proteins, e.g., kinases, involved in post-translational modification processes (as measured or predicted by RNA-seq, mass spectrometry, or other methods).
[0247] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.
[0248] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual in question (e.g., as measured by RNA-seq or mass spectrometry): peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.
[0249] Allelic interaction information can also include the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals expressing that particular MHC allele.
[0250] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and thus presentation of a peptide by HLA-C is a priori less likely than presentation by HLA-A or HLA-B.
[0251] The allele interaction information can also include the protein sequence of a particular MHC allele.
[0252] Any of the MHC allele non-interacting information listed in the section below can also be modeled as MHC allele interacting information.
[0253] VII.B.2. Allelic non-interaction information The allele non-interaction information can include a C-terminal sequence adjacent to the neoantigen-encoded peptide in its source protein sequence. The C-terminal flanking sequence can affect the proteasome processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters the MHC allele on the surface of the cell. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and therefore the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include the C-terminal flanking sequence FOEIFNDKSLDKFJI of the presented peptide FJIEJFOESS, identified from the source protein of the peptide.
[0254] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same sample that provides mass spectrometry training data. As described later with respect to Figure 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, mRNA quantification measurements are identified from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).
[0255] The allele-non-interacting information can also include N-terminal sequences adjacent to the peptide within its source protein sequence.
[0256] The allele non-interaction information can also include the presence of protease cleavage motifs in peptides, which are optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and therefore less stable in cells, and therefore less likely to be presented.
[0257] Allele non-interaction information can also include the turnover rate of the source protein when measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation; however, this characteristic has low predictive power when measured in dissimilar cell types.
[0258] The allelic non-interaction information can also optionally include the length of the source protein, taking into account the specific splice variant ("isoform") that is most highly expressed in the tumor cells, as measured by RNA-seq or proteomic mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.
[0259] Allele non-interaction information can also include the expression level of proteasome, immunoproteasome, thymoproteasome, or other proteases in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Different proteasomes have different cleavage site preferences. More weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.
[0260] Allele non-interaction information can also include the expression of the peptide's source gene (e.g., as measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in tumor samples. Peptides from genes with higher expression are more likely to be presented. Peptides from genes with undetectable levels of expression can be eliminated from consideration.
[0261] Allelic non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to nonsense-mediated decay as predicted by a model of nonsense-mediated decay, e.g., a model from Rivas et al, Science 2015.
[0262] Allele-free interaction information can also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle. Genes that are expressed at low levels overall (as measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are presented than genes that are stably expressed at very low levels.
[0263] The allelic non-interacting information can also include a comprehensive catalog of properties of the source protein, such as those provided in uniProt or the PDB http: / / www.rcsb.org / pdb / home / home.do. These properties can include, among others, the secondary and tertiary structure of the protein, subcellular localization,11, Gene Ontology (GO) terms. Specifically, this information can contain annotations that operate at the level of the protein, such as the 5'UTR length, and annotations that operate at the level of specific residues, such as the helix motif at residues 300-310. These properties can also include turn motifs, sheet motifs, and disordered residues.
[0264] Allelic non-interacting information can also include features describing the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (eg, alpha helices versus beta sheets); alternative splicing.
[0265] The allelic non-interaction information can also include features describing the presence or absence of a presentation hotspot at the peptide's position in its source protein.
[0266] Allelic non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression levels of the source protein in those individuals and the effects of the various HLA types of those individuals).
[0267] Allelic non-interaction information can also include the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias.
[0268] Expression of various gene modules / pathways (not necessarily containing the source protein of the peptide) as measured by gene expression assays such as targeted panels such as RNASeq, microarrays, Nanostring, etc., or single / multiple genes representing gene modules measured by assays such as RT-PCR, that inform on the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs).
[0269] Allele non-interaction information can also include the copy number of the peptide's source gene in the tumor cell. For example, a peptide derived from a gene that is subject to homozygous deletion in the tumor cell can be assigned a presentation probability of zero.
[0270] The allele non-interaction information can also include the probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP or that bind with higher affinity to TAP are more likely to be presented.
[0271] Allele non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, immunohistochemistry). Higher TAP expression levels increase the probability of presentation of all peptides.
[0272] Allelic non-interaction information can also include the presence or absence of tumor mutations, including but not limited to: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3. ii. in genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery that are affected by loss-of-function mutations in the tumor have a reduced probability of presentation.
[0273] Presence or absence of functional germline polymorphisms, including but not limited to: i. in genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome).
[0274] The allelic non-interaction information can also include tumor type (eg, NSCLC, melanoma).
[0275] The allele non-interaction information can also include the known functionality of the HLA allele, e.g., as reflected by the HLA allele suffix. * The N suffix in 24:09N indicates a null allele that is not expressed and therefore unlikely to present an epitope; the complete HLA allele suffix nomenclature is described in https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.
[0276] Allelic non-interaction information can also include clinical tumor subtype (e.g., squamous cell lung cancer vs. non-squamous).
[0277] The allele non-interaction information can also include smoking history.
[0278] Allele non-interaction information can also include a history of sunburn, sun exposure, or exposure to other mutagens.
[0279] The allelic non-interaction information can also include regional expression of the peptide source genes in the relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in the relevant tumor types are more likely to be represented.
[0280] The allelic non-interaction information can also include the frequency of a mutation in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.
[0281] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation can also include the mutation annotation (e.g., missense, read-through, frameshift, fusion, etc.) or whether the mutation is predicted to result in nonsense-mediated decay (NMD). For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in a decrease in mRNA translation, which reduces the probability of presentation.
[0282] VII.C. Presentation Identification System 3 is a high-level block diagram illustrating computer logic components of a presentation identification system 160, according to one embodiment. In this exemplary embodiment, the presentation identification system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. The presentation identification system 160 is also comprised of a training data store 170 and a presentation model store 175. Some embodiments of the model management system 160 have different modules than those described herein. Similarly, functionality may be distributed among the modules in a manner different than that described herein.
[0283] VII.C.1. Data Management Module The data management module 312 generates sets of training data 170 from the representation information 165. Each training data set contains a number of data examples, each of which includes at least a representation of a peptide sequence p i and the peptide sequence p i one or more relevant MHC alleles in combination with i and the dependent variable y , which represents information that the presentation identification system 160 is interested in predicting new values of the independent variables. i and the independent variable z i Contains a set of:
[0284] In one particular embodiment, which will be referred to throughout the remainder of this specification, the dependent variable y i is peptide p i but one or more associated MHC alleles i However, in other embodiments, the dependent variable y i is a function of the presentation identification system 160 for determining the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that one is interested in predicting. For example, in another embodiment, the dependent variable y i may also be a numerical value indicating the mass analysis ion current determined for the example data.
[0285] Peptide sequence p for data example i i is k i is a sequence of amino acids, k i can vary within a range among data examples i. For example, the range can be 8 to 15 for MHC class I, or 9 to 30 for MHC class II. In one specific embodiment of the system 160, all peptide sequences p in the training data set are i can have the same length, e.g., 9. The number of amino acids in a peptide sequence can vary depending on the type of MHC allele (e.g., MHC allele in humans, etc.). MHC allele a for data example i iis the peptide sequence p i Indicates whether it exists in combination with
[0286] The data management module 312 also calculates the peptide sequences p i and bound MHC allele a i Together with the binding affinity b i and stability i For example, the training data 170 may include a set of allele interaction variables, such as a predicted value of i And, a i The predicted binding affinity b between each of the bound MHC molecules shown in i As another example, the training data 170 may include a i The predicted stability value s for each of the MHC alleles shown in i may contain
[0287] The data management module 312 also receives the peptide sequence p i along with allele-free interacting variables such as C-terminal flanking sequences and mRNA quantification measurements. i It may also include.
[0288] The data management module 312 also identifies peptide sequences that are not presented by MHC alleles to generate the training data 170. Generally, this involves identifying a "longer" sequence of the source protein that contains the peptide sequence to be presented prior to presentation. If the presentation information contains an engineered cell line, the data management module 312 identifies a set of peptide sequences in the synthetic protein to which the cell was exposed that were not presented on the MHC alleles of the cell. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein from which the presented peptide sequence originated and identifies a set of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.
[0289] The data management module 312 also artificially generates peptides with random sequences of amino acids and identifies the generated sequences as peptides that are not presented on MHC alleles. This can be accomplished by randomly generating peptide sequences, allowing the data management module 312 to easily generate a large amount of synthetic data for peptides that are not presented on MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, it is highly likely that synthetically generated peptide sequences are not presented by MHC alleles, even if they are included in proteins processed by the cell.
[0290] FIG. 4 illustrates an exemplary set of training data 170A, according to one embodiment. Specifically, the first three data examples in training data 170A are alleles HLA-C * Monoallelic cell lines containing 01:03 and three different peptide sequences The peptide presentation information from TIFF2025065197000007.tif6128 is shown. The fourth data example in training data 170A shows the allele HLA-B * 07:02, HLA-C * 01:03, HLA-A * Multi-allelic cell lines containing 01:01 and peptide sequences The peptide information from TIFF2025065197000008.tif4128 is shown. The first data example shows the peptide sequence TIFF2025065197000009.tif4128 is the allele HLA-C * 01:03 indicates that the peptide sequences were not presented by the data management module 312. As discussed in the previous two paragraphs, the peptide sequences may be randomly generated by the data management module 312 or may be identified from the source protein of the presented peptide. The training data 170A also includes a predicted binding affinity of 1000 nM and a predicted half-life stability of 1 hour for the peptide sequence-allele pair. The C-terminal flanking sequence of TIFF2025065197000010.tif3128, and 10 2 It also includes allele-noninteracting variables, such as FPKM mRNA quantification measurements. The fourth data example is a peptide sequence. TIFF2025065197000011.tif4128 is the allele HLA-B * 07:02, HLA-C * 01:03, or HLA-A * 01:01. The training data 170A also includes predicted binding affinity and stability values for each of the alleles, as well as the C flanking sequences of the peptides and mRNA quantification measurements for the peptides.
[0291] VII.C.2. Coding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more representation models. In one embodiment, the encoding module 314 one-hot encodes sequences (e.g., peptide sequences or C-terminal flanking sequences) for a predetermined 20-letter amino acid alphabet. Specifically, k i A peptide sequence having amino acids p i is 20·k i p i 20·(j-1)+1 , p i 20·(j-1)+2 , ..., p i 20·j A single element in has a value of 1. The remaining elements have a value of 0. As an example, for a given alphabet {A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y}, the 3 amino acid peptide sequence EAF of data example i is a 60-element row vector The C-terminal flanking sequence c i, and the protein sequence for the MHC allele d h , and other sequence data in the presentation information can similarly be coded as above.
[0292] If the training data 170 contains sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equivalent length by adding PAD characters to extend the predetermined alphabet. For example, this may be done by left padding the peptide sequences with PAD characters until the length of the peptide sequence reaches the peptide sequence with the longest length in the training data 170. Thus, if the peptide sequence with the longest length is k 最大 If the sequence has amino acids, the encoding module 314 encodes each sequence as (20+1) k 最大 We express it numerically as a row vector of elements. For example, let us consider the extended alphabet {PAD, A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y} and k 最大 For a maximum amino acid length of =5, the same exemplary peptide sequence EAF of 3 amino acids is represented as a 105-element row vector The C-terminal flanking sequence c i or other sequence data can be similarly encoded as above. Thus, the peptide sequence p i or c i Each argument or row in represents the occurrence of a particular amino acid at a particular position in the sequence.
[0293] Although the above method for encoding sequence data has been described with respect to sequences having amino acid sequences, the method can be extended to other types of sequence data as well, such as, for example, DNA or RNA sequence data.
[0294] The encoding module 314 also encodes one or more MHC alleles a for data instance i. iinto a row vector of m elements, where each element h=1, 2, ..., m corresponds to a uniquely specified MHC allele. The element corresponding to the MHC allele specified for data instance i has a value of 1. The remaining elements have a value of 0. As an example, * 01:01, HLA-C * 01:08, HLA-B * 07:02, HLA-C * 01:03}, allele HLA-B for data example i corresponding to a multi-allele cell line * 07:02 and HLA-C * 01:03 is a four-element row vector a i = [0 0 1 1], and a3 i =1 and a4 i = 1. An example with four identified MHC allele types is described herein, but the number of MHC allele types can be hundreds or thousands in practice. As previously discussed, each data instance i typically contains a peptide sequence p i Contains up to six different MHC allele types associated with
[0295] The encoding module 314 also generates a label y for each data instance i. i We code,x,as a binary variable with values from the set {0, 1}, where a value of 1 indicates that the peptide x i However, the associated MHC allele a i A value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele a i The dependent variable y i If represents the mass analysis ion current, the encoding module 314 may additionally scale the value using various functions, such as a log function having a range of [-∞, ∞] for ion current values between [0, ∞].
[0296] The coding module 314 encodes the peptide p iand the allelic interaction variable x for the associated MHC allele h h i A pair of alleles x may be represented as a row vector in which the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], and b h i is peptide p i and the predicted binding affinity for the associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).
[0297] In one example, the encoding module 314 encodes the measured or predicted values for binding affinity as a function of the allele interaction variable x h i The binding affinity information is expressed by incorporating
[0298] In one example, the encoding module 314 encodes the measured or predicted values for binding stability as allele interaction variables x h i By incorporating the information into the
[0299] In one example, the encoding module 314 encodes the measured or predicted values for the binding on-rate as a function of the allele interaction variable x h i The binding on-rate information is represented by incorporating
[0300] In one example, the encoding module 314 encodes the peptide length in a vector TIFF2025065197000014.tif4128(in the formula, TIFF2025065197000015.tif3128 is the index function, L k is peptide p k The vector T k Let x be the allele interaction variable h i can be included in
[0301] In one example, the encoding module 314 encodes the RNA-seq-based expression levels of the MHC alleles as a function of the allele interaction variable x h i By incorporating this information into the RNA expression information of MHC alleles,
[0302] Similarly, the encoding module 314 encodes the allele non-interacting variable w i can be represented as a row vector in which the numerical representations of the allele-noninteraction variables are concatenated one after the other. For example, w i is [c i ] or [c i m i w i ], or a row vector equivalent to w i is peptide p i Quantitative measurements of mRNA related to C-terminal flanking sequences and peptides i Alternatively, one or more combinations of allele-non-interacting variables may be stored individually (e.g., as individual vectors or matrices).
[0303] In one example, the encoding module 314 encodes the turnover rate or half-life as a function of the allele non-interacting variable w i represents the turnover rate of the source protein for the peptide sequence.
[0304] In one example, the encoding module 314 encodes the protein length in terms of the allele non-interacting variable w i represents the length of the source protein or isoform.
[0305] In one example, the encoding module 314 i , β2 i , β5 i The mean expression of immunoproteasome-specific proteasome subunits, including the subunits, was calculated using the allele-free interaction variable w i Incorporation into the IL-1 protein indicates activation of the immunoproteasome.
[0306] In one example, the encoding module 314 may generate RNA-seq abundances of the peptides (quantified in units of FPKM, TPM by techniques such as RSEM) or source proteins of the genes or transcripts of the peptides, and correlate the abundances of the source proteins with the allele-non-interaction variable w i This is expressed by incorporating it into
[0307] In one example, the coding module 314 calculates the probability that the transcript from which the peptide originates will undergo nonsense-mediated decay (NMD), e.g., as estimated by the model in Rivas et al. Science, 2015, and calculates this probability as a function of an allele-noninteraction variable w i This is expressed by incorporating it into
[0308] In one example, the encoding module 314 represents the activation status of a gene module or pathway assessed via RNA-seq, for example, by quantifying the expression of genes in the pathway, e.g., in units of TPM using RSEM, for each of the genes in the pathway, and then computing a summary statistic, e.g., the mean, across the genes in the pathway. The mean is calculated using the allele non-interaction variable w i can be incorporated into.
[0309] In one example, the encoding module 314 encodes the copy number of the source gene by dividing the copy number by the allele non-interacting variable w i This is expressed by incorporating it into
[0310] In one example, the encoding module 314 encodes a measured or predicted TAP binding affinity (e.g., in nanomolar units) relative to the allele-non-interacting variable w i The TAP binding affinity is expressed by including
[0311] In one example, the encoding module 314 encodes the TAP expression levels measured by RNA-seq (and quantified, for example, by RSEM in units of TPM) into the allele-non-interacting variable w i The expression level of TAP is represented by including
[0312] In one example, the encoding module 314 encodes the tumor mutation as a non-allele interaction variable w i A vector of indicator variables in (i.e., peptide p k d if derived from a sample with a KRAS G12D mutation k = 1, otherwise 0).
[0313] In one example, the encoding module 314 encodes germline polymorphisms in antigen presenting genes as vectors of indicator variables (i.e., peptide p k d if derived from a sample with a specific germline polymorphism in TAP k = 1). We use these indicator variables as the allele non-interaction variables w i can be included in
[0314] In one example, the encoding module 314 represents the tumor types as a one-hot coded vector of length 1 over an alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot coded variables are then combined into an allele-non-interaction variable w i can be included in
[0315] In one example, the encoding module 314 represents MHC allele suffixes by processing four-digit HLA alleles with various suffixes. For example, HLA-A * 24:09N is a model for HLA-A * 24:09 is considered a different allele. Alternatively, since HLA alleles ending in an N suffix are not expressed, the probability of presentation by an MHC allele with an N suffix can be set to zero for all peptides.
[0316] In one example, the encoding module 314 represents the tumor subtype as a one-hot coded vector of length 1 for an alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot coded variables are combined with the allele non-interaction variables w i can be included in
[0317] In one example, the encoding module 314 encodes smoking history as a function of the allele non-interaction variable w i A binary indicator variable (d if the patient has a smoking history) can be included in k = 1, 0 otherwise). Alternatively, smoking history can be coded as a one-hot coded variable of length 1 for the alphabet of smoking severity. For example, smoking status can be assessed on a scale of 1 to 5, with 1 indicating non-smoker and 5 indicating current heavy smoker. Because smoking history is primarily relevant for lung tumors, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a smoking history and the tumor type is lung tumor, and zero otherwise.
[0318] In one example, the encoding module 314 encodes sunburn history as a function of the allele non-interacting variable w i A binary indicator variable (d if the patient has a history of severe sunburn) can be included in k= 1, otherwise 0). Because severe sunburn is primarily associated with melanoma, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and zero otherwise.
[0319] In one example, the coding module 314 represents the distribution of expression levels of a particular gene or transcript for each gene or transcript in the human genome as a summary statistic (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, the expression level of peptide p in samples with tumor type melanoma is expressed as k Regarding peptide p k The measured gene or transcript expression levels of the genes or transcripts of origin of alleles are compared with the allele-non-interacting variable w i Not only can it be included in the peptide p in melanoma as measured by TCGA, k The mean and / or median gene or transcript expression of the genes or transcripts of a given source may also be included.
[0320] In one example, the encoding module 314 represents the mutation types as one-hot coded variables of length 1 for an alphabet of mutation types (e.g., missense, frameshift, NMD-induced, etc.). These one-hot coded variables are grouped together into a set of allele-non-interacting variables w i can be included in
[0321] In one example, the encoding module 314 encodes the protein level characteristics of the protein as values of the source protein annotation (e.g., 5' UTR length) and the allele non-interacting variable w i In another example, the coding module 314 encodes the peptide p k The residue-level annotation of the source protein for peptide p kis equal to 1 if it overlaps with the helical motif, and 0 otherwise; or k We define an indicator variable, which is equal to 1 if is completely contained within the helical motif, as the allele non-interaction variable w i In another example, the peptide p k The property that represents the proportion of residues in the allele-free variable w i can be included in
[0322] In one example, the encoding module 314 encodes the types of proteins or isoforms in the human proteome into an index vector o having a length equal to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is peptide p k is 1 if comes from protein i, and 0 otherwise.
[0323] The coding module 314 also encodes the peptide p i and the variable z for the associated MHC allele h i The entire set of alleles is treated as the allele interaction variable x i and the allele non-interaction variable w i For example, the encoding module 314 may represent z h i [x h i w i ] or [w i x h i ] can be represented as a row vector equivalent to
[0324] VIII. Training Module The training module 316 constructs one or more presentation models that generate a likelihood that a peptide sequence will be presented by an MHC allele associated with the peptide sequence. kand the peptide sequence p k MHC alleles associated with a k Given a set of peptide sequences, each proposed model k However, the associated MHC allele a k , which indicates the likelihood that one or more of k Generate.
[0325] VIII.A. Overview The training module 316 constructs one or more representation models based on a training data set stored in storage 170, which is generated from the representation information stored in 165. Generally, regardless of the specific type of representation model, all of the representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function TIFF2025065197000016.tif4128 represents the dependent variable y for one or more data examples S in the training data 170. i∈S and the estimated likelihood u for a data example S generated by the proposed model. i∈S In one particular embodiment, which will be mentioned throughout the remainder of this specification, the loss function (y i∈S , u i∈S ; θ) is the negative log-likelihood function given by equation (1a) as follows: However, in practice, another loss function may be used. For example, if a prediction is made for mass spectrometry ion current, the loss function is the mean squared loss given by Equation 1b as follows: TIFF2025065197000018.tif10128
[0326] The proposed model may be a parametric model, where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function (y i∈S , u i∈SThe various parameters of the proposed model of parametric type that minimizes θ(θ;θ) are determined through a gradient-based numerical optimization algorithm, such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the proposed model can be a non-parametric model, where the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.
[0327] VIII.B. Allele-by-Allele Model The training module 316 may build a presentation model to predict the presentation likelihood of a peptide on a per-allele basis. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele.
[0328] In one embodiment, the training module 316 includes: TIFF2025065197000019.tif7128 identifies peptide p for a specific allele h k The estimated presentation likelihood u k where the peptide sequence x h k is peptide p k and the coded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, which for convenience of description will be referred to as a transformation function throughout this specification. h (·) is an arbitrary function, which for ease of description will be referred to throughout this specification as the dependence function, and the parameter θ determined for the MHC allele h h Based on the set of allele interaction variables x h k Generate a dependency score for each MHC allele h using the parameters θ h The set of values is θ h where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele h.
[0329] Dependency function g h (x h k ;θ h ) is the output of the MHC allele h that has at least one allele interaction characteristic x h k Based on, and in particular, the peptide p k The dependency score for MHC allele h indicates whether MHC allele h presents the corresponding neoantigen based on the amino acid position of the peptide sequence of p. k The transformation function f(·) transforms the input, more specifically, g in this example. h (x h k ;θ h ) is used to calculate the dependency score for peptide p k into an appropriate value indicating the likelihood that a gene will be presented by an MHC allele.
[0330] In one particular embodiment referred to throughout the remainder of this specification, f(·) is a function with range in [0, 1] for the appropriate domain range. In one example, f(·) is As another example, f(·) is also given by the expit function given by TIFF2025065197000020.tif10128 for values in the domain z greater than or equal to 0: It can also be the hyperbolic tangent function given by TIFF2025065197000021.tif4128. Alternatively, if the prediction is made for mass analysis ion currents with values outside the range [0, 1], f(·) can be any function, for example the identity function, the exponential function, the log function, etc.
[0331] Therefore, the peptide sequence p k The allele-specific likelihood that ���� will be presented by MHC allele h is the dependence function g h (·) is the peptide sequence p kto generate a corresponding dependency score. k may be transformed by a transformation function f(·) to produce the per-allele likelihood that h will be presented by MHC allele h.
[0332] VIII.B.1 Dependence Functions for Allelic Interaction Variables In one particular embodiment mentioned throughout this specification, the dependence function g h (·) is x h k Each allele interaction variable in is expressed as a function of the parameter θ determined for the relevant MHC allele h. h with the corresponding parameters in the set This is an affine function given by TIFF2025065197000022.tif5128.
[0333] In another specific embodiment mentioned throughout this specification, the dependence function g h (·) is a network model NN with a set of nodes arranged in one or more layers. h (·), The network function is given by TIFF2025065197000023.tif5128. The nodes are connected with the parameters θ h A node may be connected to other nodes through connections, each having an associated parameter in a set of activation functions. The value at one particular node may be expressed as the sum of the values of the nodes connected to the particular node, weighted by the associated parameters mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and process data with amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how this interaction affects peptide presentation.
[0334] In general, the network model NN h (·) may be structured as a feedforward network, such as an artificial neural network (ANN), a convolutional neural network (CNN), or a deep neural network (DNN), and / or a recurrent network, such as a long short-term memory network (LSTM), a bidirectional recurrent network, or a deep bidirectional recurrent network.
[0335] In one example, which will be mentioned throughout the remainder of this specification, each MHC allele in h=1,2, ...,m is associated with a separate network model, NN h (·) denotes the output from the network model related to MHC allele h.
[0336] FIG. 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h=3. As shown in FIG. 5, the network model NN3(·) for MHC allele h=3 includes three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2), ..., θ3(10). The network model NN3(·) includes three allele interaction variables x3 for MHC allele h=3. k (1), x3 k (2), and x3 k Receive input values for (3) (individual data examples, including encoded polypeptide sequence data and any other training data used) and generate the value NN3(x3 k ) to output.
[0337] In another example, the identified MHC alleles h=1,2, ..., m are modeled using a single network model NN H (·), and NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θh may correspond to a set of parameters for a single network model, and thus the parameters θ h A set of can be shared by all MHC alleles.
[0338] FIG. 6A shows an exemplary network model NN shared by MHC alleles h=1,2, ...,m. H As shown in FIG. 6A, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. The network model NN3(·) has an allele interaction variable x3 for MHC allele h=3. k , and the value NN3(x3 k ) to output m values.
[0339] In yet another example, a single network model NN H (·) is the allele interaction variable x for MHC allele h h k and the encoded protein sequence d h In such an example, the parameter θ h The set of parameters θ may again correspond to the set of parameters for a single network model, and thus the parameters θ h The set of NN may be shared by all MHC alleles. h (·) is the input [x h k d h ], a single network model NN H Such a network model is advantageous because it can correctly predict peptide presentation probabilities for MHC alleles that were unknown in the training data simply by identifying their protein sequences.
[0340] FIG. 6B shows an exemplary network model NN shared by MHC alleles. H As shown in FIG. 6B, the network model NN H (·) takes as input the allele interaction variables and protein sequence of MHC allele h=3 and calculates the dependency score NN3 (x3 k ) to output.
[0341] In yet another example, the dependency function g h (·)teeth, TIFF2025065197000024.tif5128, where g' h (x h k ;θ' h ) is the parameter θ' h a bias parameter θ in the set of parameters for allele interaction variables of the MHC alleles, which represents the baseline probability of presentation for the MHC allele h. h 0 This is accompanied by:
[0342] In another embodiment, the bias parameter θ h 0 may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ for the MHC allele h h 0 is θ 遺伝子(h) 0 and gene (h) is the gene family for MHC allele h. For example, MHC allele HLA-A * 02:01, HLA-A * 02:02, and HLA-A * 02:03 may be assigned to the "HLA-A" gene family, and the bias parameters θ for each of these MHC alleles h 0 may be shared.
[0343] As an example, going back to equation (2), the affine dependency function gh (·) was used to identify peptide p by MHC allele h = 3 among m = 4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000025.tif5128 can be generated by k are the allele interaction variables identified for MHC allele h=3, and θ3 is the set of parameters determined for MHC allele h=3 through loss function minimization.
[0344] As another example, a separate network transformation function g h (·) was used to identify peptide p by MHC allele h = 3 among m = 4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000026.tif5128 can be generated by k are the allele interaction variables identified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.
[0345] FIG. 7 shows the correlation coefficients of peptide p associated with MHC allele h=3 using the exemplary network model NN3(·). k As shown in Figure 7, the network model NN3(·) is a set of allele interaction variables x3 for MHC allele h=3. k Receives and outputs NN3(x3 k The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0346] VIII.B.2. Per-allele with non-allele interaction variables In one embodiment, the training module 316 incorporates allele non-interacting variables to: Peptide p by TIFF2025065197000027.tif7128 k The estimated presentation likelihood u kwhere w k is peptide p k means the coded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interaction variable w Based on the set of allele non-interaction variables w k Specifically, the parameter θ h and the parameters θ for the allele-noninteracting variables w The set of values of θ h and θ w where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele.
[0347] Dependency function g w (w k ;θ w ) output is the probability of peptide p by one or more MHC alleles based on the influence of allele-non-interacting variables. k For example, the dependency score for an allele-non-interacting variable is k The C-terminal flanking sequences and peptide p, which are known to positively affect the presentation of k If peptide p is bound, it may have a high value k The C-terminal flanking sequences and peptide p, which are known to negatively affect the presentation of k If bound, it may have a low value.
[0348] According to equation (8), the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is the function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. w(·) is also applied to the coded version of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. Both scores are combined, and the combined score is the probability of a peptide sequence p being associated with MHC allele h. k are transformed by a transformation function f(·) to produce the allele-specific likelihoods that will be presented.
[0349] Alternatively, the training module 316 may include a training module 316 that performs a training on the allele non-interaction variable w k Let x be the allele interaction variable. h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF2025065197000028.tif7128.
[0350] VIII.B.3 Dependence Functions for Allelic Non-Interacting Variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g w (·) is an affine function, or a separate network model for the allele non-interaction variables w k It may be a network function related to
[0351] Specifically, the dependency function g w (·) is w k The allele-free interaction variables in are the parameters θ w with the corresponding parameters in the set This is an affine function given by TIFF2025065197000029.tif5128.
[0352] Dependency function g w (·) also represents the parameter θ w The network model NN has relevant parameters in the set w (·), This is the network function given by TIFF2025065197000030.tif5128.
[0353] In another example, the dependence function g for the allele-non-interacting variables w (·)teeth, TIFF2025065197000031.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of m k is peptide p k is the mRNA quantification measurement for θ(·), h(·) is a function that transforms the quantification measurement, and θ w m is a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measurement to generate a dependency score for the mRNA quantification measurement. In one particular embodiment mentioned throughout the remainder of this specification, h(·) is a log function, although in practice h(·) can be any one of a variety of different functions.
[0354] In yet another example, the dependence function g for the allele-non-interacting variables w (·)teeth, TIFF2025065197000032.tif6128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w affine functions, network functions, etc. with a set of k is peptide p k is the above indicator vector representing proteins and isoforms in the human proteome, and θ w o is the set of parameters in the set of parameters for the allele non-interaction variables that are combined with the indicator vector.k and parameter θ w o If the dimension of the set of is significantly higher, A parameter regularization term such as TIFF2025065197000033.tif4128 (where ||·|| represents the L1 norm, L2 norm, a combination, etc.) can be added to the loss function when determining the value of the parameter. The optimal value of the hyperparameter λ can be determined through an appropriate method.
[0355] As an example, going back to equation (8), the affine transformation function g h (·), g w (·) was used to identify peptide p by MHC allele h = 3 among m = 4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000034.tif5128, where w k is peptide p k are allele-free interaction variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0356] As another example, the network transformation function g h (·), g w (·) was used to identify peptide p by MHC allele h = 3 among m = 4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000035.tif5128, where w k is peptide p k are allele interaction variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0357] FIG. 8 shows exemplary network models NN3(·) and NN w Peptide p associated with MHC allele h=3 using (·) kAs shown in Figure 8, the network model NN3(·) is a set of allele interaction variables x3 for MHC allele h=3. k Receives and outputs NN3(x3 k ) is generated. w (·) is peptide p k The allele non-interaction variable w k Receives and outputs NN w (w k The outputs are then combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0358] VIII.C. Multi-allele models The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting where more than one MHC allele is present. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof.
[0359] VIII.C.1. Example 1: Maximum per allele model In one embodiment, the training module 316 trains peptides p associated with a set of MHC alleles H. k The estimated presentation likelihood u k is the presentation likelihood determined for each of the MHC alleles h in set H determined based on cells expressing a single allele, as described above in conjunction with equations (2)-(11). We model the likelihood of the proposed feature as a function of TIFF2025065197000036.tif5128. Specifically, k teeth, In one embodiment, the function is a maximum function, as shown in equation (12), where the presented likelihood u k can be determined as the maximum of the presentation likelihood for each MHC allele h in set H. TIFF2025065197000038.tif5128
[0360] VIII.C.2. Example 2.1: Sum Function Model In one embodiment, the training module 316 trains peptide p k The estimated presentation likelihood u k of, Modeled by TIFF2025065197000039.tif15128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles H associated with x h k is peptide p k and the coded allele interaction variables for the corresponding MHC alleles. h The set of values is θ h The dependence function g can be determined by minimizing a loss function for each example in a subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependence function g introduced above in Section VIII.B.1. h The composition may be in any of the following forms:
[0361] According to equation (13), the peptide sequence p k The likelihood that a gene will be presented by one or more MHC alleles h is given by the dependence function g h (·) for each of the MHC alleles H, k to generate a corresponding score for the allele interaction variables. The scores for each MHC allele h are combined to generate a corresponding score for the peptide sequence p k is transformed by a transformation function f(·) to produce the presentation likelihood that α will be presented by the set of MHC alleles H.
[0362] The model presented in equation (13) is that for each peptide p k It differs from the allele-by-allele model of equation (2) in that the number of associated alleles for a can be greater than 1. In other words, h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0363] As an example, the affine transformation function g h (·) was used to identify peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000040.tif5128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0364] As another example, the network transformation function g h (·), g w (·) was used to identify peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000041.tif5128, where NN2(·), NN3(·) are the network models specified for MHC alleles h=2, h=3, and θ2, θ3 are the sets of parameters determined for MHC alleles h=2, h=3.
[0365] FIG. 9 shows the correlation between peptide p associated with MHC alleles h=2, h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) is a set of allele interaction variables x2 for MHC allele h=2.k Receives and outputs NN2(x2 k ) and the network model NN3(·) is the allele interaction variable x3 for MHC allele h=3. k Receives and outputs NN3(x3 k The outputs are then combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0366] VIII.C.3. Example 2.2: Sum Function Model with Allelic Non-Interacting Variables In one embodiment, the training module 316 incorporates allele non-interacting variables to: Peptide p by TIFF2025065197000042.tif13130 k The estimated presentation likelihood u k where w k is peptide p k Specifically, the parameters θ for each MHC allele h h and the parameters θ for the allele-noninteracting variables w The set of values of θ h and θ w The dependence function g can be determined by minimizing a loss function for each example in a subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependence function g introduced above in Section VIII.B.3. w The composition may be in any of the following forms:
[0367] Thus, according to equation (14), one or more MHC alleles H can bind to a peptide sequence p k The likelihood of being presented is the function g h (·) for each of the MHC alleles H, kto generate corresponding dependency scores for allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores are combined and the combined score is used to estimate the association between the peptide sequence p and the MHC allele H. k is transformed by a transformation function f(·) to produce the presentation likelihood that
[0368] In the model presented in equation (14), each peptide p k The number of related alleles for a can be greater than 1. h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0369] As an example, the affine transformation function g h (·), gw(·), and peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000043.tif5128, where w k is peptide p k are allele-free interaction variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0370] As another example, the network transformation function g h (·), g w (·) was used to identify peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000044.tif5128, where w k is peptide p k are allele interaction variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0371] FIG. 10 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h=2 and h=3 using (·) k As shown in Figure 10, the network model NN2(·) is a set of allele interaction variables x2 for MHC allele h=2. k Receives and outputs NN2(x2 k The network model NN3(·) generates the allele interaction variable x3 for MHC allele h=3. k Receives and outputs NN3(x3 k ) is generated. w (·) is peptide p k The allele non-interaction variable w k Receives and outputs NN w (w k The outputs are then combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0372] Alternatively, the training module 316 may include the allele non-interaction variable w k Let x be the allele interaction variable. h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF2025065197000045.tif15128.
[0373] VIII.C.4. Example 3.1: Model with Implicit Per-Allele Likelihood In another embodiment, the training module 316k The estimated presentation likelihood u k of, Modeled by TIFF2025065197000046.tif9132, where element a h k is the peptide sequence p k 1 for multiple MHC alleles h ∈ H associated with u' k h is the implicit allele-wise presentation likelihood for MHC allele h, and vector v has elements v h A h k ·u' k h where v is a vector corresponding to, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values of the input into a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, although it will be appreciated that in other embodiments s(·) may be any function, such as a maximum function. A set of values for the parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i are each example in a subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.
[0374] The presentation likelihood in the presentation model of equation (17) is the likelihood that each peptide p is presented by an individual MHC allele h. k The implicit allele presentation likelihood u' corresponds to the likelihood that k h The implicit per-allele likelihood differs from the per-allele presentation likelihood of Section VIII.B in that parameters for the implicit per-allele likelihood can be learned from a multi-allelic setting, in addition to a single-allelic setting, where the direct association between the presented peptide and the corresponding MHC allele is unknown. Thus, in the multi-allelic setting, the presentation model is k Not only can we estimate whether peptide p is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are presented by peptide p kThe individual likelihoods that TIFF2025065197000047.tif5128 can also be provided. The advantage of this is that the presented model can generate implicit likelihoods without training data for cells expressing a single MHC allele.
[0375] In one particular embodiment that will be mentioned throughout the remainder of this specification, r(·) is a function having the range [0, 1]. For example, r(·) is the clip function: r(z)=min(max(z,0), 1) and the minimum value between z and 1 is the presented likelihood u k In another embodiment, r(·) is selected as: r(z)=tanh(z) where the domain z is 0 or greater.
[0376] VIII.C.5. Example 3.2: Sum of Functions Model In one particular embodiment, s(·) is a summation function, and the presentation likelihood is given by summing the implied per-allele presentation likelihoods. TIFF2025065197000048.tif18128
[0377] In one embodiment, the implicit per-allele presentation likelihood for MHC allele h is calculated as: The proposed likelihood is generated by TIFF2025065197000049.tif7128. As estimated by TIFF2025065197000050.tif15128.
[0378] According to equation (19), one or more MHC alleles H bind to a peptide sequence p k The likelihood of being presented is the function g h (·) for each of the MHC alleles H, kto generate the corresponding dependency scores for the allele interaction variables. Each dependency score can be generated by first applying k h The allele likelihood u' is transformed by the function f(·) to generate k h are combined and a clipping function is applied to the combined likelihood to clip the values into the range [0, 1] to obtain the peptide sequence p k A presentation likelihood can be generated that g will be presented by a set of MHC alleles H. h is the dependence function g introduced above in Section VIII.B.1. h The composition may be in any of the following forms:
[0379] As an example, the affine transformation function g h (·) was used to identify peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000051.tif7128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0380] As another example, the network transformation function g h (·), g w (·) was used to identify peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000052.tif7128, where NN2(·), NN3(·) are the network models specified for MHC alleles h=2, h=3, and θ2, θ3 are the sets of parameters determined for MHC alleles h=2, h=3.
[0381] FIG. 11 shows the correlation between peptide p associated with MHC alleles h=2, h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) is a set of allele interaction variables x2 for MHC allele h=2. k Receives and outputs NN2(x2 k ) and the network model NN3(·) is the allele interaction variable x3 for MHC allele h=3. k Receives and outputs NN3(x3 k Each output is then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.
[0382] In another embodiment, if the prediction is made for the log of the mass spectrometric ion current, then r(·) is the log function and f(·) is the exponential function.
[0383] VIII.C.6. Example 3.3: Sum of Functions Model with Allelic Non-Interacting Variables In one embodiment, the implied per-allele presentation likelihood for MHC allele h is calculated as: The proposed likelihood is generated by TIFF2025065197000053.tif7128. Incorporates the influence of allele-non-interacting variables into peptide presentation as generated by TIFF2025065197000054.tif16133.
[0384] According to equation (21), the peptide sequence p k The likelihood of being presented is the function g h (·) for each of the MHC alleles H, k to generate corresponding dependency scores for allele interaction variables for each MHC allele h. w(·) is also applied to the coded version of the allele non-interaction variables to generate a dependency score for the allele non-interaction variables. The scores of the allele non-interaction variables are combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by a function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined and a clipping function is applied to the combined output to clip the values into the range [0,1] to estimate the likelihood of presentation of the peptide sequence p by the MHC allele H. k A likelihood of being proposed may be generated. w is the dependence function g introduced above in Section VIII.B.3. w The composition may be in any of the following forms:
[0385] As an example, the affine transformation function g h (·), gw(·), and peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000055.tif7128, where w k is peptide p k are allele-free interaction variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0386] As another example, the network transformation function g h (·), g w (·) was used to identify peptide p by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025065197000056.tif7128, where w k is peptide p k are allele interaction variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0387] FIG. 12 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h=2 and h=3 using (·) k As shown in FIG. 12, the network model NN2(·) is a model of the MHC allele h=2 with an allele interaction variable x2 k Receives and outputs NN2(x2 k ) is generated. w (·) is peptide p k The allele non-interaction variable w k Receives and outputs NN w (w k The outputs are combined and mapped by a function f(·). The network model NN3(·) generates a set of allele interaction variables x3 for MHC allele h=3. k Receives and outputs NN3(x3 k ) is generated, which is also the same network model NN w (·) Output NN w (w k ) and mapped by a function f(·). Both outputs are combined to give the estimated presentation likelihood u k Generate.
[0388] In another embodiment, the implied per-allele presentation likelihood for MHC allele h is calculated as: The proposed likelihood is generated by TIFF2025065197000057.tif7128. It should be generated by TIFF2025065197000058.tif13128.
[0389] VIII.C.7. Example 4: Quadratic Model In one embodiment, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k teeth, TIFF2025065197000059.tif13135, where element u'k h is the implicit per-allele presentation likelihood for MHC allele h. Values for a set of parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit per-allele presentation likelihood can be of any of the forms shown in equations (18), (20), and (22) above.
[0390] In one aspect, the model of Equation (23) is a peptide sequence p k The possibility exists that a given antigen may be simultaneously presented by two MHC alleles, which may imply that the presentation by the two HLA alleles is statistically independent.
[0391] According to equation (23), one or more MHC alleles H bind to a peptide sequence p k The presentation likelihood that p will be presented is calculated by combining the implicit per-allele presentation likelihoods and the likelihood that p will be presented by the MHC allele H. k Each pair of MHC alleles is presented with a peptide p such that it generates a presentation likelihood that k can be generated by subtracting from the sum the likelihood that
[0392] As an example, the affine transformation function g h (·) was used to identify peptide p by HLA alleles h=2 and h=3 among m=4 different identified HLA alleles. k The likelihood that will be presented is TIFF2025065197000060.tif5128, where x2 k , x3 k are the allele interaction variables specified for HLA alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for HLA alleles h=2, h=3.
[0393] As another example, the network transformation function g h (·), g w (·) was used to identify peptide p by HLA alleles h=2 and h=3 among m=4 different identified HLA alleles. k The likelihood that will be presented is TIFF2025065197000061.tif5128, where NN2(·), NN3(·) are the network models specified for HLA alleles h=2, h=3, and θ2, θ3 are the sets of parameters determined for HLA alleles h=2, h=3.
[0394] IX. Example 5: Prediction Module The prediction module 320 receives sequence data and selects candidate neoantigens in the sequence data using the proposed model. Specifically, the sequence data may be DNA sequences, RNA sequences, and / or protein sequences extracted from the patient's tumor tissue cells. The prediction module 320 divides the sequence data into a number of peptide sequences having 8-15 amino acids p k For example, the prediction module 320 processes a given sequence TIFF2025065197000062.tif3128 was divided into three peptide sequences each having nine amino acids. TIFF2025065197000063.tif6128. In one embodiment, the prediction module 320 may identify candidate neoantigens that are mutated peptide sequences by comparing sequence data extracted from normal tissue cells of a patient with sequence data extracted from tumor tissue cells of the patient to identify portions that contain one or more mutations.
[0395] The presentation module 320 applies one or more of the presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 may select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation model to the candidate neoantigens. In one embodiment, the presentation module 320 selects the candidate neoantigen sequences that have an estimated presentation likelihood above a pre-determined threshold. In another embodiment, the presentation model selects the N candidate neoantigen sequences with the highest estimated presentation likelihood (N is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the selected candidate neoantigens for a given patient can be injected into the patient to induce an immune response.
[0396] X. Example 6: Experimental results showing the performance of the exemplary proposed model The validity of the various proposed models above was tested on test data T, which was either a subset of the training data 170 that was not used to train the proposed models, or a separate data set from the training data 170 having similar variables and data structure as the training data 170.
[0397] Relevant metrics that indicate the performance of the proposed models are: TIFF2025065197000064.tif13128, which shows the ratio of the number of peptide examples correctly predicted to be presented on the relevant HLA allele to the number of peptide examples predicted to be presented on the HLA allele. i is the corresponding likelihood estimate u i is predicted to be presented on one or more relevant HLA alleles if it is equal to or greater than a given threshold t. Another relevant metric of the performance of the presentation model is TIFF2025065197000065.tif12128, which shows the ratio of the number of peptide examples that were correctly predicted to be presented on the relevant HLA allele to the number of peptide examples that were known to be presented on the HLA allele. Another relevant metric of the performance of the presentation model is the area under the curve (AUC) of the receiver operating characteristic (ROC). ROC is Plot recall against false positive rate (FPR), given by TIFF2025065197000066.tif12128.
[0398] XA Comparison of the performance of the proposed model based on mass spectrometry data against previous state-of-the-art models Figure 13A compares the performance results of the exemplary display model as presented herein and the conventional state-of-the-art model for peptide display prediction based on multi-allele mass spectrometry data.The results show that the exemplary display model performs significantly better in predicting peptide display than the conventional state-of-the-art model based on affinity and stability prediction.
[0399] Specifically, the exemplary proposed model shown in FIG. 13A as “MS” has an affine dependency function g h The maximum of the allele-by-allele presentation model shown in equation (12) with (·) and expit function f(·). An exemplary presentation model was the single-allelic HLA-A * A subset of the 02:01 mass spectrometry data (dataset "D1") (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip) and monoallelic HLA-B from the IEDB dataset * 07:02 Trained on a subset of the mass spectrometry (dataset "D2") (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). All peptides from the source protein that contained the displayed peptides in the test set were removed from the training data so that the exemplary display model cannot simply memorize the sequence of the displayed antigen.
[0400] The model shown in FIG. 13A as "Affinity" was a model similar to the current state-of-the-art model for predicting peptide presentation based on affinity prediction NETMHCpan. http: / / www.cbs.dtu.dk / services / NetMHCpan / The model shown in FIG. 13A as "Stability" was a model similar to the current state-of-the-art model for predicting peptide presentation based on the stability prediction NETMHCstab. http: / / www.cbs.dtu.dk / services / NetMHCstab-1.0 / The study data are presented in detail in. The study data are multi-allelic JY cell line HLA-A from the Bassani-Sternberg dataset. * 02:01 and HLA-B * A subset of the 07:02 mass spectrometry data (dataset "D3") (data can be found at www.ebi.ac.uk / pride / archive / projects / PXD000394). Error bars (shown as solid lines) indicate 95% confidence intervals.
[0401] As shown in the results of Figure 13A, the exemplary presentation model trained on mass spectrometry data had a significantly higher PPV value at 10% recall than the conventional state-of-the-art models predicting peptide presentation based on MHC binding affinity prediction or MHC binding stability prediction. Specifically, the exemplary presentation model had a PPV approximately 14% higher than the model based on affinity prediction and approximately 12% higher than the model based on stability prediction.
[0402] These results demonstrate that the exemplary display models performed significantly better than prior state-of-the-art models that predict peptide presentation based on predicted MHC binding affinity or stability, even though they were not trained on the protein sequence that contained the displayed peptide.
[0403] XB Comparison of the performance of presentation models based on T cell epitope data against conventional state-of-the-art models FIG. 13B compares the performance results of another exemplary presentation model as presented herein and a conventional state-of-the-art model for peptide presentation prediction based on T cell epitope data. T cell epitope data contains peptide sequences presented by MHC alleles on cell surface and recognized by T cells. The results showed that the exemplary presentation model, even though trained based on mass spectrometry data, performed significantly better in predicting T cell epitopes than the conventional state-of-the-art model based on affinity and stability prediction. In other words, the results of FIG. 13B showed that the exemplary presentation model not only performed better than the conventional state-of-the-art model in predicting peptide presentation based on mass spectrometry test data, but also performed significantly better than the conventional state-of-the-art model in predicting epitopes actually recognized by T cells. This is an indication that various presentation models as presented herein can provide improved identification of antigens that are likely to induce immunogenic responses in the immune system.
[0404] Specifically, the exemplary proposed model, shown as “MS” in FIG. 13B, is trained on a subset of the dataset D1, and is a function of the affine transformation function g h The allele-by-allele presentation model was shown in equation (2) with α(·) and expit function f(·). All peptides from the source protein containing the presented peptides in the test set were removed from the training data so that the presentation model could not simply memorize the sequence of the presented antigen.
[0405] Each of the models was *02:01 The model was applied to test data that is a subset of the mass spectrometry data (dataset "D4") for T cell epitope data (data can be found at www.iedb.org / doc / tcell full v3.zip). The model shown in FIG. 13B as "affinity" was a model similar to the current state-of-the-art model predicting peptide presentation based on affinity prediction NETMHCpan, and the model shown in FIG. 13B as "stability" was a model similar to the current state-of-the-art model predicting peptide presentation based on stability prediction NETMHCstab. Error bars (shown as solid lines) indicate 95% confidence intervals.
[0406] As shown in the results of Figure 13A, the allele-by-allele presentation model trained on mass spectrometry data had a PPV value at 10% recall rate that was significantly higher than the conventional state-of-the-art model predicting peptide presentation based on prediction of MHC binding affinity or MHC binding stability, even though the presentation model was not trained on the protein sequence that contained the presented peptide. Specifically, the allele-by-allele presentation model had a PPV value approximately 9% higher than the model based on affinity prediction, and approximately 8% higher than the model based on stability prediction.
[0407] These results demonstrated that the exemplary representation models trained on mass spectrometry data performed significantly better than the conventional state-of-the-art models for predicting epitopes recognized by T cells.
[0408] Comparison of the performance of different proposed models based on XC mass spectrometry data Figure 13C compares the performance results of an exemplary sum function model (Equation (13)), an exemplary sum of functions model (Equation (19)), and an exemplary quadratic model (Equation (23)) for peptide presentation prediction based on multi-allele mass spectrometry data. The results showed that the sum of functions model and the quadratic model performed better than the sum function model. This is because the sum function model implies that alleles in a multi-allele setting may interfere with each other for peptide presentation, when in fact the presentation of peptides is effectively independent.
[0409] Specifically, the exemplary presented model labeled “Sum of Sigmoids” in FIG. 13C is a representation of the network dependency function g h (·), the identity function f(·), and the expit function r(·). An exemplary model, labeled “sum of sigmoids”, is a sum function model with a network dependency function g h (·), the expit function f(·), and the identity function r(·). An exemplary model, labeled “hyperbolic tangent”, is the sum of the functions in equation (19) with the network dependence function g h (·), the expit function f(·), and the hyperbolic tangent function r(·). An exemplary model, labeled “quadratic”, is the sum of the functions in equation (19) with the network dependence function g h The quadratic model in Equation (23) uses the implicit per-allele representation likelihood form shown in Equation (18) with (·) and expit function f(·). Each model was trained on a subset of datasets D1, D2, and D3. The exemplary representation models were applied to the test data, which was a random subset of dataset D3 that did not overlap with the training data.
[0410] As shown in Figure 13C, the first column refers to the AUC of the ROC when each proposed model was applied to the test set, the second column refers to the negative log likelihood loss value, and the third column refers to the PPV at 10% recall. As shown in Figure 13C, the performance of the proposed models "Sum of Sigmoids", "Hyperbolic Tangent", and "Quadratic" was almost tied with a PPV at 10% recall of approximately 15-16%, while the performance of the model "Sum of Sigmoids" was slightly lower at approximately 11%.
[0411] As previously discussed in Section VIII.C.4., the results show that the presentation models "sum of sigmoids", "hyperbolic tangent", and "quadratic" have higher PPV values compared to the "sum of sigmoids" model, because the models accurately describe how peptides are presented independently by each MHC allele in a multi-allelic setting.
[0412] Comparison of the performance of presented models with and without training based on XD single-allele mass spectrometry data Figure 13D compares the performance results of two exemplary display models trained with and without single-allelic mass spectrometry data for peptide display prediction for multi-allelic mass spectrometry data.Results show that the exemplary display model trained without single-allelic data achieves performance comparable to that of the exemplary display model trained with single-allelic data.
[0413] An exemplary model "with A2 / B7 monoallelic data" is based on the network dependency function g h The model was the "sum of sigmoids" presented in equation (19) with α(·), expit function f(·), and identity function r(·). The model was trained on a subset of dataset D3 and on single-allelic mass spectrometry data for various MHC alleles from the IEDB database (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). The exemplary model "without A2 / B7 single-allelic data" was the same model, but with the addition of the allele HLA-A * 02:01 and HLA-B* The multi-allelic D3 dataset was trained on a subset of the multi-allelic D3 dataset without single-allelic mass spectrometry data for 07:02, but with single-allelic mass spectrometry data for other alleles. Within the multi-allelic training data, cell line HCC1937 was found to be HLA-B * 07:02, but expressed HLA-A * 02:01, and the cell line HCT116 does not express HLA-A * 02:01, but HLA-B * 07:02 was not expressed. The exemplary presented model was applied to the test data, which was a random subset of dataset D3 but did not overlap with the training data.
[0414] The column "Correlation" refers to the correlation between the actual label, which indicates whether the peptide was presented on the corresponding allele in the test data, and the predicted label. As shown in Figure 13D, * The prediction based on the implicit allele presentation likelihood for 02:01 is * 07:02 rather than MHC allele HLA-A * The single allele test data for 02:01 performed significantly better. Similar results were obtained for the MHC allele HLA-B * Shown at 07:02.
[0415] These results indicate that the presentation model's implicit per-allele presentation likelihood can accurately predict and identify binding motifs for individual MHC alleles, even if the direct association between peptides and each individual MHC allele was not known in the training data.
[0416] Comparison of the performance of allele-by-allele prediction without training based on XE single-allele mass spectrometry data FIG. 13E shows the HLA-A alleles provided in the analysis shown in FIG. * 02:01 and HLA-B *Based on the single-allelic mass spectrometry data for 07:02, the performance of the exemplary models "without A2 / B7 single-allelic data" and "with A2 / B7 single-allelic data" shown in Figure 13D is shown. The results show that even if the exemplary presented model is trained without the single-allelic mass spectrometry data for these two alleles, the model can learn the binding motif for each MHC allele.
[0417] As shown in FIG. 13E, the "A2 model predicting B7" shows that peptide presentation is related to the MHC allele HLA-A * Based on the implied allele-specific presentation likelihood estimates for 02:01, monoallelic HLA-B * The performance of the model when predicting for 07:02 is shown. Similarly, the "A2 model predicts A2" shows that peptide presentation is related to the MHC allele HLA-A * Based on the implied allele-specific presentation likelihood estimates for 02:01, monoallelic HLA-A * The performance of the model when predicting for 02:01 is shown. The "B7 model predicts B7" shows that peptide presentation is related to the MHC allele HLA-B * Based on the implied allele-specific presentation likelihood estimates for 07:02, monoallelic HLA-B * The performance of the model when predicting for the 07:02 data is shown. The "B7 model predicts A2" shows that peptide presentation is related to the MHC allele HLA-B * Based on the implied allele-specific presentation likelihood estimates for 07:02, monoallelic HLA-A * The performance of the model when predicting for 02:01 is shown.
[0418] As shown in Figure 13E, the predictive power of the implied allele-likelihood for the HLA alleles is significantly higher for the intended allele and significantly lower for the other HLA alleles. Similar to the results shown in Figure 13D, the exemplary presentation model was able to predict the individual alleles HLA-A, even though no direct association between peptide presentation and these alleles existed in the multi-allele training data. * 02:01 and HLA-B *It accurately learned to discriminate peptide presentation at 07:02.
[0419] Frequently occurring anchor residues in XF allele predictions match known canonical anchor motifs FIG. 13F shows the common anchor residues at positions 2 and 9 in the nonamer predicted by the exemplary model "without A2 / B7 single allele data" shown in FIG. 13D. Peptides were predicted to be presented if the estimated likelihood was above 5%. The results are shown in Table 13. * 02:01 and HLA-B * The results show that the most common anchor residues in peptides identified for presentation on 07:02 matched known anchor motifs for these MHC alleles, indicating that the exemplary presentation model accurately learned peptide binding based on the specific positions of amino acids in the peptide sequence, as expected.
[0420] As shown in FIG. 13F, the amino acids L / M at position 2 and V / L at position 9 are related to HLA-A * About 02:01 ( https: / / link.springer.com / article / 10.1186 / 1745-7580-4-2 The amino acid P at position 2 and the amino acids L / V at position 9 are known to be canonical anchor residue motifs (as shown in Table 4 of HLA-B * The most common anchor residue motifs at positions 2 and 9 for peptides identified by the model matched known canonical anchor residue motifs for both HLA alleles.
[0421] Comparison of performance of proposed models with and without XG allele non-interacting variables Figure 13G compares the performance results between the exemplary display model that incorporates C-terminal flanking sequence and N-terminal flanking sequence as allele interaction variables and the exemplary display model that incorporates C-terminal flanking sequence and N-terminal flanking sequence as allele non-interaction variables.Results show that incorporating C-terminal flanking sequence and N-terminal flanking sequence as allele non-interaction variables significantly improves model performance.More specifically, it is valuable to identify the characteristics that are appropriate for peptide display that are common across various MHC alleles, and model them so that the statistical strength for these allele non-interaction variables is shared across MHC alleles to improve the performance of display model.
[0422] An exemplary "allele interaction" model is a network dependence function g h The model was a sum-of-functions model using the form of the implicit per-allele representation likelihood in equation (22) incorporating the C-terminal and N-terminal flanking sequences as allele interaction variables, with the network dependence function g(·) and expit function f(·). An exemplary “non-allele interaction” model was h The model was a sum-of-functions model, as shown in equation (21), incorporating the C-terminal and N-terminal flanking sequences as allele-noninteracting variables, with the expit function f(·) and the C-terminal flanking sequences as allele-noninteracting variables. The allele-noninteracting variables were modeled using separate network dependence functions g w (·). Both models were trained on a subset of dataset D3 and on single-allele mass spectrometry data for various MHC alleles from the IEDB database (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Each of the presented models was applied to a test dataset, which was a random subset of dataset D3 that did not overlap with the training data.
[0423] As shown in Figure 13G, the incorporation of the C-terminal flanking sequence and the N-terminal flanking sequence as allele-non-interaction variables in the exemplary model achieved an improvement of approximately 3% in PPV values compared to modeling them as allele-interaction variables. This is because the exemplary model of "allele-non-interaction" was able to share the statistical strength of the allele-non-interaction variables across MHC alleles by modeling their effects in separate network dependency functions with very little additional computational power.
[0424] Dependence between XH-presented peptides and mRNA quantification Figure 13H illustrates the dependence between gene based mRNA quantification and fraction of peptides presented for mass spectrometry data for tumor cells. The results show a strong dependence between mRNA expression and peptide presentation.
[0425] Specifically, the horizontal axis in FIG. 13G shows mRNA expression in terms of transcripts per million (TPM) quartiles. The vertical axis in FIG. 13G shows the fraction of epitopes presented from genes in the corresponding mRNA expression quartiles. Each solid line is a plot of two measurements from a tumor sample related to the corresponding mass spectrometry data and mRNA expression measurements. As shown in FIG. 13G, there is a strong positive correlation between mRNA expression and the fraction of peptides in the corresponding genes. Specifically, peptides from genes in the top quartile of RNA expression are more than 20 times more likely to be presented than those in the bottom quartile. Furthermore, essentially zero peptides are presented from genes that are not detected through RNA.
[0426] The results show that mRNA quantification measurements are strongly predictive of peptide presentation and thus incorporation of these measurements can greatly improve the performance of presentation models.
[0427] XI. Comparison of the performance of the proposed models with the incorporation of RNA quantification data Figure 13I shows the performance of two exemplary presentation models, one of which is trained on mass spectrometry tumor cell data, and the other incorporates mRNA quantification data and mass spectrometry tumor cell data. As expected from Figure 13H, since mRNA expression is a strong indicator of peptide presentation, the results showed that there was a significant improvement in performance by incorporating mRNA quantification measurements into the exemplary presentation model.
[0428] The "MHCflurry+RNA filter" was a model similar to the current state-of-the-art model that predicts peptide presentation based on affinity prediction. It was performed using MHCflurry with a standard gene expression filter that removed all peptides derived from proteins with mRNA quantification measurements that were below 3.2 FPKM. https: / / github.com / hammerlab / mhcflurry / and http: / / biorxiv.org / content / early / 2016 / 05 / 22 / 054775. The "exemplary model, no RNA" model uses a network dependency function g h (·), the network dependency function g w The “exemplary model, no RNA” model was the “sum of sigmoids” model shown in equation (21) with the network dependency function g w The C-terminal flanking sequence was incorporated as an allelic non-interacting variable through (·).
[0429] The “exemplary model, with RNA” model is a network dependency function g h (·), the network dependency function g in equation (10) that incorporates the mRNA quantification data through a log function w The proposed model was the “sum of sigmoids” model shown in equation (19) with the network dependency function g w We incorporated C-terminal flanking sequences as allele-noninteracting variables through (·) and mRNA quantification measurements through the log function.
[0430] Each model was trained on a combination of single-allelic mass spectrometry data from the IEDB dataset, seven cell lines from multi-allelic mass spectrometry data from the Bassani-Sternberg dataset, and 20 mass spectrometry tumor samples. Each model was applied to a test set containing 5,000 donated proteins from seven tumor samples, constituting 9,830 donated peptides from a total of 52,156,840 peptides.
[0431] As shown in the first two bars in FIG. 13I, the "exemplary model, no RNA" model has a PPV value at 20% recall of 21%, compared to approximately 3% for the conventional state-of-the-art model. This represents an 18% initial performance improvement in PPV value, despite not incorporating mRNA quantification measurements. As shown in the third bar in FIG. 13I, the "exemplary model, with RNA" model, which incorporates mRNA quantification data into the proposed model, exhibits a PPV value of approximately 30%, which is an increase in performance of nearly 10% compared to the exemplary proposed model without mRNA quantification measurements.
[0432] Thus, the results show that, as expected from the findings in Figure 13H, mRNA expression is indeed a strong predictor of peptide prediction, allowing for significant improvement in the performance of the proposed model with very little additional computational complexity.
[0433] XJ MHC allele HLA-C * Example of parameters determined for 16:04 Figure 13J compares the probability of peptide presentation for various peptide lengths between the results generated by the "exemplary model, with RNA" presentation model described with respect to Figure 13I and the results predicted by a conventional state-of-the-art model that does not account for peptide length when predicting peptide presentation. The results showed that the "exemplary model, with RNA" exemplary presentation model from Figure 13I captured the variation in likelihood across peptides of different lengths.
[0434] The horizontal axis represented the sample of peptides with lengths 8, 9, 10, and 11. The vertical axis represented the probability of peptide presentation conditioned on the peptide length. The plot of "Probability of Actual Test Data" showed the percentage of peptides presented as a function of peptide length in the sample test dataset. The presentation likelihood varied with peptide length. For example, as shown in FIG. 13J, a 10-mer peptide with a standard HLA-A2 L / V anchor motif was approximately 3 times less likely to be presented than a 9-mer with the same anchor residues. The plot of "Model Ignoring Length" showed the predicted measurements when a conventional state-of-the-art model that ignores peptide length was applied to the same test dataset for presentation prediction. These models may be NetMHC versions before version 4.0, NetMHCpan versions before version 3.0, and MHCflurry, which do not take into account the variation in peptide presentation as a function of peptide length. As shown in FIG. 13J, the percentage of peptides presented will be constant across different values of peptide length, indicating that these models will not be able to capture the variation in peptide presentation as a function of length. The "Gritstone with RNA" plot showed the measurements generated from the "Gritstone with RNA" presentation model. As shown in Figure 13J, the measurements generated by the "Gritstone with RNA" model closely tracked those shown in "Probabilities of Actual Test Data" and accurately accounted for the different degrees of peptide presentation for lengths 8, 9, 10, and 11.
[0435] Thus, the results showed that the exemplary presentation model as presented herein generated improved predictions not only for 9-mer peptides, but also for peptides of 8 to 15 other lengths that account for up to 40% of peptides presented in HLA class I alleles.
[0436] Example of parameters determined for the XK MHC allele HLA-C*16:04 The following are the MHC alleles represented by h: HLA-C* For 16:04, the set of parameters determined for the variation of the allele-by-allele presentation model (equation (2)) is shown: TIFF2025065197000067.tif5128In the formula, relu(·) is the rectified linear unit (RELU) function, and W h 1 , b h 1 , W h 2 , and b h 2 is the set of parameters θ determined for the model. The allele interaction variables x h k consists of a peptide sequence. h 1 The dimensions of are (231 x 256), and b h 1 has dimensions (1 x 256), and W h 2 has dimensions (256 x 1) and b h 2 is a scalar. For the purposes of proof, b h 1 , b h 2 , W h 1 , and W h 2 The values are listed below. TIFF2025065197000068.tif222135TIFF2025065197000069.tif178128
[0437] XI. Exemplary Computer Figure 14 illustrates an exemplary computer 1400 for implementing the entities shown in Figures 1 and 3. The computer 1400 includes at least one processor 1402 coupled to a chipset 1404. The chipset 1404 includes a memory controller hub 1420 and an input / output (I / O) controller hub 1422. A memory 1406 and a graphics adapter 1412 are coupled to the memory controller hub 1420, and a display 1418 is coupled to the graphics adapter 1412. A storage device 1408, an input device 1414, and a network adapter 1416 are coupled to the I / O controller hub 1422. Other embodiments of the computer 1400 have different architectures.
[0438] The storage device 1408 is a non-transitory computer-readable storage medium, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 1406 holds instructions and data used by the processor 1402. The input interface 1414 is a touch screen interface, a mouse, a trackball, or other type of pointing device, a keyboard, or some combination thereof, and is used to input data into the computer 1400. In some embodiments, the computer 1400 may be configured to receive input (e.g., commands) from the input interface 1414 via gestures from a user. The graphics adapter 1412 displays images and other information on the display 1418. The network adapter 1416 couples the computer 1400 to one or more computer networks.
[0439] The computer 1400 is adapted to execute computer program modules for providing the functionality described herein. As used herein, the term "module" refers to computer program logic used to provide a particular functionality. Thus, a module can be implemented in hardware, firmware, and / or software. In one embodiment, the program module is stored in the storage device 1408, loaded into the memory 1406, and executed by the processor 1402.
[0440] 1 can vary depending on the aspect and processing power required by the entity. For example, the presentation specification system 160 can run on a single computer 1400 or on multiple computers 1400 that communicate with each other over a network, such as in a server farm. The computer 1400 can lack some of the components described above, such as a graphics adapter 1412 and a display 1418.
[0441] References TIFF2025065197000070.tif205145TIFF2025065197000071.tif218145TIFF2025065197000072.tif218146TIFF2025065197000073.tif151146
[0442] TIFF2025065197000074.tif230164TIFF2025065197000075.tif228156TIFF2025065197000076.tif228155TIFF2025065197000077.tif228155TIFF2025065197000078.tif228155TIFF2025065197000079.tif228156TIFF2025065197000080.tif228156TIFF2025065197000081.tif228155TIFF2025065197000082.tif228155TIFF2025065197000083.tif228156TIFF2025065197000084.tif228156TIFF2025065197000085.tif228156TIFF2025065197000086.tif228156TIFF2025065197000087.tif228155TIFF2025065197000088.tif228156TIFF2025065197000089.tif228156TIFF2025065197000090.tif228156TIFF2025065197000091.tif228155TIFF2025065197000092.tif228156TIFF2025065197000093.tif228155TIFF2025065197000094.tif228155TIFF2025065197000095.tif228156TIFF2025065197000096.tif228155TIFF2025065197000097.tif228155TIFF2025065197000098.tif228155TIFF2025065197000099.tif228155TIFF2025065197000100.tif228155TIFF2025065197000101.tif228156TIFF2025065197000102.tif228156TIFF2025065197000103.tif228156TIFF2025065197000104.tif228155TIFF2025065197000105.tif228155TIFF2025065197000106.tif228156TIFF2025065197000107.tif228155TIFF2025065197000108.tif228155TIFF2025065197000109.tif228156TIFF2025065197000110.tif228156TIFF2025065197000111.tif228155TIFF2025065197000112.tif228156TIFF2025065197000113.tif228156TIFF2025065197000114.tif228155TIFF2025065197000115.tif228155TIFF2025065197000116.tif228155TIFF2025065197000117.tif228155TIFF2025065197000118.tif228156TIFF2025065197000119.tif228156TIFF2025065197000120.tif228155TIFF2025065197000121.tif228155TIFF2025065197000122.tif228155TIFF2025065197000123.tif228156TIFF2025065197000124.tif228155TIFF2025065197000125.tif228156TIFF2025065197000126.tif228155TIFF2025065197000127.tif228155TIFF2025065197000128.tif228156TIFF2025065197000129.tif228155TIFF2025065197000130.tif228155TIFF2025065197000131.tif228155TIFF2025065197000132.tif228155TIFF2025065197000133.tif228156TIFF2025065197000134.tif228156TIFF2025065197000135.tif228155TIFF2025065197000136.tif228156TIFF2025065197000137.tif228156TIFF2025065197000138.tif228155TIFF2025065197000139.tif228156TIFF2025065197000140.tif228155TIFF2025065197000141.tif228155TIFF2025065197000142.tif228156TIFF2025065197000143.tif228155TIFF2025065197000144.tif228155TIFF2025065197000145.tif228156TIFF2025065197000146.tif228155TIFF2025065197000147.tif228155TIFF2025065197000148.tif228155TIFF2025065197000149.tif228156TIFF2025065197000150.tif228155TIFF2025065197000151.tif228155TIFF2025065197000152.tif228155TIFF2025065197000153.tif228155TIFF2025065197000154.tif228156TIFF2025065197000155.tif228156TIFF2025065197000156.tif228155TIFF2025065197000157.tif228155TIFF2025065197000158.tif228155TIFF2025065197000159.tif228155TIFF2025065197000160.tif228155TIFF2025065197000161.tif228155TIFF2025065197000162.tif228155TIFF2025065197000163.tif228156TIFF2025065197000164.tif228156TIFF2025065197000165.tif228156TIFF2025065197000166.tif228156TIFF2025065197000167.tif228156TIFF2025065197000168.tif228156TIFF2025065197000169.tif228156TIFF2025065197000170.tif228156TIFF2025065197000171.tif228156.
[0443] Array information SEQUENCE LISTING <110> GRITSTONE BIO, INC. <120> NEOANTIGEN IDENTIFICATION, MANUFACTURE, AND USE <150> US 62 / 425,995 <151> 2016-11-23 <150> US 62 / 394,074 <151> 2016-09-13 <150> US 62 / 379,986 <151> 2016-08-26 <150> US 62 / 317,823 <151> 2016-04-04 <150> US 62 / 268,333 <151> 2015-12-16 <160> 20 <170> PatentIn version 3.5 <210> 1 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 1 Tyr Val Tyr Val Ala Asp Val Ala Ala Lys 1 5 10 <210> 2 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 2 Tyr Glu Met Phe Asn Asp Lys Ser 1 5 <210> 3 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 3 Tyr Glu Met Phe Asn Asp Lys Ser Phe 1 5 <210> 4 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (3)..(3) <223> Pyrrolysine <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <400> 4 His Arg Xaa Glu Ile Phe Ser His Asp Phe Xaa 1 5 10 <210> 5 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Ile or Leu <220> <221> MOD_RES <222> (5)..(5) <223> Ile or Leu <220> <221> MOD_RES <222> (7)..(7) <223> Pyrrolysine <400> 5 Phe Xaa Ile Glu Xaa Phe Xaa Glu Ser Ser 1 5 10 <210> 6 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Pyrrolysine <400> 6 Asn Glu Ile Xaa Arg Glu Ile Arg Glu Ile 1 5 10 <210> 7 <211> 15 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (1)..(1) <223> Ile or Leu <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <220> <221> MOD_RES <222> (15)..(15) <223> Selenocysteine <400> 7 Xaa Phe Lys Ser Ile Phe Glu Met Met Ser Xaa Asp Ser Ser Xaa 1 5 10 15 <210> 8 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (11)..(11) <223> Pyrrolysine <400> 8 Lys Asn Phe Leu Glu Asn Phe Ile Glu Ser Xaa Phe Ile 1 5 10 <210> 9 <211> 15 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Pyrrolysine <220> <221> MOD_RES <222> (14)..(14) <223> Ile or Leu <400> 9 Phe Xaa Glu Ile Phe Asn Asp Lys Ser Leu Asp Lys Phe Xaa Ile 1 5 10 15 <210> 10 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <400> 10 Gln Cys Glu Ile Xaa Trp Ala Arg Glu 1 5 <210> 11 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Selenocysteine <400> 11 Phe Ile Glu Xaa His Phe Trp Ile 1 5 <210> 12 <211> 12 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (7)..(7) <223> Ile or Leu <220> <221> MOD_RES <222> (10)..(10) <223> Selenocysteine <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <400> 12 Phe Glu Trp Arg His Arg Xaa Thr Arg Xaa Xaa Arg 1 5 10 <210> 13 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Ile or Leu <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (8)..(8) <223> Ile or Leu <400> 13 Gln Ile Glu Xaa Xaa Glu Ile Xaa Glu 1 5 <210> 14 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Ile or Leu <220> <221> MOD_RES <222> (9)..(9) <223> Pyrrolysine <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <400> 14 Phe Xaa Glu Leu Phe Ile Ser Asx Xaa Ser Xaa Phe Ile Glu 1 5 10 <210> 15 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (9)..(9) <223> Ile or Leu <400> 15 Ile Glu Phe Arg Xaa Glu Ile Phe Xaa Glu Phe 1 5 10 <210> 16 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (9)..(9) <223> Ile or Leu <400> 16 Ile Glu Phe Arg Xaa Glu Ile Phe Xaa 1 5 <210> 17 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Pyrrolysine <220> <221> MOD_RES <222> (8)..(8) <223> Ile or Leu <400> 17 Glu Phe Arg Xaa Glu Ile Phe Xaa Glu 1 5 <210> 18 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (3)..(3) <223> Pyrrolysine <220> <221> MOD_RES <222> (7)..(7) <223> Ile or Leu <400> 18 Phe Arg Xaa Glu Ile Phe Xaa Glu Phe 1 5 <210> 19 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (6)..(6) <223> Selenocysteine <220> <221> MOD_RES <222> (7)..(8) <223> Pyrrolysine <400> 19 Phe Glu Gly Arg Lys Xaa Xaa Xaa Ile 1 5 <210> 20 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Ile or Leu <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (7)..(7) <223> Ile or Leu <220> <221> MOD_RES <222> (8)..(8) <223> Pyrrolysine <220> <221> MOD_RES <222> (10)..(10) <223> Ile or Leu <220> <221> MOD_RES <222> (14)..(14) <223> Pyrrolysine <400> 20 Pro Xaa Phe Ile Xaa Glu Xaa Xaa Ile Xaa Gly Glu Ile Xaa 1 5 10
Claims
1. 1. A method for identifying one or more neoantigens from a subject's tumor cells that are presented on the surface of the tumor cells, comprising the steps of: obtaining a peptide sequence for each of a set of neoantigens; inputting the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each of the neoantigens will be presented by one or more MHC alleles on the tumor cell surface of tumor cells of the subject, wherein the one or more presentation models: training peptide sequences, at least one MHC allele derived from a sample comprising a cell line engineered to express a single MHC allele; and a label derived from the mass spectrometry data indicating whether one or more training peptide sequences were presented by said at least one MHC allele; and Selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens.
2. 2. The method of claim 1, wherein the training dataset comprises one or more synthetically generated peptide sequences and a label indicating that the one or more synthetically generated peptide sequences were not presented by at least one MHC allele.
3. The method described in claim 1 or 2, wherein the peptide sequence of the neoantigen has a length of 8, 9, 10, 11, 12, 13, 14, or 15 amino acids.
4. The presented model is the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in the peptide sequence; the likelihood of presentation on the surface of tumor cells of such peptide sequences containing particular amino acids at particular positions by a particular one of the MHC alleles of the pair; The method according to any one of claims 1 to 3, which expresses a dependency between
5. The step of inputting a peptide sequence comprises: applying one or more presentation models to the corresponding peptide sequences of the neoantigens to generate a dependency score for each of the one or more MHC alleles that indicates whether the MHC allele presents the corresponding neoantigen based on at least the position of an amino acid in the peptide sequence of the corresponding neoantigen; The method according to any one of claims 1 to 4, comprising:
6. transforming the dependency scores to generate a corresponding per-allele likelihood for each MHC allele, which indicates the likelihood that the corresponding MHC allele will present the corresponding neoantigen; and Combining the allele likelihoods to generate a numerical likelihood 6. The method of claim 5, further comprising:
7. 7. The method of claim 6, wherein the step of transforming the dependency scores models the presentation of peptide sequences of corresponding neoantigens as mutually exclusive.
8. transforming the combination of dependency scores to generate a numerical likelihood The method of any one of claims 5 to 7, further comprising:
9. 9. The method of claim 8, wherein the step of transforming the combination of dependency scores models the presentation of peptide sequences of corresponding neoantigens as interference between MHC alleles.
10. The set of numerical likelihoods is further specified by at least an allele non-interaction characteristic, and applying an allele non-interaction model of the one or more presentation models to the allele non-interaction feature to generate a dependency score for the allele non-interaction feature that indicates whether the corresponding neoantigen peptide sequence will be presented based on the allele non-interaction feature. The method of any one of claims 5 to 9, further comprising:
11. combining the dependency score for each MHC allele in the one or more MHC alleles with the dependency score for the allele-non-interacting property; transforming the combined dependency scores for each MHC allele to generate a corresponding per-allele likelihood for the MHC allele, the likelihood being that the corresponding MHC allele will present the corresponding neoantigen; and Combining the allele likelihoods to generate a numerical likelihood 11. The method of claim 10, further comprising:
12. transforming the combination of the dependency scores for each of the MHC alleles and the dependency scores for the allele-non-interacting trait to generate a numerical likelihood.
12. The method of claim 10 or 11, further comprising:
13. The method of claim 1, wherein the training dataset further comprises data on mRNA expression levels of tumor cells.
14. 2. The method of claim 1, wherein at least one MHC allele is derived from a sample comprising a cell line engineered to express multiple MHC class I or class II alleles.
15. 15. The method of claim 1 or 14, wherein the samples are obtained from or comprise human cell lines derived from multiple patients.
16. 15. The method of claim 1 or 14, wherein the samples comprise fresh or frozen tumor samples obtained from multiple patients.
17. 15. The method of claim 1 or 14, wherein the samples comprise fresh or frozen tissue samples obtained from multiple patients.
18. 18. The method of any one of claims 14 to 17, wherein the training dataset is generated based on obtaining at least one of normal nucleotide sequencing data of the exome, transcriptome, and whole genome from normal tissue samples.
19. 19. The method according to claim 1, wherein the method comprises carrying out any of the steps according to claims 1 to 18, and A process for producing or having produced a tumor vaccine comprising a selected set of neoantigens. The method for producing a tumor vaccine further comprises: