Neoantigen identification using hot spots

Optimized exome and transcriptome analysis with deep learning models enhance the prediction of neoantigen presentation, addressing inefficiencies in current methods and improving the effectiveness of personalized cancer vaccines and T cell therapies.

JP2025143480APending Publication Date: 2025-10-01GRITSTONE BIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025116924
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-10-10
Filing Date
2025-07-11
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

Current methods for identifying neoantigens and neoantigen-recognizing T cells are time-consuming, labor-intensive, and have low positive predictive value (PPV), leading to ineffective neoantigen-based vaccines and T cell therapies due to inaccurate prediction of peptide presentation on tumor surfaces.

Method used

Utilizing optimized tumor exome and transcriptome analysis combined with nonlinear deep learning models to predict peptide presentation likelihood, incorporating k-mer blocks and MHC alleles, to identify neoantigens with high PPV for personalized cancer vaccines and T cell therapies.

Benefits of technology

The models significantly improve the positive predictive value of neoantigen identification, enabling more efficient and cost-effective personalized cancer therapies by accurately selecting neoantigens likely to be presented on tumor cells and induce anti-tumor immunity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025143480000001_ABST
    Figure 2025143480000001_ABST
Patent Text Reader

Abstract

To provide methods for identifying neoantigens that are likely to be presented on a surface of tumor cells of a subject.SOLUTION: Peptide sequences of tumor neoantigens are obtained by sequencing the tumor cells of the subject. The peptide sequence of each neoantigen is associated with one or more k-mer blocks of a plurality of k-mer blocks of the nucleotide sequencing data of the subject. The peptide sequences and the associated k-mer blocks are input into a machine-learned presentation model to generate presentation likelihoods for the tumor neoantigens, each presentation likelihood representing the likelihood that a neoantigen is presented by an MHC allele on the surfaces of the tumor cells of the subject. A subset of the neoantigens is selected based on the presentation likelihood to return the set of selected neoantigens.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Therapeutic vaccines and T cell therapies based on tumor-specific neoantigens hold great promise as the next generation of personalized cancer immunotherapy. 1~3 Cancers with high mutational burden, such as non-small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies due to their relatively high likelihood of generating neoantigens. 4,5 Early evidence suggests that neoantigen-based vaccination induces T cell responses. 6 , T cell therapy targeting neoantigens can induce tumor regression in selected patients 7 Both MHC class I and MHC class II influence T cell responses. 70~71 .

[0002] However, the identification of neoantigens and neoantigen-recognizing T cells is crucial for assessing tumor response. 77,110 , examining tumor evolution 111 , designing the next generation of personalized therapies 112 Current methods for identifying neoantigens are time-consuming and labor-intensive. 84,96 , or the accuracy is not sufficient 87,91-93 Neoantigen-recognizing T cells are the main component of TILs. 84,96,113,114 circulating in the peripheral blood of cancer patients. 107 Although it has been shown recently that neoantigen-reactive T cells can be identified using the TIL method, the current methods are as follows: 97,98 or leukapheresis 107 (2) they require screening of impractically large peptide libraries; or (3) they rely on MHC alleles that may be practically available for only a small number of MHC alleles.

[0003] Furthermore, early methods incorporating mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides have been proposed.8 However, these proposed methods involve many steps other than gene expression and MHC binding (e.g., TAP transport, proteasomal cleavage, MHC binding, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-I; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsins), competition with CLIP peptides for HLA binding catalyzed by HLA-DM, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-II). 9 The entire epitope generation process cannot be modeled, and therefore existing methods tend to suffer from low positive predictive value (PPV) (Figure 1A).

[0004] Indeed, analyses of peptides presented by tumor cells conducted by multiple groups have shown that less than 5% of peptides predicted to be presented using gene expression and MHC binding affinity are found on tumor surface MHC. 10,11 (Figure 1B). This low correlation between binding prediction and MHC presentation is further supported by the lack of improved prediction accuracy for binding-restricted neoantigens in response to checkpoint inhibitors relative to the number of mutations alone. 12 .

[0005] Such low positive predictive values ​​(PPV) of existing methods for predicting presentation present a challenge in the design of neoantigen-based vaccines and in neoantigen-based T cell therapies. If vaccines are designed using low PPV predictions, most patients are unlikely to receive therapeutic neoantigens, and even fewer patients will receive multiple neoantigens (even assuming all presented peptides are immunogenic). Similarly, if therapeutic T cells are designed based on low PPV predictions, most patients are unlikely to receive T cells with reactivity against tumor neoantigens, and the time and resource costs of identifying predicted neoantigens using downstream testing methods after prediction may be unnecessarily high. Therefore, neoantigen vaccination and T cell therapy using current methods are unlikely to be effective in a significant number of tumor-bearing subjects (Figure 1C).

[0006] Furthermore, previous approaches have only used cis-acting mutations to generate candidate neoantigens, including splicing factor mutations that occur in multiple tumor types and lead to aberrant splicing of many genes. 13 , and additional sources of nascent ORFs, including mutations that create or remove protease cleavage sites, were not considered in most cases.

[0007] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard tumor analysis approaches may erroneously promote sequence artifacts or germline polymorphisms as neoantigens, which can lead to inefficient utilization of vaccine doses or the risk of autoimmunity, respectively. Summary of the Invention

[0008] Disclosed herein are optimized approaches for identifying and selecting neoantigens for personalized cancer vaccines, T cell therapy, or both. First, we address optimized tumor exome and transcriptome analysis approaches to identify neoantigen candidates using next-generation neoantigen (NGS). These methods build on standard approaches of tumor analysis by NGS so that the most sensitive and specific neoantigen candidates are developed across all classes of genomic alterations. Second, novel approaches to high PPV neoantigen selection are provided to overcome specificity issues and ensure that neoantigens developed for vaccine addition and / or as targets for T cell therapy are more likely to elicit anti-tumor immunity. Depending on the embodiment, these approaches include trained statistical regression or nonlinear deep learning models that jointly model peptide-allele mapping and per-allele motifs for peptides of multiple lengths that share statistical power across peptides of different lengths. These deep learning models also utilize parameters describing the presence or absence of presentation hotspots within k-mer blocks associated with peptide sequences to determine the presentation likelihood of the peptide. In particular, nonlinear deep learning models can be designed and trained to treat different MHC alleles within the same cell as independent, thus resolving the problem associated with linear models where linear models interfere with each other.Finally, this addresses additional concerns regarding the design and production of personalized neoantigen-based vaccines and the production of personalized neoantigen-specific T cells for T cell therapy.

[0009] The models disclosed herein outperform state-of-the-art predictive tools trained on binding affinity and earlier predictive tools based on MS peptide data by up to an order of magnitude. By predicting peptide presentation with greater confidence, the models enable more time- and cost-effective identification of neoantigen- or tumor antigen-specific T cells for personalized therapy using a clinically practical process that uses limited amounts of patient peripheral blood, screens fewer peptides per patient, and does not necessarily rely on MHC multimers. However, in another embodiment, the models disclosed herein enable more time- and cost-effective identification of tumor antigen-specific T cells using MHC multimers by reducing the number of MHC multimer-bound peptides that need to be screened to identify neoantigen- or tumor antigen-specific T cells.

[0010] The predictive performance of the model disclosed herein on the TIL neoepitope dataset and the task of identifying predicted neoantigen-reactive T cells demonstrates that by modeling HLA processing and presentation, it is now possible to obtain predictions of therapeutically useful neoepitopes. In summary, this work will accelerate progress toward patient cures by enabling actionable in silico antigen identification for antigen-targeted immunotherapy. [The present invention 1001] 1. A method for identifying one or more neoantigens derived from one or more tumor cells of a subject that are likely to be presented on the surface of the tumor cells, comprising: obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing data from the tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen includes at least one alteration that makes the peptide sequence different from a corresponding wild-type peptide sequence identified from the normal cells of the subject; encoding each peptide sequence of the neoantigen into a corresponding numerical vector, each numerical vector containing information about a set of amino acids constituting the peptide sequence and the positions of the amino acids within the peptide sequence; associating each peptide sequence of said neoantigen with one or more k-mer blocks of said plurality of k-mer blocks of said nucleotide sequencing data for said subject; using a computer processor to input the numerical vector and the one or more associated k-mer blocks into a machine-learned presentation model to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood in the set represents the likelihood that a corresponding neoantigen will be presented on the surface of the tumor cells of the subject by one or more MHC alleles, and wherein the machine-learned presentation model: a plurality of parameters determined based at least on a training data set, the training data set comprising: and determining, for each sample of a plurality of samples, the presence of a peptide bound to at least one MHC allele within the set of MHC alleles identified as being present in said sample, the label obtained by mass spectrometry. For each of the samples, a training peptide sequence encoded as a numerical vector containing information about a set of amino acids that make up the peptide and the positions of the amino acids within the peptide; and For each of the samples, for each of the training peptide sequences of the sample, an association between the training peptide sequence and one or more k-mer blocks among a plurality of k-mer blocks of the nucleotide sequencing data for the training peptide sequence. Including, a subset of the plurality of parameters representing the presence or absence of a presentation hotspot in the one or more k-mer blocks; a function representing a relationship between the numerical vector and the one or more k-mar blocks received as input, and the proposed likelihood generated as output based on the numerical vector, the one or more k-mar blocks, and the parameters; Including, the inputting step; selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens; and returning the set of selected neoantigens. The method comprising: [The present invention 1002] The step of inputting the numerical vector into the machine-learned presentation model includes: applying the machine-learned presentation model to a peptide sequence of the neoantigen to generate a dependency score for each of the one or more MHC alleles that indicates whether the MHC allele will present the neoantigen based on a particular amino acid at a particular position in the peptide sequence. The method of the present invention 1001, comprising: [The present invention 1003] The step of inputting the numerical vector into the machine-learned presentation model includes: transforming the dependency scores to generate, for each MHC allele, a corresponding per-allele likelihood that indicates the likelihood that the corresponding MHC allele will present the corresponding neoantigen; combining the per-allele likelihoods to generate a presentation likelihood for the neoantigen; The method of the present invention 1002 further comprises: [The present invention 1004] 1004. The method of claim 1003, wherein transforming said dependency score models presentation of said neoantigen as mutually exclusive across said one or more MHC alleles. [The present invention 1005] The method of claim 1002, wherein the step of inputting the numerical vector into the machine-learned presentation model further comprises transforming the combination of dependency scores to generate the presentation likelihood, and transforming the combination of dependency scores models the presentation of the neoantigen as interference between the one or more MHC alleles. [The present invention 1006] the set of representation likelihoods is further specified by at least one or more allele non-interaction characteristics; applying the machine-learned presentation model to the allele-non-interacting feature to generate a dependency score for the allele-non-interacting feature that indicates whether the corresponding neoantigen peptide sequence will be presented based on the allele-non-interacting feature. Any of the methods 1002 to 1005 of the present invention, further comprising: [The present invention 1007] combining the dependency score for each MHC allele of the one or more MHC alleles with the dependency score for the allele-non-interacting trait; transforming the combined dependency scores for each MHC allele to generate a per-allele likelihood for each MHC allele, the likelihood being that the corresponding MHC allele will present the corresponding neoantigen; combining the per-allele likelihoods to generate the representation likelihood. The method of the present invention 1006 further comprising: [The present invention 1008] combining the dependency scores for each of the MHC alleles with the dependency scores for the allele-non-interacting trait; transforming the combined dependency scores to generate the presentation likelihood; The method of the present invention 1006 further comprising: [The present invention 1009] Any of the methods of claims 1006 to 1008, wherein the at least one or more allele non-interaction characteristics comprises an association between a peptide sequence of the neoantigen and one or more k-mer blocks among a plurality of k-mer blocks of the nucleotide sequencing data of the neoantigen. [The present invention 1010] 1009. The method of any of claims 1001 to 1009, wherein said one or more MHC alleles comprises two or more different MHC alleles. [The present invention 1011] The method of any of claims 1001 to 1010, wherein said peptide sequence comprises a peptide sequence having a length other than 9 amino acids. [The present invention 1012] 1012. The method of any of claims 1001 to 1011, wherein encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme. [The present invention 1013] The plurality of samples (a) one or more cell lines engineered to express a single MHC allele; (b) one or more cell lines engineered to express multiple MHC alleles; (c) one or more human cell lines obtained or derived from multiple patients; (d) fresh or frozen tumor samples obtained from multiple patients; and (e) fresh or frozen tissue samples obtained from multiple patients; Any of the methods of the present invention 1001 to 1012, including at least one of the following. [The present invention 1014] the training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of said peptides; and (b) data relating to measurements of peptide-MHC binding stability for at least one of said peptides; The method of any one of claims 1001 to 1013, further comprising at least one of the following: [The present invention 1015] 1015. The method of any of claims 1001 to 1014, wherein said set of presentation likelihoods is further identified by at least the expression level of said one or more MHC alleles in said subject, as measured by RNA-seq or mass spectrometry. [The present invention 1016] The set of presentation likelihoods is: (a) the predicted affinity between neoantigens in the set of neoantigens and the one or more MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex; The method of any of claims 1001 to 1015, further characterized by at least one of the following characteristics: [The present invention 1017] The set of numerical likelihoods is: (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; The method of any of claims 1001 to 1016, further characterized by at least one of the following characteristics: [The present invention 1018] Any of the methods of present inventions 1001 to 1017, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of being presented on the surface of the tumor cells compared to non-selected neoantigens based on the machine-learned presentation model. [The present invention 1019] Any of the methods of present inventions 1001 to 1018, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of inducing a tumor-specific immune response in the subject compared to non-selected neoantigens based on the machine-learned presentation model. [The present invention 1020] Any of the methods of inventions 1001 to 1019, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) compared to unselected neoantigens based on the presentation model, and optionally the APCs are dendritic cells (DCs). [The present invention 1021] Any of the methods of present inventions 1001 to 1020, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of being inhibited by central tolerance or peripheral tolerance compared to non-selected neoantigens based on the machine-learned presentation model. [The present invention 1022] Any of the methods of present inventions 1001 to 1021, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of inducing an autoimmune response against normal tissue in the subject compared to non-selected neoantigens based on the machine-learned presentation model. [The present invention 1023] 10. The method of any one of claims 1001 to 1022, wherein the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer. [The present invention 1024] The method of any of claims 1001 to 1023, further comprising generating an output for constructing a personalized cancer vaccine from said set of selected neoantigens. [The present invention 1025] 1025. The method of claim 1024, wherein said output for the personalized cancer vaccine comprises at least one peptide sequence or at least one nucleotide sequence encoding said selected set of neoantigens. [The present invention 1026] The method according to any one of claims 1001 to 1025, wherein the presentation model trained by machine learning is a neural network model. [The present invention 1027] The method of claim 1026, wherein the neural network model comprises a plurality of network models for MHC alleles, each network model being assigned to a corresponding MHC allele among the plurality of MHC alleles and comprising a series of nodes arranged in one or more layers. [The present invention 1028] The method of claim 1027, wherein the neural network models are trained by updating parameters of the neural network models, and the parameters of at least two network models are updated together for at least one training iteration. [The present invention 1029] The method of any one of claims 1026 to 1028, wherein the machine-learned presentation model is a deep learning model including one or more layers of nodes. [The present invention 1030] 1029. The method of any of claims 1001 to 1029, wherein said one or more MHC alleles are class I MHC alleles. [The present invention 1031] A computer processor; When executed by the computer processor, the computer processor: obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing data from tumor cells and normal cells of a subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen includes at least one alteration that makes the peptide sequence different from a corresponding wild-type peptide sequence identified from the normal cells of the subject; encoding each peptide sequence of the neoantigen into a corresponding numerical vector, each numerical vector containing information about a set of amino acids constituting the peptide sequence and positions of the amino acids within the peptide sequence; associating each peptide sequence of said neoantigen with one or more k-mer blocks of said plurality of k-mer blocks of said nucleotide sequencing data of said subject; inputting the numerical vector and the one or more associated k-mer blocks into a machine-learned presentation model to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood in the set represents the likelihood that a corresponding neoantigen will be presented on the surface of the tumor cells of the subject by one or more MHC alleles, and wherein the machine-learned presentation model: a plurality of parameters determined based at least on a training data set, the training data set comprising: and determining, for each sample of a plurality of samples, the presence of a peptide bound to at least one MHC allele within the set of MHC alleles identified as being present in said sample, the label obtained by mass spectrometry. For each of the samples, a training peptide sequence encoded as a numerical vector containing information about a set of amino acids that make up the peptide and the positions of the amino acids within the peptide; and For each of the samples, for each of the training peptide sequences of the sample, an association between the training peptide sequence and one or more of the k-mer blocks of the nucleotide sequencing data for the training peptide sequence. Including, a subset of the plurality of parameters representing the presence or absence of a presentation hotspot in the one or more k-mer blocks; a function representing a relationship between the numerical vector and the one or more k-mar blocks received as input, and the proposed likelihood generated as output based on the numerical vector, the one or more k-mar blocks, and the parameters; Including, said inputting; selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens; and returning the set of selected neoantigens. a memory storing computer program instructions for causing the 2. A computer system comprising: [Brief explanation of the drawings]

[0011] These and other features, aspects, and aspects of the present invention will become better understood with regard to the following description and accompanying drawings.

[0012] [Figure 1A] Current clinical approaches to neoantigen identification are presented. [Figure 1B] It shows that less than 5% of the predicted binding peptides are displayed on tumor cells. [Figure 1C] Illustrates the impact of specificity issues on neoantigen prediction. [Figure 1D] This shows that binding prediction is not sufficient to identify neoantigens. [Figure 1E] Probability of MHC-I presentation as a function of peptide length. [Figure 1F] An exemplary peptide spectrum generated from a Promega dynamic range standard is shown. [Figure 1G] We show how adding features increases the positive predictive value of the model. [Figure 2A] 1 is a schematic of an environment for identifying the likelihood of peptide presentation in a patient, according to one embodiment. [Figure 2B] A method for obtaining presentation information according to one embodiment will be described. [Figure 2C] A method for obtaining presentation information according to one embodiment will be described. [Figure 3]FIG. 1 is a high-level block diagram illustrating computer logic components of a presentation identification system, according to one embodiment. [Figure 4] 1 illustrates an exemplary set of training data, according to one embodiment. [Figure 5] 1 illustrates an exemplary network model related to MHC alleles. [Figure 6A] An exemplary network model NNH(·) shared by MHC alleles according to one embodiment [Figure 6B] 1. An exemplary network model NNH(·) shared by MHC alleles according to another embodiment [Figure 7] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 8] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 9] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 10] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 11] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 12] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 13A] 1 shows the sample frequency distribution of mutation burden in NSCLC patients. [Figure 13B] 1 shows the number of presented neoantigens in a simulated vaccine for patients selected based on the selection criterion of whether the patient meets a minimum mutational load, according to one embodiment. [Figure 13C]10A-10C show a comparison of the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model, according to one embodiment. [Figure 13D] 10A-B compares the number of neoantigens presented in simulated vaccines between selected patients associated with a vaccine containing therapeutic subsets identified based on a single allele-per-presentation model for HLA-A*02:01 and selected patients associated with a vaccine containing therapeutic subsets identified based on both an allele-per-presentation model for HLA-A*02:01 and HLA-B*07:02. According to one embodiment, the vaccine volume is set at v=20 epitopes. [Figure 13E] FIG. 10 compares the number of presented neoantigens in simulated vaccines between patients selected based on mutational burden and patients selected by expected utility score, according to one embodiment. [Figure 14] Figure 1 compares the positive predictive value (PPV) at 40% recall of different versions of the MS model and earlier approaches29 for modeling HLA-presented peptides in human tumors when each model is tested on a test set consisting of five different excluded test samples, including an excluded tumor sample, where each test sample has a ratio of presented to unpresented peptides of 1:2500. [Figure 15A] This compares the average positive predictive value (PPV) across recall rates for the proposed model with proposed hotspot parameters and the proposed model without proposed hotspot parameters when each model was tested on five leave-out test samples. [Figure 15B] This shows a comparison of the precision-recall curves of the proposed model with and without the proposed hotspot parameters when each model was tested on leave-out test sample 0. [Figure 15C]10 shows a comparison of the precision-recall curves of the proposed model with and without the proposed hotspot parameters when each model was tested on the excluded test sample 1. [Figure 15D] This figure compares the precision-recall curves of the proposed model with the proposed hotspot parameters and the proposed model without the proposed hotspot parameters when each model was tested on the excluded test sample 2. [Figure 15E] 10 shows a comparison of the precision-recall curves of the proposed model with and without the proposed hotspot parameters when each model was tested on the excluded test sample 3. [Figure 15F] 10 shows a comparison of precision-recall curves for the proposed model with and without the proposed hotspot parameters when each model was tested on the excluded test sample 4. [Figure 16] The percentage of peptides spanning somatic mutations recognized by T cells in the top 5, 10, 20, and 30 ranked peptides identified by the presentation model with and without the presentation hotspot parameters was compared for a test set consisting of test samples taken from patients with at least one pre-existing T cell response. [Figure 17A] Detection of T cell responses to patient-specific neoantigen peptide pools in nine patients is shown. [Figure 17B] Detection of T cell responses to individual patient-specific neoantigen peptides in four patients is shown. [Figure 17C] An exemplary image of an ELISpot well for patient CU04 is shown. [Figure 18A-1] 1 shows the results of a control experiment using neoantigens in HLA-matched healthy donors. [Figure 18A-2] This is a continuation of Figure 18A-1. [Figure 18B-1]1 shows the results of a control experiment using neoantigens in HLA-matched healthy donors. [Figure 18B-2] This is a continuation of Figure 18B-1. [Figure 19] Detection of T cell responses against the PHA positive control is shown for each donor and each in vitro expansion shown in Figure 17A. [Figure 20A] 1 shows the detection of T cell responses to each individual patient-specific neoantigen peptide in pool #2 in patient CU04. [Figure 20B] 1 shows detection of T cell responses to individual patient-specific neoantigen peptides at each of three visits for patient CU04 and at each of two visits for patient 1-024-002 (each visit occurring at a different time point). [Figure 20C] 1 shows detection of T cell responses to individual patient-specific neoantigen peptides and to patient-specific neoantigen peptide pools at each of two visits for patient CU04 and at each of two visits for patient 1-024-002 (each visit occurring at a different time point). [Figure 21] Detection of T cell responses to two patient-specific neoantigen peptide pools and a DMSO negative control for the patient in Figure 17A is shown. [Figure 22] The predictive performance of a presentation model using presentation hotspot parameters in predicting neoantigen presentation by MHC class II molecules was compared with that of a presentation model without presentation hotspot parameters. [Figure 23-1] A method for sequencing the TCR of neoantigen-specific memory T cells from the peripheral blood of NSCLC patients is presented. [Figure 23-2] This is a continuation of Figure 23-1. [Figure 24] 1 shows an exemplary embodiment of a TCR construct for introducing a TCR into a recipient cell. [Figure 25-1] 1 shows the nucleotide sequence of an exemplary P526 construct backbone for cloning TCRs into expression systems for therapeutic development. [Figure 25-2] This is a continuation of Figure 25-1. [Figure 26-1] 1 shows an exemplary construct sequence for cloning a patient's neoantigen-specific TCR clonotype 1 into an expression system for therapeutic development. [Figure 26-2] This is a continuation of Figure 26-1. [Figure 27-1] 1 shows an exemplary construct sequence for cloning a patient's neoantigen-specific TCR clonotype 3 into an expression system for therapeutic development. [Figure 27-2] This is a continuation of Figure 27-1. [Figure 28] 1 is a flowchart of a method for providing personalized neoantigen-specific therapy to a patient, according to one embodiment. [Figure 29] An exemplary computer for implementing the entities shown in FIGS. 1 and 3 will now be described. DETAILED DESCRIPTION OF THE INVENTION

[0013] I. Definition In general, terms used in the claims and the specification shall be interpreted as having their ordinary meaning as understood by one of ordinary skill in the art. Certain terms are defined below to provide further clarity. If there is a conflict between the ordinary meaning and a given definition, the given definition shall control.

[0014] As used herein, the term "antigen" refers to a substance that induces an immune response.

[0015] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes it different from its corresponding wild-type parent antigen, for example, due to a tumor cell mutation or tumor cell-specific post-translational modification. Neoantigens may include polypeptide or nucleotide sequences. Mutations can include frameshift or non-frameshift insertion / deletions (indels), missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression change that results in a new ORF. Mutations can also include splice variants. Tumor cell-specific post-translational modifications can include aberrant phosphorylation. Tumor cell-specific post-translational modifications can also include splice antigens generated by the proteasome. See Liepe et al., "A large fraction of HLA class I ligands are proteasome-generated spliced ​​peptides"; Science. 2016 Oct 21;354(6310):354-358.

[0016] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in tumor cells or tissues of a subject, but is not present in the corresponding normal cells or tissues of the subject.

[0017] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct that is based on one or more neoantigens, eg, multiple neoantigens.

[0018] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.

[0019] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.

[0020] As used herein, the term "coding mutation" refers to a mutation that occurs in a coding region.

[0021] As used herein, the term "ORF" means open reading frame.

[0022] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that arises due to mutation or other abnormalities such as splicing.

[0023] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.

[0024] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon.

[0025] As used herein, the term "frameshift mutation" is a mutation that causes an alteration in the frame of a protein.

[0026] As used herein, the term "indel" refers to the insertion or deletion of one or more nucleic acids.

[0027] As used herein, the term "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences in which a certain percentage of nucleotides or amino acid residues are the same when compared and aligned for maximum correspondence, as determined using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN, or other algorithms available to those of skill in the art), or by visual inspection. Depending on the application, the "percent identity" can exist over a region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.

[0028] In sequence comparison, generally, one sequence serves as a reference sequence to which test sequences are compared.When using a sequence comparison algorithm, test sequences and reference sequences are input into a computer, subsequence coordinates are designated if necessary, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the percent sequence identity (%) of the test sequence to the reference sequence based on the designated program parameters.Alternatively, sequence similarity or difference can also be established by the combination of the presence or absence of a specific nucleotide at a selected sequence position (e.g., sequence motif) or an amino acid in a translated sequence.

[0029] Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (see generally Ausubel et al., infra).

[0030] One example of an algorithm that is suitable for determining percent sequence identity and percent sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.

[0031] As used herein, the term "non-stop or read-through" refers to a mutation that results in the removal of the natural stop codon.

[0032] As used herein, the term "epitope" refers to a specific portion of an antigen that is typically bound by an antibody or T-cell receptor.

[0033] As used herein, the term "immunogenic" refers to the ability to elicit an immune response, for example, via T cells, B cells, or both.

[0034] As used herein, the terms "HLA binding affinity" and "MHC binding affinity" refer to the affinity of binding between a specific antigen and a specific MHC allele.

[0035] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.

[0036] As used herein, the term "mutation" is a difference between the nucleic acid of a subject and a reference human genome used as a control.

[0037] As used herein, the term "variant calling" is the algorithmic determination, typically from sequencing, of the presence of a mutation.

[0038] As used herein, the term "polymorphism" refers to a germline mutation, ie, a mutation found in all DNA-bearing cells of an individual.

[0039] As used herein, the term "somatic mutation" is a mutation that occurs in a non-germline cell of an individual.

[0040] As used herein, the term "allele" refers to one version of a gene or one version of a gene sequence or one version of a protein.

[0041] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.

[0042] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the degradation of mRNA by the cell due to a premature stop codon.

[0043] As used herein, the term "truncal mutation" is a mutation that occurs early in the development of a tumor and is present in the majority of the cells of the tumor.

[0044] As used herein, the term "subclonal mutation" is a mutation that occurs late in the development of a tumor and is present in only a portion of the cells of the tumor.

[0045] As used herein, the term "exome" refers to the subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.

[0046] As used herein, the term "logistic regression" is a regression model for binary data from statistics in which the logit of the probability that the dependent variable is equal to 1 is modeled as a linear function of the dependent variable.

[0047] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinear transformations typically trained by stochastic gradient descent and backpropagation.

[0048] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.

[0049] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome can also refer to the properties of a cell or a collection of cells (e.g., a tumor peptidome refers to the union of the peptidomes of all cells that comprise a tumor).

[0050] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.

[0051] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.

[0052] As used herein, the term "MHC multimer" is a peptide-MHC complex consisting of multiple peptide-MHC monomer units.

[0053] As used herein, the term "MHC tetramer" is a peptide-MHC complex consisting of four peptide-MHC monomer units.

[0054] As used herein, the term "tolerance or immune tolerance" refers to a state of immune unresponsiveness to one or more antigens, eg, self-antigens.

[0055] As used herein, the term "central tolerance" is tolerance conferred in the thymus by either deleting autoreactive T cell clones or promoting their differentiation into immunosuppressive regulatory T cells (Tregs).

[0056] As used herein, the term "peripheral tolerance" refers to tolerance conferred in the peripheral system by downregulating or anergizing autoreactive T cells that survive central tolerance or by promoting the differentiation of these T cells into Tregs.

[0057] The term "sample" can include a single cell, or multiple cells, or fragments of cells, or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.

[0058] The term "subject" includes cells, tissues, or organisms, human or non-human, whether male or female, in vivo, ex vivo, or in vitro. The term subject includes mammals, including humans.

[0059] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.

[0060] The term "clinical factor" refers to a measurement of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values ​​that can be obtained from assessing a subject or a sample (or a population of samples) from a subject under a given condition. A clinical factor can also be predicted by other parameters, such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.

[0061] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.

[0062] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0063] Terms not directly defined herein should be understood to have the meanings generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide further guidance to the practitioner in describing the compositions, devices, methods, etc. of embodiments of the present invention, as well as how to make or use them. It will be recognized that multiple ways of saying the same thing may be used. Accordingly, alternative terms and synonyms may be used for any one or more of the terms discussed herein. No weight should be placed on whether a term is detailed or discussed herein. Several synonyms or alternative methods, materials, etc. are provided. The recitation of one or more synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the inventive embodiments herein.

[0064] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.

[0065] II. Methods for identifying neoantigens Disclosed herein is a method for identifying neoantigens derived from tumor cells of a subject that are likely to be displayed on the surface of tumor cells. The method includes obtaining nucleotide sequencing data of the exome, transcriptome, and / or whole genome from tumor cells and normal cells of the subject. The nucleotide sequencing data is used to obtain a peptide sequence for each neoantigen in a set of neoantigens. The set of neoantigens is identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from normal cells. Specifically, the peptide sequence of each neoantigen in the set of neoantigens contains at least one change that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells. The method further includes encoding the peptide sequence of each neoantigen in the set of neoantigens into a corresponding numeric vector. Each numeric vector contains information describing the amino acids that make up the peptide sequence and the position of the amino acids within the peptide sequence. The method includes associating each peptide sequence of the neoantigen with one or more k-mer blocks of a plurality of k-mer blocks of the nucleotide sequencing data of the subject. The method further includes generating a presentation likelihood for each neoantigen in the set of neoantigens by inputting the numerical vector and associated k-mer blocks into a machine-learned presentation model. Each presentation likelihood represents the likelihood that the corresponding neoantigen will be presented by an MHC allele on the surface of tumor cells of the subject. The machine-learned presentation model includes a plurality of parameters and functions. The plurality of parameters are identified based on a training dataset.The training dataset includes, for each sample among the plurality of samples, labels obtained by mass spectrometry measuring the presence of peptides bound to at least one MHC allele within a set of MHC alleles identified as present in that sample; training peptide sequences encoded as a numerical vector containing information describing the amino acids constituting the peptide and / or the positions of amino acids within the peptide; and, for each training peptide sequence of the sample, an association between the training peptide sequence and one or more k-mer blocks among a plurality of k-mer blocks of nucleotide sequencing data for the training peptide sequence. The function represents a relationship between the numerical vector and associated k-mer block received as input by the machine-learned presentation model, and a presentation likelihood generated as output by the machine-learned presentation model based on the numerical vector, the associated k-mer block, and a plurality of parameters. The method further includes selecting a subset of the set of neoantigens based on the presentation likelihood to generate a set of selected neoantigens, and returning the set of selected neoantigens.

[0066] In some embodiments, inputting the numerical vectors into the machine-learned presentation model includes applying the machine-learned presentation model to a peptide sequence of the neoantigen to generate a dependency score for each MHC allele. The dependency score for an MHC allele indicates whether the MHC allele will present the neoantigen based on a particular amino acid at a particular position in the peptide sequence. In further embodiments, inputting the numerical vectors into the machine-learned presentation model further includes, for each MHC allele, transforming the dependency scores to generate a corresponding per-allele likelihood indicating the likelihood that the corresponding MHC allele will present the corresponding neoantigen, and combining the per-allele likelihoods to generate a presentation likelihood for the neoantigen. In some embodiments, transforming the dependency scores models presentation of the neoantigen as mutually exclusive across MHC alleles. In alternative embodiments, inputting the numerical vectors into the machine-learned presentation model further includes transforming a combination of dependency scores to generate a presentation likelihood. In such embodiments, transforming a combination of dependency scores models presentation of the neoantigen as interference between MHC alleles.

[0067] In some embodiments, the set of presentation likelihoods is further specified by one or more allele-non-interacting features. In such embodiments, the method further includes generating a dependency score for the allele-non-interacting feature by applying a machine-learned presentation model to the allele-non-interacting feature. The dependency score indicates whether the corresponding neoantigen peptide sequence is presented based on the allele-non-interacting feature. In some embodiments, the one or more allele-non-interacting features include a value indicating the presence or absence of a presentation hotspot in each k-mer block of each neoantigen peptide sequence.

[0068] In some embodiments, the method further comprises combining the dependency scores for each MHC allele with the dependency scores for the allele-non-interacting trait, transforming the combined dependency scores for each MHC allele to generate a per-allele likelihood for each MHC allele, and combining the per-allele likelihoods to generate a presentation likelihood. The per-allele likelihood for an MHC allele indicates whether that MHC allele will present the corresponding neoantigen. In alternative embodiments, the method further comprises combining the dependency scores for the MHC allele with the dependency scores for the allele-non-interacting trait, and transforming the combined dependency scores to generate a presentation likelihood.

[0069] In some embodiments, the MHC alleles comprise two or more different MHC alleles.

[0070] In some embodiments, the peptide sequence comprises a peptide sequence having a length other than 9 amino acids.

[0071] In some embodiments, encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme.

[0072] In certain embodiments, the plurality of samples comprises at least one of cell lines engineered to express a single MHC allele, cell lines engineered to express multiple MHC alleles, human cell lines obtained or derived from multiple patients, fresh or frozen tissue samples obtained from multiple patients.

[0073] In some embodiments, the training dataset further comprises at least one of data related to a measure of peptide-MHC binding affinity for at least one of said peptides and data related to a measure of peptide-MHC binding stability for at least one of said peptides.

[0074] In some embodiments, the set of representation likelihoods is further identified by the expression levels of MHC alleles in the subject, as measured by RNA-seq or mass spectrometry.

[0075] In some embodiments, the set of presentation likelihoods is further specified by properties including at least one of the predicted affinity between neoantigens and MHC alleles within the set of neoantigens and the predicted stability of the neoantigen-encoded peptide-MHC complex.

[0076] In some embodiments, the set of numerical likelihoods is further identified by a feature that includes at least one of a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence and an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence.

[0077] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of being presented on the tumor cell surface compared to non-selected neoantigens based on a machine-learned presentation model.

[0078] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of inducing a tumor-specific immune response in a subject compared to non-selected neoantigens based on a machine-learned representation model.

[0079] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood, based on a presentation model, of being able to be presented to naive T cells by professional antigen-presenting cells (APCs) compared to unselected neoantigens. In such embodiments, the APCs are optionally dendritic cells (DCs).

[0080] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of being inhibited by central or peripheral tolerance compared to non-selected neoantigens based on a machine learning-driven presentation model.

[0081] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of inducing an autoimmune response against normal tissue in a subject compared to non-selected neoantigens based on a machine learning-based representation model.

[0082] In some embodiments, the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0083] In some embodiments, the method further comprises generating output for constructing a personalized cancer vaccine from the set of selected neoantigens, in such embodiments, the output for the personalized cancer vaccine can comprise at least one peptide sequence or at least one nucleotide sequence encoding the set of selected neoantigens.

[0084] In some embodiments, the machine-learned presentation model is a neural network model. In such embodiments, the neural network model may include a plurality of network models for MHC alleles, each network model being assigned to a corresponding MHC allele among the plurality of MHC alleles and including a series of nodes arranged in one or more layers. In such embodiments, the neural network model may be trained by updating parameters of the neural network model, and parameters of at least two network models are updated together for at least one training iteration. In some embodiments, the machine-learned presentation model may be a deep learning model including one or more layers of nodes.

[0085] In some embodiments, the MHC allele is a class I MHC allele.

[0086] Also disclosed herein is a computer system including a computer processor and a memory having stored thereon computer program instructions that, when executed by the computer processor, cause the computer processor to perform any of the methods described above.

[0087] III. Identification of tumor-specific mutations in neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells). In particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject with cancer, but may not be present in normal tissues from the subject.

[0088] Genetic mutations in tumors can be considered useful for immunological targeting of tumors if they result in changes in the amino acid sequence of proteins exclusively in tumors. Useful mutations include: (1) non-synonymous mutations that result in different amino acids in proteins; (2) read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; (3) splice site mutations that result in the inclusion of an intron in mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) frameshift mutations or deletions that result in new open reading frames with new tumor-specific protein sequences. Mutations can also include one or more of non-frameshift insertions / deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in new ORFs.

[0089] For example, mutated peptides or mutated polypeptides resulting from splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.

[0090] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.

[0091] Various methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. Several techniques have been described, including dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize amplification of target gene regions, typically by PCR. Still other methods rely on the generation of small signal molecules by invasive cleavage followed by mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.

[0092] PCR-based detection means can involve the multiplex amplification of multiple markers simultaneously.For example, it is well known in the art to select PCR primers so as to generate PCR products that do not overlap in size and can be analyzed simultaneously.Alternatively, it is possible to amplify different markers with primers that are differentially labeled and therefore can be differentially detected.Of course, hybridization-based detection means allows the differential detection of multiple PCR products in a sample.Other techniques that allow multiplex analysis of multiple markers are known in the art.

[0093] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA.For example, single nucleotide polymorphisms can be detected by using special exonuclease-resistant nucleotides, as disclosed in Mundy, CR (US Patent No. 4,656,127).According to this method, a primer complementary to the allele sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a specific animal or human.If the polymorphic site on the target molecule contains a nucleotide that is complementary to the specific exonuclease-resistant nucleotide derivative present, this derivative will be incorporated onto the end of the hybridized primer.This incorporation makes the primer resistant to exonucleases, thereby enabling its detection.Since the identity of the exonuclease-resistant derivative of the sample is known, the knowledge that the primer has become resistant to exonucleases reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of exogenous sequence data.

[0094] To determine the identity of the nucleotide at a polymorphic site, a solution-based method can be used (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087)). As in the method of Mundy, U.S. Pat. No. 4,656,127, a primer is used that is complementary to the allelic sequence immediately 3' to the polymorphic site. This method uses a labeled dideoxynucleotide derivative that becomes incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site to determine the identity of the nucleotide at that site. An alternative method, known as Genetic Bit Analysis or GBA, has been described by Goelet, P. et al. (PCT Application No. 92 / 15712). The Goelet, P. et al. method uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' to the polymorphic site. Goelet, P. et al. The method of Goelet, P. et al. uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which either the primer or the target molecule is immobilized on a solid phase.

[0095] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. et al., Nucl. Acids. Res. 17:7779-7784 (1989); Sokolov, B. P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M. et al., Proc. Natl. Acad. Sci. (USA) 88:1143-1147 (1991); Prezant, T. R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al. al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, signal is proportional to the number of incorporated deoxynucleotides, so that polymorphisms occurring in runs of the same nucleotide can result in a signal proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).

[0096] Numerous initiatives obtain sequence information directly from millions of individual molecules of DNA or RNA in parallel. Real-time single-molecule sequencing by synthesis techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent strands of DNA complementary to the template being sequenced. In one method, oligonucleotides 30–50 bases in length are covalently anchored at their 5′ ends to glass coverslips. These anchored strands serve two functions. First, they act as capture sites for the target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis for sequence reading. The capture primers serve as fixed-location sites for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and dye cleavage. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. As the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing-by-synthesis techniques also exist.

[0097] Any suitable sequencing-by-synthesis platform can be used to identify mutations. As mentioned above, four major sequencing-by-synthesis platforms are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing-by-synthesis platforms have also been described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the multiple nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. A capture sequence (also called a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to a support that can double as a universal primer.

[0098] As an alternative to capture sequences, a member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair, e.g., as described in U.S. Patent Application Publication No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.

[0099] Following capture, the sequence can be analyzed by single-molecule detection / sequencing, including, for example, template-dependent sequencing by synthesis, as described, for example, in the Examples and U.S. Patent No. 7,283,337. In sequencing by synthesis, surface-bound molecules are exposed to a multitude of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time, in a step-and-repeat mode. For real-time analysis, a different optical label can be incorporated for each nucleotide, and multiple lasers can be utilized for stimulation of the incorporated nucleotides.

[0100] Sequencing can also include other massively parallel sequencing or next-generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, ThermoPGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel sequencing technologies, and future generations of these technologies, can be used.

[0101] Any cell type or tissue can be used to obtain nucleic acid samples for use in the methods described herein.For example, DNA or RNA samples can be obtained from tumor or body fluids, for example, blood obtained by known techniques (for example, venipuncture) or saliva.Alternatively, nucleic acid testing can be performed on dry samples (for example, hair or skin).In addition, a sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal tissue is of the same tissue type as tumor.A sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal sample is of a different tissue type from tumor.

[0102] The tumor may include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0103] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. Peptides can be acid-eluted from tumor cells or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.

[0104] IV. Neoantigens Neoantigens can comprise nucleotides or polynucleotides. For example, neoantigens can be RNA sequences that encode polypeptide sequences. Neoantigens useful in vaccines can therefore comprise nucleotide sequences or polypeptide sequences.

[0105] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the neoantigen comprises nucleotide sequences (e.g., DNA or RNA) that encode the associated polypeptide sequence.

[0106] The one or more polypeptides encoded by the neoantigen nucleotide sequences can comprise at least one of the following: a binding affinity to MHC with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides; the presence of a sequence motif within or near the peptide that promotes proteasomal cleavage; and the presence of a sequence motif within or near the peptide that promotes TAP transport; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II polypeptides; and the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.

[0107] One or more neoantigens can be present on the surface of a tumor.

[0108] The one or more neoantigens can be immunogenic in a tumor-bearing subject, for example, capable of eliciting a T cell or B cell response in the subject.

[0109] One or more neoantigens that induce an autoimmune response in a subject can be eliminated from consideration in the context of generating a vaccine for a tumor-bearing subject.

[0110] The size of the at least one neoantigenic peptide molecule is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35 , about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derivable therein. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.

[0111] Neoantigenic peptides and polypeptides can be 15 residues or less in length, typically between about 8 and about 11 residues, particularly 9 or 10 residues, for MHC class I; and 6 to 30 residues for MHC class II.

[0112] If desired, longer peptides can be designed in several ways. In one example, if the likelihood of peptide presentation on HLA alleles is predicted or known, the longer peptides can consist of either (1) individual presented peptides with extensions of 2-5 amino acids toward the N- and C-termini of each corresponding gene product; or (2) a concatenation of some or all of the presented peptides, each with its extended sequence. In another example, if sequencing reveals long (more than 10 residues) neo-epitope sequences present in the tumor (e.g., due to frameshifts, readthrough, or intron inclusion resulting in novel peptide sequences), the longer peptides would (3) consist of the entire novel tumor-specific stretch of amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides that are presented to the strongest HLA alleles. In either example, the use of longer peptides may allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.

[0113] Neoantigenic peptides and polypeptides can be presented on HLA proteins. In some embodiments, the neoantigenic peptides and polypeptides are presented on HLA proteins with greater affinity than wild-type peptides. In some embodiments, the neoantigenic peptide or polypeptide can have an IC50 of at least 5000 nM or less, at least 1000 nM or less, at least 500 nM or less, at least 250 nM or less, at least 200 nM or less, at least 150 nM or less, at least 100 nM or less, at least 50 nM or less, or even less.

[0114] In some embodiments, the neoantigenic peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.

[0115] Also provided are compositions comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides can be derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some embodiments, the tumor-specific mutations are driver mutations for a particular cancer type.

[0116] Neoantigenic peptides and polypeptides with desired activities or properties can be modified to confer certain desirable attributes, e.g., improved pharmacological characteristics, while enhancing or at least retaining substantially all of the biological activity of the unmodified peptide, which binds to desired MHC molecules and activates appropriate T cells. For example, neoantigenic peptides and polypeptides can be further subjected to various modifications, such as conservative or non-conservative substitutions, which may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions refer to the replacement of an amino acid residue with another that is biologically and / or chemically similar, e.g., one hydrophobic residue with another hydrophobic residue, or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effects of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2nd Ed. (1984).

[0117] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful for increasing peptide and polypeptide stability in vivo. Stability can be assayed in a number of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally follows: Pooled human serum (type AB, non-heat-inactivated) is defatted by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatography conditions.

[0118] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is typically composed of relatively small, neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that the optional spacer need not be composed of the same residues and can therefore be a hetero- or homo-oligomer. If present, the spacer will usually be at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.

[0119] The neoantigenic peptide can be linked to a T helper peptide at either the amino or carboxy terminus of the peptide, either directly or via a spacer. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxin 830-843, influenza 307-319, and malaria sporozoite peritoneal sites 382-398 and 378-389.

[0120] Proteins or peptides can be produced by any technique known to those of skill in the art, including expressing proteins, polypeptides, or peptides through standard molecular biology techniques, isolating proteins or peptides from natural sources, or chemically synthesizing proteins or peptides. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located on the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.

[0121] In a further embodiment, the neoantigen comprises a nucleic acid (e.g., a polynucleotide) encoding a neoantigenic peptide or a portion thereof. The polynucleotide can be, for example, a single-stranded and / or double-stranded polynucleotide, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), or a polynucleotide having a phosphorothioate backbone, either in a natural or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further embodiment provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host; such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.

[0122] IV. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., a tumor-specific immune response. Vaccine compositions typically include multiple neoantigens selected, e.g., using the methods described herein. Vaccine compositions may also be referred to as vaccines.

[0123] The vaccine can contain 1 to 30 different peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine may contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, It may contain 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine contains 1-30 neoantigen sequences: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 It can contain 6, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.

[0124] In one embodiment, the different peptides and / or polypeptides, or the nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides are capable of binding to different MHC molecules, such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, a vaccine composition comprises coding sequences for peptides and / or polypeptides capable of binding to the most frequently occurring MHC class I molecules and / or MHC class II molecules. Thus, the vaccine composition can comprise different fragments capable of binding to at least two preferred, at least three preferred, or at least four preferred MHC class I molecules and / or MHC class II molecules.

[0125] The vaccine composition may generate a specific cytotoxic T cell response and / or a specific helper T cell response.

[0126] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are provided herein below. The composition can be combined with a carrier, such as a protein, or an antigen-presenting cell, such as a dendritic cell (DC), which can present peptides to T cells.

[0127] An adjuvant is any substance whose incorporation into a vaccine composition enhances or otherwise modifies the immune response to a neoantigen. The carrier can be a scaffold, such as a polypeptide or polysaccharide, to which the neoantigen can be bound. Optionally, the adjuvant is covalently or non-covalently conjugated.

[0128] The ability of adjuvant to increase the immune response to antigen is typically manifested by a significant or substantial increase in immune-mediated reaction or a reduction in disease symptoms.For example, the increase in humoral immunity is typically manifested by a significant increase in the titer of antibody produced against antigen, and the increase in T cell activity is typically manifested in an increase in cell proliferation, or cellular cytotoxicity, or cytokine secretion.Adjuvant can also change immune response, for example, by changing mainly humoral or Th response to mainly cellular or Th response.

[0129] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA206, Montanide ISA 50V, and Montanide. Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila's QS21 stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponins, mycobacterial extracts and synthetic bacterial cell wall mimics, and other proprietary adjuvants such as Ribi's Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are also useful. Several immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison AC; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Several cytokines have been directly linked to influencing dendritic cell migration to lymphoid tissues (e.g., TNF-α), accelerating dendritic cell maturation into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J ImmunotherEmphasis Tumor Immunol. 1996(6):414-418).

[0130] CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in vaccine settings. Other TLR-binding molecules, such as RNA that binds to TLR 7, TLR 8, and / or TLR 9, may also be used.

[0131] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies, such as cyclophosphamide, sunitinib, bevacizumab, Celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony-stimulating factors such as granulocyte-macrophage colony-stimulating factor (GM-CSF, sargramostim).

[0132] A vaccine composition can include more than one different adjuvant. Additionally, a therapeutic composition can include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.

[0133] The carrier (or excipient) can exist independently of the adjuvant. The function of the carrier can be, for example, to increase activity or immunogenicity, to provide stability, to increase biological activity, or to increase serum half-life, particularly to increase the molecular weight of the variant. Furthermore, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, a serum protein such as keyhole limpet hemocyanin, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is tolerated and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be a dextran, such as Sepharose.

[0134] Cytotoxic T cells (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. MHC molecules themselves are located on the cell surface of antigen-presenting cells. Therefore, CTL activation is possible when a trimeric complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when peptides are used to activate CTLs, but also when APCs bearing the respective MHC molecules are added, it can enhance the immune response. Therefore, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.

[0135] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenovirus (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880), etc. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.

[0136] IV.A. Additional Considerations for Vaccine Design and Manufacturing IV.A.1. Determining a set of peptides covering all tumor subclones Truncal peptides, meaning those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides that are predicted to be highly likely to be presented and immunogenic, or if the number of truncal peptides that are predicted to be highly likely to be presented and immunogenic is small enough that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .

[0137] IV.A.2. Neoantigen Prioritization After applying all of the above neoantigen filters, it is possible that more candidate neoantigens remain available for vaccine inclusion than vaccine technology can accommodate. Additionally, uncertainty about various aspects of neoantigen analysis may remain, and trade-offs may exist between various attributes of candidate vaccine neoantigens. Therefore, instead of predetermined filters at each stage of the selection process, an integral multidimensional model can be considered, in which candidate neoantigens are placed in a space with at least the following axes, and selection is optimized using an integral approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmune risk is typically preferred) 2. Probability of sequencing artifacts (lower artifact probabilities are typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferable) 5. Gene Expression (higher expression is typically preferred) 6. HLA gene coverage (a greater number of HLA molecules involved in presenting a set of neoantigens may decrease the probability that tumors will evade immune attack through downregulation or mutation of HLA molecules) 7. HLA class coverage (covering both HLA-I and HLA-II may increase the probability of therapeutic response and decrease the probability of tumor immune evasion)

[0138] V. Methods of Treatment and Preparation Also provided are methods for inducing a tumor-specific immune response in a subject, vaccinating against a tumor, or treating and / or alleviating symptoms of cancer in a subject by administering to the subject one or more neoantigens, such as multiple neoantigens identified using the methods disclosed herein.

[0139] In some embodiments, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desired. The tumor can be any solid tumor, such as breast, ovarian, prostate, lung, kidney, stomach, colon, testicular, head and neck, pancreas, brain, melanoma, and other tissue organ tumors, as well as hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.

[0140] The neoantigen can be administered in an amount sufficient to induce a CTL response.

[0141] The neoantigen can be administered alone or in combination with other therapeutic agents, such as chemotherapeutic agents, radiation, or immunotherapy. Any suitable therapeutic treatment for the particular cancer can be administered.

[0142] In addition, the subject can be further administered with an anti-immunosuppressive / immunostimulatory substance, such as a checkpoint inhibitor.For example, the subject can be further administered with an anti-CTLA antibody, or anti-PD-1 or anti-PD-L1.Blocking CTLA-4 or PD-L1 with an antibody can enhance the immune response against cancerous cells in patients.In particular, blocking CTLA-4 has been shown to be effective when used in vaccination protocols.

[0143] The optimal amount of each neoantigen to be included in the vaccine composition and the optimal dosing regimen can be determined. For example, the neoantigen or its variants can be formulated for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Methods of injection include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administering vaccine compositions are known to those skilled in the art.

[0144] Vaccines can be edited so that the selection, number, and / or amount of neoantigens present in the composition are tissue-, cancer-, and / or patient-specific. For example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. Selection can depend on the specific type of cancer, the state of the disease, earlier treatment regimens, the patient's immune status, and, of course, the patient's HLA haplotype. Furthermore, vaccines can contain components that are personalized according to the individual needs of a particular patient. Examples include altering the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjusting for secondary treatments after a first round or scheme of treatment.

[0145] For compositions to be used as vaccines for cancer, neoantigens with similar normal self-peptides that are abundantly expressed in normal tissues can be avoided or present in low amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a high amount of a particular neoantigen, the respective pharmaceutical composition for treating that cancer can be present in high amount and / or can include more than one neoantigen specific to that particular neoantigen or pathway of that neoantigen.

[0146] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In therapeutic applications, the compositions are administered to patients in an amount sufficient to elicit an effective CTL response against the tumor antigen and cure or at least partially halt symptoms and / or complications. An amount adequate to achieve this is defined as a "therapeutically effective dose." An amount effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general health, and the judgment of the prescribing physician. It should be kept in mind that compositions can generally be used in severe disease states, i.e., life-threatening or potentially life-threatening situations, particularly when cancer has metastasized. In such instances, it is possible, and the treating physician may find it desirable, to administer substantial excesses of these compositions, taking into account the minimization of adventitious substances and the relatively non-toxic nature of the neoantigens.

[0147] For therapeutic use, administration can begin at the time of detection or surgical removal of a tumor, followed by boosting doses until at least symptoms are substantially abated, and for a period thereafter.

[0148] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against tumors. Disclosed herein are compositions for parenteral administration that contain a solution of a neoantigen, where the vaccine composition is dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. Various aqueous carriers can be used, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional, well-known sterilization techniques or sterile filtered. The resulting aqueous solutions can be packaged for use as is or lyophilized, with the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances required to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, and the like, for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.

[0149] Neoantigens can also be administered via liposomes, which target them to specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigen to be delivered is incorporated as part of the liposome, either alone or in combination with a molecule that binds to a receptor dominant among lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Liposomes filled with the desired neoantigen can thus be directed to the site of lymphoid cells, where they then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid lability, and stability of the liposomes in the bloodstream. Various methods are available for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.

[0150] For targeting to immune cells, the ligand to be incorporated into the liposome can include, for example, an antibody or fragment thereof specific for a cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at doses that vary according to, inter alia, the mode of administration, the peptide being delivered, and the stage of the disease being treated.

[0151] The peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient for therapeutic or immunization purposes. Numerous methods are conveniently used to deliver nucleic acids to a patient. For example, nucleic acids can be delivered directly as "naked DNA." This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), and U.S. Pat. Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery, as described, for example, in U.S. Pat. No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, DNA can be attached to particles, such as gold particles. Approaches for delivering nucleic acid sequences include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.

[0152] Nucleic acids can also be delivered by complexing them with cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO 96 / 18372; 9324640 WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833; 9106309 WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).

[0153] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter can also be included in a viral vector-based vaccine platform, such as the ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20 ( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.

[0154] A means of administering nucleic acids uses minigene constructs encoding one or more epitopes. To generate DNA sequences (minigenes) encoding selected CTL epitopes for expression in human cells, the amino acid sequences of the epitopes are reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are then directly adjacent to generate a continuous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the minigene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.

[0155] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) compounds can also be complexed with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.

[0156] Also disclosed herein is a method of producing a tumor vaccine, comprising performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising multiple neoantigens or a subset of multiple neoantigens.

[0157] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector (e.g., a vector comprising at least one sequence encoding one or more neoantigens) disclosed herein can include culturing host cells under conditions suitable for expression of the neoantigen or vector, wherein the host cells comprise at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatographic, electrophoretic, immunological, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.

[0158] The host cell can comprise a Chinese hamster ovary (CHO) cell, an NS0 cell, yeast, or an HEK293 cell. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to the at least one nucleic acid sequence encoding the neoantigen or vector. In certain embodiments, the isolated polynucleotide can be a cDNA.

[0159] VI. Identification of neoantigens VI.A. Identification of Candidate Neoantigens A research method for NGS analysis of tumor and normal exomes and transcriptomes is described and applied in the specific space of neoantigens. 6,14,15The examples below consider certain optimizations for greater sensitivity and specificity for identifying neoantigens in a clinical setting. These optimizations can be grouped into two areas: those related to laboratory processes and those related to NGS data analysis.

[0160] VI.A.1. Laboratory Process Optimization The process improvements presented herein build on the concepts developed for reliable assessment of cancer driver genes in targeted cancer panels. 16 This addresses the challenges in high-precision neoantigen discovery from low tumor content and small volume clinical specimens by expanding the current approach to the whole-exome and whole-transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) unique average coverage across the tumor exome to detect mutations present at low mutant allele frequency due to either low tumor content or subclonal status. 2. Fewer than 5% of bases are covered at less than 100x to minimize missed potential neoantigens, e.g. a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for areas that are not sufficiently covered 3. Targeting uniform coverage across the normal exome, with less than 5% of bases covered below 20x, to minimize the chance of potential neoantigens remaining unclassified for somatic / germline status (and therefore unusable as TSNAs). 4. To minimize the total amount of sequencing required, sequence capture probes are designed only for the coding regions of genes, since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing18 . b. Elimination of genes predicted to produce few or no candidate neoantigens due to factors such as poor expression, suboptimal digestion by the proteasome, or atypical sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples can be subjected to probe-based enrichment with the same or similar probes used to capture the exome in DNA. 19 It is extracted using

[0161] VI.A.2. Optimizing NGS Data Analysis Analytical method improvements address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically allow for customization relevant for identifying neoantigens in the clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment, as it contains multiple MHC region assemblies that better reflect population polymorphism, as opposed to earlier genome releases. 2. Various programs 5 Overcoming the limitations of single mutation callers 20 by merging results from a. Single nucleotide mutations and indels are detected in tumor DNA, tumor RNA, and normal DNA with a range of tools including: Strelka 21 and Mutec t22 and programs based on comparison of tumor and normal DNA, such as; and 23 , UNCeqR, and other programs that incorporate tumor DNA, tumor RNA, and normal DNA. b. Indels are found in Strelka and ABRA 24 This is determined by a program that performs local reassembly, such as c. Structural rearrangements are 25 or Breakseq26 It is determined using specialized tools such as 3. To detect and prevent sample swapping, mutation calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. Extensive filtering of artificial calls is performed, for example, by: a. Removal of mutations found in normal DNA, potentially with relaxed detection parameters in the case of low coverage and permissive proximity criteria in the case of indels. b. Removal of mutations due to poor mapping quality or poor base quality 27 . c. Elimination of mutations resulting from re-emerging sequencing artifacts, even if not observed in the corresponding normal 27 Examples include mutations that are detected primarily on one strand. d. Removal of mutations detected in a set of unrelated controls 27 . 5.seq2HLA 28 , ATHLATES 29 or Optitype, and also combine exome and RNA sequencing data 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing, such as long-read DNA sequencing. 30 or adaptation of methods for linking RNA fragments to maintain continuity. 31 Includes. 6. Robust Detection of Nascent ORFs Arising from Tumor-specific Splice Variants in CLASS 32 , Bayesembler 33 , StringTie 34 This is done by assembling transcripts from RNA-seq data using Cufflinks, or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate the entire transcripts from each experiment). 35Although commonly used for this purpose, it frequently produces an incredibly large number of splice variants, many of which are much shorter than the full-length gene, and may not be able to recover a simple positive control. The coding sequence and potential nonsense-mediated decay mechanisms reintroduced the mutant sequence, SpliceR. 36 and MAMBA 37 Gene expression is determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels are determined using ASE. 38 or HTSeq 39 Potential filtering steps include: a. Removal of candidate nascent ORFs that are thought to be poorly expressed. b. Removal of candidate nascent ORFs predicted to trigger nonsense-mediated decay (NMD). 7. Candidate neoantigens observed only in RNA (e.g., neo-ORFs) that cannot be directly validated as tumor-specific are classified as likely to be tumor-specific according to additional parameters, for example, by considering the following: a. Presence of supporting cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmed trans-acting mutations in splicing factors in tumor DNA only. As an example, the gene that exhibited the most differential splicing in three independently published experiments with R625 mutant SF3B1 was 1. One experiment examined patients with uveal melanoma. 40 The second experiment examined uveal melanoma cell lines. 41 , and a third study looked at breast cancer patients. 42 Nevertheless, there was agreement. c. For novel splicing isoforms, the presence of confirmatory "novel" splice-junction reads in the RNASeq data. d. For de novo rearrangements, the presence of confirmatory exon-proximal reads in tumor DNA that are not present in normal DNA. e.GTEx 43 and absence from the gene expression compendium (i.e., making germline origin less likely). 8. Complementing reference genome alignment-based analyses by comparing tumor and normal reads (or k-mers derived from such reads) of assembled DNA to directly avoid alignment- and annotation-based errors and artifacts (e.g., for somatic mutations occurring near germline mutations or repeat-context indels).

[0162] In samples with polyadenylated RNA, the presence of viral and microbial RNA in the RNA-seq data will be assessed using RNA CoMPASS44 or similar methods to identify additional factors that may predict patient response.

[0163] VI.B. HLA Peptide Isolation and Detection Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) techniques after lysis and solubilization of tissue samples. 55~58 The clarified lysates were used for HLA-specific IP.

[0164] Immunoprecipitation was performed using antibodies coupled to beads, where the antibodies are specific for HLA molecules. For pan-class I HLA immunoprecipitation, a pan-class I CR antibody is used, and for class II HLA-DR, an HLA-DR antibody is used. The antibodies are covalently attached to NHS-Sepharose beads during overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60Immunoprecipitation can also be performed using antibodies that are not covalently attached to beads. Typically, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to retain the antibody on the column. Some antibodies that can be used to selectively enrich MHC / peptide complexes are listed below. TIFF2025143480000002.tif42149

[0165] The clarified tissue lysate is added to antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is saved for further experiments, including additional IPs. The IP beads are washed to remove nonspecific binding, and the HLA / peptide complexes are eluted from the beads using standard techniques. Protein components are removed from the peptides using molecular weight spin columns or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and, in some cases, stored at -20°C prior to MS analysis.

[0166] The dried peptides were reconstituted in an HPLC buffer suitable for reversed-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution on a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of peptide mass / charge (m / z) were collected at high resolution on an Orbitrap detector, followed by MS2 low-resolution scans on an ion trap detector after HCD fragmentation of selected ions. Additionally, MS2 spectra can be acquired using either CID or ETD fragmentation methods, or any combination of the three techniques to obtain greater amino acid coverage of the peptide. MS2 spectra can also be measured with high-resolution mass accuracy on an Orbitrap detector.

[0167] The MS2 spectra from each analysis were analyzed using Comet 61、62 and peptide identifications were analyzed using Percolator63~65 Further sequencing is performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or spectral matching and de novo sequencing are performed. 75 Sequencing methods including:

[0168] VI.B.1. Investigation of MS detection limits for comprehensive HLA peptide sequencing Using the peptide YVYVADVAAK, the limit of detection was determined using various amounts of peptide loaded onto the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol (Table 1). The results are shown in Figure 1F. These results indicate that the lowest limit of detection (LoD) was in the attomolar range (10 -18 ), a dynamic range spanning five orders of magnitude, and a signal-to-noise ratio in the low femtomole range (10 -15 ) appears to be sufficient for sequencing.

[0169] TIFF2025143480000003.tif56128

[0170] VII. Presented Model VII.A. System Overview 2A is an overview of an environment 100 for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for implementing a presentation identification system 160, which itself includes a presentation information store 165.

[0171] The presentation identification system 160 is a computer model, embodied in a computational system such as that discussed below with respect to FIG. 29, that receives a peptide sequence associated with a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the set of associated MHC alleles. The presentation identification system 160 can be applied to both class I and class II MHC alleles, making it useful in a variety of contexts. One specific example application of the presentation identification system 160 is to receive the nucleotide sequence of a candidate neoantigen associated with a set of MHC alleles from tumor cells in a patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the tumor's associated MHC alleles and / or induce an immunogenic response in the patient's 110 immune system. Those candidate neoantigens with a high likelihood, as determined by the system 160, can be selected for inclusion in a vaccine 118, such that an anti-tumor immune response can be elicited from the immune system of the patient 110 that provided the tumor cells. Furthermore, T cells with TCRs that have reactivity to candidate neoantigens with a high likelihood of presentation can be generated for use in T cell therapy, which also elicits an anti-tumor immune response from the patient's 110 immune system.

[0172] The presentation identification system 160 determines presentation likelihoods through one or more presentation models. Specifically, the presentation models generate likelihoods of whether a given peptide sequence will be presented for a set of associated MHC alleles, and the likelihoods are generated based on the presentation information stored in the storage device 165. For example, the presentation models may generate likelihoods of whether the peptide sequence "YVYVADVAAK" will be presented for the set of alleles HLA-A*02:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:03, and HLA-C*01:04 on the cell surface of a sample. The presentation information 165 contains information about whether these peptides bind to various types of MHC alleles such that the peptides are presented by the MHC alleles, which is determined in the model according to the position of the amino acid in the peptide sequence. Based on the presentation information 165, the presentation models can predict whether an unrecognized peptide sequence will be presented in association with the associated set of MHC alleles. As noted above, the presentation model can be applied to both class I and class II MHC alleles.

[0173] VII.B. Presentation information 2 illustrates a method for obtaining presentation information according to one embodiment. Presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. Allele interaction information includes information that affects presentation of peptide sequences that is dependent on the type of MHC allele. Allele non-interaction information includes information that affects presentation of peptide sequences that is independent of the type of MHC allele.

[0174] VII.B.1. Allelic Interaction Information The allele interaction information primarily includes identified peptide sequences known to be presented by one or more identified MHC molecules from humans, mice, etc. Of note, this may or may not include data obtained from tumor samples. Presented peptide sequences may be identified from cells expressing a single MHC allele. In this example, presented peptide sequences are generally collected from a monoallelic cell line engineered to express a predetermined MHC allele and then exposed to a synthetic protein. Peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows an exemplary peptide presented on the predetermined MHC allele HLA-DRB1*12:01. An example of this is shown in Figure TIFF2025143480000004.tif4128, where a peptide is isolated and identified by mass spectrometry. In this situation, the direct association between the presented peptide and the MHC protein to which it binds is definitively known, since the peptide is identified through cells engineered to express a single, predetermined MHC protein.

[0175] Presented peptide sequences may also be collected from cells expressing multiple MHC alleles. Typically, in humans, six different types of MHC1 molecules and up to 12 different types of MHCII molecules are expressed by cells. Such presented peptide sequences may be identified from multi-allelic cell lines engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal or tumor tissue samples. In this particular example, MHC molecules can be immunoprecipitated from normal or tumor tissue. Peptides presented on multiple MHC alleles can similarly be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows six exemplary peptides. An example of this is shown in TIFF2025143480000005.tif11128, which is presented on the identified class I MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, and HLA-B*08:01, and the class II MHC alleles HLA-DRB1*10:01 and HLA-DRB1:11:01, isolated, and characterized by mass spectrometry. In contrast to monoallelic cell lines, the bound peptide is isolated from the MHC molecule prior to its identification, so the direct association between the presented peptide and the MHC protein to which it is bound may be unknown.

[0176] Allele interaction information can also include mass spectrometry ion currents, which depend on both the concentration of peptide-MHC molecule complexes and the ionization efficiency of the peptides. Ionization efficiency varies from peptide to peptide in a sequence-dependent manner. Generally, ionization efficiency varies from peptide to peptide over approximately two orders of magnitude, while the concentration of peptide-MHC complexes varies over an even larger range.

[0177] Allele interaction information can also include measured or predicted binding affinities between a given MHC allele and a given peptide. One or more affinity models can generate such predictions (72, 73, 74). For example, returning to the example shown in Figure 1D, the representation 165 may represent a binding affinity between the peptide YEMFNDKSF and the class I allele HLA-A. * The presentation information 165 may include a predicted binding affinity value of 1000 nM between the peptide KNFLENFIESOFI and the class II allele HLA-DRB1:11:01. Few peptides with IC50>1000 nM are presented by the MHC, with lower IC50 values ​​increasing the probability of presentation. The presentation information 165 may include a predicted binding affinity value between the peptide KNFLENFIESOFI and the class II allele HLA-DRB1:11:01.

[0178] The allele interaction information can also include measured or predicted stability values ​​for MHC complexes. One or more stability models can generate such predictions. More stable peptide-MHC complexes (i.e., complexes with longer half-lives) are more likely to be presented in high copy number on tumor cells and on antigen-presenting cells that encounter vaccine antigens. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include a predicted stability value for a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a predicted stability value for the half-life of the class II molecule HLA-DRB1:11:01.

[0179] Allele interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at a faster rate are more likely to be presented at high concentrations on the cell surface.

[0180] Allele interaction information can also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of presented peptides have a length of 9 peptides. MHC class II molecules generally tend to present peptides with a length of 6-30 peptides.

[0181] Allele interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoded peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoded peptide. The presence of kinase motifs influences the probability of post-translational modifications that may enhance or interfere with MHC binding.

[0182] Allelic interaction information can also include expression or activity levels of proteins involved in post-translational modification processes, such as kinases (as measured or predicted by RNA-seq, mass spectrometry, or other methods).

[0183] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.

[0184] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual in question (e.g., as measured by RNA-seq or mass spectrometry): peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.

[0185] Allelic interaction information can also include the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals that express that particular MHC allele.

[0186] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and therefore, presentation of peptides by HLA-C is a priori less likely than presentation by HLA-A or HLA-B II. As another example, because HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, presentation of peptides by HLA-DP is predicted to be less likely than presentation by HLA-DR or HLA-DQ.

[0187] The allele interaction information can also include the protein sequence of a particular MHC allele.

[0188] Any of the MHC allele non-interacting information listed in the section below can also be modeled as MHC allele interacting information.

[0189] VII.B.2. Allelic Non-Interaction Information The allele-non-interacting information can include the C-terminal sequence adjacent to the neoantigen-encoded peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence can affect proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters an MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and therefore, the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in Figure 2C, the presentation information 165 can include the C-terminal flanking sequence FOEIFNDKSLDKFJI of the presented peptide FJIEJFOESS, identified from the peptide's source protein.

[0190] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide mass spectrometry training data. As described later with respect to Figure 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, mRNA quantification measurements are determined from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).

[0191] The allele non-interacting information can also include N-terminal sequences adjacent to the peptide within its source protein sequence.

[0192] The allelic non-interaction information can also include a source gene for the peptide sequence. The source gene can be defined as an Ensembl protein family for the peptide sequence. In another example, the source gene can be defined as a source DNA or source RNA for the peptide sequence. The source gene can be represented, for example, as a string of nucleotides that encodes a protein, or alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode specific proteins. In another example, the allelic non-interaction information can also include a source transcript or isoform or a set of potential source transcripts or isoforms for the peptide sequence extracted from a database such as Ensembl or RefSeq.

[0193] The allelic non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.

[0194] The allele non-interaction information can also include the presence of protease cleavage motifs in peptides, optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and therefore less stable in cells, and therefore less likely to be presented.

[0195] Allelic non-interaction information can also include the turnover rate of the source protein when measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation, but this characteristic has low predictive power when measured in dissimilar cell types.

[0196] The allelic non-interaction information can also include the length of the source protein, optionally taking into account the specific splice variants ("isoforms") that are most highly expressed in tumor cells, as measured by RNA-seq or proteome mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.

[0197] Allele-free interaction information can also include the expression level of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Different proteasomes have different cleavage site preferences. More weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.

[0198] Allele-free interaction information can also include the expression of the peptide's source gene (e.g., as measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in the tumor sample. Peptides from genes with higher expression are more likely to be presented. Peptides from genes with undetectable levels of expression can be eliminated from consideration.

[0199] Allelic non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to nonsense-mediated decay as predicted by a model of nonsense-mediated decay, e.g., the model from Rivas et al., Science 2015.

[0200] Allelic non-interaction information can also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle. Genes that are expressed at low levels overall (as measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are displayed than genes that are stably expressed at very low levels.

[0201] Allelic non-interaction information can also include a comprehensive catalog of source protein properties, such as those provided in uniProt or the PDB (http: / / www.rcsb.org / pdb / home / home.do). These properties can include, among others, protein secondary and tertiary structure, subcellular localization, and Gene Ontology (GO) terms. Specifically, this information can include annotations operating at the protein level, e.g., 5'UTR length, and annotations operating at the level of specific residues, e.g., a helix motif between residues 300 and 310. These properties can also include turn motifs, sheet motifs, and disordered residues.

[0202] Allelic non-interacting information can also include features that describe the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (eg, alpha helix versus beta sheet); alternative splicing.

[0203] The allele non-interaction information can also include associations between the neoantigen peptide sequence and one or more k-mer blocks of multiple k-mer blocks of the neoantigen source gene (as present in the nucleotide sequencing data of the subject). During training of the presentation model, these associations between the neoantigen peptide sequence and the k-mer blocks of the neoantigen nucleotide sequencing data are input into the model, and the model uses some of these associations to learn model parameters that represent the presence or absence of presentation hotspots in the k-mer blocks associated with the training peptide sequence. Then, during use of the trained model, associations between the test peptide sequence and one or more k-mer blocks of the test peptide sequence's source gene are input into the model, and the parameters learned by the model during training enable the presentation model to make more accurate predictions of the presentation likelihood of the test peptide sequence.

[0204] Generally, the model parameter representing the presence or absence of a presentation hotspot in a k-mer block represents the residual tendency of the k-mer block to produce a presented peptide after controlling for all other variables (e.g., peptide sequence, RNA expression, amino acids commonly found in HLA-bound peptides). The parameter representing the presence or absence of a presentation hotspot in a k-mer block can be a binary coefficient (e.g., 0 or 1) or an analog coefficient along a particular scale (e.g., 0 to 1 inclusive). In either case, the larger the coefficient (e.g., closer to 1 or 1), the higher the likelihood of producing a presented peptide that controls for other factors, while the smaller the coefficient (e.g., closer to 0 or 0), the lower the likelihood of the k-mer block producing a presented peptide. For example, a k-mer block with a low hotspot coefficient may be from a gene that exhibits high RNA expression for amino acids commonly found in HLA-bound peptides. In this case, the source gene produces many other presented peptides, but few presented peptides are found within the k-mer block. Since other sources of peptide presence are already accounted for by other parameters (e.g., k-mer blocks or larger RNA expression commonly seen for HLA-bound peptides), these hotspot parameters provide additional information that does not "double count" information captured by other parameters.

[0205] Allelic non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression level of the source protein in those individuals and the influence of the various HLA types of those individuals).

[0206] Allelic non-interaction information can also include the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias.

[0207] Expression of various gene modules / pathways (not necessarily containing the source protein of the peptides) as measured by gene expression assays such as RNASeq, microarrays, targeted panels such as Nanostring, or single / multiple genes representing gene modules measured by assays such as RT-PCR, that inform on the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs).

[0208] Allele non-interaction information can also include the copy number of the peptide's source gene in the tumor cell. For example, a peptide derived from a gene that is subject to homozygous deletion in the tumor cell can be assigned a presentation probability of zero.

[0209] The allele-non-interaction information can also include the probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP or that bind with higher affinity to TAP are more likely to be presented by MHC-I.

[0210] Allelic non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Higher TAP expression levels at MHC-I increase the probability of presentation of all peptides.

[0211] Allelic non-interaction information can also include the presence or absence of tumor mutations, including but not limited to: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3. ii. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery affected by loss-of-function mutations in the tumor have a reduced probability of presentation.

[0212] Presence or absence of functional germline polymorphisms, including but not limited to: i. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome).

[0213] The allelic non-interaction information can also include tumor type (eg, NSCLC, melanoma).

[0214] Allele non-interaction information can also include the known functionality of the HLA allele, e.g., as reflected by the HLA allele suffix. For example, the N suffix in the allele name HLA-A*24:09N indicates a null allele that is not expressed and therefore unlikely to present an epitope; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.

[0215] Allelic non-interaction information can also include clinical tumor subtype (eg, squamous cell lung cancer vs. non-squamous).

[0216] The allele non-interaction information can also include smoking history.

[0217] Allele non-interaction information can also include a history of sunburn, sun exposure, or exposure to other mutagens.

[0218] The allelic non-interaction information can also include regional expression of the peptide's source gene in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be represented.

[0219] The allelic non-interaction information can also include the frequency of the mutation in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.

[0220] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation can also include the mutation's annotation (e.g., missense, readthrough, frameshift, fusion, etc.) or whether the mutation is predicted to result in nonsense-mediated decay (NMD). For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in reduced mRNA translation, which reduces the probability of presentation.

[0221] VII.C. Presentation Identification System 3 is a high-level block diagram illustrating the computer logic components of presentation identification system 160, according to one embodiment. In this exemplary embodiment, presentation identification system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. Presentation identification system 160 also comprises a training data store 170 and a presentation model store 175. Some embodiments of model management system 160 have different modules than those described herein. Likewise, functionality may be distributed among the modules in a manner different from that described herein.

[0222] VII.C.1. Data Management Module The data management module 312 generates sets of training data 170 from the representation information 165. Each training data set contains a number of data examples, each of which contains at least one of the represented or unrepresented peptide sequences p i and the peptide sequence p i one or more relevant MHC alleles combined with i and the dependent variable y, which represents information that the presentation identification system 160 is interested in predicting new values ​​of the independent variables. i and the independent variable z i Contains a set of:

[0223] In one particular implementation referred to throughout the remainder of this specification, the dependent variable y i is the peptide p i but one or more associated MHC alleles a i However, in other implementations, the dependent variable y i is the result of the presentation identification system 160 determining the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that one is interested in predicting. For example, in another implementation, the dependent variable y i σ may also be a numerical value indicating the mass analysis ion current determined for the example data.

[0224] Peptide sequence p for data example i i is k i is a sequence of k amino acids, i can vary within a range among data instances i. For example, the range can be 8 to 15 for MHC class I, or 6 to 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p in the training data set are i may have the same length, e.g., 9. The number of amino acids in a peptide sequence may vary depending on the type of MHC allele (e.g., MHC allele in humans). MHC allele a for data example i i is the peptide sequence p i indicates whether it existed in combination with

[0225] The data management module 312 also manages the peptide sequences p contained in the training data 170. i and bound MHC allele a i Together with the binding affinity b i and stability i For example, the training data 170 may include a predictor of the peptide p i and, a i The predicted binding affinity b between each of the bound MHC molecules shown in iAs another example, the training data 170 may contain a i The predicted stability value s for each of the MHC alleles shown in i may contain

[0226] The data management module 312 also receives the peptide sequence p i along with non-allele interacting variables such as C-terminal flanking sequences and mRNA quantification measurements. i It may also include.

[0227] The data management module 312 also identifies peptide sequences that are not presented by MHC alleles to generate the training data 170. Generally, this involves identifying a "longer" sequence of the source protein that contains the peptide sequence to be presented prior to presentation. If the presentation information contains an engineered cell line, the data management module 312 identifies a set of peptide sequences in the synthetic protein to which the cell was exposed that were not presented on the MHC alleles of the cell. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein from which the presented peptide sequence originated and identifies a set of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.

[0228] The data management module 312 also artificially generates peptides with random sequences of amino acids and identifies the generated sequences as peptides that are not presented on MHC alleles. This can be achieved by randomly generating peptide sequences, allowing the data management module 312 to easily generate large amounts of synthetic data for peptides that are not presented on MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, synthetically generated peptide sequences are very likely not presented by MHC alleles, even if they are included in proteins processed by cells.

[0229] 4 illustrates an exemplary set of training data 170A, according to one embodiment. Specifically, the first three data examples in training data 170A are a monoallelic cell line containing the allele HLA-C*01:03, and three peptide sequences: TIFF2025143480000006.tif4128 shows peptide presentation information from TIFF2025143480000006.tif4128. The fourth data example in training data 170A shows peptide information from a multi-allelic cell line containing alleles HLA-B*07:02, HLA-C*01:03, and HLA-A*01:01, and the peptide sequence QIEJOEIJE. The first data example shows that the peptide sequence QCEIOWARE was not presented by the allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, negatively labeled peptide sequences may be randomly generated by the data management module 312 or may be identified from the source protein of the presented peptide. Training data 170A also includes a predicted binding affinity of 1000 nM and a predicted stability value of 1 hour half-life for the peptide sequence-allele pair. Training data 170A also includes the C-terminal flanking sequence of peptide FJELFISBOSJFIE, and the C-terminal flanking sequence of peptide FJELFISBOSJFIE. 2 It also includes allele-non-interacting variables, such as mRNA quantification measurements of TPM. A fourth data example shows that the peptide sequence QIEJOEIJE was presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. Training data 170A also includes predicted binding affinity and stability values ​​for each of the alleles, as well as the C-terminal flanking sequence of the peptide and mRNA quantification measurements for the peptide.

[0230] VII.C.2. Coding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more representation models. In one implementation, the encoding module 314 one-hot encodes sequences (e.g., peptide sequences or C-terminal flanking sequences) for a predetermined 20-letter amino acid alphabet. Specifically, k i Peptide sequence p having amino acids i is 20·k i p, which is represented as a row vector of elements, corresponding to the alphabet of the amino acid at the jth position of the peptide sequence. i 20·(j-1)+1 ,p i 20·(j-1)+2 ,...,p i 20·j A single element in has a value of 1. The remaining elements have a value of 0. As an example, for a given alphabet {A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y}, the three amino acid peptide sequence EAF of data example i is a 60-element row vector TIFF2025143480000007.tif18147. The C-terminal flanking sequence c i , and the protein sequence for the MHC allele d h , and other sequence data in the presentation information can be similarly coded as above.

[0231] If the training data 170 contains sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equivalent length by adding PAD characters to extend the predetermined alphabet. For example, this may be done by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the peptide sequence with the longest length in the training data 170. Thus, if the peptide sequence with the longest length is k 最大 amino acids, the encoding module 314 encodes each sequence as (20+1) k 最大It is represented numerically as a row vector of elements. For example, consider the extended alphabet {PAD,A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y} and k 最大 For a maximum amino acid length of ≡5, the same exemplary peptide sequence EAF of 3 amino acids is represented as a 105-element row vector The C-terminal flanking sequence c i or other sequence data can be similarly encoded as above. Thus, the peptide sequence p i or c i Each argument or row in represents the occurrence of a particular amino acid at a particular position in the sequence.

[0232] Although the above method for encoding sequence data has been described with respect to sequences having amino acid sequences, the method can be similarly extended to other types of sequence data, such as, for example, DNA or RNA sequence data.

[0233] The encoding module 314 also encodes one or more MHC alleles a for data instance i. i is encoded into an m-element row vector, where each element h=1,2,...,m corresponds to a uniquely identified MHC allele. The element corresponding to the identified MHC allele for data instance i has a value of 1. The remaining elements have a value of 0. As an example, among the m=4 uniquely identified MHC allele types {HLA-A*01:01, HLA-C*01:08, HLA-B*07:02, HLA-DRB1*10:01}, the alleles HLA-B*07:02 and HLA-DRB1*10:01 for data instance i corresponding to a multi-allelic cell line are encoded into a four-element row vector a i = [0 0 1 1], and a3 i =1 and a4 i = 1. An example with four identified MHC allele types is described herein, but the number of MHC allele types can actually be hundreds or thousands. As noted above, each data instance i typically contains a peptide sequence pi It contains up to six different MHC allele types associated with

[0234] The encoding module 314 also generates a label y for each data instance i. i We code, as a binary variable with values ​​from the set {0,1}, where a value of 1 indicates that the peptide x i However, the associated MHC allele a i a value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele a i The dependent variable y i If represents the mass analysis ion current, the encoding module 314 may additionally scale the value using various functions, such as a log function with a range of [-∞,∞] for ion current values ​​between [0,∞].

[0235] The coding module 314 encodes the peptide p i and the allele interaction variable x for the associated MHC allele h h i The pairs of alleles x may be represented as row vectors in which the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], and b h i is the predicted binding affinity for peptide p and associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0236] In one example, the encoding module 314 encodes the measured or predicted values ​​for binding affinity as a function of the allele interaction variable x h i The binding affinity information is expressed by incorporating the

[0237] In one example, the encoding module 314 encodes the measured or predicted values ​​for binding stability as allele interaction variables x h i By incorporating it into

[0238] In one example, the encoding module 314 encodes the measured or predicted values ​​for the binding on-rate as a function of the allele interaction variable x h i The combined on-rate information is expressed by incorporating

[0239] In one example, for peptides presented by class I MHC molecules, the encoding module 314 encodes the peptide length in the vector TIFF2025143480000009.tif4128 (However, TIFF2025143480000010.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i In another example, for peptides presented by class II MHC molecules, the encoding module 314 may include the peptide length in the vector TIFF2025143480000011.tif18146 (However, TIFF2025143480000012.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x hi can be included in

[0240] In one example, the encoding module 314 represents the RNA expression information of MHC alleles by incorporating the RNA-seq-based expression levels of the MHC alleles into an allele interaction variable xhi.

[0241] Similarly, the encoding module 314 encodes the allele non-interacting variable w i can be represented as a row vector in which the numerical representations of the allele-non-interacting variables are concatenated one after the other. For example, w i is [c i ] or [c i m i w i ], and w i is the C-terminal flanking sequence of peptide pi and the mRNA quantification measurement m associated with the peptide i Alternatively, one or more combinations of allele non-interacting variables may be stored individually (e.g., as individual vectors or matrices).

[0242] In one example, the encoding module 314 encodes the turnover rate or half-life as a function of the allele non-interacting variable w i represents the turnover rate of the source protein for the peptide sequence.

[0243] In one example, the encoding module 314 encodes the protein length as a function of the allele non-interacting variable w i represents the length of the source protein or isoform by incorporating

[0244] In one example, the encoding module 314 generates β1 i , β2 i , β5 i The mean expression of immunoproteasome-specific proteasome subunits, including the subunits, was calculated using the allele-noninteracting variable w iIncorporation into the IL-1 protein results in activation of the immunoproteasome.

[0245] In one example, the encoding module 314 encodes the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM) or a source protein of a gene or transcript of the peptide, by correlating the source protein abundance with an allele-non-interacting variable w i This is expressed by incorporating it into

[0246] In one example, the encoding module 314 calculates the probability that the transcript of the peptide's origin will undergo nonsense-mediated decay (NMD), for example, as estimated by the model in Rivas et al. Science, 2015, and calculates this probability as a function of the allele non-interaction variable w i This is expressed by incorporating it into

[0247] In one example, encoding module 314 represents the activation status of a gene module or pathway assessed via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM, for each of the genes in the pathway, and then computing a summary statistic, such as a mean, across the genes in the pathway. The mean is calculated using the allele-non-interaction variable w i can be incorporated into

[0248] In one example, the encoding module 314 encodes the copy number of the source gene by dividing the copy number by the allele non-interacting variable w i This is expressed by incorporating it into

[0249] In one example, the encoding module 314 encodes the measured or predicted TAP binding affinity (e.g., in nanomolar units) relative to the allele-non-interacting variable w i The TAP binding affinity is expressed by including

[0250] In one example, the encoding module 314 encodes TAP expression levels measured by RNA-seq (and quantified, for example, by RSEM in units of TPM) as a function of the allele-non-interacting variable w i The expression level of TAP is represented by the inclusion of

[0251] In one example, the encoding module 314 encodes the tumor mutations as allele-non-interacting variables w i vector of indicator variables in (i.e., peptide p k is derived from a sample with a KRAS G12D mutation, d k = 1, otherwise 0).

[0252] In one example, the encoding module 314 encodes germline polymorphisms in antigen-presenting genes as a vector of indicator variables (i.e., peptide p k If is derived from a sample with a specific germline polymorphism in TAP, then d k = 1). These indicator variables are expressed as the allele non-interaction variables w i can be included in

[0253] In one example, the encoding module 314 represents tumor types as one-hot encoding vectors of length 1 for an alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot encoding variables are combined into an allelic non-interaction variable w i can be included in

[0254] In one example, the encoding module 314 represents MHC allele suffixes by processing four-digit HLA alleles with various suffixes. For example, HLA-A*24:09N is considered a different allele from HLA-A*24:09 for purposes of the model. Alternatively, because HLA alleles ending in an N suffix are not expressed, the probability of presentation by an MHC allele with an N suffix can be set to zero for all peptides.

[0255] In one example, the encoding module 314 represents tumor subtypes as one-hot encoding vectors of length 1 for an alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot encoding variables are combined into an allelic non-interaction variable w i can be included in

[0256] In one example, the encoding module 314 encodes smoking history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a smoking history) can be included in k = 1 otherwise 0). Alternatively, smoking history can be coded as a one-hot encoding variable of length 1 for the smoking severity alphabet. For example, smoking status can be assessed on a 1-5 scale, with 1 indicating non-smoker and 5 indicating current heavy smoker. Because smoking history is primarily relevant for lung tumors, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a smoking history and the tumor type is lung tumor, and zero otherwise.

[0257] In one example, the encoding module 314 encodes sunburn history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a history of severe sunburn) can be included in k = 1 otherwise 0). Because severe sunburn is primarily associated with melanoma, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and zero otherwise.

[0258] In one example, the encoding module 314 represents the distribution of expression levels of a particular gene or transcript for each gene or transcript in the human genome as a summary statistic (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, the expression of peptide p in samples with the tumor type melanoma is expressed as a summary statistic (e.g., mean, median) of the distribution of expression levels. k Regarding peptide p k The measured gene or transcript expression levels of the genes or transcripts of origin are compared with the allele-non-interacting variable w i Not only can it be included in the peptide p in melanoma as measured by TCGA, k The mean and / or median gene or transcript expression of the genes or transcripts of a given source may also be included.

[0259] In one example, the encoding module 314 represents the variant types as one-hot encoding variables of length 1 for an alphabet of variant types (e.g., missense, frameshift, NMD-induced, etc.). These one-hot encoding variables are referred to as allele-non-interaction variables w i can be included in

[0260] In one example, the encoding module 314 encodes the protein-level characteristics of the protein as values ​​of the source protein annotation (e.g., 5′ UTR length) and the allele-non-interacting variable w i In another example, the encoding module 314 encodes the peptide p i The residue-level annotation of the source protein for peptide p i is equal to 1 if overlaps with the helical motif, otherwise it is equal to 0, or i The allele non-interaction variable wi represents an indicator variable that is equal to 1 if p is completely contained within the helix motif. In another example, the allele non-interaction variable wi represents an indicator variable that is equal to 1 if p is completely contained within the helix motif. i The property that represents the proportion of residues ini can be included in

[0261] In one example, the encoding module 314 encodes the types of proteins or isoforms in the human proteome into an index vector o having a length equivalent to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is the peptide p k is 1 if comes from protein i, and 0 otherwise.

[0262] In one example, the encoding module 314 encodes the peptide p i Source gene G = gene(p i ) as a categorical variable with L possible categories (where L denotes the upper bound 1, 2, ..., L on the number of subscripted source genes).

[0263] In one example, the encoding module 314 encodes the peptide p i T = tissue type, cell type, tumor type, or tumor histology type of T = tissue (p i ) as a categorical variable with M possible categories (where M denotes an upper limit on the number of subscripted types 1, 2, ..., M). Tissue types can include, for example, lung tissue, cardiac tissue, intestinal tissue, and neural tissue. Cell types can include, for example, dendritic cells, macrophages, and CD4 T cells. Cancers can include, for example, lung adenocarcinoma, lung squamous cell carcinoma, melanoma, and non-Hodgkin's lymphoma.

[0264] The encoding module 314 also encodes the peptide p i and the variable z for the associated MHC allele h i The entire set of alleles is expressed as the allele interaction variable x i and the allele non-interaction variable w i For example, the encoding module 314 may represent z h i[x h i w i ] or [w i x h i ] can be represented as a row vector equivalent to

[0265] VIII. Training Module The training module 316 constructs one or more presentation models that generate a likelihood of whether a peptide sequence will be presented by an MHC allele associated with the peptide sequence. k and peptide sequence p k MHC alleles associated with a k Given a set of peptide sequences, each proposed model k However, the associated MHC allele a k the estimate u, which indicates the likelihood that one or more of k Generate.

[0266] VIII.A. Overview The training module 316 constructs one or more representation models based on a training data set stored in storage 170, which is generated from the representation information stored in 165. Generally, regardless of the specific type of representation model, all representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function TIFF2025143480000013.tif4128 is a graph of the dependent variable y for one or more data examples S in the training data 170. i∈S and the estimated likelihood u for the data example S generated by the proposed model. i∈S In one particular implementation, which will be mentioned throughout the remainder of this document, the loss function TIFF2025143480000014.tif4128 is the negative log likelihood function given by equation (1a) as follows: However, in practice, another loss function may be used. For example, if a prediction is made for mass spectrometry ion current, the loss function is the mean square loss given by Equation 1b as follows: TIFF2025143480000016.tif10128

[0267] The proposed model may be a parametric model, where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function The various parameters of the proposed parametric model that minimizes TIFF2025143480000017.tif4128 are determined through a gradient-based numerical optimization algorithm, such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the proposed model may be a non-parametric model, in which the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.

[0268] VIII.B. Allele-by-Allele Model The training module 316 may build a presentation model to predict the presentation likelihood of a peptide on an allele-by-allele basis. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele.

[0269] In one implementation, the training module 316: TIFF2025143480000018.tif7128 identifies peptides for specific alleles. k The estimated presentation likelihood u k where the peptide sequence x h k is the peptide p k and the coded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, which for convenience of description will be referred to as a transformation function throughout this specification. h(·) is an arbitrary function, which for convenience of description will be referred to as the dependence function throughout this specification, and the parameter θ determined for the MHC allele h h Based on the set of allele interaction variables x h k Generate a dependency score for the parameter θ for each MHC allele h. h The set of values ​​is θ h where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele h.

[0270] Dependence function g h (x h k ;θ h ) output is the MHC allele h with at least the allele interaction characteristic x h k and in particular the peptide p k The dependency score for MHC allele h indicates whether the corresponding neoantigen is presented based on the amino acid position of the peptide sequence of p. For example, the dependency score for MHC allele h is determined by the relationship between the MHC allele h and the peptide p. k The transformation function f(·) transforms the input, more specifically, g in this example. h (x h k ;θ h ) is used to calculate the dependency score for peptide p k is converted to an appropriate value indicating the likelihood that it will be presented by the MHC allele.

[0271] In one particular implementation referred to throughout the remainder of this specification, f(·) is a function with range in [0,1] for the appropriate domain range. In one example, f(·) is The expit function is given by TIFF2025143480000019.tif10128. As another example, f(·) also has the following meaning: for values ​​in the domain z greater than or equal to 0, It can also be the hyperbolic tangent function given by TIFF2025143480000020.tif4128. Alternatively, if the prediction is made for mass analysis ion currents with values ​​outside the range [0,1], f(·) can be any function, for example, the identity function, the exponential function, the log function, etc.

[0272] Therefore, the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is given by the dependency function g h (·) is the peptide sequence p k to generate a corresponding dependency score. The dependency score can be generated by applying k may be transformed by a transformation function f(·) to generate the allele-specific likelihood that h will be presented by MHC allele h.

[0273] VIII.B.1 Dependence Functions for Allelic Interaction Variables In one particular implementation mentioned throughout this specification, the dependency function g h (·) is x h k Each allele interaction variable in is compared with the parameter θ determined for the relevant MHC allele h. h with the corresponding parameters in the set This is an affine function given by TIFF2025143480000021.tif5128.

[0274] In another specific implementation mentioned throughout this specification, the dependency function g h (·) is a network model NN with a set of nodes arranged in one or more layers. h (·), The network function is given by TIFF2025143480000022.tif5128. The nodes are connected by parameters θh A node may be connected to other nodes through connections, each having an associated parameter in a set of . The value at one particular node may be represented as the sum of the values ​​of the nodes connected to the particular node, weighted by the associated parameter mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and process data having amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how these interactions affect peptide presentation.

[0275] Generally speaking, the network model NN h (·) can be structured as feedforward networks such as artificial neural networks (ANNs), convolutional neural networks (CNNs), deep neural networks (DNNs), and / or recurrent networks such as long short-term memory networks (LSTMs), bidirectional recurrent networks, and deep bidirectional recurrent networks.

[0276] In one example, which will be mentioned throughout the remainder of this specification, each MHC allele in h=1, 2,..., m is associated with a separate network model, NN h (·) denotes the output from the network model related to MHC allele h.

[0277] FIG. 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h=3. As shown in FIG. 5, the network model NN3(·) for MHC allele h=3 includes three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2), ..., θ3(10). The network model NN3(·) includes three allele interaction variables x3 for MHC allele h=3. k (1), x3 k (2) and x3 k (3) receives input values ​​(individual data examples, including encoded polypeptide sequence data and any other training data used) and generates the value NN3(x3 k ) The network function may include one or more network models, each taking a different allele interaction variable as input.

[0278] In another example, the identified MHC alleles h=1,2,...,m are used to model the MHC alleles in a single network NN H (·) and NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θ h may correspond to the set of parameters for a single network model, and thus the parameters θ h The set of can be shared by all MHC alleles.

[0279] Figure 6A shows an exemplary network model NN shared by MHC alleles h = 1, 2, ..., m. H (·). As shown in Figure 6A, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. The network model NN3(·) has an allele interaction variable x3 for MHC allele h=3. k , and the value NN3(x3k ) and outputs m values.

[0280] In yet another example, a single network model NN H (·) is the allele interaction variable x for MHC allele h h k and the encoded protein sequence d h In such an example, the parameter θ h may again correspond to the set of parameters for a single network model, and thus the parameters θ h The set of MHC alleles can be shared by all MHC alleles. Therefore, in such an example, NNh(·) is a single network model with inputs [x h k d h ], a single network model NN H Such a network model is advantageous because it can correctly predict peptide presentation probabilities for MHC alleles that were unknown in the training data simply by identifying their protein sequences.

[0281] Figure 6B shows an exemplary network model NN shared by MHC alleles. H (·). As shown in Figure 6B, the network model NN H (·) takes as input the allele interaction variables and protein sequence of MHC allele h=3 and calculates the dependency score NN3 (x3 k ) is output.

[0282] In yet another example, the dependency function g h (·)teeth, TIFF2025143480000023.tif5128, where g' h (x h k ;θ' h) is an affine function, network function, etc., with a set of parameters θ'h, which represents the baseline probability of presentation for MHC allele h, and the bias parameter θ in the set of parameters for the allele interaction variables of the MHC alleles. h 0 accompanied by.

[0283] In another implementation, the bias parameter θ h 0 may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ for the MHC allele h h 0 is θ 遺伝子(h) 0 where gene (h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family "HLA-A," and the bias parameter θ for each of these MHC alleles may be h 0 As another example, if the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 are assigned to the "HLA-DRB" gene family, and the bias parameters θ for each of these MHC alleles are h 0 can be shared.

[0284] As an example, returning to equation (2), the affine dependency function g h Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000024.tif5128 can be generated, where x3 k is the allele interaction variable identified for MHC allele h=3, and θ3 is the set of parameters determined for MHC allele h=3 through loss function minimization.

[0285] As another example, let us consider the MHC allele h = 3 for peptide p among m = 4 different identified MHC alleles using separate network transformation functions gh(·). k The likelihood that will be presented is TIFF2025143480000025.tif5128 can be generated, where x3 k is the allele interaction variable identified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.

[0286] Figure 7 shows the correlation coefficients of peptide p associated with MHC allele h=3 using the exemplary network model NN3(·). k As shown in Figure 7, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0287] VIII.B.2. Per allele with allele-noninteracting variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF2025143480000026.tif7128, peptide p k We model the estimated presentation likelihood uk, where w k is the peptide p k means the coded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interacting variable w Based on the set of allele-non-interacting variables w k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values ​​of θ hand θ w where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele.

[0288] Dependence function g w (w k ;θ w ) output is a measure of the peptide p expression by one or more MHC alleles based on the influence of allele-non-interacting variables. k represents a dependency score for an allele-non-interacting variable, indicating whether peptide p k The C-terminal flanking sequences and peptide p, which are known to positively influence the presentation of k If peptide p is bound, it may have a high value k The C-terminal flanking sequences and peptide p k If bound, it may have a low value.

[0289] According to equation (8), the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is the function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. w (·) is also applied to the coded version of the allele-non-interacting variable to generate a dependency score for the allele-non-interacting variable. Both scores are combined, and the combined score is used to estimate the association of MHC allele h with peptide sequence p k is transformed by a transformation function f(·) to produce the allele-specific likelihoods that will be presented.

[0290] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (2) k the allele interaction variable xh k One may include the allele non-interacting variable wk in the prediction by adding The image can be given by TIFF2025143480000027.tif7128.

[0291] VIII.B.3 Dependence Functions for Allelic Non-Interacting Variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g for allelic non-interacting variables w (·) is an affine function, or a separate network model for the allelic non-interaction variables w k It can be a network function related to

[0292] Specifically, the dependency function g w (·) is w k The allele non-interaction variables in w with the corresponding parameters in the set This is an affine function given by TIFF2025143480000028.tif5128.

[0293] Dependence function g w (·) also corresponds to the parameter θ w the network model NN with relevant parameters in the set w (·), TIFF2025143480000029.tif5128. The network function may include one or more network models, each taking different allele-non-interacting variables as input.

[0294] In another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF2025143480000030.tif5128, where g' w (w k ;θ'w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of m k is the peptide p k is the mRNA quantitative measurement for , h(·) is a function that transforms the quantitative measurement, and θ w m is a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measurement to generate a dependency score for the mRNA quantification measurement. In one particular embodiment mentioned throughout the remainder of this specification, h(·) is a log function, although in practice h(·) can be any one of a variety of different functions.

[0295] In yet another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF2025143480000031.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of k is the peptide p k is the indicator vector described in Section VII.C.2, which represents proteins and isoforms in the human proteome, and θ w o is the set of parameters in the set of parameters for the allele non-interacting variables that are combined with the indicator vector. k and parameter θ w o If the dimension of the set is significantly higher, TIFF2025143480000032.tif4128( Parameter regularization terms such as λ (representing L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined through an appropriate method.

[0296] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025143480000034.tif13128However, g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF2025143480000035.tif4128 is peptide p k is an indicator function equal to 1 if θ is derived from the source gene l as described above for allele-non-interacting variables, and θ w l is a parameter that indicates the "antigenicity" of the source gene l. In one variation, L is sufficiently large, and therefore the number of parameters θ w l=1, 2,...,L If is large enough, Parameter regularization term such as TIFF2025143480000036.tif5128 (where, TIFF2025143480000037.tif4128 can add L1 norm, L2 norm, combinations, etc. to the loss function when determining the parameter value. The optimal value of the hyperparameter λ can be determined by an appropriate method.

[0297] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025143480000038.tif13146However, g' w (w k ;θ'w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF2025143480000039.tif4128 is a peptide p k is derived from source gene l, and peptide p k is an indicator function that is equal to 1 if originates from tissue type m, and θ w lm is a parameter indicating the antigenicity of the combination of source gene l and tissue type m. Specifically, the antigenicity of gene l in tissue type m may indicate the residual tendency of cells of tissue type m to present peptides derived from gene l after adjustment for RNA expression and peptide sequence context.

[0298] In one variation, L or M is sufficiently large, so that the number of parameters θ w lm=1, 2,...,LM If is large enough, Parameter regularization term such as TIFF2025143480000040.tif5128 (where, TIFF2025143480000041.tif4128 can be added to the loss function when determining the parameter values ​​(such as L1 norm, L2 norm, or combination). The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the parameter values ​​so that the coefficients for the same source gene do not vary significantly between tissue types. For example, a penalty term such as: TIFF2025143480000042.tif18128 (However, TIFF2025143480000043.tif5128 is the average antigenicity across tissue types for source gene l) can add a penalty to the standard deviation of antigenicity across different tissue types in the loss function.

[0299] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025143480000044.tif29128 where g' w (w k ;θ' w ) is an affine function, and the set of allelic non-interaction parameters θ' w Network functions with TIFF2025143480000045.tif4128 is peptide p k is an indicator function that is equal to 1 if θ is from the source gene l as described above for allelic non-interacting variables, and θ w l is a parameter indicating the "antigenicity" of the source gene l, TIFF2025143480000046.tif4128 is peptide p k is an indicator function that is equal to 1 if m is from proteome position m, and θ w m is a parameter indicating the degree to which proteome position m is a presentation "hot spot." In one embodiment, a proteome position may contain a block of n adjacent peptides from the same protein, where n is a hyperparameter of the model determined by a suitable method such as grid search cross-validation.

[0300] In practice, the dependence function g on the allelic non-interacting variables can be calculated by combining any of the additional terms in equations (10), (11), (12a), (12b), and (12c). w For example, the term h(·) representing the mRNA quantification measurement in equation (10) and the term representing the antigenicity of the source gene in equation (12) can be added together along with any other affine or network functions to generate a dependency function for the allele-non-interacting variables.

[0301] As an example, returning to equation (8), the affine transformation function g h (·), gw Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000047.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0302] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000048.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0303] FIG. 8 shows exemplary network models NN3(·) and NN w Peptide p associated with MHC allele h=3 using (·) k As shown in Figure 8, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k The allele non-interaction variable w k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0304] VIII.C. Multi-Allele Models The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof. [Example]

[0305] VIII.C.1. Example 1: Maximum Per Allele Model In one implementation, the training module 316 trains peptides p associated with a set of MHC alleles H. k The estimated presentation likelihood u k is the presentation likelihood determined for each of the MHC alleles h in set H determined based on cells expressing a single allele, as explained above in conjunction with equations (2)-(11). TIFF2025143480000049.tif4128. Specifically, the presentation likelihood u k teeth, TIFF2025143480000050.tif4128. In one implementation, the function is a maximum function, as shown in equation (12), where the proposed likelihood u k can be determined as the maximum of the presentation likelihood for each MHC allele h in set H. TIFF2025143480000051.tif5128

[0306] VIII.C.2. Example 2.1: Sum Function Model In one implementation, the training module 316 trains peptides p k The estimated presentation likelihood u k of, TIFF2025143480000052.tif13128, where element a h k is the peptide sequence pk 1 for multiple MHC alleles H associated with x h k is the peptide p k and the coded allele interaction variables for the corresponding MHC alleles. The parameter θ for each MHC allele h h The set of values ​​is θ h The dependence function g can be determined by minimizing a loss function for i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:

[0307] According to equation (13), the peptide sequence p k The likelihood that a given allele will be presented by one or more MHC alleles h is given by the dependency function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate a corresponding score for the allele interaction variable. The scores for each MHC allele h are combined to generate a corresponding score for the peptide sequence p k is transformed by a transformation function f(·) to produce the presentation likelihood that MHC allele H will be presented by the set of MHC alleles H.

[0308] The model presented in equation (13) is that for each peptide p k It differs from the allele-by-allele model of equation (2) in that the number of relevant alleles for a can be greater than 1. In other words, h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with

[0309] For example, the affine transformation function g hUsing (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000053.tif5128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.

[0310] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000054.tif5128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.

[0311] FIG. 9 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0312] VIII.C.3. Example 2.2: Sum Function Model with Allelic Non-Interacting Variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF2025143480000055.tif13130, peptide p k The estimated presentation likelihood u k where w k is the peptide p k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values ​​of θ h and θ w The dependence function g can be determined by minimizing a loss function for i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:

[0313] Therefore, according to equation (14), one or more MHC alleles H can bind to a peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores are combined and the combined score is used to estimate the association of the peptide sequence p with the MHC allele H. k is transformed by a transformation function f(·) to produce the presentation likelihood that

[0314] In the model presented in equation (14), each peptide p k The number of relevant alleles for a can be greater than 1. h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with

[0315] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000056.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0316] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000057.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0317] FIG. 10 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) kAs shown in Figure 10, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. The network model NN3(·) generates the allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k The allele non-interaction variable w k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0318] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (15) k the allele interaction variable x h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF2025143480000058.tif13128.

[0319] VIII.C.4. Example 3.1: Model with Implicit Allele-by-Allele Likelihood In another implementation, the training module 316 k The estimated presentation likelihood u k of, TIFF2025143480000059.tif7128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles h∈H associated with u' k h is the implicit allele-specific presentation likelihood for MHC allele h, and vector v has elements v h But, a hk ·u' k h where s(·) is a vector corresponding to v, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values ​​of the input within a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, although it will be recognized that in other embodiments, s(·) may be any function, such as a maximum function. A set of values ​​for the parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.

[0320] The presentation likelihood in the presentation model of equation (17) is the likelihood that each peptide p is presented by an individual MHC allele h. k The implicit allele-specific presentation likelihood u' corresponds to the likelihood that k h The implicit per-allele likelihood differs from the per-allele presentation likelihood of Section VIII.B in that the parameters for the implicit per-allele likelihood can be learned from a multi-allelic setting, in addition to a single-allelic setting, where the direct association between the presented peptide and the corresponding MHC allele is unknown. Thus, in a multi-allelic setting, the presentation model is based on the likelihood of the peptide p k Not only can we estimate whether peptide p is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are present in peptide p k The individual likelihood u' indicates which person is most likely to have presented k h∈H The advantage of this is that the presented model can generate implicit likelihoods without training data for cells expressing a single MHC allele.

[0321] In one particular implementation that will be mentioned throughout the remainder of this specification, r(·) is a function with range [0,1]. For example, r(·) is a clip function: r(z)=min(max(z,0),1) may be the minimum value between z and 1, which represents the likelihood u k In another implementation, r(·) is chosen as r(z)=tanh(z) where the domain z is greater than or equal to 0.

[0322] VIII.C.5. Example 3.2: Sum of Functions Model In one particular implementation, s(·) is a summation function, and the presentation likelihood is given by summing the implicit per-allele presentation likelihoods. TIFF2025143480000060.tif15128

[0323] In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: Generated by TIFF2025143480000061.tif7128, the presented likelihood is Let it be estimated by TIFF2025143480000062.tif13128.

[0324] According to equation (19), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. Each dependency score can be generated by first applying the implicit per-allele presentation likelihood u' k h The allele likelihood u' is transformed by the function f(·) to generate k h are combined and a clipping function is applied to the combined likelihood to clip the values ​​into the range [0,1] to produce a peptide sequence p k A presentation likelihood can be generated that g will be presented by a set of MHC alleles H. The dependency function g his the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:

[0325] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000063.tif7128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.

[0326] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000064.tif7128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.

[0327] FIG. 11 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k) are then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.

[0328] In another implementation, if the prediction is made in terms of the log of the mass analysis ion current, then r(·) is the log function and f(·) is the exponential function.

[0329] VIII.C.6. Example 3.3: Sum of Functions Model with Allelic Non-Interacting Variables In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: Generated by TIFF2025143480000065.tif7128, the presented likelihood is As generated by TIFF2025143480000066.tif13129, the effects of allelic non-interacting variables are incorporated into peptide presentation.

[0330] According to equation (21), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele non-interaction variables to generate dependency scores for the allele non-interaction variables. The scores of the allele non-interaction variables are combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by the function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined, and a clipping function is applied to the combined output to clip values ​​into the range [0,1] to determine the likelihood of presentation of the peptide sequence p by the MHC allele H. k A likelihood of being presented can be generated. wis the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:

[0331] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000067.tif7128, where w k is the peptide p k are the allele-non-interacting variables identified for θw, and θw is the set of parameters determined for the allele-non-interacting variables.

[0332] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025143480000068.tif7128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0333] FIG. 12 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 12, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. w (·) indicates peptide p kThe allele non-interaction variable w k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·). The network model NN3(·) generates an allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated, which is also the same network model NN w (·) output NN w (w k ) and mapped by a function f(·). Both outputs are combined to give the estimated presentation likelihood u k Generate.

[0334] In another implementation, the implicit per-allele presentation likelihood for MHC allele h can be calculated as: Generated by TIFF2025143480000069.tif7128, the presented likelihood is Generated by TIFF2025143480000070.tif13128.

[0335] VIII.C.7. Example 4: Quadratic Model In one implementation, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k teeth, TIFF2025143480000071.tif13128, where the element u' k h is the implicit per-allele presentation likelihood for MHC allele h. A set of values ​​for parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit per-allele presentation likelihood can be in any of the forms shown in equations (18), (20), and (22) above.

[0336] In one embodiment, the model of equation (23) is k However, there is a possibility that a given antigen may be simultaneously presented by two MHC alleles, which may imply that presentation by the two HLA alleles is statistically independent.

[0337] According to equation (23), one or more MHC alleles H bind to the peptide sequence p k The presentation likelihood is calculated by combining the implicit per-allele presentation likelihood and the MHC allele H k Each pair of MHC alleles is assigned to a peptide p such that it generates a presentation likelihood that p will be presented. k can be generated by subtracting from the sum the likelihood that

[0338] For example, the affine transformation function g h The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF2025143480000072.tif5128 can be generated by k , x3 k are the allele interaction variables identified for HLA alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for HLA alleles h=2, h=3.

[0339] As another example, the network transformation function g h (·), g w The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF2025143480000073.tif5128, where NN2(·) and NN3(·) are the network models specified for HLA alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for HLA alleles h=2 and h=3.

[0340] IX Example 5: Prediction Module The prediction module 320 receives sequence data and selects candidate neoantigens in the sequence data using the proposed model. Specifically, the sequence data may be DNA sequences, RNA sequences, and / or protein sequences extracted from tumor tissue cells of a patient. The prediction module 320 converts the sequence data into a plurality of peptide sequences p having 8-15 amino acids for MHC-I or 6-30 amino acids for MHC-II. k For example, the prediction module 320 processes a given sequence IEFROEIFJEF into three peptide sequences having nine amino acids: TIFF2025143480000074.tif4128. In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by comparing sequence data extracted from a patient's normal tissue cells with sequence data extracted from the patient's tumor tissue cells to identify segments that have one or more mutations.

[0341] The prediction module 320 applies one or more presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation models to the candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences with an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences with the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the selected candidate neoantigens for a given patient can be injected into the patient to induce an immune response.

[0342] X. Example 6: Patient Selection Module The patient selection module 324 selects a subset of patients for vaccine therapy and / or T cell therapy based on whether the patients meet the selection criteria. In one embodiment, the selection criteria are determined based on the patient's likelihood of presentation of neoantigen candidates generated by the presentation model. By adjusting the selection criteria, the patient selection module 324 can adjust the number of patients who receive vaccine administration and / or T cell therapy based on the patient's likelihood of presentation of neoantigen candidates. Specifically, strict selection criteria may result in a smaller number of patients being treated with the vaccine and / or T cell therapy, but a higher proportion of vaccine and / or T cell therapy-treated patients who receive effective treatment (e.g., one or more tumor-specific neoantigens (TSNAs) and / or one or more neoantigen-reactive T cells). In contrast, looser selection criteria may result in a larger number of patients being treated with the vaccine and / or T cell therapy, but a lower proportion of vaccine and / or T cell therapy-treated patients who receive effective treatment. The patient selection module 324 alters the selection criteria based on a desired balance between a target proportion of patients receiving treatment and the proportion of patients receiving effective treatment.

[0343] In some embodiments, the selection criteria for selecting patients to receive vaccine therapy are the same as the selection criteria for selecting patients to receive T cell therapy. However, in alternative embodiments, the selection criteria for selecting patients to receive vaccine therapy may differ from the selection criteria for selecting patients to receive T cell therapy. Sections XA and XB below discuss the selection criteria for selecting patients to receive vaccine therapy and T cell therapy, respectively.

[0344] Selection of patients for XA vaccine treatment In one embodiment, a patient is associated with a corresponding therapeutic subset of v neoantigen candidates that can potentially be included in a personalized vaccine for that patient, having a vaccine volume v. In one embodiment, the therapeutic subset for a patient is the neoantigen candidate with the highest likelihood of presentation as determined by the presentation model. For example, if a vaccine can include v=20 epitopes, the vaccine can include a therapeutic subset for each patient with the highest likelihood of presentation as determined by the presentation model. However, it will be recognized that in other embodiments, the therapeutic subset for a patient can be determined based on other methods. For example, the therapeutic subset for a patient can be randomly selected from the set of neoantigen candidates for that patient, or can be determined based in part on a combination of factors including prior art models that model the binding affinity or stability of peptide sequences, or presentation likelihood obtained from a presentation model and affinity or stability information for those peptide sequences.

[0345] In one embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the patient's tumor mutation burden is equal to or higher than a minimum mutation burden. A patient's tumor mutation burden (TMB) indicates the total number of nonsynonymous mutations in the tumor exome. In one embodiment, the patient selection module 324 selects a patient for vaccine treatment if the patient's absolute TMB number is equal to or higher than a predetermined threshold. In another implementation, the patient selection module 324 selects a patient for vaccine treatment if the patient's TMB is within a threshold percentile among the TMBs determined for the set of patients.

[0346] In another embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the patient's utility score based on the patient's therapeutic subset is equal to or greater than the minimum utility score. In one embodiment, the utility score is a measure of the estimated number of presented antigens from the therapeutic subset.

[0347] The estimated number of presented antigens can be predicted by modeling the presentation of neoantigens as random variables with one or more probability distributions. In one implementation, the utility score for patient i is the expected number of presented neoantigen candidates from the treatment subset, or a specific function thereof. As an example, the presentation of each neoantigen can be modeled as a Bernoulli random variable, where the probability of presentation (success) is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation (success) of each neoantigen candidate is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation of each neoantigen candidate is given by the probability of presentation of each neoantigen candidate with the highest presentation likelihood, u i1 , u i2 , …, u iv v neoantigen candidates p i1 , p i2 , …, p iv Treatment subset S i Regarding neoantigen candidate p ij The presentation of random variable A ij where: TIFF2025143480000075.tif5128. The expected number of neoantigens presented is given by the sum of the likelihoods of presentation of each neoantigen candidate. In other words, the utility score for patient i is given by Represented as TIFF2025143480000076.tif15128. The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility score for the vaccine treatment.

[0348] In another implementation, the utility score for patient i is the probability that at least a threshold number of neoantigens k are presented. In one example, a therapeutic subset S of neoantigen candidates is i The number of presented antigens in is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the likelihood of presentation of each of the epitopes. In particular, the number of presented antigens for patient i is determined by the random variable N i where: TIFF2025143480000077.tif13128, where PBD(·) denotes the Poisson binomial distribution. The probability that at least a threshold number of neoantigens, k, are presented is given by iis given by the probability that k is equal to or greater than k. In other words, the utility score for patient i is given by Represented as TIFF2025143480000078.tif13128. The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility score for the vaccine treatment.

[0349] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of neoantigen candidates that have binding affinities or predicted binding affinities below a fixed threshold (e.g., 500 nM) for one or more of the patient's HLA alleles. i The threshold is the number of neoantigens in the target region. In one example, the fixed threshold is in the range of 1000 nM to 10 nM. Optionally, the utility score may only count neoantigens detected as expressed by RNA-seq.

[0350] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of candidate neoantigens whose binding affinity to one or more HLA alleles for that patient is less than or equal to a threshold percentile of the binding affinity of a random peptide to that HLA allele. i The threshold percentile is the number of neoantigens in the target region. In one example, the threshold percentile is between the 10th percentile and the 0.1th percentile. Optionally, the utility score may only count neoantigens detected as expressed by RNA-seq.

[0351] It will be appreciated that the example utility scores described with respect to equations (25) and (27) are for illustrative purposes only, and the patient selection module 324 may use other statistics or probability distributions to generate utility scores.

[0352] Patient selection for XB T-cell therapy In another embodiment, instead of or in addition to receiving vaccine therapy, a patient can receive T cell therapy. Similar to vaccine therapy, in embodiments in which a patient receives T cell therapy, the patient can be associated with a corresponding therapeutic subset of the v candidate neoantigens as described above. This therapeutic subset of the v candidate neoantigens can be used to identify in vitro T cells from the patient that are reactive to one or more of the v candidate neoantigens. These identified T cells can then be expanded and infused into the patient in a personalized T cell therapy.

[0353] Patients can be selected for T cell therapy at two different time points: the first time point after the model predicts therapeutic subsets of v neoantigen candidates for the patient but before in vitro screening of T cells specific for the predicted therapeutic subsets of v neoantigen candidates; and the second time point after in vitro screening of T cells specific for the predicted therapeutic subsets of v neoantigen candidates.

[0354] First, patients can be selected for T cell therapy after a therapeutic subset of v neoantigen candidates for the patient has been predicted, but before in vitro identification of T cells from the patient that are specific for the predicted therapeutic subset of v neoantigen candidates has been performed. Specifically, because in vitro screening of neoantigen-specific T cells from patients can be costly, it may be desirable to select patients for screening for neoantigen-specific T cells only if the patient is likely to have neoantigen-specific T cells. Patient selection prior to the in vitro T cell screening step can use the same criteria as those used to select patients for vaccine therapy. Specifically, in some embodiments, the patient selection module 324 can select patients for T cell therapy if the patient's tumor mutational burden is equal to or higher than a minimum mutational burden, as described above. In another embodiment, the patient selection module 324 can select patients for T cell therapy if the patient's utility score based on the therapeutic subset of v neoantigen candidates for the patient is equal to or higher than a minimum utility score, as described above.

[0355] Second, in addition to or instead of selecting patients for T cell therapy before in vitro identification of T cells from the patient that are specific for a therapeutic subset of the predicted v neoantigen candidates, patients can also be selected for T cell therapy after in vitro identification of T cells that are specific for a therapeutic subset of the predicted v neoantigen candidates. Specifically, a patient can be selected for T cell therapy if at least a threshold amount of neoantigen-specific TCRs are identified for the patient in an in vitro screen of the patient's T cells for neoantigen recognition. For example, a patient can be selected for T cell therapy only if at least two neoantigen-specific TCRs have been identified for the patient, or only if neoantigen-specific TCRs have been identified for two different neoantigens.

[0356] In another embodiment, a patient can be selected to receive T cell therapy only if a threshold amount of neoantigens from a therapeutic subset of v neoantigen candidates for that patient is recognized by the patient's TCR. For example, a patient can be selected to receive T cell therapy only if at least one neoantigen from a therapeutic subset of v neoantigen candidates for that patient is recognized by the patient's TCR. In a further embodiment, a patient can be selected to receive T cell therapy only if at least a threshold amount of TCRs for that patient are identified as neoantigen-specific for a particular HLA-restricted class of neoantigen peptide. For example, a patient can be selected to receive T cell therapy only if at least one TCR for that patient is identified as a neoantigen-specific HLA class I-restricted neoantigen peptide.

[0357] In yet further embodiments, a patient can be selected to receive T cell therapy only if at least a threshold amount of neoantigen peptides of a particular HLA-restricted class is recognized by the patient's TCR. For example, a patient can be selected to receive T cell therapy only if at least one HLA class I-restricted neoantigen peptide is recognized by the patient's TCR. In another example, a patient can be selected to receive T cell therapy only if at least two HLA class II-restricted neoantigen peptides are recognized by the patient's TCR. Any combination of the above criteria can also be used to select patients for T cell therapy after in vitro identification of T cells specific for the therapeutic subset of v predicted neoantigen candidates for the patient.

[0358] XI. Example 7: Experimental Results Demonstrating Exemplary Patient Selection Performance The validity of the patient selection described in Section X is validated by selecting patients from a set of simulated patients, each associated with a test set of simulated neoantigen candidates, for which a subset of the simulated neoantigens is known to be represented in the mass spectrometry data. Specifically, each simulated neoantigen candidate in the test set is associated with a label indicating whether that neoantigen is represented in the mass spectrometry data set of the multi-allelic JY cell line HLA-A*02:01 and HLA-B*07:02 from the Bassani-Sternberg dataset (dataset "D1") (data available at www.ebi.ac.uk / pride / archive / projects / PXD0000394). As described in more detail below in conjunction with Figure 13A, a number of neoantigen candidates for the simulated patients are sampled from the human proteome based on known frequency distributions of mutation burden in non-small cell lung cancer (NSCLC) patients.

[0359] Allele-specific presentation models for the same HLA alleles are trained using a training set that is a subset of the mass spectrometry data for the single alleles HLA-A*02:01 and HLA-B*07:02 from the IEDB dataset (dataset "D2") (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Specifically, the presentation model for each allele is trained using the network dependency function g h (·) and g wThe allele-specific models were modeled as shown in equation (8), incorporating the allele-specific expression (·) and exponent function f(·). The presentation model for the HLA-A*02:01 allele generates the presentation likelihood of a particular peptide being presented on the HLA-A*02:01 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables. The presentation model for the HLA-B*07:02 allele generates the presentation likelihood of a particular peptide being presented on the HLA-B*07:02 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables.

[0360] As disclosed in the following examples with reference to Figures 13A-13E, different models, such as a presentation model trained for peptide binding prediction and a prior art model, are applied to a test set of neoantigen candidates for each simulated patient to identify different therapeutic subsets for the patient based on the predictions. Patients who meet selection criteria for vaccine treatment are selected and associated with personalized vaccines containing epitopes in the patient's therapeutic subset. The size of the therapeutic subsets varies depending on different vaccine doses. No overlap is introduced between the training set used to train the presentation model and the test set of simulated neoantigen candidates.

[0361] In the following example, the proportion of selected patients with at least a certain number of presented neoantigens among the epitopes included in the vaccine is analyzed. This statistic indicates the effectiveness of the simulated vaccine in delivering potential neoantigens that will elicit an immune response in patients. Specifically, simulated neoantigens in a test set are presented if the neoantigen is presented in mass spectrometry dataset D2. A high proportion of patients with presented neoantigens indicates the likelihood of successful treatment with the neoantigen vaccine by inducing an immune response.

[0362] XI.A. Example 7A: Frequency Distribution of Mutation Burden in NSCLC Cancer Patients Figure 13A shows the sample frequency distribution of mutation burden in NSCLC patients. Mutation burden and mutations in different tumor types, including NSCLC, can be found, for example, in the Cancer Genome Atlas (TCGA) (https: / / cancergenome.nih.gov). The X-axis represents the number of nonsynonymous mutations for each patient, and the Y-axis represents the proportion of sample patients with a specific number of nonsynonymous mutations. The sample frequency distribution in Figure 13A shows a range of 3 to 1786 mutations, with 30% of patients having fewer than 100 mutations. Although not shown in Figure 13A, studies have shown that mutation burden is higher in smokers compared to nonsmokers, and that mutation burden can be a strong indicator of neoantigen burden in patients.

[0363] As introduced at the beginning of Section XI above, each of the simulated patient populations is associated with a test set of neoantigen candidates. Each patient's test set is determined by selecting the mutation load m from the frequency distribution shown in Figure 13A for each patient. i The D1 dataset is generated by sampling the D1 sequence. For each mutation, a 21-mer peptide sequence from the human proteome is randomly selected to represent the mutant sequence to be simulated. A test set of candidate neoantigen sequences is generated for patient i by identifying each (8, 9, 10, 11)-mer peptide sequence across the mutations in the 21-mer. Each candidate neoantigen is associated with a label indicating whether the candidate neoantigen sequence is present in the mass spectrometry D1 dataset. For example, candidate neoantigen sequences present in dataset D1 can be associated with the label "1," and sequences not present in dataset D1 can be associated with the label "0." As described in more detail below, Figures 13B-13E show experimental results of patient selection based on the neoantigens presented by patients in the test set.

[0364] XI.B. Example 7B: Proportion of Selected Patients with Neoantigen Presentation Based on Selection Criteria for Mutational Burden Figure 13B shows the number of presented neoantigens in the simulated vaccine for patients selected based on the selection criteria of whether the patient met a minimum mutational load. The proportion of selected patients who had at least a certain number of presented neoantigens in the corresponding study was identified.

[0365] In Figure 13B, the x-axis shows the proportion of patients excluded from vaccine treatment based on tumor mutation burden, as indicated by the label "Minimum Number of Mutations." For example, the data point at "Minimum Number of Mutations" 200 indicates that the patient selection module 324 selected only a subset of simulated patients with a mutation burden of at least 200 mutations. As another example, the data point at "Minimum Number of Mutations" 300 indicates that the patient selection module 324 selected a lower proportion of simulated patients with at least 300 mutations. The y-axis shows the proportion of selected patients associated with at least a certain number of presented neoantigens in the test set without vaccine dose v. Specifically, the top plot shows the proportion of selected patients presenting at least one neoantigen, the middle plot shows the proportion of selected patients presenting at least two antigens, and the bottom plot shows the proportion of selected patients presenting at least three antigens.

[0366] As shown in Figure 13B, the proportion of patients with presented neoantigens significantly increased with increasing mutational burden, indicating that mutational burden as a selection criterion can be effective in selecting patients in whom neoantigen vaccines are likely to induce an effective immune response.

[0367] XI.C. Example 7C: Comparison of Neoantigen Presentation in Vaccines Identified by Presentation Models and Prior Art Models Figure 13C compares the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on the presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model. The plot on the left assumes a limiting vaccine volume of v=10, and the plot on the right assumes a limiting vaccine volume of v=20. Patients are selected based on a utility score indicating the expected number of presented neoantigens.

[0368] In Figure 13C, the solid lines indicate patients associated with a vaccine containing therapeutic subsets identified based on presentation models for the alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient is identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted lines indicate patients associated with a vaccine containing therapeutic subsets identified based on the prior art model NETMHCpan for the single allele HLA-A*02:01. Implementation details for NETMHCpan are provided at http: / / www.cbs.dtu.dk / services / NetMHCpan. The therapeutic subset for each patient is identified by applying the NETMHCpan model to the sequences in the test set and identifying the v neoantigen candidates with the highest estimated binding affinity. The x-axis of both graphs indicates the proportion of patients excluded from vaccine treatment based on the expected utility score, which indicates the expected number of presented neoantigens in the therapeutic subsets identified based on the presentation model. The expected utility score is determined as described in Section X in connection with equation (25). The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens) included in the vaccine.

[0369] As shown in Figure 13C, a significantly higher proportion of patients associated with vaccines containing therapeutic subsets based on the presentation model receive vaccines containing presented neoantigens than patients associated with vaccines containing therapeutic subsets based on the prior art model. For example, as shown in the graph on the right, 80% of selected patients associated with vaccines based on the presentation model receive at least one presented neoantigen in the vaccine, compared to only 40% of selected patients associated with vaccines based on the prior art model. These results demonstrate that the presentation model described herein is effective in selecting neoantigen candidates for vaccines that are likely to elicit an immune response to treat tumors.

[0370] XI.D. Example 7D: Effect of HLA Coverage on Neoantigen Presentation of Vaccines Identified by Presentation Models Figure 13D compares the number of neoantigens presented in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a single allele-per-presentation model for HLA-A*02:01 and selected patients associated with vaccines containing therapeutic subsets identified based on both HLA-A*02:01 and HLA-B*07:02 allele-per-presentation models. Vaccine volume is set to v = 20 epitopes. For each experiment, patients are selected based on expected utility scores determined based on different therapeutic subsets.

[0371] In Figure 13D, the solid line indicates patients associated with a vaccine containing therapeutic subsets based on both presentation models for HLA alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient was identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with a vaccine containing therapeutic subsets based on a single presentation model for HLA allele HLA-A*02:01. The therapeutic subset for each patient was identified by applying a presentation model for only a single HLA allele to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. In the solid plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility scores for the therapeutic subsets identified by both presentation models. In the dotted plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility scores for the therapeutic subsets identified by a single presentation model. The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens).

[0372] As shown in Figure 13D, patients associated with vaccines containing therapeutic subsets identified by presentation models for both HLA alleles presented neoantigens at a significantly higher rate than patients associated with vaccines containing therapeutic subsets identified by a single presentation model. These results demonstrate the importance of establishing presentation models with high HLA coverage.

[0373] XI.E. Example 7E: Comparison of Neoantigen Presentation in Patients Selected by Mutation Burden and Expected Number of Presented Neoantigens Figure 13E compares the number of presented neoantigens in simulated vaccines between patients selected based on mutation burden and patients selected by expected utility score, which is determined based on the therapeutic subset identified by the presentation model with a size of v = 20 epitopes.

[0374] In Figure 13E, the solid lines indicate patients selected based on the expected utility score associated with a vaccine containing the therapeutic subset identified by the proposed model. The therapeutic subset for each patient is identified by applying each of the proposed models to the sequences in the test set and identifying the v = 20 neoantigen candidates with the highest likelihood of presentation. The therapeutic utility score is determined based on the likelihood of presentation of the therapeutic subset identified in Section X based on Equation (25). The dotted lines indicate patients selected based on the mutational load associated with a vaccine containing the therapeutic subset identified by the proposed model. The x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score in the solid plot and the proportion of patients excluded based on the mutational load in the dotted plot. The y-axis indicates the proportion of selected patients who receive a vaccine containing at least a specific number of presented neoantigens (1, 2, or 3 neoantigens).

[0375] As shown in Figure 13E, patients selected based on expected utility score receive a vaccine containing a higher proportion of presented neoantigens than patients selected based on mutation load. However, patients selected based on mutation load receive a vaccine containing a higher proportion of presented neoantigens than patients not selected. Thus, while mutation load is an effective patient selection criterion for delayed delivery of an effective neoantigen vaccine, expected utility score is more effective.

[0376] XII. Example 8: Evaluation of Mass Spectrometry Training Models on Held-Out Mass Spectrometry Data HLA peptide presentation by tumor cells is a key requirement for antitumor immunity 91,96,97 We used a large (N = 74 patients) integrated dataset of human tumor and normal tissue samples, including paired class I HLA peptide sequences, HLA typing, and transcriptome RNA-seq (Methods), to compare these and published data to predict antigen presentation in human cancers. 92,98,99 Train a novel deep learning model using 100The purpose of this study was to generate peptides for immunotherapy development. Samples were selected based on tissue availability across several tumor types of interest for immunotherapy development. Mass spectrometry identified an average of 3704 peptides per sample (range 344–11301) with a peptide-level FDR of less than 0.1. These peptides ranged in length from 8–15 amino acids and followed a characteristic class I HLA length distribution with a modal length of 9 (56% of peptides). Consistent with previous reports, the majority of peptides (median 79%) were predicted by MHC flurry to bind to at least one patient HLA allele at a standard affinity threshold of 500 nM. 90 However, considerable variability was observed between samples (e.g., in one sample, 33% of the peptides had predicted affinities above 500 nM). 101 A "strong binder" threshold of 50 nM captured a median of only 42% of presented peptides. Transcriptome sequencing yielded an average of 131M unique reads per sample, with 68% of genes expressed at levels of at least 1 TPM (transcript per million) in at least one sample, highlighting the value of large, diverse sample sets designed to observe expression of the maximum number of genes. HLA-mediated peptide presentation was strongly correlated with mRNA expression. Significant and reproducible differences in peptide presentation rates were observed between genes, greater than those explained by differences in RNA expression or sequence alone. The observed HLA types were consistent with expectations for samples derived from a patient population of primarily European ancestry.

[0377] These and published HLA peptide data 92,98,99We used the method to train a neural network (NN) model to predict HLA antigen presentation. To learn an allele-specific model from tumor mass spectrometry data, in which each peptide could be presented by one of six HLA alleles, we developed a novel network architecture (Methods) capable of jointly learning allele-peptide mapping and allele-specific presentation motifs. For each patient, data points designated as positive were peptides detected by mass spectrometry, and data points designated as negative were peptides from a reference proteome (SwissProt) that were not detected by mass spectrometry in that sample. The data were divided into training, validation, and test sets (Methods). The training set consisted of 142,844 HLA-presented peptides (FDR < 0.02) from 101 samples (69 newly described in this study and 32 previously published). The validation set (used for early termination) consisted of 18,004 presented peptides from the same 101 samples. Two mass spectrometry datasets were used for testing: (1) a tumor sample test set consisting of 571 presented peptides from five additional tumor samples (two lung, two colon, and one ovarian) that were omitted from the training data, and (2) a monoallelic cell line test set consisting of 2,128 presented peptides from genomic location windows (blocks) adjacent to (but distinct from) the locations of the monoallelic peptides included in the training data (see Methods for further details regarding the training / test split).

[0378] The training data identified a predictive model for 53 HLA alleles. 92,104Unlike previous studies, these models captured the dependence of HLA presentation on each sequence position in peptides of multiple lengths. The model properly learned a critical dependence on gene RNA expression and gene-specific presentation propensity, and combining mRNA abundance and learned presentation propensity per gene independently yielded up to a 60-fold difference in presentation rate between the lowest expressed, least likely presented genes and the highest expressed, most likely presented genes. The model also predicted the measured stability of HLA / peptide complexes in the IEDB, even after controlling for predicted binding affinity. 88 (p<1e-10 for 10 alleles) and was further observed to be predictive (p<0.05 for 8 of 10 alleles tested). Taken together, these properties form the basis for improved prediction of immunogenic HLA class I peptides.

[0379] The performance of this NN model as a predictor of HLA presentation on a left-out mass spectrometry test set was evaluated. Specifically, Figure 14 compares the positive predictive value (PPV) at 40% recall for different versions of the MS model and a recently published approach (MixMHCPred) that models peptides eluting from mass spectrometry when each model was tested on five different left-out test samples. Figure 14 also shows the average PPV at 40% recall for each model across the five test samples.

[0380] The models tested in Figure 14 (from left to right) are: "Full MS model" (the full NN model described in Methods); "MS model, no flanking sequences" (same as the full NN model except for the removal of flanking sequence features); "MS model, no flanking sequences or per-gene parameters" (same as the full NN model except for the removal of flanking sequence and per-gene parameter features); "Peptide-only MS model, all lengths trained together" (same as the full NN model except for the removal of peptide sequence and HLA type features); "Peptide-only MS model, each length trained separately" (model structure is the same as the peptide-only MS model except that in this model, separate models were trained for 9-mers and 10-mers); "Linear peptide-only MS model with ensemble" (same as the peptide-only MS model, except that instead of using a neural network to model the peptide sequence, an ensemble of linear models was used, trained using the same optimization procedure as used in the full model and described in Methods); and "MixMHCPred" "1.1" is MixMHCPred with default settings; "Binding affinity" is the same MHCflurry 1.2.0.

[0381] The "Full MS model," "MS model without flanking sequences," "MS model without flanking sequences or per-gene parameters," "Peptide-only MS model, trained on all lengths together," "Peptide-only MS model, trained on all lengths separately," and "Linear peptide-only MS model" are all neural network models trained on the mass spectrometry data described above. However, each model is trained and tested using different characteristics of the sample. The "MixMHCPred 1.1" model and the "Binding Affinity" model are early approaches to model HLA-presented peptides. 104MixMHCPred does not currently model peptides of lengths other than 9 and 10, so only 9-mers and 10-mers were used in the comparison. The last five models ("Peptide-only MS model, train all lengths together" through "Binding affinity") have the same input: peptide sequence and HLA type only. Specifically, none of the last five models use RNA abundance to make predictions.

[0382] The best-performing peptide-only model ("peptide-only MS model, all lengths trained together") gave an average PPV of 0.41 at 40% recall, whereas the poorest-performing peptide-only model trained on mass spectrometry data ("linear peptide-only MS model") had an average PPV of only 28% (slightly higher than the 18% average PPV of MixMHCpred1.1), highlighting the value of improved NN modeling of peptide sequences. Note that MixMHCpred1.1 is trained on different data than the linear peptide-only MS model, but has many of the same modeling properties (e.g., it is a linear model in which each peptide length is trained separately).

[0383] Overall, the NN model achieved significantly improved prediction of HLA peptide presentation, with PPV up to 9-fold higher than standard binding affinity + gene expression in the tumor test set. The significant PPV advantage of the NN model with MS persisted across different recall thresholds and was statistically significant (p<10 for all tumor samples). -6 The positive predictive value of standard binding affinity plus gene expression for HLA peptide presentation reached a low value of 6%, consistent with previous estimates. 87,93 However, it should be noted that since only a small percentage of peptides are detected as presented (e.g., approximately 1 in 2500 in the tumor MS test dataset), this approximately 6% PPV represents a greater than 100-fold improvement over baseline abundance.

[0384] By comparing a reduced model trained on mass spectrometry data using only HLA type and peptide sequence as input to the full MS model, it was determined that approximately 30% of the increase in PPV for binding affinity prediction was due to modeling of extrinsic properties of the peptide (RNA abundance, flanking sequences, per-gene parameters) that can be captured by mass spectrometry but not by the binding affinity assay. The remaining approximately 70% of the increase was due to improved modeling of the peptide sequence. This modeling is an early approach in modeling HLA-presented peptides in human tumors. 104 The improved performance was due to the overall model architecture, not just the nature of the training dataset (HLA-presented peptides), as this new model architecture allows for better binding affinity prediction or hard clustering approaches. 104~106 Using this approach, we were able to train allele-specific models through an end-to-end training process that does not require prior assignment of peptides to potential presenting alleles. Importantly, this model architecture does not impose accuracy-decreasing constraints on the allele-specific submodels as a prerequisite for convolution, such as linear convolution, or consider each peptide length separately. 104 The full model outperforms several simplified models and previously published approaches that are subject to these limitations.

[0385] XIII. Example 9: Experimental Results Involving Proposed Hot Spot Modeling To specifically evaluate the benefit of using presentation hotspot parameters in modeling HLA presentation, we compared the performance of a neural network presentation model incorporating presentation hotspot parameters with that of a neural network presentation model without presentation hotspot parameters. The basic neural network architecture was the same for both models and was the same as the presentation model described above in Section VII. Briefly, the models included peptide and flanking amino acid sequence parameters, RNA sequencing transcriptome data (TPM), protein family data, sample-specific identifiers, and HLA-A, B, and C types. An ensemble of five networks was used for each model. For the model including presentation hotspot parameters, we used Equation 12c described above in Section VIII.B.3., with a proteome block size of 10 per gene and peptide lengths of 8–12.

[0386] The two models were compared by conducting experiments using the mass spectrometry dataset described above in Section XII. Specifically, five samples were held out from model training and evaluation to fairly evaluate the competing models. The remaining samples were randomly divided into 90% for model training and 10% for training validation.

[0387] Figure 15A compares the average positive predictive value (PPV) across recall rates for the proposed models with and without the proposed hotspot parameters when each model was tested on five leave-out samples. Models incorporating the proposed hotspot parameters outperformed models without the proposed hotspot parameters for each sample individually, with an average precision of 0.82 with the proposed hotspot parameters and 0.77 without the proposed hotspot parameters.

[0388] Figures 15B-F compare the precision-recall curves of the proposed model with and without the proposed hotspot parameters when each model was tested on each of the five leave-out test samples.

[0389] XIV. Example 10: Evaluation of Presentation Hotspot Parameters to Identify T Cell Epitopes We further directly tested the benefit of using presentation hotspot parameters to model HLA presentation to identify human tumor CD8 T cell epitopes (i.e., targets for immunotherapy). Defining an appropriate test dataset for this evaluation is challenging, as it must contain peptides that are recognized by T cells and presented on the tumor cell surface by HLA. Furthermore, formal performance evaluation requires not only positively labeled (i.e., recognized by T cells) peptides, but also a sufficient number of negatively labeled (i.e., tested but not recognized) peptides. The mass spectrometry dataset corresponds to tumor presentation but not to T cell recognition; conversely, post-vaccination priming or T cell assays correspond to T cell recognition but not to tumor presentation.

[0390] To obtain a suitable dataset, we collected published CD8 T cell epitopes from five recent studies that met the necessary criteria: Study A 96 investigated TILs in nine patients with gastrointestinal tumors and reported T cell recognition of 12 of 1,053 somatic SNV mutations tested by IFN-γ ELISPOT using the tandem minigene (TMG) method in autologous DCs. Study B 84 also reported T cell recognition of 6 of 574 SNVs by CD8+PD-1+ circulating lymphocytes from 5 melanoma patients using TMG. Study C 97 evaluated TILs from three melanoma patients using pulsed peptide stimulation and found responses to five of 381 SNV mutations tested. 108reported the recognition of 2 of 62 SNVs using a combination of TMG assays and minimal epitope peptide-pulsed TILs from a single breast cancer patient. 160 evaluated TILs in 17 patients from the National Cancer Institute with 52 TSNAs. The combined dataset contained 4,843 assayed SNVs from 33 patients, including 75 TSNAs indicative of pre-existing T cell responses. Importantly, because the dataset consists largely of neoantigen recognition by tumor-infiltrating lymphocytes, effective predictions on this dataset demonstrate the model's ability to identify neoantigens presented to T cells by tumors, as well as neoantigens that can prime T cells as described in the section above.

[0391] To stimulate antigen selection for personalized immunotherapy, somatic mutations were ranked in order of their probability of presentation using two methods: (1) an MS model including hotspot features (described in Equation 12c as a block size n = 10), and (2) a traditional MS model without hotspot features. Because the power of antigen-specific immunotherapy is limited in the number of specificities that can be targeted (e.g., current personalized vaccines encode approximately 10–20 mutations), 6、81~82 ), prediction methods were compared by counting the number of pre-existing T cell responses to peptides ranked in the top 5, 10, 20, or 30 for each patient. The results are shown in Table 16.

[0392] Specifically, Figure 16 compares the proportion of peptides spanning somatic mutations recognized by T cells in the top 5, 10, 20, and 30 ranked peptides identified by the proposed model with and without the proposed hotspot parameters for a test set consisting of test samples from patients with at least one pre-existing T cell response. As shown in Figure 16, the model including the hotspot feature performed comparable to the model without the hotspot feature, with both models predicting 45 and 31 T cell responses to the top 20 and 10 ranked peptides, respectively. However, the hotspot model showed improvement in predicting the top 30 and top 5 peptides, with the hotspot model including 6 and 4 more T cell responses, respectively.

[0393] XIII.A. Data The present inventors have followed Gros et al. 84 ,Tran et al. 140 ,Stronen et al. 141 and Zacharakis et al., and Kosaloglu-Yalcin et al. 160 Mutation calling, HLA typing, and T cell recognition data were obtained from supplementary information.

[0394] In the mutation level analysis (Fig. 16), Gros et al., Tran et al., Zacharakis et al. 108 , and Kosaloglu-Yalcin et al. 160Data points designated as positive in the study were mutations recognized by patient T cells in both the TMG assay and the minimal epitope peptide-pulsed assay. Data points designated as negative were all other mutations tested in the TMG assay. In Stronen et al., mutations designated as positive were mutations that spanned at least one recognized peptide, and negative data points were all mutations tested but not recognized in the tetramer assay. Because the mutated 25-mer TMG assay tests T cell recognition of all peptides spanning the mutation, for the Gros, Tran, and Zacharakis data, mutations were ranked by summing the probability of presentation across all peptides spanning the mutation or by taking the lowest binding affinity. For the Stronen data, mutations were ranked by summing the probability of presentation across all peptides spanning the mutation tested in the tetramer assay or by taking the lowest binding affinity. A complete list of mutations and their characteristics is provided in Supplementary Table 1.

[0395] For epitope-level analyses, data points designated as positive were all minimal epitopes recognized by patient T cells in peptide pulse or tetramer assays, and negative data points were all minimal epitopes not recognized by T cells in peptide pulse or tetramer assays, and all peptides spanning mutations from the TMGs tested that were not recognized by patient T cells. In the cases of Gros et al., Tran et al., and Zacharakis et al., minimal epitope peptides spanning mutations recognized in the TMG analysis that were not tested by peptide pulse assays were excluded from the analysis because the T cell recognition status of these peptides was not experimentally investigated.

[0396] XV. Example 11: Identification of neoantigen-responsive T cells in cancer patients In this example, we demonstrate that improved prediction enables the identification of neoantigens from normal patient samples. To do so, archived FFPE tumor biopsies and 5–30 ml of peripheral blood were analyzed from nine patients with metastatic NSCLC undergoing anti-PD(L)1 therapy (Supplementary Table 2: Patient demographics and treatment information for N = 9 patients examined in Figures 17A–C. Key fields include tumor stage and subtype, anti-PD1 therapy administered, and a summary of NGS results). Tumor whole-exome sequencing, tumor transcriptome sequencing, and matched normal exome sequencing yielded an average of 198 somatic mutations (SNVs and short indels) per patient, of which an average of 118 were expressed (Methods, Supplementary Table 2). We applied the full MS model to prioritize 20 neoantigens per patient for testing against pre-existing anti-tumor T cell responses. To focus our analysis on likely CD8 responses, prioritized peptides were synthesized as 8- to 11-mer minimal epitopes ("Methods"). Peripheral blood mononuclear cells (PBMCs) were then cultured with the synthesized peptides in short-term in vitro stimulation (IVS) cultures to expand neoantigen-responsive T cells (Supplementary Table 3). Two weeks later, the presence of antigen-specific T cells was assessed using IFN-γ ELISpot against the prioritized neoepitopes. Separate experiments were further performed in seven patients for whom sufficient PBMCs were available to fully or partially identify the specific antigens recognized. These results are shown in Figures 17A-C and 18A-21.

[0397] Figure 17A shows the detection of T cell responses to patient-specific neoantigen peptide pools in nine patients. For each patient, predicted neoantigens were combined into two pools of 10 peptides each according to model ranking and arbitrary sequence homology (homologous peptides were divided into different pools). Then, for each patient, in vitro expanded PBMCs for that patient were stimulated with the two patient-specific neoantigen peptide pools in an IFN-γ ELISpot. The data in Figure 17A are based on the number of seeded cells 10 minus the background (corresponding DMSO negative control). 5Results are shown as spot-forming units (SFU) per cell. Background measurements (DMSO negative control) are shown in Figure 21. For patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, and CU05, responses of single wells (patients 1-038-001, CU02, CU03, and 1-050-001) or replicates (all other patients) including mean and standard deviation are shown to cognate peptide pools #1 and #2. For patients CU02 and CU03, cell numbers allowed testing only against specific peptide pool #1. Samples with fold increase values ​​greater than 2-fold above background were considered positive and are indicated by an asterisk (responders included patients 1-038-001, CU04, 1-024-001, 1-024-002, and CU02). Non-responders included patients 1-050-001, 1-001-002, CU05, and CU03. Figure 17C shows photographs of ELISpot wells containing in vitro-expanded PBMCs from patient CU04 stimulated with the DMSO negative control, PHA positive control, CU04-specific neoantigen peptide pool #1, CU04-specific peptide 1, CU04-specific peptide 6, and CU04-specific peptide 8 in an IFN-γ ELISpot.

[0398] Figures 18A-B show the results of control experiments using patient neoantigens in HLA-matched healthy donors. The results of these experiments demonstrate that the in vitro culture conditions did not allow for de novo priming in vitro, but rather expanded only pre-existing in vivo primed memory T cells.

[0399] Figure 19 shows the detection of T cell responses to the PHA positive control for each donor and in vitro expansion shown in Figure 17A. For each donor and in vitro expansion shown in Figure 17A, in vitro expanded patient PBMCs were stimulated with PHA for maximal T cell activation. The data in Figure 19 are based on 10 seeded cells minus background (corresponding DMSO negative control). 5Results are shown as spot-forming units (SFU) per donor. Responses of single wells or biological replicates are shown for patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, CU05, and CU03. Patient CU02 was not tested with PHA. Cells from patient CU02 were included in the analysis because their positive response to peptide pool #1 (Figure 17A) indicated viable and functional T cells. As shown in Figure 17A, donors who responded to the peptide pools include patients 1-038-001, CU04, 1-024-001, and 1-024-002. Also shown in Figure 17A, donors who did not respond to the peptide pool include patients 1-050-001, 1-001-002, CU05, and CU03.

[0400] Figure 20A shows the detection of T cell responses to each individual patient-specific neoantigen peptide in pool #2 in patient CU04. Figure 20A also shows the detection of T cell responses to the PHA positive control in patient CU04. (This positive control data is also shown in Figure 19.) For patient CU04, the patient's in vitro expanded PBMCs were stimulated in an IFN-γ ELISpot with the patient-specific individual neoantigen peptides from pool #2 for patient CU04. The patient's in vitro expanded PBMCs were also stimulated in an IFN-γ ELISpot with PHA as a positive control. Data are based on seeded cells 10 minus background (corresponding DMSO negative control). 5 Shown as spot-forming units (SFU) per cell.

[0401] Figure 20B shows the detection of T cell responses to individual patient-specific neoantigen peptides at each of three visits for patient CU04 and at each of two visits for patient 1-024-002 (each visit occurring at a different time point). In both patients, the patient's in vitro expanded PBMCs were stimulated with the patient-specific individual neoantigen peptide in an IFN-γ ELISpot. For each patient, data from each visit are calculated based on the number of seeded cells 10 minus background (corresponding DMSO control).5 The data are presented as cumulative (sum) spot-forming units (SFU) per individual. Data for patient CU04 are presented as background-subtracted cumulative SFU over three visits. For patient CU04, background-subtracted SFU are shown for the first visit (T0) and subsequent visits 2 months (T0+2 months) and 14 months (T0+14 months) after the first visit (T0). Data for patient 1-024-002 are presented as background-subtracted cumulative SFU over two visits. For patient 1-024-002, background-subtracted SFU are shown for the first visit (T0) and subsequent visit 1 month (T0+1 month) after the first visit (T0). Samples with a fold increase value greater than 2-fold above background were considered positive and are indicated by an asterisk.

[0402] Figure 20C shows detection of T cell responses to individual patient-specific neoantigen peptides and to patient-specific neoantigen peptide pools at each of two visits for patient CU04 and patient 1-024-002 (each visit occurring at a different time point). For both patients, the patient's in vitro expanded PBMCs were stimulated with the patient-specific individual neoantigen peptides and the patient-specific neoantigen peptide pool in an IFN-γ ELISpot. Specifically, for patient CU04, in vitro expanded PBMCs from patient CU04 were stimulated in IFN-γ ELISpot with CU04-specific individual neoantigen peptides 6 and 8 and the CU04-specific neoantigen peptide pool, and for patient 1-024-002, in vitro expanded PBMCs from patient 1-024-002 were stimulated in IFN-γ ELISpot with 1-024-002-specific individual neoantigen peptide 16 and the 1-024-002-specific neoantigen peptide pool. Data in Figure 20C are for each technical replicate with mean and range, and represent the number of seeded cells 10 minus the background (corresponding DMSO control). 5Data are presented as spot-forming units (SFU) per individual. Data for patient CU04 are presented as background-subtracted SFU across two visits. For patient CU04, background-subtracted SFU are presented for the first visit (T0, technical triplicate) and a subsequent visit two months after the first visit (T0) (T0 + 2 months, technical triplicate). Data for patient 1-024-002 are presented as background-subtracted SFU across two visits. For patient 1-024-002, background-subtracted SFU are presented for the first visit (T0, technical triplicate) and a subsequent visit one month after the first visit (T0) (T0 + 1 month, technical duplicate except for the sample stimulated with patient 1-024-002-specific neoantigen peptide pool).

[0403] Figure 21 shows the detection of T cell responses to two patient-specific neoantigen peptide pools and a DMSO negative control for the patients in Figure 17A. For each patient, in vitro expanded PBMCs for that patient were stimulated with two patient-specific neoantigen peptide pools in an IFN-γ ELISpot. For each donor and each in vitro expansion, in vitro expanded patient PBMCs were also stimulated with DMSO as a negative control in an IFN-γ ELISpot. The data in Figure 21 are based on 10 seeded cells, including background (corresponding DMSO negative control), for the patient-specific neoantigen peptide pools and corresponding DMSO control. 5Results are shown as spot-forming units (SFU) per well. For patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, and CU05, responses are shown for single wells (patients 1-038-001, CU02, CU03, and 1-050-001) or averages (all other samples) with standard deviations of biological duplicates to cognate peptide pools #1 and #2. For patients CU02 and CU03, cell numbers allowed testing only against specific peptide pool #1. Samples with fold-increase values ​​>2-fold above background were considered positive and are indicated by an asterisk (responding donors included patients 1-038-001, CU04, 1-024-001, 1-024-002, and CU02). Non-responsive donors include patients 1-050-001, 1-001-002, CU05, and CU03.

[0404] As briefly described above with respect to Figures 18A-B, to confirm that the in vitro culture conditions only expanded pre-existing in vivo primed memory T cells rather than allowing de novo priming in vitro, a series of control experiments using neoantigens in HLA-matched healthy donors were performed. The results of these experiments are shown in Figures 18A-B and Supplementary Table 5. The results of these experiments confirmed that de novo priming and detectable neoantigen-specific T cell responses did not occur in healthy donors using the IVS culture method.

[0405] In contrast, pre-existing neoantigen-reactive T cells were identified in the majority of patients (5 / 9, 56%) tested with patient-specific peptide pools using IFN-γ ELISpot (Figure 17A and Figures 19-21). Of the seven patients whose cell counts allowed complete or partial testing of individual neoantigen cognate peptides, four patients responded to at least one of the tested neoantigen peptides, and all of these patients demonstrated responses to the corresponding pools (Figure 17B). The remaining three patients tested with individual neoantigens (patients 1-001-002, 1-050-001, and CU05) did not demonstrate detectable responses to single peptides (data not shown), confirming the lack of responses seen in these patients to the neoantigen pools (Figure 17A). Of the four responding patients, samples from a single visit were obtained for two responding patients (patients 1-024-001 and 1-038-001), and samples from multiple visits were obtained for the remaining two responding patients (CU04 and 1-024-002). For the two patients with samples from multiple visits, cumulative (summed) spot-forming units (SFU) from three visits (patient CU04) and two visits (patient 1-024-002) are shown in Figure 17B, with a breakdown by visit shown in Figure 20B. Additional PBMC samples from the same visits were also obtained for patients 1-024-002 and CU04, and responses to patient-specific neoantigens were confirmed by repeat IVS culture and ELISpot (Figure 20C).

[0406] Overall, among patients in whom at least one T cell-recognized neoepitopes was identified, as shown by responses to the pool of 10 peptides in Figure 17A, the average number of recognized neoepitopes per patient was at least two (counting the non-deconvolutable recognized pool as one recognized peptide, with a minimum of 10 epitopes identified in five patients). In addition to testing for IFN-γ responses by ELISpot, culture supernatants were also tested for granzyme B by ELISA and for TNF-α, IL-2, and IL-5 by MSD cytokine multiplex assay. Of the five patients with positive ELISpots, cells from four secreted three or more test compounds, including granzyme B (Supplementary Table 4), demonstrating polyfunctionality of neoantigen-specific T cells. Importantly, because the combined prediction and IVS method does not rely on a limited set of available MHC multimers, responses were broadly tested across restricted HLA alleles. Furthermore, this approach directly identifies the minimal epitope, unlike tandem minigene screening, which requires a separate deconvolution step to identify the recognized mutation and identify the minimal epitope. Overall, the yield of neoantigen identification is significantly higher than the previous best method of testing TILs for all mutations using apheresis samples. 96 While the results were comparable to those of the previous method, it only required screening of 20 synthetic peptides using the usual 5-30 mL of whole blood.

[0407] XV.A. Peptides Custom-made recombinant lyophilized peptides were purchased from JPT Peptide Technologies (Berlin, Germany) or Genscript (Piscataway, NJ, USA), reconstituted at 10–50 mM in sterile DMSO (VWR International, Pittsburgh, PA, USA), and stored in aliquots at −80°C.

[0408] XV.B.Human peripheral blood mononuclear cells (PBMC) Lyophilized, HLA-typed PBMCs from healthy donors (confirmed to be seronegative for HIV, HCV, and HBV) were purchased from Precision for Medicine (Gladstone, NJ, USA) or Cellular Technology, Ltd. (Cleveland, OH, USA) and stored in liquid nitrogen until use. Fresh blood samples were purchased from Research Blood Components (Boston, MA, USA), and leukopaks were purchased from AllCells (Boston, MA, USA). PBMCs were isolated using a Ficoll-Paque density gradient (GE Healthcare Bio, Marlborough, MA, USA) and then cryopreserved. Patient PBMCs were processed at a local clinical processing center according to local clinical standard operating procedures (SOPs) and IRB-approved protocols. The approving IRBs were the Quorum Review IRB, Comitato Etico Interaziendale AOU San Luigi Gonzaga di Orbassano, and Comite Etico de la Investigacion del Grupo Hospitalario Quiron en Barcelona.

[0409] Briefly, PBMCs were isolated by density gradient centrifugation, washed, counted, and cultured at 5x10 in CryoStor CS10 (STEMCELL Technologies, Vancouver, BC, V6A 1B6, Canada). 6Cells were cryopreserved at 1000 cells / ml. Cryopreserved cells were shipped via cryoport and transferred and stored in LN2 upon arrival. Patient demographics are shown in Supplementary Table 2. Cryopreserved cells were thawed and washed twice in OpTmizer T-cell Expansion Basal Medium (Gibco, Gaithersburg, MD, USA) with Benzonase (EMD Millipore, Billerica, MA, USA) and once without Benzonase. Cell count and viability were assessed using Guava ViaCount reagent and a module on a Guava easyCyte HT cytometer (EMD Millipore). Cells were then resuspended in the appropriate medium at the appropriate concentration for subsequent assays (see next section).

[0410] XV.C. In vitro stimulation (IVS) culture Pre-existing T cells obtained from healthy donors or patient samples were analyzed using the approach applied by Ott et al. 81 The same approach was used to expand PBMCs in the presence of cognate peptides and IL-2. Briefly, thawed PBMCs were rested overnight and stimulated in 24-well tissue culture plates in the presence of peptide pools (10 μM per peptide, 10 peptides per pool) in ImmunoCult™-XF T-cell Expansion Medium (STEMCELL Technologies) supplemented with 10 IU / ml rhIL-2 (R&D Systems Inc., Minneapolis, MN) for 14 days. Cells were grown at 2x10 6 Cells were seeded at 2 x 10 cells / well and cultured by changing 2 / 3 of the medium every 2-3 days. One patient sample demonstrated protocol deviations and should be considered a potential false negative. Patient CU03 did not yield sufficient numbers of cells after thawing, so cells were cultured at 2 x 10 cells / peptide pool. 5 cells (10-fold fewer than stated in the protocol).

[0411] XV.D. IFNγ Enzyme-Linked Immunospot (ELISpot) Assay ELISpot assay for detection of IFNγ-producing T cells 142 Briefly, PBMCs (ex vivo or after in vitro expansion) were harvested, washed in serum-free RPMI (VWR International), and cultured in ELISpot Multiscreen plates (EMD Millipore) coated with anti-human IFNγ capture antibody (Mabtech, Cincinatti, OH, USA) in OpTmizer T-cell Expansion Basal Medium (ex vivo) or ImmunoCult™-XF T-cell Expansion Medium (expanded cultures) in the presence of control or cognate peptide. After 18 h of incubation in a humidified incubator at 37°C with 5% CO2, cells were removed from the plates, and membrane-bound IFNγ was detected using anti-human IFNγ detection antibody (Mabtech), Vectastain Avidin peroxidase conjugate (Vector Labs, Burlingame, CA, USA), and AEC Substrate (BD Biosciences, San Jose, CA, USA). ELISpot plates were dried, stored protected from light, and sent to Zellnet Consulting, Inc., Fort Lee, NJ, USA) for standardized evaluation. Data are presented as spot-forming units (SFU) per number of cells plated.

[0412] XV.E. Granzyme B ELISA and MSD Multiplex Assay Detection of secreted IL-2, IL-5, and TNF-α in ELISpot supernatants was performed using a triplex assay, the MSD U-PLEX Biomarker assay (catalog number K15067L-2). The assay was performed according to the manufacturer's instructions. Test article concentrations (pg / ml) were calculated using serial dilutions of known standards for each cytokine. For data graphing, values ​​below the minimum range of the standard curve were set equal to zero. Detection of granzyme B in ELISpot supernatants was performed using the GranzymeB DuoSet® ELISA (R&D Systems, Minneapolis, MN) according to the manufacturer's instructions. Briefly, ELISpot supernatants were diluted 1:4 in sample diluent and run alongside serial dilutions of granzyme B standard to calculate concentrations (pg / ml). For data graphing, values ​​below the minimum range of the standard curve were set equal to zero.

[0413] Negative control experiment of XV.F.IVS assay - neoantigens derived from tumor cell lines tested in healthy donors Figure 18A shows a negative control experiment of the IVS assay for neoantigens derived from tumor cell lines tested in healthy donors. PBMCs from healthy donors were stimulated during IVS culture with a peptide pool containing a positive control peptide (pre-exposed to an infection), HLA-matched neoantigens from tumor cell lines (unexposed), and peptides derived from pathogens to which the donor was seronegative. Expanded cells were then stimulated with DMSO (negative control, black circles), PHA and common infectious disease peptides (positive control, red circles), neoantigens (unexposed, light blue circles), or HIV and HCV peptides (donor was seronegative; dark blue, A and B), followed by IFNγ ELISpot (10 5 The data were analyzed by seeding cells 10 5 Shown as spot-forming units (SFU) per donor. Means and biological replicates with SEM are shown. No responses were observed to neoantigens or peptides derived from pathogens to which the donor had not been exposed (seronegative).

[0414] Negative control experiment of XV.G.IVS assay - patient-derived neoantigens tested in healthy donors Figure 18A shows a negative control experiment of the IVS assay for patient-derived neoantigens tested for responsiveness in healthy donors. Evaluation of T cell responses in healthy donors to HLA-matched neoantigen peptide pools. Left panel: Healthy donor PBMCs were stimulated with control (DMSO, CEF, and PHA) or HLA-matched patient-derived neoantigen peptides in an ex vivo IFNγ ELISpot. Data are from 2x10 cells plated in triplicate wells. 5 Shown as spot-forming units (SFU) per cell. Right panel: PBMCs from healthy donors after IVS culture expanded in the presence of neoantigen pools or CEF pools were stimulated with control (DMSO, CEF, and PHA) or neoantigen peptide pools from HLA-matched patients in IFNγ ELISpot. Data are for 1 x 10 cells plated for triplicate wells. 5 Shown as SFU per cell. No response to neoantigens was observed in healthy donors.

[0415] XV.H. Supplementary Table 3: Peptides tested for T cell recognition in NSCLC patients Details of neoantigen peptides tested in N=9 patients examined in Figure 17A-C (identification of neoantigen-reactive T cells from NSCLC patients). Key fields include source mutation, peptide sequence, and observed pool and individual peptide sequences. The column "most_probable_restriction" indicates which allele the model predicted was most likely to present each peptide. The rank of these peptides among all mutant peptides for each patient, as calculated by binding affinity prediction ("Methods"), is also included.

[0416] Four peptides were highly ranked by the full MS model and had low predicted binding affinity or were recognized by CD8 T cells that were low ranked by the binding affinity prediction.

[0417] For three of these peptides, this is due to differences in HLA coverage between the model and MHCflurry1.2.0. Peptide YEHEDVKEA is predicted to be presented by HLA-B*49:01, which is not covered by MHCflurry1.2.0. Similarly, peptides SSAAAPFPL and FVSTSDIKSM are predicted to be presented by HLA-C*03:04, which is also not covered by MHCflurry1.2.0. The online NetMHCpan 4.0 (BA) prediction tool, a pan-allele binding affinity prediction tool that in principle covers all alleles, ranks SSAAAPFPL as a strong binder to HLA-C*03:04 (23.2 nM, ranked 2nd in patient 1-024-002), predicts weak binding of FVSTSDIKSM to HLA-C*03:04 (943.4 nM, ranked 39th in patient 1-024-002) and weak binding of YEHEDVKEA to HLA-B*49:01 (3387.8 nM), and stronger binding to HLA-B*41:01 (208.9 nM, ranked 11th in patient 1-038-001), which is also present in this patient but not covered by the model. Thus, of these three peptides, FVSTSDIKSM would have been leaked according to binding affinity prediction, SSAAAPFPL would have been captured, and the HLA restriction of YEHEDVKEA is unknown.

[0418] The remaining five peptides for which peptide-specific T cell responses were deconvolved were derived from patients whose most likely presenting alleles, as determined by the model, were also covered by MHCflurry1.2.0. Four of these five peptides (4 / 5) had predicted binding affinities stronger than the standard 500 nM threshold and were ranked in the top 20, although somewhat lower than their model rankings (peptides (TIFF2025143480000079.tif4128 was ranked 2, 14, 7, and 9 by MHCflurry, whereas it was ranked 0, 4, 5, and 7, respectively, by the model.) The peptide GTKKDVDVLK was recognized by CD8 T cells and ranked 1 by the model, but 70th by MHCflurry, with a predicted binding affinity of 2169 nM.

[0419] Overall, six of the eight individually recognized peptides (6 / 8) that were highly ranked by the full MS model were also highly ranked using binding affinity prediction and had predicted binding affinities below 500 nM, whereas two of the eight individually recognized peptides (2 / 8) would have been missed if binding affinity prediction had been used instead of the full MS model.

[0420] XV.I. Supplementary Table 4: MSD Cytokine Multiplex and ELISA Assays for ELISpot Supernatants Obtained from NSCLC Neoantigen Peptides Test substances detected in supernatants from positive ELISpot (IFNγ) wells are shown for Granzyme B (ELISA), TNFα, IL-2, and IL-5 (MSD). Values ​​are shown as the mean pg / ml from technical replicates. Positive values ​​are shown in italics. Granzyme B ELISA: values ​​≥ 1.5-fold above DMSO background were considered positive. U-Plex MSD assay: values ​​≥ 1.5-fold above DMSO background were considered positive.

[0421] XV.J. Supplementary Table 5: Neoantigens and infectious disease epitopes in IVS control experiments Details of tumor cell line neoantigens and viral peptides tested in IVS control experiments shown in Figures 18A-B. Key fields include source cell line or virus, peptide sequence, and predicted presenting HLA alleles.

[0422] XV.K.Data The MS peptide dataset used to train and test the predictive models (Figure 16) is available at the MassIVE archive (massive.ucsd.edu), accession number MSV000082648. Neoantigen peptides tested by ELISpot (Figures 17A-C and Figures 18A-B) are included with the manuscript (Supplementary Tables 3 and 5).

[0423] XVI. Methods of Examples 8-11 XVI.A. Mass spectrometry XVI.A.1. Sample Archived frozen tissue samples for mass spectrometry analysis were obtained from commercial sources, including BioServe (Beltsville, MD), ProteoGenex (Culver City, CA), iSpecimen (Lexington, MA), and Indivumed (Hamburg, Germany). A subset of samples was also prospectively collected from patients at the Hôpital Marie Lannelongue (Le Plessis-Robinson, France) under a research protocol approved by the Comite de Protection des Personnes, Ile-de-France VII.

[0424] XVI.A.2. HLA Immunoprecipitation Isolation of HLA peptide molecules after lysis and solubilization of tissue samples 87,124-126 Immunoprecipitation (IP) was performed using the immunoprecipitation (IP) method. Fresh frozen tissue was pulverized (CryoPrep; Covaris, Woburn, MA), and lysis buffer (1% CHAPS, 20 mM Tris-HCl, 150 mM NaCl, protease and phosphatase inhibitors, pH = 8) was added to solubilize the tissue. The resulting solution was centrifuged at 4°C for 2 hours to pellet debris. The clarified lysate was used for HLA-specific IP. Immunoprecipitation was performed using antibody W6 / 32 as previously described. 127The lysate was added to antibody beads and rotated overnight at 4°C for immunoprecipitation. After immunoprecipitation, the beads were removed from the lysate. The IP beads were washed to remove nonspecific binding, and the HLA / peptide complexes were eluted from the beads with 2N acetic acid. Protein components were removed from the peptides using molecular weight spin columns. The resulting peptides were dried by SpeedVac evaporation and stored at -20°C until MS analysis.

[0425] XVI.A.3 Peptide sequencing The dried peptides were reconstituted in HPLC buffer A and loaded onto a C-18 microcapillary HPLC column for gradient elution into a mass spectrometer. A 180-minute gradient of 0-40% B (solvent A: 0.1% formic acid, solvent B: 0.1% formic acid in 80% acetonitrile) was used to elute the peptides into a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of the peptide mass / charge (m / z) were collected in an Orbitrap detector at a resolution of 120,000, followed by 20 MS2 low-resolution scans in an Orbitrap or ion trap detector after HCD fragmentation of selected ions. MS2 ion selection was performed using data-dependent acquisition mode and 30-second dynamic exclusion after MS2 ion selection. The automatic gain control (AGC) for the MS1 scan was set to 4x10. 5 and 1x10 for MS2 scans. 4 For sequencing of HLA peptides, charge states of +1, +2 and +3 can be selected for MS2 fragmentation.

[0426] MS2 spectra from each analysis were collected using Comet 128,129 Search against protein databases using Percolator 130~132 was scored using

[0427] XVI.B. Machine Learning XVI.B.1. Data Coding For each sample, training data points were all 8-11mer (encompassing) peptides from the reference proteome that mapped to exactly one gene expressed in the sample. The entire training dataset was generated by concatenating the training datasets from each training sample. The 8-11 length range was chosen because it captures approximately 95% of all HLA class I-presented peptides, but lengths of 12-15 can also be achieved using the same method at the expense of a moderate increase in computational demand. Peptides and flanking sequences were vectorized using a one-hot encoding scheme. Peptides of multiple lengths (8-11) were represented as fixed-length vectors by increasing the amino acid alphabet with pad characters and padding all peptides to a maximum length of 11. The RNA abundance of the source proteins for the training peptides was calculated using RSEM. 133 The data were expressed as the logarithm of the isoform-level TPM (transcripts per million) estimates obtained from [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119

[0428] XVI.B.2. Model Architecture Specification The full representation model has the following functional form: TIFF2025143480000080.tif5128In the formula, k is the subscript of the HLA allele in the dataset from 1 to m, TIFF2025143480000081.tif5128 is an indicator variable whose value is 1 if allele k is present in the sample from which peptide i is derived, and 0 otherwise. For a particular peptide i, All but a maximum of six of TIFF2025143480000082.tif5128 (six corresponding to the HLA type of the sample from which peptide i was derived) are 0. The sum of the probabilities is, for example, TIFF2025143480000083.tif4128, Clip to TIFF2025143480000084.tif3128.

[0429] The per-allele probability of presentation is modeled as follows: Pr(peptide i presented by allele a) = sigmoid{NN a (peptide i )+NN フランキング (Franking i )+NN RNA (log(TPM i ))+α 試料(i) +β タンパク質(i)} where the variables have the following meanings: sigmoid is the sigmoid (also known as expit) function, and i is the one-hot coded intermediate padded amino acid sequence of peptide i, and NN a is a neural network with linear final layer activations that models the contribution of peptide sequences to the probability of presentation, and flanking i is the one-hot coded flanking sequence of peptide i in the source protein, and N フランキング is a neural network with linear final layer activations that models the contribution of flanking sequences to the probability of presentation, and is the TPM i is the expression of the source mRNA of peptide i in TPM units, sample (i) is the sample (i.e., patient) from which peptide i is derived, and α 試料(i) is the sample section, protein (i) is the source protein of peptide i, and β タンパク質(i) is the per-protein intercept (also known as the per-gene trend of representation).

[0430] In the model described in the Results section, each component neural network has the following architecture: NN a Each of these is a single output node of a single hidden layer multilayer perceptron (MLP) with input dimensionality of 231 (11 residues × 21 possible characters per residue, including pad characters), width 256, rectified linear unit (ReLU) activation in the hidden layer, linear activation in the output layer, and one output node for each HLA allele a in the training dataset. NN フランキング is a single-hidden-layer MLP with input dimensionality of 210 (5 residues of N-terminal flanking sequence + 5 residues of C-terminal flanking sequence × 21 possible characters per residue (including pad characters)), width 32, rectified linear unit (ReLU) activations in the hidden layer, and linear activations in the output layer. NN RNA is a single-hidden-layer MLP with input dimensionality 1, width 16, rectified linear unit (ReLU) activations in the hidden layer, and linear activations in the output layer.

[0431] Some components of the model (e.g., NN a ) depends on the specific HLA, but many components (NN フランキング , N.N. RNA , α 試料(i) , β タンパク質(i) ) are independent. The former is called "allele interactivity" and the latter is called "allele non-interaction." The properties to model as allele interactivity and allele non-interaction are selected based on conventional biological knowledge. That is, since HLA alleles see peptides, peptide sequences should be modeled as allele interactivity, but information about source protein, RNA expression, or flanking sequences is not conveyed to HLA alleles (because peptides are separated from their source proteins in the endoplasmic reticulum before encountering HLA), and therefore these properties should be modeled as allele non-interaction. This model was implemented using Keras v2.0.4. 134 and Theano v0.9.0 135 It was carried out at.

[0432] The peptide MS model uses the same deconvolution procedure as the full MS model (Equation 1), but the allele-by-allele probabilities of presentation were generated using a reduced allele-by-allele model that only considered the peptide sequence and HLA alleles. Pr(peptide i presented by allele a) = sigmoid{NN a (peptide i )}

[0433] Although the peptide MS model uses the same features as the binding affinity prediction, the model weights are trained on different data types (i.e., mass spectrometry data versus HLA peptide binding affinity data). Thus, comparing the predictive performance of the peptide MS model with the full MS model reveals the contribution of non-peptide features (i.e., RNA abundance, flanking sequences, gene ID) to overall predictive performance, and comparing the predictive performance of the peptide MS model with the binding affinity model reveals the importance of improved modeling of the peptide sequence to overall predictive performance.

[0434] XVI.B.3. Training / Validation / Testing Split We used the following procedure to ensure that no peptides appeared in more than one of the training / validation / test sets. First, all peptides that appeared in more than one protein were removed from the reference proteome. Then, the proteome was partitioned into blocks of 10 contiguous peptides. Each block was uniquely assigned to the training, validation, and test sets. This ensured that no peptides appeared in more than one of the training, validation, or test datasets. The validation set was used exclusively for early stopping. The tumor sample test data in Figures 14-16 represent test set peptides from five tumor samples that were completely excluded from the training and validation sets (i.e., peptides from blocks of contiguous peptides uniquely assigned to the test set).

[0435] XVI.B.4. Model Training To train the model, we modeled every peptide independently, with the per-peptide loss being the negative Bernoulli log-likelihood loss function (also known as log-loss). Formally, the contribution of peptide i to the overall loss is Loss(i)=-log(Bernoulli(y i |Pr(presenting peptide i))) where y i is the label of peptide i (i.e., when peptide i is presented, y i = 1 otherwise 0, and given an iid binomial observation vector y, Bernoulli(y|P) denotes the Bernoulli likelihood with parameters p∈[0,1]. The model was trained by minimizing the loss function:

[0436] To reduce training time, class balance was adjusted by randomly removing 90% of the negatively labeled training data, resulting in an overall training set class balance of approximately one presented peptide for every 2000 unpresented peptides. Model weights were initialized using the Glorot uniform procedure61 and trained on an Nvidia Maxwell TITAN X GPU using the ADAM62 stochastic optimizer with standard parameters. A validation set consisting of 10% of the total data was used for early stopping. The model was evaluated on the validation set every ¼ epoch, and model training was terminated after the first ¼ epoch when the validation loss (i.e., the negative Bernoulli log-likelihood on the validation set) no longer decreased.

[0437] The full representation model was an ensemble of 10 model replicates, each trained independently on shuffled copies of the same training data with different random initializations of model weights for all models in the ensemble. At test time, predictions were generated by averaging the probabilities output by the model replicates.

[0438] XVI.B.5. Motif logo Motif logo, weblogolib Python API v3.5.0 138 To generate the binding affinity logo, the mhc_ligand_full.csv file was submitted to the Immune Epitope Database (IEDB) in July 2017. 88 Peptides were downloaded from the IEDB and retained for the following criteria: measurements in nanomolar (nM), a reference date after 2000, an object type equal to "linear peptide," and all residues in the peptide drawn from the standard 20-letter amino acid alphabet. Logos were generated using a filtered subset of peptides with measured binding affinities below the conventional binding threshold of 500 nM. Logos were not generated for allele pairs with too few binders in the IEDB. To generate logos representing the trained display model, model predictions were made for 2,000,000 random peptides at each allele and peptide length. For each allele and length, peptides ranked in the top 1% (i.e., the top 20,000) by the trained display model were used to generate logos. Importantly, this binding affinity data from the IEDB was not used for model training or testing; it was used only to compare the learned motifs.

[0439] XVI.B.6. Prediction of Binding Affinity We used MHCflurry v1.2.0, an open-source, GPU-compatible HLA class I binding affinity prediction tool with performance comparable to the NetMHC family of models. 139Peptide-MHC binding affinities were predicted using a binding affinity-only prediction tool from [link removed]. To combine binding affinity predictions for a single peptide across multiple HLA alleles, the minimum binding affinity was selected. To combine binding affinities across multiple peptides (i.e., to rank mutations across multiple mutant peptides, as shown in Figure 16), the minimum binding affinity across peptides was selected. To determine the RNA expression threshold for the T cell dataset, tumor-type-matched RNA-seq data from TCGA up to a threshold of TPM > 1 was used. Because all of the initial T cell datasets were filtered for TPM > 0 during initial publication, TCGA RNA-seq data filtered for TPM > 0 were not used.

[0440] XVI.B.7. Predictions of Presentation To sum the probability of presentation for a single peptide across multiple HLA alleles, the sum of probabilities was determined as shown in Equation 1. To sum the probability of presentation across multiple peptides (i.e., to rank mutations across multiple mutant peptides, as shown in Figure 16), the sum of probabilities of presentation was determined. Probabilistically, if peptide presentation is considered as an i.i.d. Bernoulli random variable, the sum of probabilities corresponds to the expected number of m...

Claims

[Claim 1] The invention described herein.