Identifying neoantigens for T cell therapy
The use of next-generation sequencing and nonlinear deep learning models enhances the identification of neoantigens for personalized cancer therapies, addressing inefficiencies in current methods by improving predictive performance and ensuring effective neoantigen-based vaccines and T cell therapies.
Patent Information
- Application Number
- JP2020534819
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-25
- Filing Date
- 2018-09-05
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2038-09-05
AI Technical Summary
Current methods for identifying neoantigens and neoantigen-recognizing T cells are time-consuming, labor-intensive, and have low positive predictive value (PPV), leading to ineffective neoantigen-based vaccines and T cell therapies due to insufficient accuracy in predicting peptide presentation on tumor surfaces.
An optimized approach using next-generation sequencing and nonlinear deep learning models to identify neoantigens, modeling peptide-allele mapping and MHC processing, enhancing the predictive performance by up to an order of magnitude, enabling more efficient and accurate selection of neoantigens for personalized cancer vaccines and T cell therapies.
The method accelerates the development of personalized cancer therapies by improving the identification of therapeutically useful neoepitopes, reducing time and cost, and ensuring higher accuracy in selecting neoantigens that elicit anti-tumor immunity.
Smart Images

Figure 0007763588000110 
Figure 0007763588000111 
Figure 0007763588000112
Abstract
Description
[Background technology]
[0001] Therapeutic vaccines and T cell therapies based on tumor-specific neoantigens hold great promise as the next generation of personalized cancer immunotherapy. 1~3 Cancers with high mutational burden, such as non-small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies due to their relatively high likelihood of generating neoantigens. 4,5 Early evidence suggests that neoantigen-based vaccination induces T cell responses. 6 , T cell therapy targeting neoantigens can induce tumor regression in selected patients 7 Both MHC class I and MHC class II influence T cell responses. 70~71 .
[0002] However, the identification of neoantigens and neoantigen-recognizing T cells is crucial for assessing tumor response. 77,110 , examining tumor evolution 111 , designing the next generation of personalized therapies 112 Current methods for identifying neoantigens are time-consuming and labor-intensive. 84,96 , or the accuracy is not sufficient 87,91-93 Neoantigen-recognizing T cells are the main component of TILs. 84,96,113,114 circulating in the peripheral blood of cancer patients. 107 Although it has been shown recently that neoantigen-reactive T cells can be identified using the TIL method, the current methods are as follows: 97,98 or leukapheresis 107 (2) they require screening of impractically large peptide libraries; or (3) they rely on MHC alleles that may be practically available for only a small number of MHC alleles.
[0003] Furthermore, early methods incorporating mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides have been proposed.8 However, these proposed methods involve many steps other than gene expression and MHC binding (e.g., TAP transport, proteasomal cleavage, MHC binding, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-I; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsins), competition with CLIP peptides for HLA binding catalyzed by HLA-DM, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-II). 9 The entire epitope generation process cannot be modeled, and therefore existing methods tend to suffer from low positive predictive value (PPV) (Figure 1A).
[0004] Indeed, analyses of peptides presented by tumor cells conducted by multiple groups have shown that less than 5% of peptides predicted to be presented using gene expression and MHC binding affinity are found on tumor surface MHC. 10,11 (Figure 1B). This low correlation between binding prediction and MHC presentation is further supported by the lack of improved prediction accuracy for binding-restricted neoantigens in response to checkpoint inhibitors relative to the number of mutations alone. 12 .
[0005] Such low positive predictive values (PPV) of existing methods for predicting presentation present a challenge in the design of neoantigen-based vaccines and in neoantigen-based T cell therapies. If vaccines are designed using low PPV predictions, most patients are unlikely to receive therapeutic neoantigens, and even fewer patients will receive multiple neoantigens (even assuming all presented peptides are immunogenic). Similarly, if therapeutic T cells are designed based on low PPV predictions, most patients are unlikely to receive T cells with reactivity against tumor neoantigens, and the time and resource costs of identifying predicted neoantigens using downstream testing methods after prediction may be unnecessarily high. Therefore, neoantigen vaccination and T cell therapy using current methods are unlikely to be effective in a significant number of tumor-bearing subjects (Figure 1C).
[0006] Furthermore, previous approaches have only used cis-acting mutations to generate candidate neoantigens, including splicing factor mutations that occur in multiple tumor types and lead to aberrant splicing of many genes. 13 , and additional sources of nascent ORFs, including mutations that create or remove protease cleavage sites, were not considered in most cases.
[0007] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard tumor analysis approaches may erroneously promote sequence artifacts or germline polymorphisms as neoantigens, which can lead to inefficient utilization of vaccine doses or risk of autoimmunity, respectively. Summary of the Invention
[0008] Disclosed herein are optimized approaches for identifying and selecting neoantigens for personalized cancer vaccines, T cell therapy, or both. First, we address optimized tumor exome and transcriptome analysis approaches to identify neoantigen candidates using next-generation neoantigen (NGS). These methods build on standard approaches of tumor analysis by NGS so that the most sensitive and specific neoantigen candidates are developed across all classes of genomic alterations. Second, novel approaches for high PPV neoantigen selection are provided to overcome specificity issues and ensure that neoantigens developed for vaccine addition and / or as targets for T cell therapy are more likely to elicit anti-tumor immunity. Depending on the embodiment, these approaches include trained statistical regression or nonlinear deep learning models that jointly model peptide-allele mapping, as well as per-allele motifs for peptides of multiple lengths that share statistical power across peptides of different lengths. In particular, nonlinear deep learning models can be designed and trained to treat different MHC alleles within the same cell as independent, thereby overcoming the problem associated with linear models where linear models interfere with each other. Finally, further concerns regarding the design and production of neoantigen-based personalized vaccines and in the production of personalized neoantigen-specific T cells for T cell therapy are resolved.
[0009] The models disclosed herein outperform state-of-the-art predictive tools trained on binding affinity and earlier predictive tools based on MS peptide data by up to an order of magnitude. By predicting peptide presentation with greater confidence, the models enable more time- and cost-effective identification of neoantigen- or tumor antigen-specific T cells for personalized therapy using a clinically practical process that uses limited amounts of patient peripheral blood, screens fewer peptides per patient, and does not necessarily rely on MHC multimers. However, in another embodiment, the models disclosed herein enable more time- and cost-effective identification of tumor antigen-specific T cells using MHC multimers by reducing the number of MHC multimer-bound peptides that need to be screened to identify neoantigen- or tumor antigen-specific T cells.
[0010] The predictive performance of the model disclosed herein on the TIL neoepitope dataset and the task of identifying predicted neoantigen-reactive T cells demonstrates that by modeling HLA processing and presentation, it is now possible to obtain predictions of therapeutically useful neoepitopes. In summary, this work will accelerate progress toward patient cures by enabling actionable in silico antigen identification for antigen-targeted immunotherapy. [The present invention 1001] 1. A method for identifying one or more T cells that are antigen-specific for at least one neoantigen that is likely to be derived from and presented on the surface of one or more tumor cells in a subject, comprising: obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing data from the tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen comprises at least one alteration that makes the peptide sequence different from a corresponding wild-type peptide sequence identified from the normal cells of the subject; encoding the peptide sequence of each of the neoantigens into a corresponding numerical vector, each numerical vector containing information about a set of amino acids constituting the peptide sequence and the positions of the amino acids in the peptide sequence; inputting the numerical vectors into a machine-learned presentation model using a computer processor to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood in the set represents the likelihood that a corresponding neoantigen will be presented on the surface of the tumor cells of the subject by one or more MHC alleles, and wherein the machine-learned presentation model: a label obtained by mass spectrometry that measures, for each sample in the plurality of samples, the presence of a peptide bound to at least one MHC allele within the set of MHC alleles identified as being present in said sample; and For each of the samples, a training peptide sequence is coded as a numerical vector containing information about a set of amino acids that make up the peptide and the positions of the amino acids in the peptide. a plurality of parameters identified based at least on a training dataset including: a function representing a relationship between the numerical vector received as an input and the proposed likelihood generated as an output based on the numerical vector and the parameters; the process comprising: selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens; identifying one or more T cells in said subset that are antigen-specific for at least one of said neoantigens; and returning the one or more identified T cells The method comprising: [The present invention 1002] inputting the numerical vector into the machine-learned presentation model, applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate, for each of the one or more MHC alleles, a dependency score indicating whether the MHC allele will present the neoantigen based on a particular amino acid at a particular position in the peptide sequence. The method of the present invention 1001, comprising: [The present invention 1003] inputting the numerical vector into the machine-learned presentation model, transforming the dependency scores to generate, for each MHC allele, a corresponding per-allele likelihood that indicates the likelihood that the corresponding MHC allele will present the corresponding neoantigen; and combining the per-allele likelihoods to generate the presentation likelihood of the neoantigen. The method of the present invention 1002 further comprises: [The present invention 1004] 1004. The method of claim 1003, wherein transforming said dependency score models presentation of said neoantigen as mutually exclusive across said one or more MHC alleles. [The present invention 1005] inputting the numerical vector into the machine-learned presentation model, transforming the combination of dependency scores to generate the presentation likelihood, wherein transforming the combination of dependency scores models presentation of the neoantigen as interference between the one or more MHC alleles. The method of the present invention 1002 further comprises: [The present invention 1006] The set of representation likelihoods is further specified by at least one or more allele non-interaction characteristics; and applying the machine-learned presentation model to the allele-non-interacting trait to generate a dependency score for the allele-non-interacting trait that indicates whether the corresponding neoantigen peptide sequence will be presented based on the allele-non-interacting trait. Any of the methods 1002 to 1005 of the present invention, further comprising: [The present invention 1007] combining the dependency score for each MHC allele in the one or more MHC alleles with the dependency score for the allele-non-interacting trait; transforming the combined dependency scores for each MHC allele to generate a per-allele likelihood for each MHC allele, the likelihood being that the corresponding MHC allele will present the corresponding neoantigen; and combining the per-allele likelihoods to generate the representation likelihood. The method of the present invention 1006 further comprising: [The present invention 1008] combining the dependency scores for each of the MHC alleles with the dependency scores for the allele-non-interacting trait; and transforming the combined dependency scores to generate the presentation likelihood. The method of the present invention 1006 further comprising: [The present invention 1009] The method of any of claims 1001 to 1008, wherein said one or more MHC alleles comprises two or more different MHC alleles. [The present invention 1010] 1009. The method of any of claims 1001 to 1009, wherein said peptide sequence comprises a peptide sequence having a length other than 9 amino acids. [The present invention 1011] 1010. The method of any of claims 1001 to 1010, wherein the step of encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme. [The present invention 1012] The plurality of samples (a) one or more cell lines engineered to express a single MHC allele; (b) one or more cell lines engineered to express multiple MHC alleles; (c) one or more human cell lines obtained or derived from multiple patients; (d) fresh or frozen tumor samples obtained from multiple patients; and (e) fresh or frozen tissue samples obtained from multiple patients; Any of the methods of the present invention 1001 to 1011, including at least one of the following. [The present invention 1013] the training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of said peptides; and (b) data relating to measurements of peptide-MHC binding stability for at least one of said peptides; The method of any one of 1001 to 1012, further comprising at least one of the following: [The present invention 1014] 1014. The method of any of claims 1001 to 1013, wherein said set of presentation likelihoods is further identified by at least the expression level of said one or more MHC alleles in said subject, as measured by RNA-seq or mass spectrometry. [The present invention 1015] The set of presentation likelihoods is: (a) the predicted affinity between neoantigens in the set of neoantigens and the one or more MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex; The method of any of claims 1001 to 1014, further characterized by at least one of the following characteristics: [The present invention 1016] The set of numerical likelihoods is: (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; The method of any of claims 1001 to 1015, further characterized by at least one of the following characteristics: [The present invention 1017] Any of the methods of present inventions 1001 to 1016, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of being presented on the surface of the tumor cells compared to non-selected neoantigens based on the machine-learned presentation model. [The present invention 1018] Any of the methods of present inventions 1001 to 1017, wherein the step of selecting the set of selected neoantigens comprises selecting, based on the machine-learned presentation model, neoantigens that have an increased likelihood of inducing a tumor-specific immune response in the subject compared to non-selected neoantigens. [The present invention 1019] Any of the methods of inventions 1001 to 1018, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that, based on the presentation model, have an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) compared to unselected neoantigens, and optionally the APCs are dendritic cells (DCs). [The present invention 1020] Any of the methods of present inventions 1001 to 1019, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of being inhibited by central tolerance or peripheral tolerance compared to non-selected neoantigens based on the machine-learned presentation model. [The present invention 1021] Any of the methods of present inventions 1001 to 1020, wherein the step of selecting the set of selected neoantigens comprises selecting, based on the machine-learned presentation model, neoantigens that have a reduced likelihood of inducing an autoimmune response against normal tissue in the subject compared to non-selected neoantigens. [The present invention 1022] 1022. The method of any one of claims 1001 to 1021, wherein said one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer. [The present invention 1023] The method of any of claims 1001 to 1022, further comprising generating an output for constructing a personalized cancer vaccine from said set of selected neoantigens. [The present invention 1024] 1024. The method of claim 1023, wherein the output for said personalized cancer vaccine comprises at least one peptide sequence or at least one nucleotide sequence encoding said selected set of neoantigens. [The present invention 1025] The method according to any one of claims 1001 to 1024, wherein the presentation model trained by machine learning is a neural network model. [The present invention 1026] The method of claim 1025, wherein the neural network model comprises a plurality of network models for MHC alleles, each network model being assigned to a corresponding MHC allele among the plurality of MHC alleles and comprising a series of nodes arranged in one or more layers. [The present invention 1027] The method of claim 1026, wherein the neural network models are trained by updating parameters of the neural network models, and the parameters of at least two network models are updated together for at least one training iteration. [The present invention 1028] The method of any one of claims 1025 to 1027, wherein the machine-learned presentation model is a deep learning model including one or more layers of nodes. [The present invention 1029] Any of the methods of claims 1001 to 1028, wherein the step of identifying the one or more T cells comprises co-culturing the one or more T cells with one or more of the neoantigens in the subset under conditions that allow the one or more T cells to proliferate. [The present invention 1030] Any of the methods of claims 1001 to 1029, wherein the step of identifying the one or more T cells comprises contacting the one or more T cells with an MHC multimer comprising one or more of the neoantigens in the subset under conditions that allow binding of the T cells to the MHC multimer. [The present invention 1031] The method of any of claims 1001 to 1030, further comprising identifying one or more T cell receptors (TCRs) of said one or more identified T cells. [The present invention 1032] 1032. The method of claim 1031, wherein identifying said one or more T cell receptors comprises sequencing T cell receptor sequences of said one or more identified T cells. [The present invention 1033] An isolated T cell antigen-specific to at least one selected neoantigen in said subset of any of 1001 to 1032 of the present invention. [The present invention 1034] genetically engineering a plurality of T cells to express at least one of said one or more identified T cell receptors; culturing the plurality of T cells under conditions that expand the plurality of T cells; and injecting the expanded T cells into the subject. The method of the present invention 1032 further comprises: [This invention 1035] genetically engineering said plurality of T cells to express at least one of said one or more identified T cell receptors; cloning the T cell receptor sequence of said one or more identified T cells into an expression vector; and transfecting each of said plurality of T cells with said expression vector. The method of the present invention 1034, comprising: [The present invention 1036] culturing the one or more identified T cells under conditions that expand the one or more identified T cells; and injecting the expanded T cells into the subject. Any of the methods of claims 1001 to 1035, further comprising: [This invention 1037] Any of the methods of claims 1001 to 1036, wherein the one or more T cells in the subset that are antigen-specific to at least one of the neoantigens are identified using 5 to 30 mL of whole blood from the subject. [The present invention 1038] Any of the methods of present inventions 1001 to 1037, wherein the subset of neoantigens includes up to 20 types of neoantigens, and the one or more identified T cells recognize at least two types of neoantigens in the subset of neoantigens. [This invention 1039] The method of any of claims 1001 to 1038, wherein said one or more MHC alleles are class I MHC alleles. [Brief explanation of the drawings]
[0011] These and other features, aspects, and aspects of the present invention will become better understood with regard to the following description and accompanying drawings.
[0012] [Figure 1A] Current clinical approaches to neoantigen identification are presented. [Figure 1B] It shows that less than 5% of the predicted binding peptides are displayed on tumor cells. [Figure 1C] 1 illustrates the impact of the specificity problem on neoantigen prediction. [Figure 1D] This shows that binding prediction is not sufficient to identify neoantigens. [Figure 1E] Probability of MHC-I presentation as a function of peptide length. [Figure 1F] Figure 1F shows an exemplary peptide spectrum generated from a dynamic range standard from Promega. Figure 1F discloses SEQ ID NO: 1. [Figure 1G] We show how adding features increases the positive predictive value of the model. [Figure 2A] 1 is a schematic of an environment for identifying the likelihood of peptide presentation in a patient, according to one embodiment. [Figure 2B] A method for obtaining presentation information according to one embodiment is described below. Figure 2B discloses SEQ ID NO:28. [Figure 2C] A method for obtaining presentation information according to one embodiment is illustrated in Figure 2C, which discloses SEQ ID NOs: 3 to 8, respectively, in order of appearance. [Figure 3] FIG. 1 is a high-level block diagram illustrating computer logic components of a presentation identification system, according to one embodiment. [Figure 4] 4 illustrates an exemplary set of training data, according to one embodiment. Figure 4 discloses, in order of appearance, the peptide sequences as SEQ ID NOS: 10-13 and the C-terminal flanking sequences as SEQ ID NOS: 15, 29-30, and 30, respectively. [Figure 5] 1 illustrates an exemplary network model related to MHC alleles. [Figure 6A] 1 illustrates an exemplary network model NNH(·) shared by MHC alleles, according to one embodiment. [Figure 6B] 1 illustrates an exemplary network model NNH(·) shared by MHC alleles, according to another embodiment. [Figure 7] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 8] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 9] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 10] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 11] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 12] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 13A] 1 shows the sample frequency distribution of mutation burden in NSCLC patients. [Figure 13B] 1 shows the number of presented neoantigens in a simulated vaccine for patients selected based on the selection criterion of whether the patient meets a minimum mutational load, according to one embodiment. [Figure 13C] 10A-10C show a comparison of the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model, according to one embodiment. [Figure 13D]10A-B compares the number of neoantigens presented in simulated vaccines between selected patients associated with a vaccine containing therapeutic subsets identified based on a single allele-per-presentation model for HLA-A*02:01 and selected patients associated with a vaccine containing therapeutic subsets identified based on both an allele-per-presentation model for HLA-A*02:01 and HLA-B*07:02. According to one embodiment, the vaccine volume is set at v=20 epitopes. [Figure 13E] FIG. 10 compares the number of presented neoantigens in simulated vaccines between patients selected based on mutational burden and patients selected by expected utility score, according to one embodiment. [Figure 14A] The positive predictive value (PPV) at 40% recall is compared for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on a test set consisting of five different test samples, including an excluded tumor sample, where each test sample has a ratio of presented peptides to unpresented peptides of 1:2500. [Figure 14B] The PPV at 40% recall is compared for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on a test set consisting of 15 different test samples, each containing an excluded peptide from the single-allele cell line test dataset with a ratio of presented to unpresented peptides of 1:10,000. [Figure 14C]For a test set consisting of 12 different test samples, each from a patient with at least one pre-existing T cell response, the percentage of somatic mutations recognized by T cells (e.g., pre-existing T cell responses) was compared for the top 5, 10, and 20 ranked somatic mutations identified by the "full MS model," the "peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2. [Figure 14D] For a test set consisting of 12 different test samples, each from a patient with at least one pre-existing T cell response, the percentage of minimal nascent epitopes recognized by T cells (e.g., pre-existing T cell responses) was compared for the top 5, 10, and 20 ranked minimal nascent epitopes identified by the "full MS model," the "peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2. [Figure 15A] Detection of T cell responses to patient-specific neoantigen peptide pools in nine patients is shown. [Figure 15B] Detection of T cell responses to individual patient-specific neoantigen peptides in four patients is shown. [Figure 15C] An exemplary image of an ELISpot well for patient CU04 is shown. [Figure 16] The positive predictive value (PPV) at a recall rate of 40% was compared for the "full MS model" and the "anchor residue-only MS model" when each model was tested on a test set consisting of five different test samples, including an excluded tumor sample, where each test sample had a ratio of presented peptides to unpresented peptides of 1:2500. [Figure 17A] Shown are the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 0 from Figure 14A. [Figure 17B]The PPV at 40% recall is compared for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on a test set consisting of 15 different test samples, each containing an excluded peptide from the single-allele cell line test dataset with a ratio of presented to unpresented peptides of 1:5,000. [Figure 17C] Shown are the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 0 from Figure 14A. [Figure 17D] Shown are the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 1 from Figure 14A. [Figure 17E] Shown are the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 2 from Figure 14A. [Figure 17F] Shown are the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 3 from Figure 14A. [Figure 17G] Shown are the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 4 from Figure 14A. [Figure 17H]Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*01:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17I] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*02:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17J] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*02:03 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17K] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*02:07 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17L] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*03:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17M]Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*24:02 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17N] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*29:02 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17O] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*31:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17P] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*68:02 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17Q] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*35:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17R]Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*44:02 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17S] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*44:03 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17T] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*51:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17U] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*54:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 17V] Figure 14B shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model with omitted peptides from the test dataset of the HLA-A*57:01 cell line in Figure 14B with a presented to unpresented peptide ratio of 1:10,000. [Figure 18]Figure 14A compares the positive predictive value (PPV) at 40% recall of different versions of the MS model and earlier approaches29 for modeling HLA-presented peptides in human tumors when each model is tested on the test set of 5 different test samples in Figure 14A, where each test sample has a ratio of presented to unpresented peptides of 1:2500, including an excluded tumor sample. [Figure 19A-1] 1 shows the results of a control experiment using neoantigens in HLA-matched healthy donors. [Figure 19A-2] 1 shows the results of a control experiment using neoantigens in HLA-matched healthy donors. [Figure 19B-1] Figure 19B shows the results of a control experiment using neoantigens in HLA-matched healthy donors: 27, 24, 21-22, 31-36, 21, 37-45, respectively, in order of appearance. [Figure 19B-2] Figure 19B shows the results of a control experiment using neoantigens in HLA-matched healthy donors: 27, 24, 21-22, 31-36, 21, 37-45, respectively, in order of appearance. [Figure 20] Detection of T cell responses against the PHA positive control is shown for each donor and each in vitro expansion shown in Figure 15A. [Figure 21A] 1 shows the detection of T cell responses to each individual patient-specific neoantigen peptide in pool #2 in patient CU04. [Figure 21B] 1 shows detection of T cell responses to individual patient-specific neoantigen peptides at each of three visits for patient CU04 and at each of two visits for patient 1-024-002 (each visit occurring at a different time point). [Figure 21C] 1 shows detection of T cell responses to individual patient-specific neoantigen peptides and to patient-specific neoantigen peptide pools at each of two visits for patient CU04 and at each of two visits for patient 1-024-002 (each visit occurring at a different time point). [Figure 22]Detection of T cell responses to two patient-specific neoantigen peptide pools and a DMSO negative control for the patient in Figure 15A is shown. [Figure 23] The figure compares the predictive performance of the "MS model," which adopts the lowest NetMHCIIpan percentile rank across HLA-DRB1*15:01 and HLA-DRB5*01:01 ("NetMHCIIpan rank": NetMHCIIpan 3.177), and the strongest affinity in nM across HLA-DRB1*15:01 and HLA-DRB5*01:01 ("NetMHCIIpan nM": NetMHCIIpan 3.1), in ranking peptides within the HLA-DRB1*15:01 / HLA-DRB5*01:01 test dataset. [Figure 24-1] Figure 24 shows a method for sequencing the TCRs of neoantigen-specific memory T cells derived from the peripheral blood of NSCLC patients. Figure 24 discloses 46 to 48 in order of appearance, respectively. [Figure 24-2] Figure 24 shows a method for sequencing the TCRs of neoantigen-specific memory T cells derived from the peripheral blood of NSCLC patients. Figure 24 discloses 46 to 48 in order of appearance, respectively. [Figure 25] 1 shows an exemplary embodiment of a TCR construct for introducing a TCR into a recipient cell. [Figure 26] Figure 26 discloses SEQ ID NO:49, which shows the nucleotide sequence of an exemplary P526 construct backbone for cloning TCRs into expression systems for therapeutic development. [Figure 27] Figure 27 discloses SEQ ID NO:50, an exemplary construct sequence for cloning a patient's neoantigen-specific TCR clonotype 1 into an expression system for therapeutic development. [Figure 28] Figure 28 discloses SEQ ID NO:51, which shows an exemplary construct sequence for cloning a patient's neoantigen-specific TCR clonotype 3 into an expression system for therapeutic development. [Figure 29]1 is a flowchart of a method for providing personalized neoantigen-specific therapy to a patient, according to one embodiment. [Figure 30] An exemplary computer for implementing the entities shown in FIGS. 1 and 3 will now be described. DETAILED DESCRIPTION OF THE INVENTION
[0013] I. Definition In general, terms used in the claims and the specification shall be interpreted as having their ordinary meaning as understood by one of ordinary skill in the art. Certain terms are defined below to provide further clarity. If there is a conflict between the ordinary meaning and a given definition, the given definition shall control.
[0014] As used herein, the term "antigen" refers to a substance that induces an immune response.
[0015] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes it different from its corresponding wild-type parent antigen, for example, due to a tumor cell mutation or tumor cell-specific post-translational modification. Neoantigens may include polypeptide or nucleotide sequences. Mutations can include frameshift or non-frameshift insertion / deletions (indels), missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression change that results in a new ORF. Mutations can also include splice variants. Tumor cell-specific post-translational modifications can include aberrant phosphorylation. Tumor cell-specific post-translational modifications can also include splice antigens generated by the proteasome. See Liepe et al., "A large fraction of HLA class I ligands are proteasome-generated spliced peptides"; Science. 2016 Oct 21;354(6310):354-358.
[0016] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in tumor cells or tissues of a subject, but is not present in the corresponding normal cells or tissues of the subject.
[0017] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct that is based on one or more neoantigens, eg, multiple neoantigens.
[0018] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.
[0019] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.
[0020] As used herein, the term "coding mutation" refers to a mutation that occurs in a coding region.
[0021] As used herein, the term "ORF" means open reading frame.
[0022] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that arises due to mutation or other abnormalities such as splicing.
[0023] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.
[0024] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon.
[0025] As used herein, the term "frameshift mutation" is a mutation that causes an alteration in the frame of a protein.
[0026] As used herein, the term "indel" refers to the insertion or deletion of one or more nucleic acids.
[0027] As used herein, the term "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences in which a certain percentage of nucleotides or amino acid residues are the same when compared and aligned for maximum correspondence, as determined using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN, or other algorithms available to those of skill in the art), or by visual inspection. Depending on the application, the "percent identity" can exist over a region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.
[0028] In sequence comparison, generally, one sequence serves as a reference sequence to which test sequences are compared.When using a sequence comparison algorithm, test sequences and reference sequences are input into a computer, subsequence coordinates are designated if necessary, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the percent sequence identity (%) of the test sequence to the reference sequence based on the designated program parameters.Alternatively, sequence similarity or difference can also be established by the combination of the presence or absence of a specific nucleotide at a selected sequence position (e.g., sequence motif) or an amino acid in a translated sequence.
[0029] Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (see generally Ausubel et al., infra).
[0030] One example of an algorithm that is suitable for determining percent sequence identity and percent sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.
[0031] As used herein, the term "non-stop or read-through" refers to a mutation that results in the removal of the natural stop codon.
[0032] As used herein, the term "epitope" refers to a specific portion of an antigen that is typically bound by an antibody or T-cell receptor.
[0033] As used herein, the term "immunogenic" refers to the ability to elicit an immune response, for example, via T cells, B cells, or both.
[0034] As used herein, the terms "HLA binding affinity" and "MHC binding affinity" refer to the affinity of binding between a specific antigen and a specific MHC allele.
[0035] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.
[0036] As used herein, the term "mutation" is a difference between the nucleic acid of a subject and a reference human genome used as a control.
[0037] As used herein, the term "variant calling" is the algorithmic determination, typically from sequencing, of the presence of a mutation.
[0038] As used herein, the term "polymorphism" refers to a germline mutation, ie, a mutation found in all DNA-bearing cells of an individual.
[0039] As used herein, the term "somatic mutation" is a mutation that occurs in a non-germline cell of an individual.
[0040] As used herein, the term "allele" refers to one version of a gene or one version of a gene sequence or one version of a protein.
[0041] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.
[0042] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the degradation of mRNA by the cell due to a premature stop codon.
[0043] As used herein, the term "truncal mutation" is a mutation that occurs early in the development of a tumor and is present in the majority of the cells of the tumor.
[0044] As used herein, the term "subclonal mutation" is a mutation that occurs late in the development of a tumor and is present in only a portion of the cells of the tumor.
[0045] As used herein, the term "exome" refers to the subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.
[0046] As used herein, the term "logistic regression" is a regression model for binary data from statistics in which the logit of the probability that the dependent variable is equal to 1 is modeled as a linear function of the dependent variable.
[0047] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinear transformations typically trained by stochastic gradient descent and backpropagation.
[0048] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.
[0049] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome can also refer to the properties of a cell or a collection of cells (e.g., a tumor peptidome refers to the union of the peptidomes of all cells that comprise a tumor).
[0050] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.
[0051] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.
[0052] As used herein, the term "MHC multimer" is a peptide-MHC complex consisting of multiple peptide-MHC monomer units.
[0053] As used herein, the term "MHC tetramer" is a peptide-MHC complex consisting of four peptide-MHC monomer units.
[0054] As used herein, the term "tolerance or immune tolerance" refers to a state of immune unresponsiveness to one or more antigens, eg, self-antigens.
[0055] As used herein, the term "central tolerance" is tolerance conferred in the thymus by either deleting autoreactive T cell clones or promoting their differentiation into immunosuppressive regulatory T cells (Tregs).
[0056] As used herein, the term "peripheral tolerance" refers to tolerance conferred in the peripheral system by downregulating or anergizing autoreactive T cells that survive central tolerance or by promoting the differentiation of these T cells into Tregs.
[0057] The term "sample" can include a single cell, or multiple cells, or fragments of cells, or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.
[0058] The term "subject" includes cells, tissues, or organisms, human or non-human, whether male or female, in vivo, ex vivo, or in vitro. The term subject includes mammals, including humans.
[0059] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.
[0060] The term "clinical factor" refers to a measurement of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values that can be obtained from assessing a subject or a sample (or a population of samples) from a subject under a given condition. A clinical factor can also be predicted by other parameters, such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.
[0061] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.
[0062] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0063] Terms not directly defined herein should be understood to have the meanings generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide further guidance to the practitioner in describing the compositions, devices, methods, etc. of embodiments of the present invention, as well as how to make or use them. It will be recognized that multiple ways of saying the same thing may be used. Accordingly, alternative terms and synonyms may be used for any one or more of the terms discussed herein. No weight should be placed on whether a term is detailed or discussed herein. Several synonyms or alternative methods, materials, etc. are provided. The recitation of one or more synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the inventive embodiments herein.
[0064] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.
[0065] II. Methods for identifying neoantigens Disclosed herein is a method for identifying T cells that are antigen-specific for neoantigens derived from tumor cells of a subject that have a high probability of being presented on the surface of the tumor cells. The method includes obtaining nucleotide sequencing data of the exome, transcriptome, and / or whole genome from tumor cells and normal cells of the subject. The nucleotide sequencing data is used to obtain a peptide sequence for each neoantigen in a set of neoantigens. The set of neoantigens is identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from normal cells. Specifically, the peptide sequence of each neoantigen in the set of neoantigens contains at least one change that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells. The method further includes encoding the peptide sequence of each neoantigen in the set of neoantigens into a corresponding numeric vector. Each numeric vector contains information describing the amino acids that make up the peptide sequence and the position of the amino acids within the peptide sequence. The method further includes generating a presentation likelihood for each neoantigen in the set of neoantigens by inputting the numeric vector into a machine-learned presentation model. Each presentation likelihood represents the likelihood that the corresponding neoantigen is presented by an MHC allele on the surface of a subject's tumor cells. The machine-learned presentation model includes a plurality of parameters and a function. The plurality of parameters are identified based on a training dataset. The training dataset includes, for each sample among a plurality of samples, labels obtained by mass spectrometry to measure the presence of peptides bound to at least one MHC allele from a set of MHC alleles identified as being present in that sample, and training peptide sequences encoded as a numerical vector containing information describing the amino acids that make up the peptide and / or the positions of amino acids within the peptide. The function represents the relationship between the numerical vector received as input by the machine-learned presentation model and the presentation likelihood generated as output by the machine-learned presentation model based on the numerical vector and the plurality of parameters.The method further includes selecting a subset of the set of neoantigens based on presentation likelihood to generate a set of selected neoantigens. The method further includes identifying T cells that are antigen-specific for at least one of the neoantigens in the subset and returning the identified T cells.
[0066] In some embodiments, inputting the numerical vectors into the machine-learned presentation model includes applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate a dependency score for each MHC allele. The dependency score for an MHC allele indicates whether that MHC allele will present the neoantigen based on a particular amino acid at a particular position in the peptide sequence. In further embodiments, inputting the numerical vectors into the machine-learned presentation model further includes, for each MHC allele, transforming the dependency scores to generate a corresponding per-allele likelihood indicating the likelihood that the corresponding MHC allele will present the corresponding neoantigen, and combining the per-allele likelihoods to generate a presentation likelihood for the neoantigen. In some embodiments, transforming the dependency scores models presentation of the neoantigen as mutually exclusive across MHC alleles. In alternative embodiments, inputting the numerical vectors into the machine-learned presentation model further includes transforming a combination of dependency scores to generate a presentation likelihood. In such embodiments, transforming a combination of dependency scores models presentation of the neoantigen as interference between MHC alleles.
[0067] In some embodiments, the set of presentation likelihoods is further specified by one or more allele non-interaction features. In such embodiments, the method further includes generating a dependency score for the allele non-interaction feature by applying a machine-learned presentation model to the allele non-interaction feature. The dependency score indicates whether the peptide sequence of the corresponding neoantigen will be presented based on the allele non-interaction feature. In some embodiments, the method further includes combining the dependency score for each MHC allele with the dependency score for the allele non-interaction feature, transforming the combined dependency scores for each MHC allele to generate a per-allele likelihood for each MHC allele, and combining the per-allele likelihoods to generate a presentation likelihood. The per-allele likelihood for a given MHC allele indicates whether that MHC allele will present the corresponding neoantigen. In an alternative embodiment, the method further includes combining the dependency score for the MHC allele with the dependency score for the allele non-interaction feature and transforming the combined dependency score to generate a presentation likelihood.
[0068] In some embodiments, the MHC alleles comprise two or more different MHC alleles.
[0069] In some embodiments, the peptide sequence comprises a peptide sequence having a length other than 9 amino acids.
[0070] In some embodiments, encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme.
[0071] In certain embodiments, the plurality of samples comprises at least one of cell lines engineered to express a single MHC allele, cell lines engineered to express multiple MHC alleles, human cell lines obtained or derived from multiple patients, fresh or frozen tissue samples obtained from multiple patients.
[0072] In some embodiments, the training dataset further comprises at least one of data related to a measure of peptide-MHC binding affinity for at least one of the peptides and data related to a measure of peptide-MHC binding stability for at least one of the peptides.
[0073] In some embodiments, the set of representation likelihoods is further identified by the expression level of the MHC allele in the subject, as measured by RNA-seq or mass spectrometry.
[0074] In some embodiments, the set of presentation likelihoods is further specified by properties including at least one of the predicted affinity between neoantigens and MHC alleles in the set of neoantigens and the predicted stability of the neoantigen-encoded peptide-MHC complex.
[0075] In some embodiments, the set of numerical likelihoods is further identified by a feature that includes at least one of a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence and an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence.
[0076] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of being presented on the tumor cell surface compared to non-selected neoantigens based on a machine-learned presentation model.
[0077] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of inducing a tumor-specific immune response in a subject compared to non-selected neoantigens based on a machine-learned representation model.
[0078] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood, based on a presentation model, of being able to be presented to naive T cells by professional antigen-presenting cells (APCs) compared to unselected neoantigens. In such embodiments, the APCs are optionally dendritic cells (DCs).
[0079] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of being inhibited by central or peripheral tolerance compared to non-selected neoantigens based on a machine learning-driven presentation model.
[0080] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have a reduced likelihood of inducing an autoimmune response against normal tissue in a subject compared to non-selected neoantigens based on a machine learning-based representation model.
[0081] In some embodiments, the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0082] In some embodiments, the method further comprises generating output for constructing a personalized cancer vaccine from the set of selected neoantigens, in such embodiments, the output for the personalized cancer vaccine can comprise at least one peptide sequence or at least one nucleotide sequence encoding the set of selected neoantigens.
[0083] In some embodiments, the machine-learned presentation model is a neural network model. In such embodiments, the neural network model may include multiple network models for MHC alleles, each network model being assigned to a corresponding MHC allele among the multiple MHC alleles and including a series of nodes arranged in one or more layers. In such embodiments, the neural network model may be trained by updating parameters of the neural network model, and parameters of at least two network models are updated together for at least one training iteration. In some embodiments, the machine-learned presentation model may be a deep learning model including one or more layers of nodes.
[0084] In some embodiments, identifying the T cells comprises co-culturing the T cells with one or more of the neoantigens in the subset under conditions that allow the T cells to expand.
[0085] In some embodiments, identifying the T cell comprises contacting the T cell with an MHC multimer comprising one or more of the neoantigens in the subset under conditions that allow binding of the T cell to the MHC multimer.
[0086] In some embodiments, the method further includes identifying the T cell receptor (TCR) of the identified T cells. In such embodiments, identifying the T cell receptor includes sequencing the T cell receptor sequence of the identified T cells. In such embodiments, the method can further include genetically engineering the T cells to express at least one of the one or more identified T cell receptors, culturing the T cells under conditions for T cell expansion, and infusing the expanded T cells into a subject. In such embodiments, genetically engineering the T cells to express at least one of the identified T cell receptors can include cloning the T cell receptor sequence of the identified T cells into an expression vector and transfecting each of the T cells with the expression vector.
[0087] In some embodiments, the method can further include culturing the identified T cells under conditions that expand the identified T cells, and injecting the expanded T cells into a subject.
[0088] In some embodiments, T cells that are antigen-specific for at least one of the neoantigens in the subset are identified using 5-30 mL of whole blood from the subject.
[0089] In some embodiments, the subset of neoantigens includes up to 20 neoantigens, and the specified T cells recognize at least two neoantigens in the subset of neoantigens.
[0090] In some embodiments, the MHC allele is a class I MHC allele.
[0091] Also disclosed herein are isolated T cells that are antigen-specific for at least one selected neoantigen from the subset of neoantigens mentioned above.
[0092] III. Identification of tumor-specific mutations in neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells). In particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject with cancer, but may not be present in normal tissues from the subject.
[0093] Genetic mutations in tumors can be considered useful for immunological targeting of tumors if they result in changes in the amino acid sequence of proteins exclusively in tumors. Useful mutations include: (1) non-synonymous mutations that result in different amino acids in proteins; (2) read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; (3) splice site mutations that result in the inclusion of an intron in mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) frameshift mutations or deletions that result in new open reading frames with new tumor-specific protein sequences. Mutations can also include one or more of non-frameshift insertions / deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in new ORFs.
[0094] For example, mutated peptides or mutated polypeptides resulting from splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.
[0095] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0096] Various methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. Several techniques have been described, including dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize amplification of target gene regions, typically by PCR. Still other methods rely on the generation of small signal molecules by invasive cleavage followed by mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.
[0097] PCR-based detection means can involve the multiplex amplification of multiple markers simultaneously.For example, it is well known in the art to select PCR primers so as to generate PCR products that do not overlap in size and can be analyzed simultaneously.Alternatively, it is possible to amplify different markers with primers that are differentially labeled and therefore can be differentially detected.Of course, hybridization-based detection means allows the differential detection of multiple PCR products in a sample.Other techniques that allow multiplex analysis of multiple markers are known in the art.
[0098] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA.For example, single nucleotide polymorphisms can be detected by using special exonuclease-resistant nucleotides, as disclosed in Mundy, CR (US Patent No. 4,656,127).According to this method, a primer complementary to the allele sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a specific animal or human.If the polymorphic site on the target molecule contains a nucleotide that is complementary to the specific exonuclease-resistant nucleotide derivative present, this derivative will be incorporated onto the end of the hybridized primer.This incorporation makes the primer resistant to exonucleases, thereby enabling its detection.Since the identity of the exonuclease-resistant derivative of the sample is known, the knowledge that the primer has become resistant to exonucleases reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of exogenous sequence data.
[0099] To determine the identity of the nucleotide at a polymorphic site, a solution-based method can be used (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087)). As in the method of Mundy, U.S. Pat. No. 4,656,127, a primer is used that is complementary to the allelic sequence immediately 3' to the polymorphic site. This method uses a labeled dideoxynucleotide derivative that becomes incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site to determine the identity of the nucleotide at that site. An alternative method, known as Genetic Bit Analysis or GBA, has been described by Goelet, P. et al. (PCT Application No. 92 / 15712). The Goelet, P. et al. method uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' to the polymorphic site. Goelet, P. et al. The method of Goelet, P. et al. uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which either the primer or the target molecule is immobilized on a solid phase.
[0100] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. et al., Nucl. Acids. Res. 17:7779-7784 (1989); Sokolov, B. P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M. et al., Proc. Natl. Acad. Sci. (USA) 88:1143-1147 (1991); Prezant, T. R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al. al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, signal is proportional to the number of incorporated deoxynucleotides, so that polymorphisms occurring in runs of the same nucleotide can result in a signal proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).
[0101] Numerous initiatives obtain sequence information directly from millions of individual molecules of DNA or RNA in parallel. Real-time single-molecule sequencing by synthesis techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent strands of DNA complementary to the template being sequenced. In one method, oligonucleotides 30–50 bases in length are covalently anchored at their 5′ ends to glass coverslips. These anchored strands serve two functions. First, they act as capture sites for the target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis for sequence reading. The capture primers serve as fixed-location sites for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and dye cleavage. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. As the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing-by-synthesis techniques also exist.
[0102] Any suitable sequencing-by-synthesis platform can be used to identify mutations. As mentioned above, four major sequencing-by-synthesis platforms are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing-by-synthesis platforms have also been described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the multiple nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. A capture sequence (also called a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to a support that can double as a universal primer.
[0103] As an alternative to capture sequences, a member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair, e.g., as described in U.S. Patent Application Publication No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.
[0104] Following capture, the sequence can be analyzed by single-molecule detection / sequencing, including, for example, template-dependent sequencing by synthesis, as described, for example, in the Examples and U.S. Patent No. 7,283,337. In sequencing by synthesis, surface-bound molecules are exposed to a multitude of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time, in a step-and-repeat mode. For real-time analysis, a different optical label can be incorporated for each nucleotide, and multiple lasers can be utilized for stimulation of the incorporated nucleotides.
[0105] Sequencing can also include other massively parallel sequencing or next-generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, ThermoPGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel sequencing technologies, and future generations of these technologies, can be used.
[0106] Any cell type or tissue can be used to obtain nucleic acid samples for use in the methods described herein.For example, DNA or RNA samples can be obtained from tumor or body fluids, for example, blood obtained by known techniques (for example, venipuncture) or saliva.Alternatively, nucleic acid testing can be performed on dry samples (for example, hair or skin).In addition, a sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal tissue is of the same tissue type as tumor.A sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal sample is of a different tissue type from tumor.
[0107] The tumor may include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0108] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. Peptides can be acid-eluted from tumor cells or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.
[0109] IV. Neoantigens Neoantigens can comprise nucleotides or polynucleotides. For example, neoantigens can be RNA sequences that encode polypeptide sequences. Neoantigens useful in vaccines can therefore comprise nucleotide sequences or polypeptide sequences.
[0110] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the neoantigen comprises nucleotide sequences (e.g., DNA or RNA) that encode the associated polypeptide sequence.
[0111] The one or more polypeptides encoded by the neoantigen nucleotide sequences can comprise at least one of the following: a binding affinity to MHC with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides; the presence of a sequence motif within or near the peptide that promotes proteasomal cleavage; and the presence of a sequence motif within or near the peptide that promotes TAP transport; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II polypeptides; and the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.
[0112] One or more neoantigens can be present on the surface of a tumor.
[0113] The one or more neoantigens can be immunogenic in a tumor-bearing subject, for example, capable of eliciting a T cell or B cell response in the subject.
[0114] One or more neoantigens that induce an autoimmune response in a subject can be eliminated from consideration in the context of generating a vaccine for a tumor-bearing subject.
[0115] The size of the at least one neoantigenic peptide molecule is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35 , about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derivable therein. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.
[0116] Neoantigenic peptides and polypeptides can be 15 residues or less in length, typically between about 8 and about 11 residues, particularly 9 or 10 residues, for MHC class I; and 6 to 30 residues for MHC class II.
[0117] If desired, longer peptides can be designed in several ways. In one example, if the likelihood of peptide presentation on HLA alleles is predicted or known, the longer peptides can consist of either (1) individual presented peptides with extensions of 2-5 amino acids toward the N- and C-termini of each corresponding gene product; or (2) a concatenation of some or all of the presented peptides, each with its extended sequence. In another example, if sequencing reveals long (more than 10 residues) neo-epitope sequences present in the tumor (e.g., due to frameshifts, readthrough, or intron inclusion resulting in novel peptide sequences), the longer peptides would (3) consist of the entire novel tumor-specific stretch of amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides that are presented to the strongest HLA alleles. In either example, the use of longer peptides may allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.
[0118] Neoantigenic peptides and polypeptides can be presented on HLA proteins. In some embodiments, the neoantigenic peptides and polypeptides are presented on HLA proteins with greater affinity than wild-type peptides. In some embodiments, the neoantigenic peptide or polypeptide can have an IC50 of at least 5000 nM or less, at least 1000 nM or less, at least 500 nM or less, at least 250 nM or less, at least 200 nM or less, at least 150 nM or less, at least 100 nM or less, at least 50 nM or less, or even less.
[0119] In some embodiments, the neoantigenic peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.
[0120] Also provided are compositions comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides can be derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some embodiments, the tumor-specific mutations are driver mutations for a particular cancer type.
[0121] Neoantigenic peptides and polypeptides with desired activities or properties can be modified to confer certain desirable attributes, e.g., improved pharmacological characteristics, while enhancing or at least retaining substantially all of the biological activity of the unmodified peptide, which binds to desired MHC molecules and activates appropriate T cells. For example, neoantigenic peptides and polypeptides can be further subjected to various modifications, such as conservative or non-conservative substitutions, which may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions refer to the replacement of an amino acid residue with another that is biologically and / or chemically similar, e.g., one hydrophobic residue with another hydrophobic residue, or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effects of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2nd Ed. (1984).
[0122] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful for increasing peptide and polypeptide stability in vivo. Stability can be assayed in a number of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally follows: Pooled human serum (type AB, non-heat-inactivated) is defatted by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatography conditions.
[0123] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is typically composed of relatively small, neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that the optional spacer need not be composed of the same residues and can therefore be a hetero- or homo-oligomer. If present, the spacer will usually be at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.
[0124] The neoantigenic peptide can be linked to a T helper peptide at either the amino or carboxy terminus of the peptide, either directly or via a spacer. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxin 830-843, influenza 307-319, and malaria sporozoite peritoneal sites 382-398 and 378-389.
[0125] Proteins or peptides can be produced by any technique known to those of skill in the art, including expressing proteins, polypeptides, or peptides through standard molecular biology techniques, isolating proteins or peptides from natural sources, or chemically synthesizing proteins or peptides. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located on the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.
[0126] In a further embodiment, the neoantigen comprises a nucleic acid (e.g., a polynucleotide) encoding a neoantigenic peptide or a portion thereof. The polynucleotide can be, for example, a single-stranded and / or double-stranded polynucleotide, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), or a polynucleotide having a phosphorothioate backbone, either in a natural or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further embodiment provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host; such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.
[0127] IV. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., a tumor-specific immune response. Vaccine compositions typically include multiple neoantigens selected, e.g., using the methods described herein. Vaccine compositions may also be referred to as vaccines.
[0128] The vaccine can contain 1 to 30 different peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine may contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, It may contain 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine contains 1-30 neoantigen sequences: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 It can contain 6, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.
[0129] In one embodiment, the different peptides and / or polypeptides, or the nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides are capable of binding to different MHC molecules, such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, a vaccine composition comprises coding sequences for peptides and / or polypeptides capable of binding to the most frequently occurring MHC class I molecules and / or MHC class II molecules. Thus, the vaccine composition can comprise different fragments capable of binding to at least two preferred, at least three preferred, or at least four preferred MHC class I molecules and / or MHC class II molecules.
[0130] The vaccine composition may generate a specific cytotoxic T cell response and / or a specific helper T cell response.
[0131] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are provided herein below. The composition can be combined with a carrier, such as a protein, or an antigen-presenting cell, such as a dendritic cell (DC), which can present peptides to T cells.
[0132] An adjuvant is any substance whose incorporation into a vaccine composition enhances or otherwise modifies the immune response to a neoantigen. The carrier can be a scaffold, such as a polypeptide or polysaccharide, to which the neoantigen can be bound. Optionally, the adjuvant is covalently or non-covalently conjugated.
[0133] The ability of adjuvant to increase the immune response to antigen is typically manifested by a significant or substantial increase in immune-mediated reaction or a reduction in disease symptoms.For example, the increase in humoral immunity is typically manifested by a significant increase in the titer of antibody produced against antigen, and the increase in T cell activity is typically manifested in an increase in cell proliferation, or cellular cytotoxicity, or cytokine secretion.Adjuvant can also change immune response, for example, by changing mainly humoral or Th response to mainly cellular or Th response.
[0134] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA206, Montanide ISA 50V, and Montanide. Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila's QS21 stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponins, mycobacterial extracts and synthetic bacterial cell wall mimics, and other proprietary adjuvants such as Ribi's Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are also useful. Several immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison AC; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Several cytokines have been directly linked to influencing dendritic cell migration to lymphoid tissues (e.g., TNF-α), accelerating dendritic cell maturation into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J ImmunotherEmphasis Tumor Immunol. 1996(6):414-418).
[0135] CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in vaccine settings. Other TLR-binding molecules, such as RNA that binds to TLR 7, TLR 8, and / or TLR 9, may also be used.
[0136] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies, such as cyclophosphamide, sunitinib, bevacizumab, Celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony-stimulating factors such as granulocyte-macrophage colony-stimulating factor (GM-CSF, sargramostim).
[0137] A vaccine composition can include more than one different adjuvant. Additionally, a therapeutic composition can include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.
[0138] The carrier (or excipient) can exist independently of the adjuvant. The function of the carrier can be, for example, to increase activity or immunogenicity, to provide stability, to increase biological activity, or to increase serum half-life, particularly to increase the molecular weight of the variant. Furthermore, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, a serum protein such as keyhole limpet hemocyanin, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is tolerated and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be a dextran, such as Sepharose.
[0139] Cytotoxic T cells (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. MHC molecules themselves are located on the cell surface of antigen-presenting cells. Therefore, CTL activation is possible when a trimeric complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when peptides are used to activate CTLs, but also when APCs bearing the respective MHC molecules are added, it can enhance the immune response. Therefore, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.
[0140] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenovirus (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880), etc. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0141] IV.A. Additional Considerations for Vaccine Design and Manufacturing IV.A.1. Determining a set of peptides covering all tumor subclones Truncal peptides, meaning those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides that are predicted to be highly likely to be presented and immunogenic, or if the number of truncal peptides that are predicted to be highly likely to be presented and immunogenic is small enough that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .
[0142] IV.A.2. Neoantigen Prioritization After applying all of the above neoantigen filters, it is possible that more candidate neoantigens remain available for vaccine inclusion than vaccine technology can accommodate. Additionally, uncertainty about various aspects of neoantigen analysis may remain, and trade-offs may exist between various attributes of candidate vaccine neoantigens. Therefore, instead of predetermined filters at each stage of the selection process, an integral multidimensional model can be considered, in which candidate neoantigens are placed in a space with at least the following axes, and selection is optimized using an integral approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmune risk is typically preferred) 2. Probability of sequencing artifacts (lower artifact probabilities are typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferable) 5. Gene Expression (higher expression is typically preferred) 6. HLA gene coverage (a greater number of HLA molecules involved in presenting a set of neoantigens may decrease the probability that tumors will evade immune attack through downregulation or mutation of HLA molecules) 7. HLA class coverage (covering both HLA-I and HLA-II may increase the probability of therapeutic response and decrease the probability of tumor immune evasion)
[0143] V. Methods of Treatment and Preparation Also provided are methods for inducing a tumor-specific immune response in a subject, vaccinating against the tumor, and treating and / or alleviating symptoms of cancer in a subject by administering to the subject one or more neoantigens, such as multiple neoantigens identified using the methods disclosed herein.
[0144] In some embodiments, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desired. The tumor can be any solid tumor, such as breast, ovarian, prostate, lung, kidney, stomach, colon, testicular, head and neck, pancreas, brain, melanoma, and other tissue organ tumors, as well as hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.
[0145] The neoantigen can be administered in an amount sufficient to induce a CTL response.
[0146] The neoantigen can be administered alone or in combination with other therapeutic agents, such as chemotherapeutic agents, radiation, or immunotherapy. Any suitable therapeutic treatment for the particular cancer can be administered.
[0147] In addition, the subject can be further administered with an anti-immunosuppressive / immunostimulatory substance, such as a checkpoint inhibitor.For example, the subject can be further administered with an anti-CTLA antibody, or anti-PD-1 or anti-PD-L1.Blocking CTLA-4 or PD-L1 with an antibody can enhance the immune response against cancerous cells in patients.In particular, blocking CTLA-4 has been shown to be effective when used in vaccination protocols.
[0148] The optimal amount of each neoantigen to be included in the vaccine composition and the optimal dosing regimen can be determined. For example, the neoantigen or its variants can be formulated for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Methods of injection include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administering vaccine compositions are known to those skilled in the art.
[0149] Vaccines can be edited so that the selection, number, and / or amount of neoantigens present in the composition are tissue-, cancer-, and / or patient-specific. For example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. Selection can depend on the specific type of cancer, the state of the disease, earlier treatment regimens, the patient's immune status, and, of course, the patient's HLA haplotype. Furthermore, vaccines can contain components that are personalized according to the individual needs of a particular patient. Examples include altering the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjusting for secondary treatments after a first round or scheme of treatment.
[0150] For compositions to be used as vaccines for cancer, neoantigens with similar normal self-peptides that are abundantly expressed in normal tissues can be avoided or present in low amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a high amount of a particular neoantigen, the respective pharmaceutical composition for treating that cancer can be present in high amount and / or can include more than one neoantigen specific to that particular neoantigen or pathway of that neoantigen.
[0151] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In therapeutic applications, the compositions are administered to patients in an amount sufficient to elicit an effective CTL response against the tumor antigen and cure or at least partially halt symptoms and / or complications. An amount adequate to achieve this is defined as a "therapeutically effective dose." An amount effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general health, and the judgment of the prescribing physician. It should be kept in mind that compositions can generally be used in severe disease states, i.e., life-threatening or potentially life-threatening situations, particularly when cancer has metastasized. In such instances, it is possible, and the treating physician may find it desirable, to administer substantial excesses of these compositions, taking into account the minimization of adventitious substances and the relatively non-toxic nature of the neoantigens.
[0152] For therapeutic use, administration can begin at the time of detection or surgical removal of the tumor, followed by boosting doses until at least symptoms are substantially abated, and for a period thereafter.
[0153] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against tumors. Disclosed herein are compositions for parenteral administration that contain a solution of a neoantigen, where the vaccine composition is dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. Various aqueous carriers can be used, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional, well-known sterilization techniques or sterile filtered. The resulting aqueous solutions can be packaged for use as is or lyophilized, with the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances required to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, and the like, for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.
[0154] Neoantigens can also be administered via liposomes, which target them to specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigen to be delivered is incorporated as part of the liposome, either alone or in combination with a molecule that binds to a receptor dominant among lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Liposomes filled with the desired neoantigen can thus be directed to the site of lymphoid cells, where they then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid lability, and stability of the liposomes in the bloodstream. Various methods are available for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.
[0155] For targeting to immune cells, the ligand to be incorporated into the liposome can include, for example, an antibody or fragment thereof specific for a cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at doses that vary according to, inter alia, the mode of administration, the peptide being delivered, and the stage of the disease being treated.
[0156] The peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient for therapeutic or immunization purposes. Numerous methods are conveniently used to deliver nucleic acids to a patient. For example, nucleic acids can be delivered directly as "naked DNA." This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), and U.S. Pat. Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery, as described, for example, in U.S. Pat. No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, DNA can be attached to particles, such as gold particles. Approaches for delivering nucleic acid sequences include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.
[0157] Nucleic acids can also be delivered by complexing them with cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO 96 / 18372; 9324640 WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833; 9106309 WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).
[0158] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter can also be included in a viral vector-based vaccine platform, such as the ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20 ( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0159] A means of administering nucleic acids uses minigene constructs encoding one or more epitopes. To generate DNA sequences (minigenes) encoding selected CTL epitopes for expression in human cells, the amino acid sequences of the epitopes are reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are then directly adjacent to generate a continuous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the minigene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.
[0160] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) compounds can also be complexed with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.
[0161] Also disclosed herein is a method of producing a tumor vaccine, comprising performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising multiple neoantigens or a subset of multiple neoantigens.
[0162] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector (e.g., a vector comprising at least one sequence encoding one or more neoantigens) disclosed herein can include culturing host cells under conditions suitable for expression of the neoantigen or vector, wherein the host cells comprise at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatographic, electrophoretic, immunological, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.
[0163] The host cell can comprise a Chinese hamster ovary (CHO) cell, an NS0 cell, yeast, or an HEK293 cell. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to the at least one nucleic acid sequence encoding the neoantigen or vector. In certain embodiments, the isolated polynucleotide can be a cDNA.
[0164] VI. Identification of neoantigens VI.A. Identification of Candidate Neoantigens A research method for NGS analysis of tumor and normal exomes and transcriptomes is described and applied in the specific space of neoantigens. 6,14,15The examples below consider certain optimizations for greater sensitivity and specificity for identifying neoantigens in a clinical setting. These optimizations can be grouped into two areas: those related to laboratory processes and those related to NGS data analysis.
[0165] VI.A.1. Laboratory Process Optimization The process improvements presented herein build on the concepts developed for reliable assessment of cancer driver genes in targeted cancer panels. 16 This addresses the challenges in high-precision neoantigen discovery from clinical specimens with low tumor content and small volumes by expanding the method to the whole-exome and whole-transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) unique average coverage across the tumor exome to detect mutations present at low mutant allele frequency due to either low tumor content or subclonal status. 2. Fewer than 5% of bases are covered at less than 100x to minimize missed potential neoantigens, e.g. a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for areas that are not sufficiently covered 3. Targeting uniform coverage across the normal exome, with less than 5% of bases covered below 20x, to minimize the chance of potential neoantigens remaining unclassified for somatic / germline status (and therefore unusable as TSNAs). 4. To minimize the total amount of sequencing required, sequence capture probes are designed only for the coding regions of genes, since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing 18. b. Elimination of genes predicted to produce few or no candidate neoantigens due to factors such as poor expression, suboptimal digestion by the proteasome, or atypical sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples can be subjected to probe-based enrichment with the same or similar probes used to capture the exome in DNA. 19 It is extracted using
[0166] VI.A.2. Optimizing NGS Data Analysis Analytical method improvements address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically allow for customization relevant for identifying neoantigens in the clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment, as it contains multiple MHC region assemblies that better reflect population polymorphism, as opposed to earlier genome releases. 2. Various programs 5 Overcoming the limitations of single mutation callers 20 by merging results from a. Single nucleotide mutations and indels are detected in tumor DNA, tumor RNA, and normal DNA with a range of tools including: Strelka 21 and Mutec t22 and programs based on comparison of tumor and normal DNA, such as; and 23 , UNCeqR, and other programs that incorporate tumor DNA, tumor RNA, and normal DNA. b. Indels are found in Strelka and ABRA 24 This is determined by a program that performs local reassembly, such as c. Structural rearrangements are 25 or Breakseq 26It is determined using specialized tools such as 3. To detect and prevent sample swapping, mutation calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. Extensive filtering of artificial calls is performed, for example, by: a. Removal of mutations found in normal DNA, potentially with relaxed detection parameters in the case of low coverage and permissive proximity criteria in the case of indels. b. Removal of mutations due to poor mapping quality or poor base quality 27 . c. Elimination of mutations resulting from re-emerging sequencing artifacts, even if not observed in the corresponding normal 27 Examples include mutations that are detected primarily on one strand. d. Removal of mutations detected in a set of unrelated controls 27 . 5.seq2HLA 28 , ATHLATES 29 or Optitype, and also combine exome and RNA sequencing data 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing, such as long-read DNA sequencing. 30 or adaptation of methods for linking RNA fragments to maintain continuity. 31 Includes. 6. Robust Detection of Nascent ORFs Arising from Tumor-specific Splice Variants in CLASS 32 , Bayesembler 33 , StringTie 34 This is done by assembling transcripts from RNA-seq data using Cufflinks, or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate the entire transcript from each experiment). 35Although commonly used for this purpose, it frequently produces an incredibly large number of splice variants, many of which are much shorter than the full-length gene, and may not be able to recover a simple positive control. The coding sequence and potential nonsense-mediated decay mechanisms reintroduced the variant sequence, SpliceR. 36 and MAMBA 37 Gene expression is determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels are determined using ASE. 38 or HTSeq 39 Potential filtering steps include: a. Removal of candidate nascent ORFs that are thought to be poorly expressed. b. Removal of candidate nascent ORFs predicted to trigger nonsense-mediated decay (NMD). 7. Candidate neoantigens observed only in RNA (e.g., neo-ORFs) that cannot be directly validated as tumor-specific are classified as likely to be tumor-specific according to additional parameters, for example, by considering the following: a. Presence of supporting cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmed trans-acting mutations in splicing factors in tumor DNA only. As an example, the gene that exhibited the most differential splicing in three independently published experiments with R625 mutant SF3B1 was 1. One experiment examined patients with uveal melanoma. 40 The second experiment examined uveal melanoma cell lines. 41 , and a third study looked at breast cancer patients. 42 Nevertheless, there was agreement. c. For novel splicing isoforms, the presence of confirmatory "novel" splice-junction reads in the RNASeq data. d. For de novo rearrangements, the presence of confirmatory exon-proximal reads in tumor DNA that are not present in normal DNA. e.GTEx 43 and absence from the gene expression compendium (i.e., making germline origin less likely). 8. Complementing reference genome alignment-based analyses by comparing tumor and normal reads (or k-mers derived from such reads) of assembled DNA to directly avoid alignment- and annotation-based errors and artifacts (e.g., for somatic mutations occurring near germline mutations or repeat-context indels).
[0167] In samples with polyadenylated RNA, the presence of viral and microbial RNA in the RNA-seq data will be assessed using RNA CoMPASS44 or similar methods to identify additional factors that may predict patient response.
[0168] VI.B. HLA Peptide Isolation and Detection Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) techniques after lysis and solubilization of tissue samples. 55~58 The clarified lysates were used for HLA-specific IP.
[0169] Immunoprecipitation was performed using antibodies coupled to beads, where the antibodies are specific for HLA molecules. For pan-class I HLA immunoprecipitation, a pan-class I CR antibody is used, and for class II HLA-DR, an HLA-DR antibody is used. The antibodies are covalently attached to NHS-Sepharose beads during overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60Immunoprecipitation can also be performed using antibodies that are not covalently attached to beads. Typically, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to retain the antibody on the column. Some antibodies that can be used to selectively enrich MHC / peptide complexes are listed below. TIFF0007763588000001.tif42149
[0170] The clarified tissue lysate is added to antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is saved for further experiments, including additional IPs. The IP beads are washed to remove nonspecific binding, and the HLA / peptide complexes are eluted from the beads using standard techniques. Protein components are removed from the peptides using molecular weight spin columns or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and, in some cases, stored at -20°C prior to MS analysis.
[0171] The dried peptides were reconstituted in an HPLC buffer suitable for reversed-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution on a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of peptide mass / charge (m / z) were collected at high resolution on an Orbitrap detector, followed by MS2 low-resolution scans on an ion trap detector after HCD fragmentation of selected ions. Additionally, MS2 spectra can be acquired using either CID or ETD fragmentation methods, or any combination of the three techniques to obtain greater amino acid coverage of the peptide. MS2 spectra can also be measured with high-resolution mass accuracy on an Orbitrap detector.
[0172] The MS2 spectra from each analysis were analyzed using Comet 61、62 and peptide identifications were analyzed using Percolator63~65 Further sequencing is performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or spectral matching and de novo sequencing are performed. 75 Sequencing methods including:
[0173] VI.B.1. Investigation of MS detection limits for comprehensive HLA peptide sequencing Peptide YVYVADVAAK (SEQ ID NO: 1) Using this method, the limit of detection was determined using various amounts of peptide loaded onto the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol (Table 1). The results are shown in Figure 1F. These results indicate that the lowest limit of detection (LoD) was in the attomole range (10 -18 ), a dynamic range spanning five orders of magnitude, and a signal-to-noise ratio in the low femtomole range (10 -15 ) appears to be sufficient for sequencing.
[0174] TIFF0007763588000002.tif56128
[0175] VII. Presented Model VII.A. System Overview 2A is an overview of an environment 100 for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for implementing a presentation identification system 160, which itself includes a presentation information store 165.
[0176] The presentation identification system 160 is a computer model, embodied in a computational system such as that discussed below with respect to FIG. 30, that receives a peptide sequence associated with a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the set of associated MHC alleles. The presentation identification system 160 can be applied to both class I and class II MHC alleles, making it useful in a variety of contexts. One specific example application of the presentation identification system 160 is to receive the nucleotide sequence of a candidate neoantigen associated with a set of MHC alleles from tumor cells of a patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the tumor's associated MHC alleles and / or induce an immunogenic response in the patient's 110 immune system. Those candidate neoantigens with a high likelihood, as determined by the system 160, can be selected for inclusion in a vaccine 118, such that an anti-tumor immune response can be elicited from the immune system of the patient 110 that provided the tumor cells. Furthermore, T cells with TCRs that have reactivity to candidate neoantigens with a high likelihood of presentation can be generated for use in T cell therapy, which also elicits an anti-tumor immune response from the patient's 110 immune system.
[0177] The presentation identification system 160 determines the presentation likelihood through one or more presentation models. Specifically, the presentation models generate a likelihood that a given peptide sequence will be presented for a set of associated MHC alleles, the likelihood being generated based on the presentation information stored in the storage device 165. For example, a presentation model may be generated for the peptide sequence "YVYVADVAAK" (SEQ ID NO: 1)" may generate a likelihood of whether a peptide sequence will be presented for the set of alleles HLA-A*02:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:03, HLA-C*01:04 on the cell surface of the sample. Presentation information 165 contains information about whether these peptides bind to various types of MHC alleles such that the peptides are presented by the MHC alleles, which is determined in the model according to the position of the amino acid in the peptide sequence. Based on presentation information 165, the presentation model can predict whether an unrecognized peptide sequence will be presented in association with the associated set of MHC alleles. As noted above, the presentation model can be applied to both class I and class II MHC alleles.
[0178] VII.B. Presentation information 2 illustrates a method for obtaining presentation information according to one embodiment. Presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. Allele interaction information includes information that affects presentation of peptide sequences that is dependent on the type of MHC allele. Allele non-interaction information includes information that affects presentation of peptide sequences that is independent of the type of MHC allele.
[0179] VII.B.1. Allelic Interaction Information The allele interaction information primarily includes identified peptide sequences known to be presented by one or more identified MHC molecules from humans, mice, etc. Of note, this may or may not include data obtained from tumor samples. Presented peptide sequences may be identified from cells expressing a single MHC allele. In this example, presented peptide sequences are generally collected from a monoallelic cell line engineered to express a predetermined MHC allele and then exposed to a synthetic protein. Peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows an exemplary peptide presented on the predetermined MHC allele HLA-DRB1*12:01. An example of this is shown in TIFF0007763588000003.tif5128, where a peptide is isolated and identified by mass spectrometry. In this situation, the direct association between the presented peptide and the MHC protein to which it binds is definitively known, since the peptide is identified through cells engineered to express a single, predetermined MHC protein.
[0180] Presented peptide sequences may also be collected from cells expressing multiple MHC alleles. Typically, in humans, six different types of MHC1 molecules and up to 12 different types of MHCII molecules are expressed by cells. Such presented peptide sequences may be identified from multi-allelic cell lines engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal or tumor tissue samples. In this particular example, MHC molecules can be immunoprecipitated from normal or tumor tissue. Peptides presented on multiple MHC alleles can similarly be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows six exemplary peptides. An example of this is shown in TIFF0007763588000004.tif21128, which is presented on the identified class I MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, and HLA-B*08:01, and the class II MHC alleles HLA-DRB1*10:01 and HLA-DRB1:11:01, isolated, and characterized by mass spectrometry. In contrast to monoallelic cell lines, the bound peptide is isolated from the MHC molecule prior to its identification, so the direct association between the presented peptide and the MHC protein to which it is bound may be unknown.
[0181] Allele interaction information can also include mass spectrometry ion currents, which depend on both the concentration of peptide-MHC molecule complexes and the ionization efficiency of the peptides. Ionization efficiency varies from peptide to peptide in a sequence-dependent manner. Generally, ionization efficiency varies from peptide to peptide over approximately two orders of magnitude, while the concentration of peptide-MHC complexes varies over an even larger range.
[0182] Allele interaction information can also include measured or predicted binding affinities between a given MHC allele and a given peptide. One or more affinity models can generate such predictions (72, 73, 74). For example, returning to the example shown in Figure 1D, representation 165 may represent a sequence of peptides YEMFNDKSF (SEQ ID NO:3) and class I alleles of HLA-A * The presentation information 165 may include a predicted binding affinity value of 1000 nM between 01:01. Few peptides with IC50 > 1000 nM are presented by the MHC, with lower IC50 values increasing the probability of presentation. (SEQ ID NO:8) and the class II allele HLA-DRB1:11:01.
[0183] The allele interaction information can also include measured or predicted stability values for MHC complexes. One or more stability models can generate such predictions. More stable peptide-MHC complexes (i.e., complexes with longer half-lives) are more likely to be presented in high copy number on tumor cells and on antigen-presenting cells that encounter vaccine antigens. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include a predicted stability value for a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a predicted stability value for the half-life of the class II molecule HLA-DRB1:11:01.
[0184] Allele interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at a faster rate are more likely to be presented at high concentrations on the cell surface.
[0185] Allele interaction information can also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of presented peptides have a length of 9 peptides. MHC class II molecules generally tend to present peptides with a length of 6-30 peptides.
[0186] Allele interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoded peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoded peptide. The presence of kinase motifs influences the probability of post-translational modifications that may enhance or interfere with MHC binding.
[0187] Allelic interaction information can also include expression or activity levels of proteins involved in post-translational modification processes, such as kinases (as measured or predicted by RNA-seq, mass spectrometry, or other methods).
[0188] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.
[0189] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual in question (e.g., as measured by RNA-seq or mass spectrometry): peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.
[0190] Allelic interaction information can also include the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals that express that particular MHC allele.
[0191] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and therefore, presentation of peptides by HLA-C is a priori less likely than presentation by HLA-A or HLA-B II. As another example, because HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, presentation of peptides by HLA-DP is predicted to be less likely than presentation by HLA-DR or HLA-DQ.
[0192] The allele interaction information can also include the protein sequence of a particular MHC allele.
[0193] Any of the MHC allele non-interacting information listed in the section below can also be modeled as MHC allele interacting information.
[0194] VII.B.2. Allelic Non-Interaction Information Non-allele-interacting information can include the C-terminal sequence adjacent to the neoantigen-encoded peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence can affect proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters an MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and therefore, the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in Figure 2C, presentation information 165 is the presented peptide FJIEJFOESS identified from the peptide's source protein. (SEQ ID NO:5) The C-terminal flanking sequence FOEIFNDKSLDKFJI (SEQ ID NO:9) may include:
[0195] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide mass spectrometry training data. As described later with respect to Figure 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, mRNA quantification measurements are determined from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).
[0196] The allele non-interacting information can also include N-terminal sequences adjacent to the peptide within its source protein sequence.
[0197] The allelic non-interaction information can also include a source gene for the peptide sequence. The source gene can be defined as an Ensembl protein family for the peptide sequence. In another example, the source gene can be defined as a source DNA or source RNA for the peptide sequence. The source gene can be represented, for example, as a string of nucleotides that encodes a protein, or alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode specific proteins. In another example, the allelic non-interaction information can also include a source transcript or isoform or a set of potential source transcripts or isoforms for the peptide sequence extracted from a database such as Ensembl or RefSeq.
[0198] The allelic non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.
[0199] The allele non-interaction information can also include the presence of protease cleavage motifs in peptides, optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and therefore less stable in cells, and therefore less likely to be presented.
[0200] Allelic non-interaction information can also include the turnover rate of the source protein when measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation, but this characteristic has low predictive power when measured in dissimilar cell types.
[0201] The allelic non-interaction information can also include the length of the source protein, optionally taking into account the specific splice variants ("isoforms") that are most highly expressed in tumor cells, as measured by RNA-seq or proteome mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.
[0202] Allele-free interaction information can also include the expression level of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Different proteasomes have different cleavage site preferences. More weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.
[0203] Allele-free interaction information can also include the expression of the peptide's source gene (e.g., as measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in the tumor sample. Peptides from genes with higher expression are more likely to be presented. Peptides from genes with undetectable levels of expression can be eliminated from consideration.
[0204] Allelic non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to nonsense-mediated decay as predicted by a model of nonsense-mediated decay, e.g., the model from Rivas et al., Science 2015.
[0205] Allelic non-interaction information can also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle. Genes that are expressed at low levels overall (as measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are displayed than genes that are stably expressed at very low levels.
[0206] Allelic non-interaction information can also include a comprehensive catalog of source protein properties, such as those provided in uniProt or the PDB (http: / / www.rcsb.org / pdb / home / home.do). These properties can include, among others, protein secondary and tertiary structure, subcellular localization, and Gene Ontology (GO) terms. Specifically, this information can include annotations operating at the protein level, e.g., 5'UTR length, and annotations operating at the level of specific residues, e.g., a helix motif between residues 300 and 310. These properties can also include turn motifs, sheet motifs, and disordered residues.
[0207] Allelic non-interacting information can also include features that describe the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (eg, alpha helix versus beta sheet); alternative splicing.
[0208] The allelic non-interaction information can also include properties that describe the presence or absence of presentation hotspots at the peptide's position in its source protein.
[0209] Allelic non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression level of the source protein in those individuals and the influence of the various HLA types of those individuals).
[0210] Allelic non-interaction information can also include the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias.
[0211] Expression of various gene modules / pathways (not necessarily containing the source protein of the peptides) as measured by gene expression assays such as RNASeq, microarrays, targeted panels such as Nanostring, or single / multiple genes representing gene modules measured by assays such as RT-PCR, that inform on the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs).
[0212] Allele non-interaction information can also include the copy number of the peptide's source gene in the tumor cell. For example, a peptide derived from a gene that is subject to homozygous deletion in the tumor cell can be assigned a presentation probability of zero.
[0213] The allele-non-interaction information can also include the probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP or that bind with higher affinity to TAP are more likely to be presented by MHC-I.
[0214] Allelic non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Higher TAP expression levels at MHC-I increase the probability of presentation of all peptides.
[0215] Allelic non-interaction information can also include the presence or absence of tumor mutations, including but not limited to: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3. ii. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery affected by loss-of-function mutations in the tumor have a reduced probability of presentation.
[0216] Presence or absence of functional germline polymorphisms, including but not limited to: i. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome).
[0217] The allelic non-interaction information can also include tumor type (eg, NSCLC, melanoma).
[0218] Allele non-interaction information can also include the known functionality of the HLA allele, e.g., as reflected by the HLA allele suffix. For example, the N suffix in the allele name HLA-A*24:09N indicates a null allele that is not expressed and therefore unlikely to present an epitope; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.
[0219] Allelic non-interaction information can also include clinical tumor subtype (eg, squamous cell lung cancer vs. non-squamous).
[0220] The allele non-interaction information can also include smoking history.
[0221] Allele non-interaction information can also include a history of sunburn, sun exposure, or exposure to other mutagens.
[0222] The allelic non-interaction information can also include regional expression of the peptide's source gene in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be represented.
[0223] The allelic non-interaction information can also include the frequency of the mutation in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.
[0224] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation can also include the mutation's annotation (e.g., missense, readthrough, frameshift, fusion, etc.) or whether the mutation is predicted to result in nonsense-mediated decay (NMD). For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in reduced mRNA translation, which reduces the probability of presentation.
[0225] VII.C. Presentation Identification System 3 is a high-level block diagram illustrating the computer logic components of presentation identification system 160, according to one embodiment. In this exemplary embodiment, presentation identification system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. Presentation identification system 160 also comprises a training data store 170 and a presentation model store 175. Some embodiments of model management system 160 have different modules than those described herein. Likewise, functionality may be distributed among the modules in a manner different from that described herein.
[0226] VII.C.1. Data Management Module The data management module 312 generates sets of training data 170 from the representation information 165. Each training data set contains a number of data examples, each of which contains at least one of the represented or unrepresented peptide sequences p i and the peptide sequence p i one or more relevant MHC alleles combined with i and the dependent variable y, which represents information that the presentation identification system 160 is interested in predicting new values of the independent variables. i and the independent variable z i Contains a set of
[0227] In one particular implementation referred to throughout the remainder of this specification, the dependent variable y i is the peptide p i but one or more associated MHC alleles a i However, in other implementations, the dependent variable y i is the result of the presentation identification system 160 determining the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that one is interested in predicting. For example, in another implementation, the dependent variable y i σ may also be a numerical value indicating the mass analysis ion current determined for the example data.
[0228] Peptide sequence p for data example i i is k i is a sequence of k amino acids, i can vary within a range among data instances i. For example, the range can be 8 to 15 for MHC class I, or 6 to 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p in the training data set are i may have the same length, e.g., 9. The number of amino acids in a peptide sequence may vary depending on the type of MHC allele (e.g., MHC allele in humans). MHC allele a for data example i i is the peptide sequence p i indicates whether it existed in combination with
[0229] The data management module 312 also manages the peptide sequences p contained in the training data 170. i and bound MHC allele a i Together with the binding affinity b i and stability i For example, the training data 170 may include a predictor of the peptide p i and a i The predicted binding affinity b between each of the bound MHC molecules shown in iAs another example, the training data 170 may contain a i The predicted stability value s for each of the MHC alleles shown in i may contain
[0230] The data management module 312 also receives the peptide sequence p i along with non-allele interacting variables such as C-terminal flanking sequences and mRNA quantification measurements. i It may also include.
[0231] The data management module 312 also identifies peptide sequences that are not presented by MHC alleles to generate the training data 170. Generally, this involves identifying a "longer" sequence of the source protein that contains the peptide sequence to be presented prior to presentation. If the presentation information contains an engineered cell line, the data management module 312 identifies a set of peptide sequences in the synthetic protein to which the cell was exposed that were not presented on the MHC alleles of the cell. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein from which the presented peptide sequence originated and identifies a set of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.
[0232] The data management module 312 also artificially generates peptides with random sequences of amino acids and identifies the generated sequences as peptides that are not presented on MHC alleles. This can be achieved by randomly generating peptide sequences, allowing the data management module 312 to easily generate large amounts of synthetic data for peptides that are not presented on MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, synthetically generated peptide sequences are very likely not presented by MHC alleles, even if they are included in proteins processed by cells.
[0233] 4 illustrates an exemplary set of training data 170A, according to one embodiment. Specifically, the first three data examples in training data 170A are a monoallelic cell line containing the allele HLA-C*01:03, and three peptide sequences: The fourth example data in training data 170A shows peptide presentation information from TIFF0007763588000005.tif9128. The fourth example data in training data 170A shows a multi-allelic cell line containing the alleles HLA-B*07:02, HLA-C*01:03, and HLA-A*01:01, and the peptide sequence QIEJOEIJE (SEQ ID NO: 13) The first data example shows peptide information from the peptide sequence QCEIOWARE (SEQ ID NO:14) was not presented by the allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, the negatively labeled peptide sequences may be randomly generated by the data management module 312 or may be identified from the source protein of the presented peptide. The training data 170A also includes a predicted binding affinity of 1000 nM and a predicted stability with a half-life of 1 hour for the peptide sequence-allele pair. The training data 170A also includes a predicted binding affinity of 1000 nM and a predicted stability with a half-life of 1 hour for the peptide FJELFISBOSJFIE. (SEQ ID NO: 15) and the C-terminal flanking sequence of 10 2 It also includes non-allele-interacting variables, such as the mRNA quantification measurements of TPM. The fourth data example is the peptide sequence QIEJOEIJE (SEQ ID NO: 13) was presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. Training data 170A also includes predicted binding affinity and stability values for each of the alleles, as well as the C-terminal flanking sequence of the peptide and mRNA quantification measurements for the peptide.
[0234] VII.C.2. Coding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more representation models. In one implementation, the encoding module 314 one-hot encodes sequences (e.g., peptide sequences or C-terminal flanking sequences) for a predetermined 20-letter amino acid alphabet. Specifically, k i Peptide sequence p having amino acids i is 20·k i p, which is represented as a row vector of elements, corresponding to the alphabet of the amino acid at the jth position of the peptide sequence. i 20·(j-1)+1 ,p i 20·(j-1)+2 ,...,p i 20·j A single element in has a value of 1. The remaining elements have a value of 0. As an example, for a given alphabet {A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y}, the three amino acid peptide sequence EAF of data example i is a 60-element row vector The C-terminal flanking sequence c i , and the protein sequence for the MHC allele d h , and other sequence data in the presentation information can be similarly coded as above.
[0235] If the training data 170 contains sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equivalent length by adding PAD characters to extend the predetermined alphabet. For example, this may be done by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the peptide sequence with the longest length in the training data 170. Thus, if the peptide sequence with the longest length is k 最大 amino acids, the encoding module 314 encodes each sequence as (20+1) k 最大It is represented numerically as a row vector of elements. For example, consider the extended alphabet {PAD,A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y} and k 最大 For a maximum amino acid length of ≡5, the same exemplary peptide sequence EAF of 3 amino acids is represented as a 105-element row vector The C-terminal flanking sequence c i or other sequence data can be similarly encoded as above. Thus, the peptide sequence p i or c i Each argument or row in represents the occurrence of a particular amino acid at a particular position in the sequence.
[0236] Although the above method for encoding sequence data has been described with respect to sequences having amino acid sequences, the method can be similarly extended to other types of sequence data, such as, for example, DNA or RNA sequence data.
[0237] The encoding module 314 also encodes one or more MHC alleles a for data instance i. i is encoded into an m-element row vector, where each element h=1,2,...,m corresponds to a uniquely identified MHC allele. The element corresponding to the identified MHC allele for data instance i has a value of 1. The remaining elements have a value of 0. As an example, among the m=4 uniquely identified MHC allele types {HLA-A*01:01, HLA-C*01:08, HLA-B*07:02, HLA-DRB1*10:01}, the alleles HLA-B*07:02 and HLA-DRB1*10:01 for data instance i corresponding to a multi-allelic cell line are encoded into a four-element row vector a i = [0 0 1 1], and a3 i =1 and a4 i = 1. An example with four identified MHC allele types is described herein, but the number of MHC allele types can actually be hundreds or thousands. As noted above, each data instance i typically contains a peptide sequence pi It contains up to six different MHC allele types associated with
[0238] The encoding module 314 also generates a label y for each data instance i. i We code, as a binary variable with values from the set {0,1}, where a value of 1 indicates that the peptide x i However, the associated MHC allele a i a value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele a i The dependent variable y i If represents the mass analysis ion current, the encoding module 314 may additionally scale the value using various functions, such as a log function with a range of [-∞,∞] for ion current values between [0,∞].
[0239] The coding module 314 encodes the peptide p i and the allele interaction variable x for the associated MHC allele h h i pairs as row vectors in which the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], and b h i is the predicted binding affinity for peptide p and associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).
[0240] In one example, the encoding module 314 encodes the measured or predicted values for binding affinity as a function of the allele interaction variable x h i The binding affinity information is expressed by incorporating the
[0241] In one example, the encoding module 314 encodes the measured or predicted values for binding stability as allele interaction variables x h i By incorporating the information into the
[0242] In one example, the encoding module 314 encodes the measured or predicted values for the binding on-rate as a function of the allele interaction variable x h i The combined on-rate information is expressed by incorporating
[0243] In one example, for peptides presented by class I MHC molecules, the encoding module 314 encodes the peptide length in the vector TIFF0007763588000008.tif11128 (However, TIFF0007763588000009.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i In another example, for peptides presented by class II MHC molecules, the encoding module 314 may include the peptide length in the vector TIFF0007763588000010.tif18146 (However, TIFF0007763588000011.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x hi can be included in
[0244] In one example, the encoding module 314 represents the RNA expression information of MHC alleles by incorporating the RNA-seq-based expression levels of the MHC alleles into an allele interaction variable xhi.
[0245] Similarly, the encoding module 314 encodes the allele non-interacting variable w i can be represented as a row vector in which the numerical representations of the allele-non-interacting variables are concatenated one after the other. For example, w i is [c i ] or [c i m i w i ], and w i is the C-terminal flanking sequence of peptide pi and the mRNA quantification measurement m associated with the peptide i Alternatively, one or more combinations of allele non-interacting variables may be stored individually (e.g., as individual vectors or matrices).
[0246] In one example, the encoding module 314 encodes the turnover rate or half-life as a function of the allele non-interacting variable w i represents the turnover rate of the source protein for the peptide sequence.
[0247] In one example, the encoding module 314 encodes the protein length as a function of the allele non-interacting variable w i represents the length of the source protein or isoform by incorporating
[0248] In one example, the encoding module 314 generates β1 i , β2 i , β5 i The mean expression of immunoproteasome-specific proteasome subunits, including the subunits, was calculated using the allele-noninteracting variable w iIncorporation into the IL-1 protein results in activation of the immunoproteasome.
[0249] In one example, the encoding module 314 encodes the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM) or a source protein of a gene or transcript of the peptide, by correlating the source protein abundance with an allele-non-interacting variable w i This is expressed by incorporating it into
[0250] In one example, the encoding module 314 calculates the probability that the transcript of the peptide's origin will undergo nonsense-mediated decay (NMD), for example, as estimated by the model in Rivas et al. Science, 2015, and calculates this probability as a function of the allele non-interaction variable w i This is expressed by incorporating it into
[0251] In one example, encoding module 314 represents the activation status of a gene module or pathway assessed via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM, for each of the genes in the pathway, and then computing a summary statistic, such as a mean, across the genes in the pathway. The mean is calculated using the allele-non-interaction variable w i can be incorporated into.
[0252] In one example, the encoding module 314 encodes the copy number of the source gene by dividing the copy number by the allele non-interacting variable w i This is expressed by incorporating it into
[0253] In one example, the encoding module 314 encodes the measured or predicted TAP binding affinity (e.g., in nanomolar units) relative to the allele-non-interacting variable w i The TAP binding affinity is expressed by including
[0254] In one example, the encoding module 314 encodes TAP expression levels measured by RNA-seq (and quantified, for example, by RSEM in units of TPM) as a function of the allele-non-interacting variable w i The expression level of TAP is represented by the inclusion of
[0255] In one example, the encoding module 314 encodes the tumor mutations as allele-non-interacting variables w i vector of indicator variables in (i.e., peptide p k is derived from a sample with a KRAS G12D mutation, d k = 1, otherwise 0).
[0256] In one example, the encoding module 314 encodes germline polymorphisms in antigen-presenting genes as a vector of indicator variables (i.e., peptide p k If is derived from a sample with a specific germline polymorphism in TAP, then d k = 1). These indicator variables are expressed as the allele non-interaction variables w i can be included in
[0257] In one example, the encoding module 314 represents tumor types as one-hot encoded vectors of length 1 for an alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot encoded variables are then combined into a set of allelic non-interaction variables w i can be included in
[0258] In one example, the encoding module 314 represents MHC allele suffixes by processing four-digit HLA alleles with various suffixes. For example, HLA-A*24:09N is considered a different allele from HLA-A*24:09 for purposes of the model. Alternatively, because HLA alleles ending in an N suffix are not expressed, the probability of presentation by an MHC allele with an N suffix can be set to zero for all peptides.
[0259] In one example, the encoding module 314 represents tumor subtypes as one-hot encoded vectors of length 1 for an alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot encoded variables are combined into an allelic non-interaction variable w i can be included in
[0260] In one example, the encoding module 314 encodes smoking history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a smoking history) can be included in k = 1 otherwise 0). Alternatively, smoking history can be coded as a one-hot encoded variable of length 1 for the smoking severity alphabet. For example, smoking status can be assessed on a 1-5 scale, with 1 indicating non-smoker and 5 indicating current heavy smoker. Because smoking history is primarily relevant for lung tumors, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a smoking history and the tumor type is lung tumor, and zero otherwise.
[0261] In one example, the encoding module 314 encodes sunburn history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a history of severe sunburn) can be included in k = 1 otherwise 0). Because severe sunburn is primarily associated with melanoma, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and zero otherwise.
[0262] In one example, the encoding module 314 represents the distribution of expression levels of a particular gene or transcript for each gene or transcript in the human genome as a summary statistic (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, the expression of peptide p in samples with the tumor type melanoma is expressed as a summary statistic (e.g., mean, median) of the distribution of expression levels. k Regarding peptide p k The measured gene or transcript expression levels of the genes or transcripts of origin are compared with the allele-non-interacting variable w i Not only can it be included in the peptide p in melanoma as measured by TCGA, k The mean and / or median gene or transcript expression of the genes or transcripts of a given source may also be included.
[0263] In one example, the encoding module 314 represents variant types as one-hot encoded variables of length 1 for an alphabet of variant types (e.g., missense, frameshift, NMD-induced, etc.). These one-hot encoded variables are referred to as allele-non-interacting variables w i can be included in
[0264] In one example, the encoding module 314 encodes the protein-level characteristics of the protein as values of the source protein annotation (e.g., 5′ UTR length) and the allele-non-interacting variable w i In another example, the encoding module 314 encodes the peptide p i The residue-level annotation of the source protein for peptide p i is equal to 1 if overlaps with the helical motif, otherwise it is equal to 0, or i The allele non-interaction variable wi represents the allele non-interaction variable wi by including an indicator variable equal to 1 if p is completely contained within the helix motif. iThe property that represents the proportion of residues in i can be included in
[0265] In one example, the encoding module 314 encodes the types of proteins or isoforms in the human proteome into an index vector o having a length equivalent to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is the peptide p k is 1 if comes from protein i, and 0 otherwise.
[0266] In one example, the encoding module 314 encodes the peptide p i Source gene G = gene(p i ) as a categorical variable with L possible categories (where L denotes the upper bound 1, 2, ..., L on the number of subscripted source genes).
[0267] In one example, the encoding module 314 encodes the peptide p i T = tissue type, cell type, tumor type, or tumor histology type of T = tissue (p i ) as a categorical variable with M possible categories (where M denotes an upper limit on the number of subscripted types 1, 2, ..., M). Tissue types can include, for example, lung tissue, cardiac tissue, intestinal tissue, and neural tissue. Cell types can include, for example, dendritic cells, macrophages, and CD4 T cells. Cancers can include, for example, lung adenocarcinoma, lung squamous cell carcinoma, melanoma, and non-Hodgkin's lymphoma.
[0268] The encoding module 314 also encodes the peptide p i and the variable z for the associated MHC allele h i The entire set of alleles is expressed as the allele interaction variable x i and the allele non-interaction variable w iFor example, the encoding module 314 may represent z h i [x h i w i ] or [w i x h i ] can be represented as a row vector equivalent to
[0269] VIII. Training Module The training module 316 constructs one or more presentation models that generate a likelihood of whether a peptide sequence will be presented by an MHC allele associated with the peptide sequence. k and peptide sequence p k MHC alleles associated with a k Given a set of peptide sequences, each proposed model k However, the associated MHC allele a k the estimate u, which indicates the likelihood that one or more of k Generate.
[0270] VIII.A. Overview The training module 316 constructs one or more representation models based on a training data set stored in storage 170, which is generated from the representation information stored in 165. Generally, regardless of the specific type of representation model, all representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function TIFF0007763588000012.tif4128 is a graph of the dependent variable y for one or more data examples S in the training data 170. i∈S and the estimated likelihood u for the data example S generated by the proposed model. i∈S In one particular implementation, which will be mentioned throughout the remainder of this document, the loss function TIFF0007763588000013.tif4128 is the negative log likelihood function given by equation (1a) as follows: However, in practice, another loss function may be used. For example, if the prediction is made for mass spectrometry ion current, the loss function is the mean square loss given by Equation 1b as follows: TIFF0007763588000015.tif10128
[0271] The proposed model may be a parametric model, where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function The various parameters of the proposed parametric model that minimizes TIFF0007763588000016.tif4128 are determined through a gradient-based numerical optimization algorithm, such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the proposed model may be a non-parametric model, in which the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.
[0272] VIII.B. Allele-by-Allele Model The training module 316 may build a presentation model to predict the presentation likelihood of a peptide on an allele-by-allele basis. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele.
[0273] In one implementation, the training module 316: TIFF0007763588000017.tif7128 estimates the likelihood of peptide pk being presented for a particular allele h. k where the peptide sequence x h k is the peptide p k and the coded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, which for convenience of description will be referred to as a transformation function throughout this specification. h(·) is an arbitrary function, which for convenience of description will be referred to as the dependence function throughout this specification, and the parameter θ determined for the MHC allele h h Based on the set of allele interaction variables x h k Generate a dependency score for the parameter θ for each MHC allele h. h The set of values is θ h where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele h.
[0274] Dependence function g h (x h k ;θ h ) output is the MHC allele h with at least the allele interaction characteristic x h k and in particular the peptide p k The dependency score for MHC allele h indicates whether the corresponding neoantigen is presented based on the amino acid position of the peptide sequence of p. For example, the dependency score for MHC allele h is determined by the relationship between the MHC allele h and the peptide p. k The transformation function f(·) transforms the input, more specifically, g in this example. h (x h k ;θ h ) is used to calculate the dependency score for peptide p k is converted to an appropriate value indicating the likelihood that it will be presented by the MHC allele.
[0275] In one particular implementation referred to throughout the remainder of this specification, f(·) is a function with range in [0,1] for the appropriate domain range. In one example, f(·) is The expit function is given by TIFF0007763588000018.tif10128. As another example, f(·) also has the following meaning: for values in the domain z greater than or equal to 0, It can also be the hyperbolic tangent function given by TIFF0007763588000019.tif4128. Alternatively, if the prediction is made for mass analysis ion currents with values outside the range [0,1], f(·) can be any function, for example, the identity function, the exponential function, the log function, etc.
[0276] Therefore, the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is given by the dependency function g h (·) is the peptide sequence p k to generate a corresponding dependency score. The dependency score can be generated by applying k may be transformed by a transformation function f(·) to generate the allele-specific likelihood that h will be presented by MHC allele h.
[0277] VIII.B.1 Dependence Functions for Allelic Interaction Variables In one particular implementation mentioned throughout this specification, the dependency function g h (·) is x h k Each allele interaction variable in is compared with the parameter θ determined for the relevant MHC allele h. h with the corresponding parameters in the set is an affine function given by TIFF0007763588000020.tif5128.
[0278] In another specific implementation mentioned throughout this specification, the dependency function g h (·) is a network model NN with a set of nodes arranged in one or more layers. h (·), The network function is given by TIFF0007763588000021.tif5128. The nodes are connected by parameters θh A node may be connected to other nodes through connections, each having an associated parameter in a set of . The value at one particular node may be represented as the sum of the values of the nodes connected to the particular node, weighted by the associated parameter mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and process data having amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how these interactions affect peptide presentation.
[0279] Generally speaking, the network model NN h (·) can be structured as feedforward networks such as artificial neural networks (ANNs), convolutional neural networks (CNNs), deep neural networks (DNNs), and / or recurrent networks such as long short-term memory networks (LSTMs), bidirectional recurrent networks, and deep bidirectional recurrent networks.
[0280] In one example, which will be mentioned throughout the remainder of this specification, each MHC allele in h=1, 2,..., m is associated with a separate network model, NN h (·) denotes the output from the network model related to MHC allele h.
[0281] FIG. 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h=3. As shown in FIG. 5, the network model NN3(·) for MHC allele h=3 includes three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2), ..., θ3(10). The network model NN3(·) includes three allele interaction variables x3 for MHC allele h=3. k (1), x3 k (2) and x3 k (3) receives input values (individual data examples, including encoded polypeptide sequence data and any other training data used) and generates the value NN3(x3 k ) The network function may include one or more network models, each taking a different allele interaction variable as input.
[0282] In another example, the identified MHC alleles h=1,2,...,m are used to model the MHC alleles in a single network NN H (·) and NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θ h may correspond to the set of parameters for a single network model, and thus the parameters θ h The set of can be shared by all MHC alleles.
[0283] Figure 6A shows an exemplary network model NN shared by MHC alleles h = 1, 2, ..., m. H (·). As shown in Figure 6A, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. The network model NN3(·) has an allele interaction variable x3 for MHC allele h=3. k , and the value NN3(x3k ) and outputs m values.
[0284] In yet another example, a single network model NN H (·) is the allele interaction variable x for MHC allele h h k and the encoded protein sequence d h In such an example, the parameter θ h may again correspond to the set of parameters for a single network model, and thus the parameters θ h The set of MHC alleles can be shared by all MHC alleles. Therefore, in such an example, NNh(·) is a single network model with inputs [x h k d h ], a single network model NN H Such a network model is advantageous because it can correctly predict peptide presentation probabilities for MHC alleles that were unknown in the training data simply by identifying their protein sequences.
[0285] Figure 6B shows an exemplary network model NN shared by MHC alleles. H (·). As shown in Figure 6B, the network model NN H (·) takes as input the allele interaction variables and protein sequence of MHC allele h=3 and calculates the dependency score NN3 (x3 k ) is output.
[0286] In yet another example, the dependency function g h (·)teeth, TIFF0007763588000022.tif5128, where g' h (x h k ;θ' h) is an affine function, network function, etc., with a set of parameters θ'h, which represents the baseline probability of presentation for MHC allele h, and the bias parameter θ in the set of parameters for the allele interaction variables of the MHC alleles. h 0 accompanied by.
[0287] In another implementation, the bias parameter θ h 0 may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ for the MHC allele h h 0 is θ 遺伝子(h) 0 where gene (h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family "HLA-A," and the bias parameter θ for each of these MHC alleles may be h 0 As another example, if the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 are assigned to the "HLA-DRB" gene family, and the bias parameters θ for each of these MHC alleles are h 0 can be shared.
[0288] As an example, returning to equation (2), the affine dependency function g h Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000023.tif5128, where x3k is the allele interaction variable identified for MHC allele h=3 and θ3 is the set of parameters determined for MHC allele h=3 through loss function minimization.
[0289] As another example, let us consider the MHC allele h = 3 for peptide p among m = 4 different identified MHC alleles using separate network transformation functions gh(·). k The likelihood that will be presented is TIFF0007763588000024.tif5128 can be generated by k is the allele interaction variable identified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.
[0290] Figure 7 shows the correlation coefficients of peptide p associated with MHC allele h=3 using the exemplary network model NN3(·). k As shown in Figure 7, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0291] VIII.B.2. Per allele with allele-noninteracting variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF0007763588000025.tif7128 k We model the estimated presentation likelihood uk, where w k is the peptide p k means the coded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interacting variable w Based on the set of allele-non-interacting variables w k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values of θ h and θw where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele.
[0292] Dependence function g w (w k ;θ w ) output is a measure of the peptide p expression by one or more MHC alleles based on the influence of allele-non-interacting variables. k represents a dependency score for an allele-non-interacting variable, indicating whether peptide p k The C-terminal flanking sequences and peptide p, which are known to positively influence the presentation of k If peptide p is bound, it may have a high value k The C-terminal flanking sequences and peptide p k If bound, it may have a low value.
[0293] According to equation (8), the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is the function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. w (·) is also applied to the coded version of the allele-non-interacting variable to generate a dependency score for the allele-non-interacting variable. Both scores are combined, and the combined score is used to estimate the association of peptide sequence p with MHC allele h. k is transformed by a transformation function f(·) to produce the allele-specific likelihoods that will be presented.
[0294] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (2) k the allele interaction variable xh k One may include the allele non-interacting variable wk in the prediction by adding The image can be given by TIFF0007763588000026.tif7128.
[0295] VIII.B.3 Dependence Functions for Allelic Non-Interacting Variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g for allelic non-interacting variables w (·) is an affine function, or a separate network model for the allelic non-interaction variables w k It can be a network function related to
[0296] Specifically, the dependency function g w (·) is w k The allele non-interaction variables in w with the corresponding parameters in the set is an affine function given by TIFF0007763588000027.tif5128.
[0297] Dependence function g w (·) also corresponds to the parameter θ w the network model NN with relevant parameters in the set w (·), TIFF0007763588000028.tif5128. A network function may include one or more network models, each taking different allele-non-interacting variables as input.
[0298] In another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF0007763588000029.tif5128, where g' w (w k ;θ'w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of m k is the peptide p k is the mRNA quantitative measurement for , h(·) is a function that transforms the quantitative measurement, and θ w m is a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measurement to generate a dependency score for the mRNA quantification measurement. In one particular embodiment mentioned throughout the remainder of this specification, h(·) is a log function, although in practice h(·) can be any one of a variety of different functions.
[0299] In yet another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF0007763588000030.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of k is the peptide p k is the indicator vector described in Section VII.C.2, which represents proteins and isoforms in the human proteome, and θ w o is the set of parameters in the set of parameters for the allele non-interacting variables that are combined with the indicator vector. k and parameter θ w o If the dimension of the set is significantly higher, TIFF0007763588000031.tif4128( Parameter regularization terms such as L1 norm, L2 norm, combination, etc. can be added to the loss function when determining the parameter value. The optimal value of the hyperparameter λ can be determined through an appropriate method.
[0300] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF0007763588000033.tif13128However, g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF0007763588000034.tif5128 is peptide p k is an indicator function equal to 1 if θ is derived from the source gene l as described above for allele-non-interacting variables, and θ w l is a parameter that indicates the "antigenicity" of the source gene l. In one variation, L is sufficiently large, and therefore the number of parameters θ w l=1, 2,...,L If is large enough, TIFF0007763588000035.tif5128 parameter regularization term (where TIFF0007763588000036.tif4128 can be added to the loss function when determining the parameter value (such as L1 norm, L2 norm, or a combination). The optimal value of the hyperparameter λ can be determined by an appropriate method.
[0301] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF0007763588000037.tif13146However, g' w (w k ;θ'w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF0007763588000038.tif5128 is a peptide p k is derived from source gene l, and peptide p k is an indicator function that is equal to 1 if originates from tissue type m, and θ w lm is a parameter indicating the antigenicity of the combination of source gene l and tissue type m. Specifically, the antigenicity of gene l in tissue type m may indicate the residual tendency of cells of tissue type m to present peptides derived from gene l after adjustment for RNA expression and peptide sequence context.
[0302] In one variation, L or M is sufficiently large, so that the number of parameters θ w lm=1, 2,...,LM If is large enough, Parameter regularization term such as TIFF0007763588000039.tif5128 (where, TIFF0007763588000040.tif4128 can be added to the loss function when determining the parameter values (such as L1 norm, L2 norm, or combination). The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the parameter values so that the coefficients for the same source gene do not differ significantly between tissue types. For example, a penalty term such as: TIFF0007763588000041.tif18128 (However, TIFF0007763588000042.tif5128 is the average antigenicity across tissue types for source gene l) can add a penalty to the standard deviation of antigenicity across different tissue types in the loss function.
[0303] In practice, the dependence function g on the allelic non-interacting variables can be calculated by combining any of the additional terms in equations (10), (11), (12a) and (12b). w For example, the term h(·) representing the mRNA quantification measurement in equation (10) and the term representing the antigenicity of the source gene in equation (12) can be added together along with any other affine or network functions to generate a dependency function for the allele-non-interacting variables.
[0304] As an example, returning to equation (8), the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000043.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0305] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000044.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0306] FIG. 8 shows exemplary network models NN3(·) and NN w Peptide p associated with MHC allele h=3 using (·) kAs shown in Figure 8, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0307] VIII.C. Multi-Allele Models The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof. [Example]
[0308] VIII.C.1. Example 1: Maximum Per Allele Model In one implementation, the training module 316 trains peptides p associated with a set of MHC alleles H. k The estimated presentation likelihood u k is the presentation likelihood determined for each of the MHC alleles h in set H determined based on cells expressing a single allele, as explained above in conjunction with equations (2)-(11). TIFF0007763588000045.tif4128. Specifically, the presentation likelihood u k teeth, In one implementation, the function is a maximum function, as shown in equation (12), where the proposed likelihood u kcan be determined as the maximum of the presentation likelihood for each MHC allele h in set H. TIFF0007763588000047.tif5128
[0309] VIII.C.2. Example 2.1: Sum Function Model In one implementation, the training module 316 trains peptides p k The estimated presentation likelihood u k of, TIFF0007763588000048.tif13128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles H associated with x h k is the peptide p k and the coded allele interaction variables for the corresponding MHC alleles. The parameter θ for each MHC allele h h The set of values is θ h The dependence function g can be determined by minimizing a loss function with respect to i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:
[0310] According to equation (13), the peptide sequence p k The likelihood that a given allele will be presented by one or more MHC alleles h is given by the dependency function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate a corresponding score for the allele interaction variable. The scores for each MHC allele h are combined to generate a corresponding score for the peptide sequence p k is transformed by a transformation function f(·) to produce the presentation likelihood that MHC allele H will be presented by the set of MHC alleles H.
[0311] The model presented in equation (13) is that for each peptide p k It differs from the allele-by-allele model of equation (2) in that the number of relevant alleles for a can be greater than 1. In other words, h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0312] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000049.tif5128 can be generated by k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0313] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000050.tif5128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.
[0314] FIG. 9 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). kAs shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0315] VIII.C.3. Example 2.2: Sum Function Model with Allelic Non-Interacting Variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF0007763588000051.tif13128, peptide p k The estimated presentation likelihood u k where w k is the peptide p k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values of θ h and θ w The dependence function g can be determined by minimizing a loss function with respect to i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:
[0316] Therefore, according to equation (14), one or more MHC alleles H can bind to a peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p kto generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores are combined and the combined score is used to estimate the association of the peptide sequence p with the MHC allele H. k is transformed by a transformation function f(·) to produce the presentation likelihood that
[0317] In the model presented in equation (14), each peptide p k The number of relevant alleles for a can be greater than 1. h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0318] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000052.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0319] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000053.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0320] FIG. 10 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 10, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. The network model NN3(·) generates the allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0321] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (15) k the allele interaction variable x h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF0007763588000054.tif13128.
[0322] VIII.C.4. Example 3.1: Model with Implicit Allele-by-Allele Likelihood In another implementation, the training module 316 k The estimated presentation likelihood u k of, TIFF0007763588000055.tif7128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles h∈H associated with u' k h is the implicit allele-specific presentation likelihood for MHC allele h, and vector v has elements v h But, a h k ·u' k h where s(·) is a vector corresponding to v, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values of the input within a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, although it will be recognized that in other embodiments, s(·) may be any function, such as a maximum function. A set of values for the parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.
[0323] The presentation likelihood in the presentation model of equation (17) is the likelihood that each peptide p is presented by an individual MHC allele h. k The implicit allele-specific presentation likelihood u' corresponds to the likelihood that k h The implicit per-allele likelihood differs from the per-allele presentation likelihood of Section VIII.B in that the parameters for the implicit per-allele likelihood can be learned from a multi-allelic setting, in addition to a single-allelic setting, where the direct association between the presented peptide and the corresponding MHC allele is unknown. Thus, in a multi-allelic setting, the presentation model is based on the likelihood of the peptide p kNot only can we estimate whether peptide p is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are present in peptide p k The individual likelihood u' indicates which person is most likely to have presented k h∈H The advantage of this is that the presented model can generate implicit likelihoods without training data for cells expressing a single MHC allele.
[0324] In one particular implementation that will be mentioned throughout the remainder of this specification, r(·) is a function with range [0,1]. For example, r(·) is a clip function: r(z)=min(max(z,0),1) may be the minimum value between z and 1, which represents the likelihood u k In another implementation, r(·) is chosen as r(z)=tanh(z) where the domain z is greater than or equal to 0.
[0325] VIII.C.5. Example 3.2: Sum of Functions Model In one particular implementation, s(·) is a summation function, and the presentation likelihood is given by summing the implicit per-allele presentation likelihoods. TIFF0007763588000056.tif15128
[0326] In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: The likelihood of the image generated by TIFF0007763588000057.tif7128 is Let it be estimated by TIFF0007763588000058.tif13128.
[0327] According to equation (19), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence pk to generate the corresponding dependency scores for the allele interaction variables. Each dependency score can be generated by first applying the implicit per-allele presentation likelihood u' k h The allele likelihood u' is transformed by the function f(·) to generate k h are combined and a clipping function is applied to the combined likelihood to clip the values into the range [0,1] to produce a peptide sequence p k A presentation likelihood can be generated that g will be presented by a set of MHC alleles H. The dependency function g h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:
[0328] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000059.tif7128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0329] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000060.tif7128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.
[0330] FIG. 11 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) are then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.
[0331] In another implementation, if the prediction is made in terms of the log of the mass analysis ion current, then r(·) is the log function and f(·) is the exponential function.
[0332] VIII.C.6. Example 3.3: Sum of Functions Model with Allelic Non-Interacting Variables In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: The likelihood of the image generated by TIFF0007763588000061.tif7128 is As generated by TIFF0007763588000062.tif13128, the effects of allelic non-interacting variables are incorporated into peptide presentation.
[0333] According to equation (21), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h(·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele non-interaction variables to generate dependency scores for the allele non-interaction variables. The scores of the allele non-interaction variables are combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by the function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined, and a clipping function is applied to the combined output to clip values into the range [0,1] to determine the likelihood of presentation of the peptide sequence p by the MHC allele H. k A likelihood of being presented can be generated. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:
[0334] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000063.tif7128, where w k is the peptide p k are the allele-non-interacting variables identified for θw, and θw is the set of parameters determined for the allele-non-interacting variables.
[0335] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007763588000064.tif7128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0336] FIG. 12 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 12, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·). The network model NN3(·) generates an allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated, which is also the same network model NN w (·) output NN w (w k ) and mapped by a function f(·). Both outputs are combined to give the estimated presentation likelihood u k Generate.
[0337] In another implementation, the implicit per-allele presentation likelihood for MHC allele h can be calculated as: The likelihood of the image generated by TIFF0007763588000065.tif7128 is Generated by TIFF0007763588000066.tif13128.
[0338] VIII.C.7. Example 4: Quadratic Model In one implementation, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k teeth, TIFF0007763588000067.tif13128, where the element u' k h is the implicit per-allele presentation likelihood for MHC allele h. A set of values for parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit per-allele presentation likelihood can be in any of the forms shown in equations (18), (20), and (22) above.
[0339] In one embodiment, the model of equation (23) is k However, there is a possibility that a given antigen may be simultaneously presented by two MHC alleles, which may imply that presentation by the two HLA alleles is statistically independent.
[0340] According to equation (23), one or more MHC alleles H bind to the peptide sequence p k The presentation likelihood is calculated by combining the implicit per-allele presentation likelihood and the MHC allele H k Each pair of MHC alleles is assigned to a peptide p such that it generates a presentation likelihood that p will be presented. k can be generated by subtracting from the sum the likelihood that
[0341] For example, the affine transformation function g h The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF0007763588000068.tif5128 can be generated byk , x3 k are the allele interaction variables identified for HLA alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for HLA alleles h=2, h=3.
[0342] As another example, the network transformation function g h (·), g w The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF0007763588000069.tif5128, where NN2(·) and NN3(·) are the network models specified for HLA alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for HLA alleles h=2 and h=3.
[0343] IX. Example 5: Prediction Module The prediction module 320 receives sequence data and selects candidate neoantigens in the sequence data using the proposed model. Specifically, the sequence data may be DNA sequences, RNA sequences, and / or protein sequences extracted from tumor tissue cells of a patient. The prediction module 320 converts the sequence data into a plurality of peptide sequences p having 8-15 amino acids for MHC-I or 6-30 amino acids for MHC-II. k For example, the prediction module 320 processes a predetermined sequence TIFF0007763588000070.tif5128, three peptide sequences with nine amino acids TIFF0007763588000071.tif9128. In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by comparing sequence data extracted from a patient's normal tissue cells with sequence data extracted from the patient's tumor tissue cells to identify segments that have one or more mutations.
[0344] The prediction module 320 applies one or more presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation models to the candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences with an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences with the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the selected candidate neoantigens for a given patient can be injected into the patient to induce an immune response.
[0345] X. Example 6: Patient Selection Module The patient selection module 324 selects a subset of patients for vaccine therapy and / or T cell therapy based on whether the patients meet the selection criteria. In one embodiment, the selection criteria are determined based on the patient's likelihood of presentation of neoantigen candidates generated by the presentation model. By adjusting the selection criteria, the patient selection module 324 can adjust the number of patients who receive vaccine administration and / or T cell therapy based on the patient's likelihood of presentation of neoantigen candidates. Specifically, strict selection criteria may result in a smaller number of patients being treated with the vaccine and / or T cell therapy, but a higher proportion of vaccine and / or T cell therapy-treated patients who receive effective treatment (e.g., one or more tumor-specific neoantigens (TSNAs) and / or one or more neoantigen-reactive T cells). In contrast, looser selection criteria may result in a larger number of patients being treated with the vaccine and / or T cell therapy, but a lower proportion of vaccine and / or T cell therapy-treated patients who receive effective treatment. The patient selection module 324 alters the selection criteria based on a desired balance between a target proportion of patients receiving treatment and the proportion of patients receiving effective treatment.
[0346] In some embodiments, the selection criteria for selecting patients to receive vaccine therapy are the same as the selection criteria for selecting patients to receive T cell therapy. However, in alternative embodiments, the selection criteria for selecting patients to receive vaccine therapy may differ from the selection criteria for selecting patients to receive T cell therapy. Sections XA and XB below discuss the selection criteria for selecting patients to receive vaccine therapy and T cell therapy, respectively.
[0347] Selection of patients for XA vaccine treatment In one embodiment, a patient is associated with a corresponding therapeutic subset of v neoantigen candidates that can potentially be included in a personalized vaccine for that patient, having a vaccine volume v. In one embodiment, the therapeutic subset for a patient is the neoantigen candidate with the highest likelihood of presentation as determined by the presentation model. For example, if a vaccine can include v=20 epitopes, the vaccine can include a therapeutic subset for each patient with the highest likelihood of presentation as determined by the presentation model. However, it will be appreciated that in other embodiments, the therapeutic subset for a patient can be determined based on other methods. For example, the therapeutic subset for a patient can be randomly selected from the set of neoantigen candidates for that patient, or can be determined in part based on a combination of factors including prior art models that model the binding affinity or stability of peptide sequences, or the likelihood of presentation obtained from the presentation model and affinity or stability information for those peptide sequences.
[0348] In one embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the patient's tumor mutation burden is equal to or higher than a minimum mutation burden. A patient's tumor mutation burden (TMB) indicates the total number of nonsynonymous mutations in the tumor exome. In one embodiment, the patient selection module 324 selects a patient for vaccine treatment if the patient's absolute TMB number is equal to or higher than a predetermined threshold. In another implementation, the patient selection module 324 selects a patient for vaccine treatment if the patient's TMB is within a threshold percentile among the TMBs determined for the set of patients.
[0349] In another embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the patient's utility score based on the patient's therapeutic subset is equal to or greater than the minimum utility score. In one embodiment, the utility score is a measure of the estimated number of presented antigens from the therapeutic subset.
[0350] The estimated number of presented antigens can be predicted by modeling the presentation of neoantigens as random variables with one or more probability distributions. In one implementation, the utility score for patient i is the expected number of presented neoantigen candidates from the treatment subset, or a specific function thereof. As an example, the presentation of each neoantigen can be modeled as a Bernoulli random variable, where the probability of presentation (success) is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation (success) of each neoantigen candidate is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation of each neoantigen candidate is given by the probability of presentation of each neoantigen candidate with the highest presentation likelihood, u i1 , u i2 , …, u iv V types of neoantigen candidates p i1 , p i2 , …, p iv Treatment subset S i Regarding neoantigen candidate p ij The presentation of random variable A ij where: The expected number of neoantigens presented is given by the sum of the likelihoods of presentation of each neoantigen candidate. In other words, the utility score for patient i is expressed as: The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility score for the vaccine treatment.
[0351] In another implementation, the utility score for patient i is the probability that at least a threshold number of neoantigens k are presented. In one example, a therapeutic subset S of candidate neoantigens is i The number of presented antigens in is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the likelihood of presentation of each of the epitopes. In detail, the number of presented antigens for patient i is determined by the random variable N i can be given by: TIFF0007763588000074.tif13128 where PBD(·) denotes the Poisson binomial distribution. The probability that at least a threshold number of neoantigens k are presented is given by the number of presented antigens N i is given by the probability that k is equal to or greater than k. In other words, the utility score for patient i is expressed as: The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility score for the vaccine treatment.
[0352] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of neoantigen candidates that have binding affinities or predicted binding affinities below a fixed threshold (e.g., 500 nM) for one or more of the patient's HLA alleles. i The threshold is the number of neoantigens in the target region. In one example, the fixed threshold is in the range of 1000 nM to 10 nM. Optionally, the utility score may only count neoantigens detected as expressed by RNA-seq.
[0353] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of candidate neoantigens whose binding affinity to one or more HLA alleles for that patient is less than or equal to a threshold percentile of the binding affinity of a random peptide to that HLA allele. i The threshold percentile is the number of neoantigens in the target region. In one example, the threshold percentile is between the 10th percentile and the 0.1th percentile. Optionally, the utility score may only count neoantigens detected as expressed by RNA-seq.
[0354] It will be appreciated that the example utility scores described with respect to equations (25) and (27) are for illustrative purposes only, and that the patient selection module 324 may use other statistics or probability distributions to generate utility scores.
[0355] Patient selection for XB T-cell therapy In another embodiment, instead of or in addition to receiving vaccine therapy, a patient can receive T cell therapy. Similar to vaccine therapy, in embodiments in which a patient receives T cell therapy, the patient can be associated with a corresponding therapeutic subset of the v candidate neoantigens described above. This therapeutic subset of the v candidate neoantigens can be used to identify in vitro T cells from the patient that are reactive to one or more of the v candidate neoantigens. These identified T cells can then be expanded and infused into the patient in a personalized T cell therapy.
[0356] Patients can be selected for T cell therapy at two different time points: the first time point after a therapeutic subset of v candidate neoantigens for the patient is predicted using the model but before in vitro screening of T cells specific for the predicted therapeutic subset of v candidate neoantigens; and the second time point after in vitro screening of T cells specific for the predicted therapeutic subset of v candidate neoantigens.
[0357] First, patients can be selected for T cell therapy after a therapeutic subset of v neoantigen candidates for the patient is predicted, but before in vitro identification of T cells from the patient specific for the predicted therapeutic subset of v neoantigen candidates is performed. Specifically, because in vitro screening of neoantigen-specific T cells from patients can be costly, it may be desirable to select patients for screening for neoantigen-specific T cells only if the patient is likely to have neoantigen-specific T cells. Patient selection prior to the in vitro T cell screening step can use the same criteria used to select patients for vaccine therapy. Specifically, in some embodiments, the patient selection module 324 can select patients for T cell therapy if the patient's tumor mutational burden is equal to or higher than a minimum mutational burden, as described above. In another embodiment, the patient selection module 324 can select patients for T cell therapy if the patient's utility score based on the therapeutic subset of v neoantigen candidates for the patient is equal to or higher than a minimum utility score, as described above.
[0358] Second, in addition to or instead of selecting patients for T cell therapy before in vitro identification of T cells from the patient that are specific for a therapeutic subset of the predicted v neoantigen candidates, patients can also be selected for T cell therapy after in vitro identification of T cells that are specific for a therapeutic subset of the predicted v neoantigen candidates. Specifically, a patient can be selected for T cell therapy if in vitro screening of the patient's T cells for neoantigen recognition identifies at least a threshold amount of neoantigen-specific TCRs for the patient. For example, a patient can be selected for T cell therapy only if at least two neoantigen-specific TCRs have been identified for the patient, or only if neoantigen-specific TCRs have been identified for two different neoantigens.
[0359] In another embodiment, a patient can be selected to receive T cell therapy only if a threshold amount of neoantigens from a therapeutic subset of v candidate neoantigens for that patient is recognized by the patient's TCR. For example, a patient can be selected to receive T cell therapy only if at least one neoantigen from a therapeutic subset of v candidate neoantigens for that patient is recognized by the patient's TCR. In a further embodiment, a patient can be selected to receive T cell therapy only if at least a threshold amount of TCRs for that patient are identified as neoantigen-specific for a particular HLA-restricted class of neoantigen peptide. For example, a patient can be selected to receive T cell therapy only if at least one TCR for that patient is identified as a neoantigen-specific HLA class I-restricted neoantigen peptide.
[0360] In yet further embodiments, a patient can be selected for T cell therapy only if at least a threshold amount of a particular HLA-restricted class of neoantigen peptide is recognized by the patient's TCR. For example, a patient can be selected for T cell therapy only if at least one HLA class I-restricted neoantigen peptide is recognized by the patient's TCR. In another example, a patient can be selected for T cell therapy only if at least two HLA class II-restricted neoantigen peptides are recognized by the patient's TCR. Any combination of the above criteria can also be used to select patients for T cell therapy after in vitro identification of T cells specific for the patient's predicted therapeutic subset of v neoantigen candidates.
[0361] XI. Example 7: Experimental Results Demonstrating Exemplary Patient Selection Performance The validity of the patient selection described in Section X is validated by selecting patients from a set of simulated patients, each associated with a test set of simulated neoantigen candidates, for which a subset of the simulated neoantigens is known to be represented in the mass spectrometry data. Specifically, each simulated neoantigen candidate in the test set is associated with a label indicating whether that neoantigen is represented in the mass spectrometry data set of the multi-allelic JY cell line HLA-A*02:01 and HLA-B*07:02 from the Bassani-Sternberg dataset (dataset "D1") (data available at www.ebi.ac.uk / pride / archive / projects / PXD0000394). As described in more detail below in conjunction with Figure 13A, a number of neoantigen candidates for the simulated patients are sampled from the human proteome based on known frequency distributions of mutation burden in non-small cell lung cancer (NSCLC) patients.
[0362] Allele-specific presentation models for the same HLA alleles are trained using a training set that is a subset of the mass spectrometry data for the single alleles HLA-A*02:01 and HLA-B*07:02 from the IEDB dataset (dataset "D2") (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Specifically, the presentation model for each allele is trained using the network dependency function g h (·) and g wThe allele-specific models were modeled as shown in equation (8), incorporating the allele-specific expression (·) and exponent function f(·). The presentation model for the HLA-A*02:01 allele generates the presentation likelihood of a particular peptide being presented on the HLA-A*02:01 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables. The presentation model for the HLA-B*07:02 allele generates the presentation likelihood of a particular peptide being presented on the HLA-B*07:02 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables.
[0363] As disclosed in the following examples with reference to Figures 13A-13E, different models, such as a presentation model trained for peptide binding prediction and a prior art model, are applied to a test set of neoantigen candidates for each simulated patient to identify different therapeutic subsets for the patient based on the predictions. Patients who meet selection criteria for vaccine treatment are selected and associated with personalized vaccines containing epitopes in the patient's therapeutic subset. The size of the therapeutic subsets varies depending on different vaccine doses. No overlap is introduced between the training set used to train the presentation model and the test set of simulated neoantigen candidates.
[0364] In the following example, the proportion of selected patients with at least a certain number of presented neoantigens among the epitopes included in the vaccine is analyzed. This statistic indicates the effectiveness of the simulated vaccine in delivering potential neoantigens that will elicit an immune response in patients. Specifically, simulated neoantigens in a test set are presented if the neoantigen is presented in mass spectrometry dataset D2. A high proportion of patients with presented neoantigens indicates the likelihood of successful treatment with the neoantigen vaccine by inducing an immune response.
[0365] XI.A. Example 7A: Frequency Distribution of Mutation Burden in NSCLC Cancer Patients Figure 13A shows the sample frequency distribution of mutation burden in NSCLC patients. Mutation burden and mutations in different tumor types, including NSCLC, can be found, for example, in the Cancer Genome Atlas (TCGA) (https: / / cancergenome.nih.gov). The X-axis represents the number of nonsynonymous mutations for each patient, and the Y-axis represents the proportion of sample patients with a specific number of nonsynonymous mutations. The sample frequency distribution in Figure 13A shows a range of 3 to 1786 mutations, with 30% of patients having fewer than 100 mutations. Although not shown in Figure 13A, studies have shown that mutation burden is higher in smokers compared to nonsmokers, and that mutation burden can be a strong indicator of neoantigen burden in patients.
[0366] As introduced at the beginning of Section XI above, each of the simulated patient populations is associated with a test set of neoantigen candidates. Each patient's test set is determined by selecting the mutation load m from the frequency distribution shown in Figure 13A for each patient. i The D1 dataset is generated by sampling the D1 sequence. For each mutation, a 21-mer peptide sequence from the human proteome is randomly selected to represent the mutant sequence to be simulated. A test set of candidate neoantigen sequences is generated for patient i by identifying each (8, 9, 10, 11)-mer peptide sequence across the mutations in the 21-mer. Each candidate neoantigen is associated with a label indicating whether the candidate neoantigen sequence is present in the mass spectrometry D1 dataset. For example, candidate neoantigen sequences present in dataset D1 can be associated with the label "1," and sequences not present in dataset D1 can be associated with the label "0." As described in more detail below, Figures 13B-13E show experimental results of patient selection based on the neoantigens presented by patients in the test set.
[0367] XI.B. Example 7B: Proportion of Selected Patients with Neoantigen Presentation Based on Selection Criteria for Mutational Burden Figure 13B shows the number of presented neoantigens in the simulated vaccine for patients selected based on the selection criteria of whether the patient met a minimum mutational load. The proportion of selected patients who had at least a certain number of presented neoantigens in the corresponding study is identified.
[0368] In Figure 13B, the x-axis shows the proportion of patients excluded from vaccine treatment based on tumor mutation burden, as indicated by the label "Minimum Number of Mutations." For example, the data point at "Minimum Number of Mutations" 200 indicates that the patient selection module 324 selected only a subset of simulated patients with a mutation burden of at least 200 mutations. As another example, the data point at "Minimum Number of Mutations" 300 indicates that the patient selection module 324 selected a lower proportion of simulated patients with at least 300 mutations. The y-axis shows the proportion of selected patients associated with at least a certain number of presented neoantigens in the test set without vaccine dose v. Specifically, the top plot shows the proportion of selected patients presenting at least one neoantigen, the middle plot shows the proportion of selected patients presenting at least two antigens, and the bottom plot shows the proportion of selected patients presenting at least three antigens.
[0369] As shown in Figure 13B, the proportion of patients with presented neoantigens significantly increased with increasing mutational burden, indicating that mutational burden as a selection criterion can be effective in selecting patients in whom neoantigen vaccines are likely to induce an effective immune response.
[0370] XI.C. Example 7C: Comparison of Neoantigen Presentation in Vaccines Identified by Presentation Models vs. Prior Art Models Figure 13C compares the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on the presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model. The plot on the left assumes a limiting vaccine volume of v=10, and the plot on the right assumes a limiting vaccine volume of v=20. Patients are selected based on a utility score indicating the expected number of presented neoantigens.
[0371] In Figure 13C, the solid lines indicate patients associated with a vaccine containing therapeutic subsets identified based on presentation models for the alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient was identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted lines indicate patients associated with a vaccine containing therapeutic subsets identified based on the prior art model NETMHCpan for the single allele HLA-A*02:01. Implementation details for NETMHCpan are provided at http: / / www.cbs.dtu.dk / services / NetMHCpan. The therapeutic subset for each patient was identified by applying the NETMHCpan model to the sequences in the test set and identifying the v neoantigen candidates with the highest estimated binding affinity. The x-axis of both graphs indicates the proportion of patients excluded from vaccine treatment based on the expected utility score, which indicates the expected number of presented neoantigens in the therapeutic subsets identified based on the presentation model. The expected utility score is determined as described in Section X in connection with Equation (25). The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens) contained in the vaccine.
[0372] As shown in Figure 13C, a significantly higher proportion of patients associated with vaccines containing therapeutic subsets based on the presentation model receive vaccines containing presented neoantigens than patients associated with vaccines containing therapeutic subsets based on the prior art model. For example, as shown in the graph on the right, 80% of selected patients associated with vaccines based on the presentation model receive at least one presented neoantigen in the vaccine, compared to only 40% of selected patients associated with vaccines based on the prior art model. These results demonstrate that the presentation model described herein is effective in selecting neoantigen candidates for vaccines that are likely to elicit an immune response to treat tumors.
[0373] XI.D. Example 7D: Effect of HLA Coverage on Neoantigen Presentation of Vaccines Identified by Presentation Models Figure 13D compares the number of neoantigens presented in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a single allele-per-presentation model for HLA-A*02:01 and selected patients associated with vaccines containing therapeutic subsets identified based on both HLA-A*02:01 and HLA-B*07:02 allele-per-presentation models. Vaccine volume is set to v = 20 epitopes. For each experiment, patients are selected based on expected utility scores determined based on different therapeutic subsets.
[0374] In Figure 13D, the solid line indicates patients associated with a vaccine containing therapeutic subsets based on both presentation models for HLA alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient was identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with a vaccine containing therapeutic subsets based on a single presentation model for HLA allele HLA-A*02:01. The therapeutic subset for each patient was identified by applying a presentation model for only a single HLA allele to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. In the solid plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility scores for the therapeutic subsets identified by both presentation models. In the dotted plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility scores for the therapeutic subsets identified by a single presentation model. The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens).
[0375] As shown in Figure 13D, patients associated with vaccines containing therapeutic subsets identified by presentation models for both HLA alleles presented neoantigens at a significantly higher rate than patients associated with vaccines containing therapeutic subsets identified by a single presentation model. These results demonstrate the importance of establishing presentation models with high HLA allele coverage.
[0376] XI.E. Example 7E: Comparison of Neoantigen Presentation in Patients Selected by Mutational Burden Versus Expected Number of Presented Neoantigens Figure 13E compares the number of presented neoantigens in simulated vaccines between patients selected based on mutation burden and patients selected by expected utility score, which is determined based on the therapeutic subset identified by the presentation model with a size of v = 20 epitopes.
[0377] In Figure 13E, the solid lines indicate patients selected based on the expected utility score associated with a vaccine containing the therapeutic subset identified by the proposed model. The therapeutic subset for each patient is identified by applying each of the proposed models to the sequences in the test set and identifying the v = 20 neoantigen candidates with the highest likelihood of presentation. The therapeutic utility score is determined based on the likelihood of presentation of the therapeutic subset identified in Section X based on Equation (25). The dotted lines indicate patients selected based on the mutational load associated with a vaccine containing the therapeutic subset identified by the proposed model. The x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score in the solid plot and the proportion of patients excluded based on the mutational load in the dotted plot. The y-axis indicates the proportion of selected patients who receive a vaccine containing at least a specified number of presented neoantigens (1, 2, or 3 neoantigens).
[0378] As shown in Figure 13E, patients selected based on expected utility score receive a vaccine containing a higher proportion of presented neoantigens than patients selected based on mutation load. However, patients selected based on mutation load receive a vaccine containing a higher proportion of presented neoantigens than patients not selected. Thus, while mutation load is an effective patient selection criterion for delayed delivery of an effective neoantigen vaccine, expected utility score is more effective.
[0379] XII. Example 8: Evaluation of Mass Spectrometry Training Models on Held-Out Mass Spectrometry Data HLA peptide presentation by tumor cells is a key requirement for antitumor immunity 91,96,97 We used a large (N = 74 patients) integrated dataset of human tumor and normal tissue samples, including paired class I HLA peptide sequences, HLA typing, and transcriptome RNA-seq (Methods), to compare these and published data to predict antigen presentation in human cancers. 92,98,99 Train a novel deep learning model using100 The purpose of this study was to generate peptides for immunotherapy development. Samples were selected based on tissue availability across several tumor types of interest for immunotherapy development. Mass spectrometry identified an average of 3704 peptides per sample (range 344–11301) with a peptide-level FDR of less than 0.1. These peptides ranged in length from 8–15 amino acids and followed a characteristic class I HLA length distribution with a modal length of 9 (56% of peptides). Consistent with previous reports, the majority of peptides (median 79%) were predicted by MHC flurry to bind to at least one patient HLA allele at a standard affinity threshold of 500 nM. 90 However, considerable variability was observed between samples (e.g., in one sample, 33% of the peptides had predicted affinities above 500 nM). 101 A "strong binder" threshold of 50 nM captured a median of only 42% of presented peptides. Transcriptome sequencing yielded an average of 131M unique reads per sample, with 68% of genes expressed at levels of at least 1 TPM (transcript per million) in at least one sample, highlighting the value of large, diverse sample sets designed to observe expression of the maximum number of genes. HLA-mediated peptide presentation was strongly correlated with mRNA expression. Significant and reproducible differences in peptide presentation rates were observed between genes, greater than those explained by differences in RNA expression or sequence alone. The observed HLA types were consistent with expectations for samples derived from a patient population of primarily European ancestry.
[0380] These and published HLA peptide data 92,98,99We used the method to train a neural network (NN) model to predict HLA antigen presentation. To learn an allele-specific model from tumor mass spectrometry data, in which each peptide could be presented by one of six HLA alleles, we developed a novel network architecture (Methods) capable of jointly learning allele-peptide mapping and allele-specific presentation motifs. For each patient, data points designated as positive were peptides detected by mass spectrometry, and data points designated as negative were peptides from a reference proteome (SwissProt) that were not detected by mass spectrometry in that sample. The data were divided into training, validation, and test sets (Methods). The training set consisted of 142,844 HLA-presented peptides (FDR < 0.02) from 101 samples (69 newly described in this study and 32 previously published). The validation set (used for early termination) consisted of 18,004 presented peptides from the same 101 samples. Two mass spectrometry datasets were used for testing: (1) a tumor sample test set consisting of 571 presented peptides from five additional tumor samples (two lung, two colon, and one ovarian) that were omitted from the training data, and (2) a monoallelic cell line test set consisting of 2,128 presented peptides from genomic location windows (blocks) adjacent to (but distinct from) the locations of the monoallelic peptides included in the training data (see Methods for further details regarding the training / test split).
[0381] The training data identified a predictive model for 53 HLA alleles. 92,104Unlike previous studies, these models captured the dependence of HLA presentation on each sequence position in peptides of multiple lengths. The model properly learned a critical dependence on gene RNA expression and gene-specific presentation propensity, and combining mRNA abundance and learned presentation propensity per gene independently yielded up to a 60-fold difference in presentation rate between the lowest expressed, least likely presented genes and the highest expressed, most likely presented genes. The model also predicted the measured stability of HLA / peptide complexes in the IEDB, even after controlling for predicted binding affinity. 88 (p<1e-10 for 10 alleles) and was further observed to be predictive (p<0.05 for 8 of 10 alleles tested). Taken together, these properties form the basis for improved prediction of immunogenic HLA class I peptides.
[0382] We evaluated the performance of this NN model as a predictor of HLA presentation on a left-out mass spectrometry test set and compared it with the state-of-the-art binding affinity predictor, MHCFlurry, a neural network tool trained on in vitro HLA binding data. 90 Based on previous reports highlighting the importance of mRNA levels in HLA presentation, we compared the results with those from the previous version (Methods). 81,92,103 Increasing thresholds for gene expression assayed by .times. ...
[0383] Figures 14A-D compare the predictive performance of the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three gene expression thresholds. Both the "Full MS Model" and the "Peptide MS Model" are neural network models trained on the mass spectrometry data described above. However, while the "Full MS Model" is trained and tested based on all characteristics of the sample, the "Peptide MS Model" is trained and tested based only on the sample's HLA type and peptide sequence. Three different versions of the MHC Flurry 1.2.0 binding affinity model are tested: the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 0, the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 1, and the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 2. Because the "peptide MS model" and the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM>1 are both trained and tested based solely on the HLA type and peptide sequence of the sample, and both have the same RNA expression threshold, a direct comparison of the performance of these two models quantifies the improvement in prediction due to differences in peptide motifs learned from the mass spectrometry data versus the binding affinity training data.
[0384] Referring first to Figure 14A, Figure 14A compares the positive predictive value (PPV) at 40% recall for the "full MS model," the "peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model on a test set consisting of five different test samples, each with an excluded tumor sample having a ratio of presented to unpresented peptides of 1:2500 (Methods). Figure 14A also shows the average PPV at 40% recall for the "full MS model," the "peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 for the five test samples. As shown in Figure 14A, the average PPV at 40% recall was 0.54 for the "full MS model," and 0.076, 0.072, and 0.061 for the MHC Flurry 1.2.0 binding affinity model with gene expression thresholds of TPM > 2, 1, and 0, respectively. All comparisons between the "full MS model" and the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 0 are statistically significant at p < 1e-6.
[0385] Referring now to Figure 14B, Figure 14B compares the PPV at 40% recall for the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model on a test set consisting of 15 different test samples, each test sample containing an omitted peptide from the monoallelic cell line test dataset with a ratio of presented to unpresented peptides of 1:10,000 (Methods). Figure 14B also shows the average PPV at 40% recall for the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 for the 15 test samples. As shown in Figure 14B, the average PPV at 40% recall was 0.37 for the "Full MS model," and 0.094, 0.090, and 0.071 for the MHC Flurry 1.2.0 binding affinity model with gene expression thresholds of TPM > 2, 1, and 0, respectively. All comparisons between the "Full MS model" and the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 0 are statistically significant at p < 1e-6, except for the test sample containing HLA-A*01:01, where p = 1.6e-4.
[0386] Figure 16 compares the positive predictive value (PPV) at 40% recall for the "full MS model" and the "anchor residues-only MS model" when each model is tested on the test set described above with respect to Figure 14A (Methods). Figure 16 also shows the average PPV at 40% recall for the "full MS model" and the "anchor residues-only MS model" for the five test samples. Like the "full MS model," the "anchor residues-only MS model" is a neural network model trained using the mass spectrometry analysis described above. However, instead of training and testing the "anchor residues-only MS model" based on the entire peptide sequence in the sample, the "anchor residues-only MS model" is trained and tested based only on the "anchor" residues (the first, second, and last residues) of the sample's peptide sequence. Therefore, the results shown in Figure 16 quantify the relative importance of anchor and non-anchor residues to the model's predictive performance. As shown in Figure 16, the performance of the "anchor residues-only MS model" is significantly lower than that of the "full MS model." The average PPV at 40% recall for the anchor residue-only MS model is 0.13 compared to 0.50 for the full MS model. Therefore, it can be inferred that training and testing models that include non-anchor sequences of the peptide sequence will lead to improved model predictions.
[0387] Figure 17A shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test sample 0 from Figure 14A (Methods). As shown in Figure 17A, the "Full MS model" and the "Peptide MS model" achieve better performance compared to the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2.
[0388] Figure 17B compares the PPV at 40% recall for the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model on a test set consisting of 15 different test samples, each test sample containing an omitted peptide from the monoallelic cell line test dataset with a ratio of presented to unpresented peptides of 1:5,000 (Methods). Figure 17B also shows the average PPV at 40% recall for the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 for the 15 test samples. By comparing the results of Figure 14B (each test sample contains an omitted peptide from the single-allelic cell line test dataset with a ratio of presented to unpresented peptides of 1:10,000) with the results of Figure 17A (each test sample contains an omitted peptide from the single-allelic cell line test dataset with a ratio of presented to unpresented peptides of 1:5,000), it can be inferred that the abundance of peptide presentation is strongly correlated with absolute PPV. In general, the lower the abundance of the predicted event (e.g., presentation), the more difficult it is to obtain a high PPV prediction. Therefore, by decreasing (increasing) the abundance in the test data, the absolute PPV of all models decreases (increases). However, the relative difference between the PPVs of different models is not affected by changes in the expected test set abundance.
[0389] Figure 17C-G shows the full precision-recall curves for the "Full MS model," the "Peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when each model is tested on test samples 0-4 from Figure 14A (Methods).
[0390] Figures 17H-V show the full precision-recall curves for the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 when testing each model on the 15 different test samples from Figure 14B, where each test sample contains an excluded peptide from the monoallelic cell line test dataset with a ratio of presented to unpresented peptides of 1:10,000 (Methods).
[0391] FIG. 18 shows different versions of the MS model and initial approaches to model HLA-presented peptides in human tumors when each model is tested on the five different test samples in FIG. 14A. 104 The positive predictive value (PPV) at 40% recall was compared for each model (Methods). Figure 18 also shows the average PPV at 40% recall for each model for the five test samples. The models tested in Figure 18 include the "Full MS Model," "MS Model without Flanking Sequences," "MS Model without Flanking Sequences or Per-Gene Coefficients," "Peptide-Only MS Model, All Lengths Trained Together," "Peptide-Only MS Model, All Lengths Trained Separately," "Linear Peptide-Only MS Model," "MixMHCPred 1.1" model, and "Binding Affinity" model. The "Full MS Model," "MS Model without Flanking Sequences," "MS Model without Flanking Sequences or Per-Gene Coefficients," "Peptide-Only MS Model, All Lengths Trained Together," "Peptide-Only MS Model, All Lengths Trained Separately," and "Linear Peptide-Only MS Model" are all neural network models trained with mass spectrometry as described above. However, each model is trained and tested using different sample characteristics. The "MixMHCPred 1.1" model and the "binding affinity" model are early approaches to model HLA-presented peptides. 104 .
[0392] Overall, the NN model achieved significantly improved prediction of HLA peptide presentation, with PPVs up to 9-fold higher than standard binding affinity + gene expression in the tumor test set (Figure 14A) and up to 5-fold higher in the single-allele dataset (Figure 14B). The large PPV advantage of the MS-based NN model persisted across different recall rates (Figure 17A) and was statistically significant (p<10 in all tumor samples in Figures 14A and 14B and in the single-allele samples except for HLA-A*01:01, where p=1.6e-4). -6 The positive predictive value of standard binding affinity plus gene expression for HLA peptide presentation reached a low value of 6%, consistent with previous estimates. 87,93 However, it should be noted that since only a small percentage of peptides are detected as presented (e.g., approximately 1 in 2500 in the tumor MS test dataset), this approximately 6% PPV represents a greater than 100-fold improvement over baseline abundance.
[0393] By comparing a reduced model trained on mass spectrometry data using only HLA type and peptide sequence as input ("Peptide MS Model," Figure 14A-B, see Methods) with the full MS model, it was determined that approximately 30% of the increase in PPV for binding affinity prediction was due to modeling of extrinsic properties of peptides (RNA abundance, flanking sequences, per-gene coefficients) that can be captured by mass spectrometry but not by binding affinity assays (Figure 14A-B, see also Figures 17A and 18). The remaining 70% of the increase was due to improved modeling of peptide sequence (Figure 14A-B). This modeling is consistent with earlier approaches in modeling HLA-presented peptides in human tumors. 104 The improved performance was due to the overall model architecture, not just the nature of the training dataset (HLA-presented peptides), as shown in Figure 18. This new model architecture allows for a more accurate prediction of binding affinities or hard clustering approaches. 104~106Using this approach, we were able to train allele-specific models through an end-to-end training process that does not require prior assignment of peptides to potential presenting alleles. Importantly, the new model architecture also considers each peptide length separately without imposing accuracy-degrading constraints on the allele-specific submodels as prerequisites for deconvolution, such as linearity. 104 The full model outperformed several simplified models and previously published approaches that impose these constraints (Figure 18).
[0394] Figure 18 shows the performance of several simplified models on the MS test set. The relative importance of the modeling refinements incorporated into the full model is quantified by removing the modeling refinements one at a time and testing their predictive performance on the MS test set. Additionally, a comparison was made between the proposed model disclosed herein and a recently published approach (MixMHCPred) for modeling peptides eluting from mass spectrometry analyses. Because MixMHCPred does not currently model peptides of lengths other than 9 and 10, only 9-mers and 10-mers were used in the comparison. The models are (from left to right): "Full MS model" (the full NN model described in Methods); "MS model, no flanking sequences" (same as the full NN model except without the flanking sequence features); "MS model, no flanking sequences or per-gene coefficients" (same as the full NN model except without the flanking sequence features and per-gene coefficients); "Peptide-only MS model, all lengths trained together" (same as the full NN model except that the only features used were peptide sequence and HLA type); "Peptide-only MS model, each length trained separately" (the model structure is the same as the peptide-only MS model except that in this model, separate models were trained for 9-mers and 10-mers); "Linear peptide-only MS model (with ensemble)" (same as the peptide-only MS model except that instead of using a neural network to model the peptide sequence, an ensemble of linear models was used as in the full model and trained using the same optimization procedure as described in Methods); and "MixMHCPred" "1.1" is MixMHCPred with default settings; "Binding Affinity" is MHCflurry 1.2.0, the same as in the main text. The last five models ("Peptide-only MS model, train all lengths together" through "Binding Affinity") have the same input of only peptide sequence and HLA type. Specifically, none of the last five models use RNA abundance to make predictions.The best-performing peptide-only model ("peptide-only MS model, all lengths trained together") gave an average PPV of 0.41 at 40% recall, whereas the poorest-performing peptide-only model trained on mass spectrometry data ("linear peptide-only MS model with ensembling") had an average PPV of only 28% (very slightly higher than the 18% average PPV of MixMHCpred), highlighting the value of improved NN modeling of peptide sequences. Note that MixMHCpred is trained on different data than the linear peptide-only MS model, but has many of the same modeling properties (e.g., it is a linear model in which each peptide length is trained separately).
[0395] XIII. Example 9: Retrospective Neoantigen T Cell Data Applicator Model Evaluation We evaluated whether such accurate prediction of HLA peptide presentation could lead to the ability to identify epitopes (i.e., targets for immunotherapy) for human tumor CD8+ T cells. A suitable test dataset for this evaluation would include peptides recognized by T cells and presented by HLA on the surface of tumor cells. Furthermore, formal performance evaluation would require not only peptides that were positively labeled (i.e., recognized by T cells) but also a sufficient number of negatively labeled peptides (i.e., tested but not recognized). Mass spectrometry datasets correspond to tumor presentation but not to T cell recognition; conversely, post-vaccination priming or T cell assays correspond to the presence of T cell precursors and T cell recognition but not to tumor presentation. For example, strongly HLA-binding peptides whose source genes are expressed at low levels in tumors may generate strong post-immunization CD8+ T cell responses without therapeutic utility because such peptides are not presented by tumors.
[0396] To obtain a suitable dataset, published CD8 T cell epitopes were collected from four recent studies that met the necessary criteria: Study A 96investigated TILs in nine patients with gastrointestinal tumors and reported T cell recognition of 12 of 1,053 somatic SNV mutations tested by IFN-γ ELISPOT using a tandem minigene (TMG) method in autologous dendritic cells (DCs). Study B 107 also reported T cell recognition of 6 of 574 SNVs by CD8+PD-1+ circulating lymphocytes from 4 melanoma patients using TMG. Study C 97 evaluated TILs from three melanoma patients using pulsed peptide stimulation and found responses to five of 381 SNV mutations tested. 108 reported the recognition of 2 of 62 SNVs using a combination of TMG assays to evaluate TILs from one breast cancer patient pulsed with minimal epitope peptides. The combined dataset consisted of 2,009 assayed SNVs from 17 patients, including 26 neoantigens indicative of pre-existing T cell responses. Importantly, because the dataset consisted largely of neoantigen recognition by tumor-infiltrating lymphocytes, effective predictions were based on previous literature data. 81,82,97 This suggests the ability to identify not only neoantigens that can prime T cells, but more specifically neoantigens that are presented to T cells by tumors, as described in.
[0397] To simulate antigen selection for personalized immunotherapy, somatic mutations were ranked in order of probability of presentation using the "full MS model," the "peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds: TPM > 0, 1, and 2. Because antigen-specific immunotherapy is technically constrained in the number of specificities targeted (e.g., current personalized vaccines encode approximately 10–20 somatic mutations), 80~82), prediction methods were compared by counting the number of pre-existing T cell responses in the top 5, 10, or 20 ranked somatic mutations for each patient with at least one pre-existing T cell response. These results are shown in Figure 14C. Specifically, Figure 14C compares the proportion of somatic mutations recognized by T cells (e.g., pre-existing T cell responses) for the top 5, 10, and 20 ranked somatic mutations identified by the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds of TPM > 0, 1, and 2 for a test set consisting of 12 different test samples, each of which was collected from a patient with at least one pre-existing T cell response. All comparisons between the "Full MS Model" and the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 0 are statistically significant at p < 0.005, except for the top 5 ranked somatic mutations, where p = 0.056.
[0398] As expected, binding affinity predictions included only a small fraction of pre-existing T cell responses among prioritized mutations, e.g., only 9 of 26 (35%) among the top 20-ranked mutations with TPM > 0 (Supplementary Table 1). In contrast, the majority of pre-existing T cell responses (19 / 26, 73%) were ranked in the top 20 by the full MS model, and the average held across different rank and gene expression thresholds (Figure 14C, Supplementary Table 1). At the patient level, the full MS model showed an average of 1.54 pre-existing neoantigen T cell responses among the top 20 predicted mutations in the 13 patients with at least one pre-existing T cell response, compared with only 0.69 (p = 1.4e-4) for binding affinities with TPM > 0.
[0399] We then evaluated mutations at the minimal neo-epitope level (i.e., 8- to 11-mers overlapping with the mutations were recognized) as potentially useful for identifying T cell / TCR targets for cell therapy. In other words, we ranked minimal neo-epitopes in order of their probability of presentation using the "full MS model," the "peptide MS model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds: TPM > 0, 1, and 2. As mentioned above, because antigen-specific immunotherapy is technically limited in the number of targeted specificities, we compared prediction methods by counting the number of pre-existing T cell responses in the top 5, 10, or 20 ranked minimal neo-epitopes for each patient who showed at least one pre-existing T cell response. Epitopes shown as positive are those confirmed as minimally immunogenic epitopes by peptide-based assays (instead of or in addition to TMG-based assays), while negative examples are all epitopes not recognized by the peptide-based assays and epitopes spanning all mutations contained in the minigene that are not recognized. These results are shown in Figure 14D.
[0400] Specifically, Figure 14D compares the percentage of minimal nascent epitopes recognized by T cells (e.g., pre-existing T cell responses) for the top 5, 10, and 20 ranked minimal nascent epitopes identified by the "Full MS Model," the "Peptide MS Model," and the MHC Flurry 1.2.0 binding affinity model with three different gene expression thresholds: TPM > 0, 1, and 2, for a test set consisting of 12 different test samples, each from a patient exhibiting at least one pre-existing T cell response. All comparisons between the "Full MS Model" and the MHC Flurry 1.2.0 binding affinity model with a gene expression threshold of TPM > 0 are statistically significant at p < 0.05, except for the top 5 ranked minimal nascent epitopes, where p = 0.082. In all panels, error bars represent 90% confidence intervals.
[0401] As shown in Figure 14D, the advantage of the NN model for binding affinity at TPM > 0 was even more pronounced than in Figure 14C, with at least four times more emerging epitopes included in the top-ranked minimal epitopes. It is noteworthy that this comparison is biased in favor of binding affinity predictions because only peptides with strong binding affinity were individually tested in Studies A, B, and D. It is possible that there are T cell-recognized peptides with weak predicted HLA binding affinity that were not assayed in these studies but would have been selected according to this model. Such peptides were observed in this study and are discussed in detail below with respect to Figure 15A and Supplementary Table 3.
[0402] Known limitations of mass spectrometry in detecting cysteine-containing peptides 92,104 Nevertheless, the NN model outperformed in binding affinity prediction for cysteine-containing T cell-recognized epitopes, ranking three of the seven cysteine-containing epitopes in the top five (43%) compared to one of the seven in the top five for binding affinity with a gene expression threshold of TPM > 0. As with the mass spectrometry test set, additional features that can be modeled based on the mass spectrometry training data (RNA, flanking sequences, per-gene coefficients) contributed significantly to the increased predictive performance. However, as with the mass spectrometry test data, the predictive performance of the peptide-only MS model improved significantly for binding affinity prediction, with most of this improvement attributed to improved modeling of the peptide sequence (Figure 14C-D, compare light blue and green bars).
[0403] Notably, this improvement was observed despite the potential increase in false negatives for neo-epitopes in the test set (i.e., cases where T cell responses were not detected at neo-epitopes presented by tumors that can be recognized by T cells) due to limitations of current TIL assays. These limitations include (a) an immunosuppressive tumor microenvironment and inefficient T cell priming, (b) depletion of neo-epitope-responsive T cells, (c) production of cytokines other than IFNg by TILs, and (d) heterogeneity of the tumor fractions used. Therefore, the absolute predictive performance of the top 5–20 immunogenic peptide counts described herein may be disappointing for other situations, such as the administration of potent neoantigen cancer vaccines.
[0404] XIII.A. Data The present inventors have reviewed the results of Gros et al. 84 , Tran et al. 140 , Stronen et al. 141 Variant calling, HLA typing, and T cell recognition data were obtained from the supplementary information in [1] and [2]. Patient-specific RNA-seq data were not available. Inferring that tumor RNA expression correlates with the same tumor type in different patients, we substituted RNA-seq data from tumor-type-matched patients obtained from TCGA, which was used in both neural network prediction and RNA expression filtering with TPM > 1 prior to binding affinity prediction. Adding tumor-type-matched RNA-seq data improved prediction performance (Figure 14C-D).
[0405] For mutation-level analyses (Figure 14C), data points designated as positive in Gros et al., Tran et al., and Zacharakis et al. were mutations recognized by patient T cells in both the TMG assay and the minimal epitope peptide-pulsed assay. Data points designated as negative were all other mutations tested in the TMG assay. In Stronen et al., mutations designated as positive were mutations spanning at least one recognized peptide, and negative data points were all mutations tested but not recognized in the tetramer assay. Because the mutated 25-mer TMG assay tests T cell recognition of all peptides spanning the mutation, for the Gros, Tran, and Zacharakis data, mutations were ranked by summing the probability of presentation across all peptides spanning the mutation or by taking the lowest binding affinity. For the Stronen data, mutations were ranked by summing the probability of presentation across all peptides spanning the mutation tested in the tetramer assay or by taking the lowest binding affinity. A complete list of mutations and their characteristics is provided in Supplementary Table 1.
[0406] For epitope-level analyses, data points designated as positive were all minimal epitopes recognized by patient T cells in peptide pulse or tetramer assays, and negative data points were all minimal epitopes not recognized by T cells in peptide pulse or tetramer assays, and all peptides spanning mutations from the TMGs tested that were not recognized by patient T cells. In the cases of Gros et al., Tran et al., and Zacharakis et al., minimal epitope peptides spanning mutations recognized in the TMG analysis that were not tested by peptide pulse assays were excluded from the analysis because the T cell recognition status of these peptides was not experimentally investigated.
[0407] XIV. Example 10: Identification of neoantigen-responsive T cells in cancer patients In this example, we demonstrate that improved prediction enables the identification of neoantigens from normal patient samples. To do so, archived FFPE tumor biopsies and 5–30 ml of peripheral blood were analyzed from nine patients with metastatic NSCLC undergoing anti-PD(L)1 therapy (Supplementary Table 2: Patient demographics and treatment information for N = 9 patients examined in Figure 15A–C. Key fields include tumor stage and subtype, anti-PD1 therapy administered, and a summary of NGS results). Tumor whole-exome sequencing, tumor transcriptome sequencing, and matched normal exome sequencing yielded an average of 198 somatic mutations (SNVs and short indels) per patient, of which an average of 118 were expressed (Methods, Supplementary Table 2). We applied the full MS model to prioritize 20 neoantigens per patient for testing against pre-existing anti-tumor T cell responses. To focus our analysis on likely CD8 responses, prioritized peptides were synthesized as 8- to 11-mer minimal epitopes ("Methods"). Peripheral blood mononuclear cells (PBMCs) were then cultured with the synthesized peptides in short-term in vitro stimulation (IVS) cultures to expand neoantigen-responsive T cells (Supplementary Table 3). Two weeks later, the presence of antigen-specific T cells was assessed using IFN-γ ELISpot against the prioritized neoepitopes. Separate experiments were further performed in seven patients for whom sufficient PBMCs were available to fully or partially identify the specific antigens recognized. These results are shown in Figures 15A-C and 19A-22.
[0408] Figure 15A shows the detection of T cell responses to patient-specific neoantigen peptide pools in nine patients. For each patient, predicted neoantigens were combined into two pools of 10 peptides according to model ranking and arbitrary sequence homology (homologous peptides were divided into different pools). Then, for each patient, in vitro expanded PBMCs for that patient were stimulated with the two patient-specific neoantigen peptide pools in an IFN-γ ELISpot. The data in Figure 15A are based on the number of seeded cells 10 minus the background (corresponding DMSO negative control). 5Results are shown as spot-forming units (SFU) per cell. Background measurements (DMSO negative control) are shown in Figure 22. For patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, and CU05, responses of single wells (patients 1-038-001, CU02, CU03, and 1-050-001) or replicates (all other patients) including means and standard deviations are shown to cognate peptide pools #1 and #2. For patients CU02 and CU03, cell numbers allowed testing only against specific peptide pool #1. Samples with fold increase values greater than 2-fold above background were considered positive and are indicated by an asterisk (responders included patients 1-038-001, CU04, 1-024-001, 1-024-002, and CU02). Non-responders included patients 1-050-001, 1-001-002, CU05, and CU03. Figure 15C shows photographs of ELISpot wells containing in vitro-expanded PBMCs from patient CU04 stimulated with the DMSO negative control, PHA positive control, CU04-specific neoantigen peptide pool #1, CU04-specific peptide 1, CU04-specific peptide 6, and CU04-specific peptide 8 in an IFN-γ ELISpot.
[0409] Figures 19A-B show the results of control experiments using patient neoantigens in HLA-matched healthy donors. The results of these experiments demonstrate that the in vitro culture conditions did not allow for de novo priming in vitro, but rather expanded only pre-existing in vivo primed memory T cells.
[0410] Figure 20 shows the detection of T cell responses to the PHA positive control for each donor and in vitro expansion shown in Figure 15A. For each donor and in vitro expansion shown in Figure 15A, in vitro expanded patient PBMCs were stimulated with PHA for maximal T cell activation. The data in Figure 20 are based on 10 seeded cells minus background (corresponding DMSO negative control). 5Results are shown as spot-forming units (SFU) per donor. Responses of single wells or biological replicates are shown for patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, CU05, and CU03. Patient CU02 was not tested with PHA. Cells from patient CU02 were included in the analysis because their positive response to peptide pool #1 (Figure 15A) indicated viable and functional T cells. As shown in Figure 15A, donors who responded to the peptide pools include patients 1-038-001, CU04, 1-024-001, and 1-024-002. Also shown in Figure 15A, donors who did not respond to the peptide pool include patients 1-050-001, 1-001-002, CU05, and CU03.
[0411] Figure 21A shows the detection of T cell responses to each individual patient-specific neoantigen peptide in pool #2 in patient CU04. Figure 21A also shows the detection of T cell responses to the PHA positive control in patient CU04. (This positive control data is also shown in Figure 20.) For patient CU04, the patient's in vitro expanded PBMCs were stimulated in an IFN-γ ELISpot with the patient-specific individual neoantigen peptides from pool #2 for patient CU04. The patient's in vitro expanded PBMCs were also stimulated in an IFN-γ ELISpot with PHA as a positive control. Data are based on seeded cells 10 minus background (corresponding DMSO negative control). 5 Shown as spot-forming units (SFU) per cell.
[0412] Figure 21B shows the detection of T cell responses to individual patient-specific neoantigen peptides at each of three visits for patient CU04 and at each of two visits for patient 1-024-002 (each visit occurring at a different time point). In both patients, the patient's in vitro expanded PBMCs were stimulated with the patient-specific individual neoantigen peptide in an IFN-γ ELISpot. For each patient, data from each visit are calculated based on the number of seeded cells 10 minus background (corresponding DMSO control).5 The data are presented as cumulative (sum) spot-forming units (SFU) per individual. Data for patient CU04 are presented as background-subtracted cumulative SFU over three visits. For patient CU04, background-subtracted SFU are shown for the first visit (T0) and subsequent visits 2 months (T0+2 months) and 14 months (T0+14 months) after the first visit (T0). Data for patient 1-024-002 are presented as background-subtracted cumulative SFU over two visits. For patient 1-024-002, background-subtracted SFU are shown for the first visit (T0) and subsequent visit 1 month (T0+1 month) after the first visit (T0). Samples with a fold increase value greater than 2-fold above background were considered positive and are indicated by an asterisk.
[0413] Figure 21C shows detection of T cell responses to individual patient-specific neoantigen peptides and to patient-specific neoantigen peptide pools at each of two visits for patient CU04 and patient 1-024-002 (each visit occurring at a different time point). For both patients, the patient's in vitro expanded PBMCs were stimulated with the patient-specific individual neoantigen peptides and the patient-specific neoantigen peptide pool in an IFN-γ ELISpot. Specifically, for patient CU04, in vitro expanded PBMCs from patient CU04 were stimulated in IFN-γ ELISpot with CU04-specific individual neoantigen peptides 6 and 8 and the CU04-specific neoantigen peptide pool, and for patient 1-024-002, in vitro expanded PBMCs from patient 1-024-002 were stimulated in IFN-γ ELISpot with 1-024-002-specific individual neoantigen peptide 16 and the 1-024-002-specific neoantigen peptide pool. Data in Figure 21C are for each technical replicate with mean and range, minus background (corresponding DMSO control) for seeded cells 10 5Data are presented as spot-forming units (SFU) per individual. Data for patient CU04 are presented as background-subtracted SFU across two visits. For patient CU04, background-subtracted SFU are presented for the first visit (T0, technical triplicate) and a subsequent visit two months after the first visit (T0) (T0 + 2 months, technical triplicate). Data for patient 1-024-002 are presented as background-subtracted SFU across two visits. For patient 1-024-002, background-subtracted SFU are presented for the first visit (T0, technical triplicate) and a subsequent visit one month after the first visit (T0) (T0 + 1 month, technical duplicate except for the sample stimulated with patient 1-024-002-specific neoantigen peptide pool).
[0414] Figure 22 shows the detection of T cell responses to two patient-specific neoantigen peptide pools and a DMSO negative control for the patients in Figure 15A. For each patient, in vitro expanded PBMCs for that patient were stimulated with two patient-specific neoantigen peptide pools in an IFN-γ ELISpot. For each donor and each in vitro expansion, in vitro expanded patient PBMCs were also stimulated with DMSO as a negative control in an IFN-γ ELISpot. The data in Figure 22 are based on 10 seeded cells, including background (corresponding DMSO negative control), for the patient-specific neoantigen peptide pools and corresponding DMSO control. 5Results are shown as spot-forming units (SFU) per well. For patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, and CU05, responses are shown for single wells (patients 1-038-001, CU02, CU03, and 1-050-001) or averages (all other samples) with standard deviations of biological duplicates to cognate peptide pools #1 and #2. For patients CU02 and CU03, cell numbers allowed testing only against specific peptide pool #1. Samples with fold-increase values >2-fold above background were considered positive and are indicated by an asterisk (responding donors included patients 1-038-001, CU04, 1-024-001, 1-024-002, and CU02). Non-responsive donors include patients 1-050-001, 1-001-002, CU05, and CU03.
[0415] As briefly described above with respect to Figures 19A-B, to confirm that the in vitro culture conditions only expanded pre-existing in vivo primed memory T cells rather than allowing de novo priming in vitro, a series of control experiments using neoantigens in HLA-matched healthy donors were performed. The results of these experiments are shown in Figures 19A-B and Supplementary Table 5. The results of these experiments confirmed that de novo priming and detectable neoantigen-specific T cell responses did not occur in healthy donors using the IVS culture method.
[0416] In contrast, pre-existing neoantigen-reactive T cells were identified in the majority of patients (5 / 9, 56%) tested with patient-specific peptide pools using IFN-γ ELISpot (Figure 15A and Figures 20-22). Of the seven patients whose cell counts allowed complete or partial testing of individual neoantigen cognate peptides, four patients responded to at least one of the tested neoantigen peptides, and all of these patients demonstrated responses to the corresponding pools (Figure 15B). The remaining three patients tested with individual neoantigens (patients 1-001-002, 1-050-001, and CU05) did not demonstrate detectable responses to single peptides (data not shown), confirming the lack of responses seen in these patients to the neoantigen pools (Figure 15A). Of the four responding patients, samples from a single visit were obtained for two responding patients (patients 1-024-001 and 1-038-001), and samples from multiple visits were obtained for the remaining two responding patients (CU04 and 1-024-002). For the two patients with samples from multiple visits, cumulative (summed) spot-forming units (SFU) from three visits (patient CU04) and two visits (patient 1-024-002) are shown in Figure 15B, with a breakdown by visit shown in Figure 21B. Additional PBMC samples from the same visits were also obtained for patients 1-024-002 and CU04, and responses to patient-specific neoantigens were confirmed by repeat IVS culture and ELISpot (Figure 21C).
[0417] Overall, among patients in whom at least one T cell-recognized neoepitopes was identified, as shown by responses to the pool of 10 peptides in Figure 15A, the average number of recognized neoepitopes per patient was at least two (counting the unrecognized pool as one recognized peptide, with a minimum of 10 epitopes identified in five patients). In addition to testing for IFN-γ responses by ELISpot, culture supernatants were also tested for granzyme B by ELISA and for TNF-α, IL-2, and IL-5 by MSD cytokine multiplex assay. Of the five patients with positive ELISpots, cells from four secreted three or more test compounds, including granzyme B (Supplementary Table 4), demonstrating polyfunctionality of neoantigen-specific T cells. Importantly, because the combined prediction and IVS method does not rely on a limited set of available MHC multimers, responses were broadly tested across restricted HLA alleles. Furthermore, this approach directly identifies the minimal epitope, unlike tandem minigene screening, which requires a separate deconvolution step to identify the recognized mutation and identify the minimal epitope. Overall, the yield of neoantigen identification is significantly higher than the previous best method of testing TILs for all mutations using apheresis samples. 96 While the results were comparable to those of the previous method, it only required screening of 20 synthetic peptides using the usual 5-30 mL of whole blood.
[0418] XIV.A. Peptides Custom-made recombinant lyophilized peptides were purchased from JPT Peptide Technologies (Berlin, Germany) or Genscript (Piscataway, NJ, USA), reconstituted at 10–50 mM in sterile DMSO (VWR International, Pittsburgh, PA, USA), and stored in aliquots at −80°C.
[0419] XIV.B.Human peripheral blood mononuclear cells (PBMC) Lyophilized, HLA-typed PBMCs from healthy donors (confirmed to be seronegative for HIV, HCV, and HBV) were purchased from Precision for Medicine (Gladstone, NJ, USA) or Cellular Technology, Ltd. (Cleveland, OH, USA) and stored in liquid nitrogen until use. Fresh blood samples were purchased from Research Blood Components (Boston, MA, USA), and leukopaks were purchased from AllCells (Boston, MA, USA). PBMCs were isolated using a Ficoll-Paque density gradient (GE Healthcare Bio, Marlborough, MA, USA) and then cryopreserved. Patient PBMCs were processed at a local clinical processing center according to local clinical standard operating procedures (SOPs) and IRB-approved protocols. The approving IRBs were the Quorum Review IRB, Comitato Etico Interaziendale AOU San Luigi Gonzaga di Orbassano, and Comite Etico de la Investigacion del Grupo Hospitalario Quiron en Barcelona.
[0420] Briefly, PBMCs were isolated by density gradient centrifugation, washed, counted, and cultured at 5x10 in CryoStor CS10 (STEMCELL Technologies, Vancouver, BC, V6A 1B6, Canada). 6Cells were cryopreserved at 1000 cells / ml. Cryopreserved cells were shipped via cryoport and transferred and stored in LN2 upon arrival. Patient demographics are shown in Supplementary Table 2. Cryopreserved cells were thawed and washed twice in OpTmizer T-cell Expansion Basal Medium (Gibco, Gaithersburg, MD, USA) with Benzonase (EMD Millipore, Billerica, MA, USA) and once without Benzonase. Cell count and viability were assessed using Guava ViaCount reagent and a module on a Guava easyCyte HT cytometer (EMD Millipore). Cells were then resuspended in the appropriate medium at the appropriate concentration for subsequent assays (see next section).
[0421] XIV.C. In vitro stimulation (IVS) culture Pre-existing T cells obtained from healthy donors or patient samples were analyzed using the approach applied by Ott et al. 81 The same approach was used to expand PBMCs in the presence of cognate peptides and IL-2. Briefly, thawed PBMCs were rested overnight and stimulated in 24-well tissue culture plates in the presence of peptide pools (10 μM per peptide, 10 peptides per pool) in ImmunoCult™-XF T-cell Expansion Medium (STEMCELL Technologies) supplemented with 10 IU / ml rhIL-2 (R&D Systems Inc., Minneapolis, MN) for 14 days. Cells were grown at 2x10 6 Cells were seeded at 2 x 10 cells / well and cultured by changing 2 / 3 of the medium every 2-3 days. One patient sample demonstrated protocol deviations and should be considered a potential false negative. Patient CU03 did not yield sufficient numbers of cells after thawing, so cells were cultured at 2 x 10 cells / peptide pool. 5 cells (10-fold fewer than stated in the protocol).
[0422] XIV.D. IFNγ enzyme-linked immunospot (ELISpot) assay ELISpot assay for detection of IFNγ-producing T cells 142 Briefly, PBMCs (ex vivo or after in vitro expansion) were harvested, washed in serum-free RPMI (VWR International), and cultured in ELISpot Multiscreen plates (EMD Millipore) coated with anti-human IFNγ capture antibody (Mabtech, Cincinatti, OH, USA) in OpTmizer T-cell Expansion Basal Medium (ex vivo) or ImmunoCult™-XF T-cell Expansion Medium (expanded cultures) in the presence of control or cognate peptide. After 18 h of incubation in a humidified incubator at 37°C with 5% CO2, cells were removed from the plates, and membrane-bound IFNγ was detected using anti-human IFNγ detection antibody (Mabtech), Vectastain Avidin peroxidase conjugate (Vector Labs, Burlingame, CA, USA), and AEC Substrate (BD Biosciences, San Jose, CA, USA). The ELISpot plates were dried, stored in the dark, and subjected to standardized evaluation. 143 The data were sent to Zellnet Consulting, Inc., Fort Lee, NJ, USA for further analysis. Data are presented as spot-forming units (SFU) per number of cells plated.
[0423] XIV.E. Granzyme B ELISA and MSD Multiplex Assay Detection of secreted IL-2, IL-5, and TNF-α in ELISpot supernatants was performed using a triplex assay, the MSD U-PLEX Biomarker assay (catalog number K15067L-2). The assay was performed according to the manufacturer's instructions. Test article concentrations (pg / ml) were calculated using serial dilutions of known standards for each cytokine. For data graphing, values below the minimum range of the standard curve were set equal to zero. Detection of granzyme B in ELISpot supernatants was performed using the GranzymeB DuoSet® ELISA (R&D Systems, Minneapolis, MN) according to the manufacturer's instructions. Briefly, ELISpot supernatants were diluted 1:4 in sample diluent and run alongside serial dilutions of granzyme B standard to calculate concentrations (pg / ml). For data graphing, values below the minimum range of the standard curve were set equal to zero.
[0424] Negative control experiment of the XIV.F.IVS assay - neoantigens derived from tumor cell lines tested in healthy donors Figure 19A shows a negative control experiment of the IVS assay for neoantigens derived from tumor cell lines tested in healthy donors. PBMCs from healthy donors were stimulated during IVS culture with a peptide pool containing a positive control peptide (pre-exposed to an infection), HLA-matched neoantigens from tumor cell lines (unexposed), and peptides derived from pathogens to which the donor was seronegative. Expanded cells were then stimulated with DMSO (negative control, black circles), PHA and common infectious disease peptides (positive control, red circles), neoantigens (unexposed, light blue circles), or HIV and HCV peptides (donor was seronegative; dark blue, A and B), followed by IFNγ ELISpot (10 5 The data were analyzed by seeding cells 10 5 Shown as spot-forming units (SFU) per donor. Means and biological replicates with SEM are shown. No responses were observed to neoantigens or peptides derived from pathogens to which the donor had not been exposed (seronegative).
[0425] Negative control experimen...
Claims
1. 1. A method for identifying one or more T cells antigen-specific for at least one neoantigen derived from and presented on the surface of one or more tumor cells in a subject, comprising: obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing data from the tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen comprises at least one alteration that makes the peptide sequence different from a corresponding wild-type peptide sequence identified from the normal cells of the subject; encoding the peptide sequence of each of the neoantigens into a corresponding numerical vector, each numerical vector containing information about a set of amino acids constituting the peptide sequence and the positions of the amino acids in the peptide sequence; inputting the numerical vectors into a machine-learned presentation model using a computer processor to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood in the set represents the likelihood that a corresponding neoantigen will be presented on the surface of the tumor cells of the subject by one or more MHC alleles, and wherein the machine-learned presentation model: For each sample in the plurality of samples, a label obtained by mass spectrometry that indicates whether the peptide was presented by at least one MHC allele within the set of MHC alleles identified as being present in the sample; and For each of the samples, a training peptide sequence is coded as a numerical vector containing information about a set of amino acids that make up the peptide and the positions of the amino acids in the peptide. a plurality of parameters identified based at least on a training dataset including: a function representing a relationship between the numerical vector received as an input and the proposed likelihood generated as an output based on the numerical vector and the parameters; the process comprising: selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens; and identifying one or more T cells in said subset that are antigen-specific for at least one of said neoantigens; The method comprising:
2. inputting the numerical vector into the machine-learned presentation model, applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate, for each of the one or more MHC alleles, a dependency score indicating whether the MHC allele will present the neoantigen based on a particular amino acid at a particular position in the peptide sequence. The method of claim 1 , comprising:
3. inputting the numerical vector into the machine-learned presentation model, (A) transforming the dependency scores to generate, for each MHC allele, a corresponding per-allele likelihood that indicates the likelihood that the corresponding MHC allele will present the corresponding neoantigen; and combining the per-allele likelihoods to generate the presentation likelihood of the neoantigen; transforming the dependency score models presentation of the neoantigen as mutually exclusive across the one or more MHC alleles; or (B) transforming the combination of dependency scores to generate the presentation likelihood, wherein transforming the combination of dependency scores models presentation of the neoantigen as interference between the one or more MHC alleles. The method of claim 2 further comprising:
4. The set of representation likelihoods is further specified by at least one or more allele non-interaction characteristics; and the method further comprises applying the machine-learned presentation model to the allele-non-interacting feature to generate a dependency score for the allele-non-interacting feature, the dependency score indicating whether the corresponding neoantigen peptide sequence is presented based on the allele-non-interacting feature; The method is (A) combining the dependency score for each MHC allele in the one or more MHC alleles with the dependency score for the allele-non-interacting trait; transforming the combined dependency scores for each MHC allele to generate a per-allele likelihood for each MHC allele, the likelihood being that the corresponding MHC allele will present the corresponding neoantigen; and combining the per-allele likelihoods to generate the representation likelihood; or (B) combining the dependency score for each of the MHC alleles with the dependency score for the allele-non-interacting trait; and transforming the combined dependency scores to generate the presentation likelihood. The method of claim 2 or claim 3, further comprising:
5. (a) the one or more MHC alleles comprise two or more different MHC alleles; (b) the peptide sequence comprises a peptide sequence having a length other than 9 amino acids; and / or (c) encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme. The method according to any one of claims 1 to 4.
6. The plurality of samples (a) one or more cell lines engineered to express a single MHC allele; (b) one or more cell lines engineered to express multiple MHC alleles; (c) one or more human cell lines obtained or derived from multiple patients; (d) fresh or frozen tumor samples obtained from multiple patients; and (e) fresh or frozen tissue samples obtained from multiple patients The method according to any one of claims 1 to 5, comprising at least one of:
7. the training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of said peptides; and (b) data relating to a measure of peptide-MHC binding stability for at least one of said peptides; The method of any one of claims 1 to 6, further comprising at least one of:
8. (A) the set of presentation likelihoods is further identified by at least the expression level of the one or more MHC alleles in the subject, as measured by RNA-seq or mass spectrometry; (B) the set of presentation likelihoods is (a) the predicted affinity between neoantigens within the set of neoantigens and the one or more MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex. and / or (C) the set of numerical likelihoods is: (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and further characterized by characteristics including at least one of: The method according to any one of claims 1 to 7.
9. selecting the set of selected neoantigens, (A) selecting neoantigens that have an increased likelihood of being presented on the surface of the tumor cells compared to non-selected neoantigens based on the machine-learned presentation model; (B) selecting, based on the machine-learned presentation model, neoantigens that have an increased likelihood of inducing a tumor-specific immune response in the subject compared to unselected neoantigens; (C) selecting neoantigens that have an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) compared to unselected neoantigens based on the presentation model, wherein the APCs are dendritic cells (DCs); (D) selecting neoantigens that have a reduced likelihood of being inhibited by central or peripheral tolerance compared to non-selected neoantigens based on the machine-learned presentation model; and / or (E) selecting, based on the machine-learned presentation model, neoantigens that have a reduced likelihood of inducing an autoimmune response against normal tissue in the subject compared to non-selected neoantigens. The method according to any one of claims 1 to 8, comprising:
10. (A) the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer; and / or (B) the method further comprises generating an output for constructing a personalized cancer vaccine from the selected set of neoantigens, wherein the output for the personalized cancer vaccine comprises at least one peptide sequence or at least one nucleotide sequence encoding the selected set of neoantigens. The method according to any one of claims 1 to 9.
11. the machine-learned presentation model (e.g., a deep learning model including one or more layers of nodes) is a neural network model; (A) the neural network model includes a plurality of network models for MHC alleles, each network model including a series of nodes assigned to a corresponding MHC allele among the plurality of MHC alleles and arranged in one or more layers; or (B) the neural network model includes a plurality of network models for MHC alleles, each network model being assigned to a corresponding MHC allele among the plurality of MHC alleles and including a series of nodes arranged in one or more layers, and the neural network models are trained by updating parameters of the neural network models, and the parameters of at least two network models are updated together for at least one training iteration. The method according to any one of claims 1 to 10.
12. identifying the one or more T cells (A) co-culturing said one or more T cells with one or more of said neoantigens in said subset under conditions that expand said one or more T cells; and / or (B) contacting the one or more T cells with an MHC multimer comprising one or more of the neoantigens in the subset under conditions that allow binding of the T cells to the MHC multimer. The method according to any one of claims 1 to 11, comprising:
13. 13. The method of any one of claims 1 to 12, wherein the method further comprises identifying one or more T cell receptors (TCRs) of the one or more identified T cells, and wherein identifying the one or more T cell receptors comprises sequencing T cell receptor sequences of the one or more identified T cells.
14. (A) identifying the one or more T cells in the subset that are antigen-specific for at least one of the neoantigens using 5 to 30 mL of whole blood from the subject; (B) the subset of neoantigens comprises up to 20 neoantigens, and the one or more specified T cells recognize at least two neoantigens in the subset of neoantigens; and / or (C) the one or more MHC alleles are class I MHC alleles. The method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Neoantigen Identification, Manufacture, and Use
US20170212984A1
Reagents and methods for identifying, enriching, and / or expanding antigen-specific t cells
WO2016044530A1
Cited By
Neoantigen identification for t-cell therapy
JP2023162369A