Identification of neoantigens with MHC class ii model
The use of next-generation sequencing and nonlinear deep learning models addresses the low PPV issue in MHC class II neoantigen prediction, enhancing the identification of tumor antigen-specific T cells for personalized cancer therapies, improving the effectiveness of immunotherapy.
Patent Information
- Application Number
- JP2025066884
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-29
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-05
AI Technical Summary
Current methods for identifying MHC class II-presented neoantigens in cancer therapy suffer from low positive predictive value (PPV), making it difficult to design effective personalized vaccines and T cell therapies, as they rely on impractical large peptide libraries, limited MHC allele availability, and fail to accurately predict peptide presentation due to structural differences in MHC class II molecules.
An optimized approach using next-generation sequencing and nonlinear deep learning models to identify MHC class II allele-presented neoantigens, which includes tumor exome and transcriptome analysis, and models peptide-MHC class II allele interactions to enhance prediction accuracy, reducing the need for impractical screening and improving PPV.
The model significantly enhances the prediction of MHC class II-presented neoantigens, enabling more efficient and cost-effective identification of tumor antigen-specific T cells for personalized cancer therapies, accelerating progress towards patient cures by improving the accuracy of immunotherapy.
Smart Images

Figure 2025114587000001_ABST
Abstract
Description
[Background technology]
[0001] Therapeutic vaccines and T cell therapies based on tumor-specific neoantigens hold great promise as the next generation of personalized cancer immunotherapy. 1~3 Cancers with high mutational burden, such as non-small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies due to their relatively high likelihood of generating neoantigens. 4,5 Early evidence suggests that neoantigen-based vaccination induces T cell responses. 6 , T cell therapy targeting neoantigens can induce tumor regression in selected patients 7 It has been shown that:
[0002] Specifically, the identification of neoantigens presented by MHC class II for use in neoantigen-based vaccination and neoantigen-targeted T cell therapy is a promising therapeutic approach, as up to 50% of neoantigen-reactive TILs contain CD4 cells that respond to neoantigens presented by MHC class II alleles. These CD4 cells have been shown to assist CD8 cells in antitumor responses and, in some cases, directly attack tumor cells. Despite this promise of MHC class II-presented neoantigens in cancer therapy, the positive predictive value (PPV) of MHC class II-presented neoantigens is lower than that of MHC class I-presented neoantigens recognized by CD8 cells.
[0003] These relatively poor presentation predictions for MHC class II-presented neoantigens may be due in part to the structure of MHC class II molecules compared to MHC class I molecules. Specifically, MHC class II molecules tend to have a more open peptide-binding groove compared to MHC class II molecules. This structural difference allows MHC class I molecules to bind peptides of 8–10 amino acids in length, whereas MHC class II molecules bind peptides of a wider range of lengths (Figure 14F). Due to the variability in the length of peptides presented by MHC class II molecules, peptides presented by MHC class II molecules may be more difficult to predict compared to peptides presented by MHC class I molecules.
[0004] Therefore, the identification of MHC class II-presented neoantigens and neoantigen-recognizing T cells is crucial for assessing tumor response. 77,110 , examining tumor evolution 111 , designing the next generation of personalized therapies 112 Current methods for identifying neoantigens are time-consuming and labor-intensive. 84,96 , or the accuracy is not sufficient 87,91-93 Furthermore, neoantigen-recognizing T cells are the main component of TILs. 84,96,113,114 circulating in the peripheral blood of cancer patients. 107 Although it has been shown recently that neoantigen-reactive T cells can be identified using the TIL method, the current methods are as follows: 97,98 or leukapheresis 107 (2) they require screening of impractically large peptide libraries; or (3) they rely on MHC alleles that may be practically available for only a small number of MHC alleles.
[0005] Furthermore, early methods incorporating mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides have been proposed. 8However, these proposed methods involve many steps other than gene expression and MHC binding (e.g., TAP transport, proteasomal cleavage, MHC binding, transport of peptide-MHC complexes to the cell surface, and / or MHC recognition by TCR; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsins), and / or competition with CLIP peptides for HLA binding catalyzed by HLA-DM). 9 The entire epitope generation process cannot be modeled, and therefore existing methods tend to suffer from low positive predictive value (PPV) (Figure 1A).
[0006] Indeed, analyses of peptides presented by tumor cells conducted by several groups have shown that less than 5% of peptides predicted to be presented using gene expression and MHC binding affinity are found on tumor surface MHC. 10,11 (Figure 1B). This low correlation between binding prediction and MHC presentation is further supported by the lack of improved prediction accuracy for binding-restricted neoantigens in response to checkpoint inhibitors relative to the number of mutations alone. 12 These failures in presentation prediction are particularly true in the case of neoantigens presented by MHC class II alleles.
[0007] Such low positive predictive values (PPV) of existing methods for predicting presentation present a challenge in the design of neoantigen-based vaccines and in neoantigen-based T cell therapies. If vaccines are designed using low PPV predictions, most patients are unlikely to receive therapeutic neoantigens, and even fewer patients will receive multiple neoantigens (even assuming all presented peptides are immunogenic). Similarly, if therapeutic T cells are designed based on low PPV predictions, most patients are unlikely to receive T cells with reactivity against tumor neoantigens, and the time and resource costs of identifying predicted neoantigens using downstream testing methods after prediction may be unnecessarily high. Therefore, neoantigen vaccination and T cell therapy using current methods are unlikely to be effective in a significant number of tumor-bearing subjects (Figure 1C).
[0008] Furthermore, previous approaches have only used cis-acting mutations to generate candidate neoantigens, including splicing factor mutations that occur in multiple tumor types and lead to aberrant splicing of many genes. 13 , and additional sources of nascent ORFs, including mutations that create or remove protease cleavage sites, were not considered in most cases.
[0009] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard tumor analysis approaches may erroneously promote sequence artifacts or germline polymorphisms as neoantigens, which can lead to inefficient utilization of vaccine doses or the risk of autoimmunity, respectively. Summary of the Invention
[0010] Disclosed herein are optimized approaches for identifying and selecting neoantigens presented by MHC class II alleles for personalized cancer vaccines, T cell therapy, or both. First, we address optimized tumor exome and transcriptome analysis approaches to identify neoantigen candidates using next-generation sequencing (NGS). These methods build on standard approaches for tumor analysis by NGS to develop the most sensitive and specific neoantigen candidates across all classes of genomic alterations. Second, we provide novel approaches for selecting MHC class II allele-presented neoantigens with high PPV to overcome specificity issues and ensure that MHC class II allele-presented neoantigens developed for vaccine administration and / or as targets for T cell therapy are more likely to elicit anti-tumor immunity. Depending on the embodiment, these approaches include trained statistical regression or nonlinear deep learning MHC class II models that jointly model peptide-MHC class II allele mapping, as well as motifs for each MHC class II allele for peptides of multiple lengths that share statistical power across peptides of different lengths. In particular, nonlinear MHC class II deep learning models can be designed and trained to treat different MHC alleles within the same cell as independent, thus resolving the problem associated with linear models where linear models interfere with each other.Finally, further concerns regarding the design and production of personalized vaccines based on MHC class II allele-presenting neoantigens and the production of personalized MHC class II allele-presenting neoantigen-specific T cells for T cell therapy are resolved.
[0011] The model disclosed herein outperforms state-of-the-art predictive tools trained on binding affinity and earlier predictive tools based on MS peptide data by up to an order of magnitude. By more reliably predicting peptide presentation by MHC class II alleles, the model enables more time- and cost-effective identification of MHC class II allele-presenting neoantigen- or tumor antigen-specific T cells for personalized therapy using a clinically practical process that uses limited amounts of patient peripheral blood, screens fewer peptides per patient, and does not necessarily rely on MHC multimers. However, in another embodiment, the model disclosed herein enables more time- and cost-effective identification of MHC class II allele-presenting tumor antigen-specific T cells using MHC multimers by reducing the number of MHC multimer-bound peptides that need to be screened to identify MHC class II allele-presenting neoantigen- or tumor antigen-specific T cells.
[0012] The predictive performance of the MHC class II model disclosed herein in the task of identifying TIL neoepitope datasets and predicted neoantigen-reactive T cells demonstrates that by modeling MHC class II allele processing and presentation, it is now possible to obtain predictions of therapeutically useful MHC class II allele-presented neoepitopes. In summary, this work will accelerate progress toward patient cures by enabling the identification of actionable in silico MHC class II allele-presented antigens for MHC class II allele-targeted immunotherapy. [The present invention 1001] 1. A method for identifying one or more T cells derived from one or more tumor cells in a subject, the T cells being antigen-specific for at least one neoantigen likely to be presented on the surface of said tumor cells by one or more class II MHC alleles, comprising: obtaining at least one of exome, transcriptome, or whole genome nucleotide sequencing data from the tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing a peptide sequence for each neoantigen of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence for each neoantigen includes at least one alteration that makes the peptide sequence different from a corresponding wild-type peptide sequence identified from the normal cells of the subject; encoding the peptide sequence of each of the neoantigens into a corresponding numerical vector, each numerical vector containing information about a set of amino acids constituting the peptide sequence and the positions of the amino acids in the peptide sequence; using a computer processor to input the numerical vectors into a machine-learned presentation model to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood in the set represents the likelihood that a corresponding neoantigen will be presented on the surface of the tumor cells of the subject by the one or more class II MHC alleles; a label obtained by mass spectrometry to measure, for each sample of the plurality of samples, the presence of a peptide bound to at least one class II MHC allele from the set of class II MHC alleles identified as being present in said sample; and a training peptide sequence, encoded as a numerical vector containing information about a set of amino acids that make up the peptide and the positions of the amino acids in the peptide, for each of the samples; a plurality of parameters identified based at least on a training dataset including: a function representing a relationship between the numerical vector received as an input and the proposed likelihood generated as an output based on the numerical vector and the parameters; including, steps; selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens; identifying one or more T cells in said subset that are antigen-specific for at least one of said neoantigens; and returning the one or more identified T cells The method comprising: [The present invention 1002] The step of inputting the numerical vector into the machine-learned presentation model includes: applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate a dependency score for each of the one or more Class II MHC alleles indicating whether the Class II MHC allele presents the neoantigen based on the particular amino acid at the particular position of the peptide sequence. The method of the present invention 1001, comprising: [The present invention 1003] The step of inputting the numerical vector into the machine-learned presentation model includes: transforming the dependency scores to generate a corresponding allele-specific likelihood for each Class II MHC allele, the likelihood indicating the likelihood that the corresponding Class II MHC allele will present the corresponding neoantigen; and combining the likelihoods for each allele to generate the presentation likelihood of the neoantigen. The method of the present invention 1002 further comprises: [The present invention 1004] 1004. The method of claim 1003, wherein transforming said dependency score models presentation of said neoantigen as mutually exclusive across said one or more class II MHC alleles. [The present invention 1005] The step of inputting the numerical vector into the machine-learned presentation model includes: transforming the combination of dependency scores to generate the presentation likelihood, wherein transforming the combination of dependency scores models presentation of the neoantigen as interference between the one or more class II MHC alleles. The method of the present invention 1002 further comprises: [The present invention 1006] The set of representation likelihoods is further specified by at least one or more allele non-interaction characteristics, and the method comprises: applying the machine-learned presentation model to the allele-non-interacting trait to generate a dependency score for the allele-non-interacting trait, the dependency score indicating whether the corresponding neoantigen peptide sequence will be presented based on the allele-non-interacting trait. Any of the methods 1002 to 1005 of the present invention, further comprising: [The present invention 1007] combining the dependency score for each class II MHC allele of the one or more class II MHC alleles with the dependency score for the allele-non-interacting trait; transforming the combined dependency scores for each Class II MHC allele to generate a per-allele likelihood for each Class II MHC allele, the likelihood being that the corresponding Class II MHC allele will present the corresponding neoantigen; and combining the likelihoods for each allele to generate the representation likelihood. The method of the present invention 1006 further comprising: [The present invention 1008] combining the dependency scores for each of the class II MHC alleles with the dependency scores for the allele-non-interacting trait; and transforming the combined dependency scores to generate the presentation likelihood. The method of the present invention 1006 further comprising: [The present invention 1009] The method of any of claims 1001 to 1008, wherein said one or more class II MHC alleles comprises two or more different class II MHC alleles. [The present invention 1010] 1009. The method of any of claims 1001 to 1009, wherein said at least one class II MHC allele comprises two or more different types of class II MHC alleles. [The present invention 1011] The method of any of claims 1001 to 1010, wherein said peptide sequence comprises a peptide sequence having a length other than 9 amino acids. [The present invention 1012] 1012. The method of any of claims 1001 to 1011, wherein said step of encoding said peptide sequence comprises encoding said peptide sequence using a one-hot coding scheme. [The present invention 1013] The plurality of samples (a) one or more cell lines engineered to express a single class II MHC allele; (b) one or more cell lines engineered to express multiple class II MHC alleles; (c) one or more human cell lines obtained or derived from multiple patients; (d) fresh or frozen tumor samples obtained from multiple patients; and (e) fresh or frozen tissue samples obtained from multiple patients; Any of the methods of the present invention 1001 to 1012, including at least one of the following. [The present invention 1014] the training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of said peptides; and (b) data relating to measurements of peptide-MHC binding stability for at least one of said peptides; The method of any one of claims 1001 to 1013, further comprising at least one of the following: [The present invention 1015] Any of the methods of claims 1001 to 1014, wherein the set of presentation likelihoods is further identified by at least the expression level of the one or more class II MHC alleles in the subject, and the expression level is measured by RNA-seq or mass spectrometry. [The present invention 1016] The set of presentation likelihoods is: (a) the predicted affinity between a neoantigen in the set of neoantigens and the one or more class II MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex; The method of any of claims 1001 to 1015, further characterized by at least one of the following characteristics: [The present invention 1017] The set of numerical likelihoods is: (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; The method of any of claims 1001 to 1016, further characterized by at least one of the following characteristics: [The present invention 1018] Any of the methods of present inventions 1001 to 1017, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that are more likely to be presented on the surface of the tumor cells than neoantigens that are not selected based on the machine-learned presentation model. [The present invention 1019] Any of the methods of present inventions 1001 to 1018, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that are more likely to be able to induce a tumor-specific immune response in the subject compared to neoantigens that are not selected based on the machine-learned presentation model. [The present invention 1020] Any of the methods of claims 1001 to 1019, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that are more likely to be presented to naive T cells by professional antigen-presenting cells (APCs) than unselected neoantigens based on the presentation model, and optionally the APCs are dendritic cells (DCs). [The present invention 1021] Any of the methods of present inventions 1001 to 1020, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that are less likely to be inhibited by central tolerance or peripheral tolerance than neoantigens that are not selected based on the machine-learned presentation model. [The present invention 1022] Any of the methods of present inventions 1001 to 1021, wherein the step of selecting the set of selected neoantigens comprises selecting neoantigens that are less likely to be able to induce an autoimmune response against normal tissue in the subject compared to neoantigens that are not selected based on the machine-learned presentation model. [The present invention 1023] 10. The method of any one of claims 1001 to 1022, wherein the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer. [The present invention 1024] The method of any of claims 1001 to 1023, further comprising generating an output for constructing a personalized cancer vaccine from said set of selected neoantigens. [The present invention 1025] 1025. The method of claim 1024, wherein the output for said personalized cancer vaccine comprises at least one peptide sequence or at least one nucleotide sequence encoding said selected set of neoantigens. [The present invention 1026] The method according to any one of claims 1001 to 1025, wherein the presentation model trained by machine learning is a neural network model. [The present invention 1027] 1026. The method of claim 1026, wherein the neural network model comprises a plurality of network models for the class II MHC alleles, each network model comprising a set of nodes assigned to a corresponding class II MHC allele among the class II MHC alleles and arranged in one or more layers. [The present invention 1028] 1027. The method of claim 1027, wherein each network model further comprises one or more convolutional neural networks, each of said one or more convolutional neural networks comprising a series of nodes arranged in one or more layers and having filters of different sizes, said filters of each of said one or more convolutional neural networks being sized to identify the position of an amino acid in the peptide sequence of each neoantigen that comprises a binding core or binding anchor of said peptide sequence. [The present invention 1029] 1029. The method of claim 1027 or 1028, wherein the neural network models are trained by updating parameters of the neural network models, and wherein the parameters of at least two network models are updated together in at least one training iteration. [The present invention 1030] The method of any one of claims 1026 to 1029, wherein the presentation model trained by machine learning is a deep learning model including one or more layers of nodes. [The present invention 1031] Any of the methods of claims 1001 to 1030, wherein the step of identifying the one or more T cells comprises co-culturing the one or more T cells with one or more of the neoantigens in the subset under conditions that expand the one or more T cells. [The present invention 1032] Any of the methods of claims 1001 to 1031, wherein the step of identifying the one or more T cells comprises contacting the one or more T cells with an MHC multimer comprising one or more of the neoantigens in the subset under conditions that allow binding of the T cells to the MHC multimer. [The present invention 1033] The method of any of claims 1001 to 1032, further comprising the step of identifying one or more T cell receptors (TCRs) of said one or more identified T cells. [The present invention 1034] 1034. The method of claim 1033, wherein said step of identifying said one or more T cell receptors comprises sequencing said T cell receptor sequences of said one or more identified T cells. [This invention 1035] An isolated T cell that is antigen-specific to at least one selected neoantigen in any of the subsets of the present invention 1001 to 1034. [The present invention 1036] genetically engineering a plurality of T cells to express at least one of the one or more identified T cell receptors; culturing the plurality of T cells under conditions that expand the plurality of T cells; and injecting the expanded T cells into the subject. The method of the present invention 1034 further comprises: [This invention 1037] said step of genetically engineering said plurality of T cells to express at least one of said one or more identified T cell receptors comprises: cloning the T cell receptor sequences of the one or more identified T cells into an expression vector; and transfecting each of said plurality of T cells with said expression vector. The method of the present invention 1036, comprising: [The present invention 1038] culturing the one or more identified T cells under conditions that expand the one or more identified T cells; and injecting the expanded T cells into the subject. Any of the methods of claims 1001 to 1037, further comprising: [Brief explanation of the drawings]
[0013] These and other features, aspects, and aspects of the present invention will become better understood with regard to the following description and accompanying drawings.
[0014] [Figure 1A] Current clinical approaches to neoantigen identification are presented.
[0015] [Figure 1B] It shows that less than 5% of the predicted binding peptides are displayed on tumor cells.
[0016] [Figure 1C] Illustrates the impact of specificity issues on neoantigen prediction.
[0017] [Figure 1D] This shows that binding prediction is not sufficient to identify neoantigens.
[0018] [Figure 1E] Probability of MHC-I presentation as a function of peptide length.
[0019] [Figure 1F] 1 shows an exemplary peptide spectrum generated from a dynamic range standard from Promega. SEQ ID NO: 1 is disclosed.
[0020] [Figure 1G] We show how adding features increases the positive predictive value of the model.
[0021] [Figure 2A] 1 is a schematic of an environment for identifying the likelihood of peptide presentation in a patient, according to one embodiment.
[0022] [Figure 2B] A method for obtaining presentation information according to one embodiment is described. SEQ ID NO: 20 is disclosed. [Figure 2C] A method for obtaining presentation information according to one embodiment will be described. SEQ ID NOs: 3 to 8 are disclosed in the order shown.
[0023] [Figure 3] FIG. 1 is a high-level block diagram illustrating computer logic components of a presentation identification system, according to one embodiment.
[0024] [Figure 4] An exemplary set of training data according to one embodiment is described below. The "peptide sequences" are disclosed as SEQ ID NOs: 10-13, and the "C-flanking sequences" are disclosed as SEQ ID NOs: 15, 21-22, and 22, in the order shown.
[0025] [Figure 5] 1 illustrates an exemplary network model related to MHC alleles.
[0026] [Figure 6A] 1 shows an exemplary network model NNH(·) shared by MHC alleles, according to one embodiment.
[0027] [Figure 6B] 1 shows an exemplary network model NNH(·) shared by MHC alleles according to another embodiment.
[0028] [Figure 7] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model.
[0029] [Figure 8] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model.
[0030] [Figure 9] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model.
[0031] [Figure 10] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model.
[0032] [Figure 11] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model.
[0033] [Figure 12] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model.
[0034] [Figure 13A] 1 shows the sample frequency distribution of mutation burden in NSCLC patients.
[0035] [Figure 13B] 1 shows the number of presented neoantigens in a simulated vaccine for patients selected based on the selection criterion of whether the patient meets a minimum mutational load, according to one embodiment.
[0036] [Figure 13C] 10A-10C show a comparison of the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model, according to one embodiment.
[0037] [Figure 13D]10A-B show a comparison of the number of neoantigens presented in simulated vaccines between selected patients associated with a vaccine containing therapeutic subsets identified based on a single allele-specific presentation model for HLA-A*02:01 and selected patients associated with a vaccine containing therapeutic subsets identified based on both an allele-specific presentation model for HLA-A*02:01 and an allele-specific presentation model for HLA-B*07:02. According to one embodiment, the vaccine volume is set to v=20 epitopes.
[0038] [Figure 13E] FIG. 10 compares the number of presented neoantigens in simulated vaccines between patients selected based on mutational burden and patients selected by expected utility score, according to one embodiment.
[0039] [Figure 14A] 1 is a histogram of peptide lengths eluted from class II MHC alleles on human tumor cells and tumor-infiltrating lymphocytes (TILs) using mass spectrometry.
[0040] [Figure 14B] The dependency of mRNA quantification and presented peptides per residue is shown for two exemplary data sets.
[0041] [Figure 14C] 1 compares the performance results of exemplary proposed models trained and tested using two exemplary datasets.
[0042] [Figure 14D] 1 is a histogram showing the amount of peptides sequenced using mass spectrometry for each sample of a total of 73 samples containing HLA class II molecules.
[0043] [Figure 14E] 1 is a histogram showing the amount of a sample in which a particular MHC class II molecule allele was identified.
[0044] [Figure 14F] 1 is a histogram showing the proportion of peptides presented by MHC class II molecules across a range of peptide lengths in a total of 73 samples.
[0045] [Figure 14G] 1 is a line graph showing the relationship between gene expression and the incidence of presentation of gene expression products by MHC class II molecules for genes present in 73 samples.
[0046] [Figure 14H] 1 is a line graph comparing the performance of the same model with different inputs in predicting the likelihood that peptides in a test dataset of peptides will be presented by MHC class II molecules.
[0047] [Figure 14I] 1 is a line graph comparing the performance of three different models in predicting the likelihood that peptides in a test dataset of peptides will be presented by MHC class II molecules.
[0048] [Figure 14J] FIG. 14I illustrates an exemplary embodiment of the Bi-LSTM model of FIG. 14I configured to predict peptide presentation by HLA-DRB (MHC class II genes).
[0049] [Figure 14K] 14B is a line graph showing the perfect precision-recall curves for the Bi-LSTM, MLP, RNN, and binding affinity models of FIG. 14I.
[0050] [Figure 14L] 1 is a line graph comparing the performance of a best-in-class conventional model using two different criteria and a presentation model disclosed herein with two different inputs in predicting the likelihood that a peptide in a test dataset of peptides will be presented by an MHC class II molecule.
[0051] [Figure 14M] 1 is a histogram showing the amount of peptides sequenced using mass spectrometry at a q value of less than 0.1 for each sample of a total of 230 samples consisting of human tumors (NSCLC, lymphoma, and ovarian cancer) and cell lines (EBV) containing HLA class II molecules.
[0052] [Figure 14N] 1 is a histogram showing the amount of a sample in which a particular MHC class II molecule allele was identified.
[0053] [Figure 14O] Peptides bound to MHC class I molecules and peptides bound to MHC class II molecules are shown.
[0054] [Figure 14P] FIG. 14Q shows an exemplary embodiment of an Inception Neural Network of the Inception model of FIG. 14Q configured to predict peptide presentation by MHC class II molecules.
[0055] [Figure 14Q] 1 is a line graph comparing the performance of the "Bi-LSTM" and "Inception" presentation models in predicting the likelihood that a peptide in a test dataset of peptides will be presented by at least one of the MHC class II molecules present in the test dataset.
[0056] [Figure 15A]Figure 15 compares the predictive performance of the "MS model," "NetMHCIIpan rank": NetMHCIIpan 3.177, which adopts the lowest NetMHCIIpan percentile rank across HLA-DRB1*15:01 and HLA-DRB5*01:01, and "NetMHCIIpan nM": NetMHCIIpan 3.1, which adopts the strongest affinity in nM across HLA-DRB1*15:01 and HLA-DRB5*01:01, in ranking peptides within the HLA-DRB1*15:01 / HLA-DRB5*01:01 test dataset. [Figure 15B] See legend to Figure 15A.
[0057] [Figure 16] 1 shows an exemplary embodiment of a TCR construct for introducing a TCR into a recipient cell.
[0058] [Figure 17-1] Figure 17 shows the nucleotide sequence of an exemplary P526 construct backbone for cloning TCRs into expression systems for therapeutic development. Figure 17 discloses SEQ ID NO: 23. [Figure 17-2] A continuation of Figure 17-1 is shown.
[0059] [Figure 18-1] Figure 18 shows an exemplary construct sequence for cloning a patient's neoantigen-specific TCR clonotype 1 into an expression system for therapeutic development. Figure 18 discloses SEQ ID NO: 24. [Figure 18-2] A continuation of Figure 18-1 is shown.
[0060] [Figure 19-1] Figure 19 shows an exemplary construct sequence for cloning a patient's neoantigen-specific TCR clonotype 3 into an expression system for therapeutic development. Figure 19 discloses SEQ ID NO: 25. [Figure 19-2] A continuation of Figure 19-1 is shown.
[0061] [Figure 20] 1 is a flowchart of a method for providing personalized neoantigen-specific therapy to a patient, according to one embodiment.
[0062] [Figure 21] An exemplary computer for implementing the entities shown in FIGS. 1 and 3 will now be described. DETAILED DESCRIPTION OF THE INVENTION
[0063] Detailed Description I. Definition In general, terms used in the claims and the specification shall be interpreted as having their ordinary meaning as understood by one of ordinary skill in the art. Certain terms are defined below to provide further clarity. If there is a conflict between the ordinary meaning and a given definition, the given definition shall control.
[0064] As used herein, the term "antigen" refers to a substance that induces an immune response.
[0065] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes it different from its corresponding wild-type parent antigen, for example, due to a tumor cell mutation or tumor cell-specific post-translational modification. Neoantigens may include polypeptide or nucleotide sequences. Mutations can include frameshift or non-frameshift insertion / deletions (indels), missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression change that results in a new ORF. Mutations can also include splice variants. Tumor cell-specific post-translational modifications can include aberrant phosphorylation. Tumor cell-specific post-translational modifications can also include splice antigens generated by the proteasome. See Liepe et al., "A large fraction of HLA class I ligands are proteasome-generated spliced peptides"; Science. 2016 Oct 21;354(6310):354-358.
[0066] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in tumor cells or tissues of a subject, but is not present in the corresponding normal cells or tissues of the subject.
[0067] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct that is based on one or more neoantigens, eg, multiple neoantigens.
[0068] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.
[0069] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.
[0070] As used herein, the term "coding mutation" refers to a mutation that occurs in a coding region.
[0071] As used herein, the term "ORF" means open reading frame.
[0072] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that arises due to mutation or other abnormalities such as splicing.
[0073] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.
[0074] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon.
[0075] As used herein, the term "frameshift mutation" is a mutation that causes an alteration in the frame of a protein.
[0076] As used herein, the term "indel" refers to the insertion or deletion of one or more nucleic acids.
[0077] As used herein, the term "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences in which a certain percentage of nucleotides or amino acid residues are the same when compared and aligned for maximum correspondence, as determined using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN, or other algorithms available to those of skill in the art), or by visual inspection. Depending on the application, the "percent identity" can exist over a region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.
[0078] In sequence comparison, generally, one sequence serves as a reference sequence to which test sequences are compared.When using a sequence comparison algorithm, test sequences and reference sequences are input into a computer, subsequence coordinates are designated if necessary, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the sequence identity (%) of the test sequence to the reference sequence based on the designated program parameters.Alternatively, sequence similarity or difference can also be established by the combination of the presence or absence of a specific nucleotide at a selected sequence position (e.g., sequence motif) or an amino acid in a translated sequence.
[0079] Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (see generally Ausubel et al., infra).
[0080] One example of an algorithm that is suitable for determining percent sequence identity and percent sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.
[0081] As used herein, the term "non-stop or read-through" refers to a mutation that results in the removal of the natural stop codon.
[0082] As used herein, the term "epitope" refers to a specific portion of an antigen that is typically bound by an antibody or T-cell receptor.
[0083] As used herein, the term "immunogenic" refers to the ability to elicit an immune response, for example, via T cells, B cells, or both.
[0084] As used herein, the terms "HLA binding affinity" and "MHC binding affinity" refer to the affinity of binding between a specific antigen and a specific MHC allele.
[0085] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.
[0086] As used herein, the term "mutation" is a difference between the nucleic acid of a subject and a reference human genome used as a control.
[0087] As used herein, the term "variant calling" is the algorithmic determination, typically from sequencing, of the presence of a mutation.
[0088] As used herein, the term "polymorphism" refers to a germline mutation, ie, a mutation found in all DNA-bearing cells of an individual.
[0089] As used herein, the term "somatic mutation" is a mutation that occurs in a non-germline cell of an individual.
[0090] As used herein, the term "allele" refers to one version of a gene or one version of a gene sequence or one version of a protein.
[0091] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.
[0092] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the degradation of mRNA by the cell due to a premature stop codon.
[0093] As used herein, the term "truncal mutation" is a mutation that occurs early in the development of a tumor and is present in the majority of the cells of the tumor.
[0094] As used herein, the term "subclonal mutation" is a mutation that occurs late in the development of a tumor and is present in only a portion of the cells of the tumor.
[0095] As used herein, the term "exome" refers to the subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.
[0096] As used herein, the term "logistic regression" is a regression model for binary data from statistics in which the logit of the probability that the dependent variable is equal to 1 is modeled as a linear function of the dependent variable.
[0097] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinear transformations typically trained by stochastic gradient descent and backpropagation.
[0098] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.
[0099] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome can also refer to the properties of a cell or a collection of cells (e.g., a tumor peptidome refers to the union of the peptidomes of all cells that comprise a tumor).
[0100] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.
[0101] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.
[0102] As used herein, the term "MHC multimer" is a peptide-MHC complex consisting of multiple peptide-MHC monomer units.
[0103] As used herein, the term "MHC tetramer" is a peptide-MHC complex consisting of four peptide-MHC monomer units.
[0104] As used herein, the term "tolerance or immune tolerance" refers to a state of immune unresponsiveness to one or more antigens, eg, self-antigens.
[0105] As used herein, the term "central tolerance" is tolerance conferred in the thymus by either deleting autoreactive T cell clones or promoting their differentiation into immunosuppressive regulatory T cells (Tregs).
[0106] As used herein, the term "peripheral tolerance" refers to tolerance conferred in the peripheral system by downregulating or anergizing autoreactive T cells that survive central tolerance or by promoting the differentiation of these T cells into Tregs.
[0107] The term "sample" can include a single cell, or multiple cells, or fragments of cells, or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.
[0108] The term "subject" includes cells, tissues, or organisms, human or non-human, whether male or female, in vivo, ex vivo, or in vitro. The term subject includes mammals, including humans.
[0109] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.
[0110] The term "clinical factor" refers to a measurement of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values that can be obtained from assessing a subject or a sample (or a population of samples) from a subject under a given condition. A clinical factor can also be predicted by other parameters, such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.
[0111] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.
[0112] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0113] Terms not directly defined herein should be understood to have the meanings generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide further guidance to the practitioner in describing the compositions, devices, methods, etc. of embodiments of the present invention, as well as how to make or use them. It will be recognized that multiple ways of saying the same thing may be used. Accordingly, alternative terms and synonyms may be used for any one or more of the terms discussed herein. No weight should be placed on whether a term is detailed or discussed herein. Several synonyms or alternative methods, materials, etc. are provided. The recitation of one or more synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the inventive embodiments herein.
[0114] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.
[0115] II. Methods for identifying neoantigens Disclosed herein is a method for identifying T cells that are antigen-specific for neoantigens derived from tumor cells of a subject that are likely to be presented by class II MHC alleles on the surface of the tumor cells. The method includes obtaining nucleotide sequencing data of the exome, transcriptome, and / or whole genome from tumor cells and normal cells of the subject. The nucleotide sequencing data is used to obtain a peptide sequence for each neoantigen in a set of neoantigens. The set of neoantigens is identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from normal cells. Specifically, the peptide sequence of each neoantigen in the set of neoantigens contains at least one change that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells. The method further includes encoding the peptide sequence of each neoantigen in the set of neoantigens into a corresponding numeric vector. Each numeric vector contains information describing the amino acids that make up the peptide sequence and the position of the amino acids in the peptide sequence. The method further includes generating a presentation likelihood for each neoantigen in the set of neoantigens by inputting a numerical vector into a machine-learned presentation model. Each presentation likelihood represents the likelihood that the corresponding neoantigen will be presented by a class II MHC allele on the surface of tumor cells in the subject. The machine-learned presentation model includes a plurality of parameters and a function. The plurality of parameters are identified based on a training dataset. The training dataset includes, for each sample in a plurality of samples, labels obtained by mass spectrometry to measure the presence of peptides bound to at least one class II MHC allele in the set of class II MHC alleles identified as present in that sample, and training peptide sequences encoded as a numerical vector containing information describing the amino acids that make up the peptides and / or the positions of amino acids in the peptides. The function represents the relationship between the numerical vector received as input by the machine-learned presentation model and the presentation likelihood generated as output by the machine-learned presentation model based on the numerical vector and the plurality of parameters.The method further includes generating a set of selected neoantigens by selecting a subset of the set of neoantigens based on likelihood of presentation, and the method further includes identifying T cells that are antigen-specific for at least one of the neoantigens in the subset and returning the identified T cells.
[0116] In some embodiments, inputting the numerical vectors into the machine-learned presentation model includes applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate a dependency score for each class II MHC allele. The dependency score for a class II MHC allele indicates whether the class II MHC allele will present the neoantigen based on a specific amino acid at a specific position in the peptide sequence. In further embodiments, inputting the numerical vectors into the machine-learned presentation model further includes transforming the dependency scores to generate, for each class II MHC allele, a corresponding allele likelihood indicating the likelihood that the corresponding class II MHC allele will present the corresponding neoantigen, and combining the allele likelihoods to generate a neoantigen presentation likelihood. In some embodiments, transforming the dependency scores models neoantigen presentation as mutually exclusive across class II MHC alleles. In alternative embodiments, inputting the numerical vectors into the machine-learned presentation model further includes transforming a combination of dependency scores to generate a presentation likelihood. In such embodiments, transforming a combination of dependency scores models neoantigen presentation as interference between class II MHC alleles.
[0117] In some embodiments, the set of presentation likelihoods is further specified by one or more allele non-interacting features. In such embodiments, the method further includes generating a dependency score for the allele non-interacting feature by applying a machine-learned presentation model to the allele non-interacting feature. The dependency score indicates whether the peptide sequence of the corresponding neoantigen will be presented based on the allele non-interacting feature. In some embodiments, the method further includes combining the dependency score for each class II MHC allele with the dependency score for the allele non-interacting feature, transforming the combined dependency scores for each class II MHC allele to generate a per-allele likelihood for each class II MHC allele, and combining the per-allele likelihoods to generate a presentation likelihood. The per-allele likelihood for a class II MHC allele indicates the likelihood that the class II MHC allele will present the corresponding neoantigen. In an alternative embodiment, the method further includes combining the dependency score for the class II MHC allele with the dependency score for the allele non-interacting feature and transforming the combined dependency score to generate a presentation likelihood.
[0118] In some embodiments, the class II MHC alleles comprise two or more different class II MHC alleles.
[0119] In some embodiments, at least one class II MHC allele in the set of class II MHC alleles identified as present in the samples of the training dataset comprises two or more different types of class II MHC alleles.
[0120] In some embodiments, the peptide sequence comprises a peptide sequence having a length other than 9 amino acids.
[0121] In some embodiments, encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme.
[0122] In certain embodiments, the plurality of samples comprises at least one of: a cell line engineered to express a single class II MHC allele; a cell line engineered to express multiple class II MHC alleles; a human cell line obtained or derived from multiple patients; or fresh or frozen tissue samples obtained from multiple patients.
[0123] In some embodiments, the training dataset further comprises at least one of data related to a measure of peptide-MHC binding affinity for at least one of the peptides and data related to a measure of peptide-MHC binding stability for at least one of the peptides.
[0124] In some embodiments, the set of representation likelihoods is further identified by the expression level of class II MHC alleles in the subject, as measured by RNA-seq or mass spectrometry.
[0125] In some embodiments, the set of presentation likelihoods is further specified by properties including at least one of the predicted affinity between neoantigens and class II MHC alleles in the set of neoantigens and the predicted stability of the neoantigen-encoded peptide-MHC complex.
[0126] In some embodiments, the set of numerical likelihoods is further identified by a feature that includes at least one of a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence and an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence.
[0127] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that are more likely to be presented on the tumor cell surface than non-selected neoantigens based on a machine learning presentation model.
[0128] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that have a higher likelihood of inducing a tumor-specific immune response in a subject compared to non-selected neoantigens based on a machine learning-driven presentation model.
[0129] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that are more likely to be presented to naive T cells by professional antigen-presenting cells (APCs) than unselected neoantigens based on a presentation model. In such embodiments, the APCs are optionally dendritic cells (DCs).
[0130] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that are less likely to be inhibited by central or peripheral tolerance compared to non-selected neoantigens based on a machine learning-driven presentation model.
[0131] In some embodiments, selecting the set of selected neoantigens comprises selecting neoantigens that are less likely to be able to induce an autoimmune response against normal tissue in a subject compared to non-selected neoantigens based on a machine learning-based presentation model.
[0132] In some embodiments, the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0133] In some embodiments, the method further comprises generating output for constructing a personalized cancer vaccine from the set of selected neoantigens, in such embodiments, the output for the personalized cancer vaccine can comprise at least one peptide sequence or at least one nucleotide sequence encoding the set of selected neoantigens.
[0134] In some embodiments, the machine-learned model is a neural network model. In such embodiments, the neural network model may include a plurality of network models for class II MHC alleles, each of which includes a series of nodes assigned to a corresponding class II MHC allele among the class II MHC alleles and arranged in one or more layers. In such embodiments, the neural network model may be trained by updating the parameters of the neural network model, and the parameters of at least two network models are updated together in at least one training iteration.
[0135] In such embodiments, each network model can further include one or more convolutional neural networks, each of which includes a set of nodes arranged in one or more layers and having filters of different sizes, and each filter of the one or more convolutional neural networks can be sized to identify the position of an amino acid in the peptide sequence of each neoantigen that comprises a binding core or a binding anchor of the peptide sequence.
[0136] In some embodiments, the machine-learned presentation model may be a deep learning model that includes one or more layers of nodes.
[0137] In some embodiments, identifying the T cells comprises co-culturing the T cells with one or more of the neoantigens in the subset under conditions that allow the T cells to expand.
[0138] In some embodiments, identifying the T cells comprises contacting the T cells with an MHC multimer comprising one or more of the neoantigens in the subset under conditions that allow binding of the T cells to the MHC multimer.
[0139] In some embodiments, the method further includes identifying the T cell receptor (TCR) of the identified T cells. In such embodiments, identifying the T cell receptor includes sequencing the T cell receptor sequence of the identified T cells. In such embodiments, the method can further include genetically engineering the T cells to express at least one of the one or more identified T cell receptors, culturing the T cells under conditions for T cell expansion, and infusing the expanded T cells into a subject. In such embodiments, genetically engineering the T cells to express at least one of the identified T cell receptors can include cloning the T cell receptor sequence of the identified T cells into an expression vector and transfecting each of the T cells with the expression vector.
[0140] In some embodiments, the method can further include culturing the identified T cells under conditions that expand the identified T cells, and injecting the expanded T cells into a subject.
[0141] Also disclosed herein are isolated T cells that are antigen-specific for at least one selected neoantigen from the subset of neoantigens mentioned above.
[0142] International Patent Application Publication Nos. WO2018 / 195357 and WO2019 / 050994 are incorporated herein by reference in their entirety. International Patent Application Publication No. WO2018 / 195357 describes a method for predicting antigen presentation by MHC class II molecules. International Patent Application Publication No. WO2019 / 050994 describes a method for identifying T cells specific for antigens presented by MHC molecules. Although these publications are mentioned in this section of the specification, the disclosures set forth in International Patent Application Publication Nos. WO2018 / 195357 and WO2019 / 050994 are incorporated by reference in their entirety into all sections of this application.
[0143] III. Identification of tumor-specific mutations in neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells). In particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject with cancer, but may not be present in normal tissues from the subject.
[0144] Genetic mutations in tumors can be considered useful for immunological targeting of tumors if they result in changes in the amino acid sequence of proteins exclusively in tumors. Useful mutations include: (1) non-synonymous mutations that result in different amino acids in proteins; (2) read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; (3) splice site mutations that result in the inclusion of an intron in mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) frameshift mutations or deletions that result in new open reading frames with new tumor-specific protein sequences. Mutations can also include one or more of non-frameshift insertions / deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in new ORFs.
[0145] For example, mutated peptides or mutated polypeptides resulting from splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.
[0146] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0147] Various methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. Several techniques have been described, including dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize amplification of target gene regions, typically by PCR. Still other methods rely on the generation of small signal molecules by invasive cleavage followed by mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.
[0148] PCR-based detection means can involve the multiplex amplification of multiple markers simultaneously.For example, it is well known in the art to select PCR primers so as to generate PCR products that do not overlap in size and can be analyzed simultaneously.Alternatively, it is possible to amplify different markers with primers that are differentially labeled and therefore can be differentially detected.Of course, hybridization-based detection means allows the differential detection of multiple PCR products in a sample.Other techniques that allow multiplex analysis of multiple markers are known in the art.
[0149] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA.For example, single nucleotide polymorphisms can be detected by using special exonuclease-resistant nucleotides, as disclosed in Mundy, CR (US Patent No. 4,656,127).According to this method, a primer complementary to the allele sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a specific animal or human.If the polymorphic site on the target molecule contains a nucleotide that is complementary to the specific exonuclease-resistant nucleotide derivative present, this derivative will be incorporated onto the end of the hybridized primer.This incorporation makes the primer resistant to exonucleases, thereby enabling its detection.Since the identity of the exonuclease-resistant derivative of the sample is known, the knowledge that the primer has become resistant to exonucleases reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of exogenous sequence data.
[0150] Solution-based methods can be used to determine the identity of the nucleotide at a polymorphic site (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087)). As in the method of Mundy, U.S. Pat. No. 4,656,127, a primer is used that is complementary to the allelic sequence immediately 3' to the polymorphic site. This method uses a labeled dideoxynucleotide derivative that becomes incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site to determine the identity of the nucleotide at that site.
[0151] An alternative method known as Genetic Bit Analysis or GBA has been described by Goelet, P. et al. (PCT Application No. 92 / 15712). The method of Goelet, P. et al. uses a mixture of labeled terminators and primers that are complementary to the sequence 3' of the polymorphic site. The method of Goelet, P. et al. uses a mixture of labeled terminators and primers that are complementary to the sequence 3' of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which the primer or target molecule is immobilized on a solid phase.
[0152] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. et al., Nucl. Acids. Res. 17:7779-7784 (1989); Sokolov, B. P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M. et al., Proc. Natl. Acad. Sci. (USA) 88:1143-1147 (1991); Prezant, T. R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al. al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, signal is proportional to the number of incorporated deoxynucleotides, so that polymorphisms occurring in runs of the same nucleotide can result in a signal proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).
[0153] Numerous initiatives obtain sequence information directly from millions of individual molecules of DNA or RNA in parallel. Real-time single-molecule sequencing by synthesis techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent strands of DNA complementary to the template being sequenced. In one method, oligonucleotides 30–50 bases in length are covalently anchored at their 5′ ends to glass coverslips. These anchored strands serve two functions. First, they act as capture sites for the target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis for sequence reading. The capture primers serve as fixed-location sites for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and dye cleavage. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. As the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing-by-synthesis techniques also exist.
[0154] Any suitable sequencing-by-synthesis platform can be used to identify mutations. As mentioned above, four major sequencing-by-synthesis platforms are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing-by-synthesis platforms have also been described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the multiple nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. A capture sequence (also called a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to a support that can double as a universal primer.
[0155] As an alternative to capture sequences, a member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair, e.g., as described in U.S. Patent Application Publication No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.
[0156] Following capture, the sequence can be analyzed by single-molecule detection / sequencing, including, for example, template-dependent sequencing by synthesis, as described, for example, in the Examples and U.S. Patent No. 7,283,337. In sequencing by synthesis, surface-bound molecules are exposed to a multitude of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time, in a step-and-repeat mode. For real-time analysis, a different optical label can be incorporated for each nucleotide, and multiple lasers can be utilized for stimulation of the incorporated nucleotides.
[0157] Sequencing can also include other massively parallel sequencing or next-generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, ThermoPGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel sequencing technologies, and future generations of these technologies, can be used.
[0158] Any cell type or tissue can be used to obtain nucleic acid samples for use in the methods described herein.For example, DNA or RNA samples can be obtained from tumor or body fluids, for example, blood obtained by known techniques (for example, venipuncture) or saliva.Alternatively, nucleic acid testing can be performed on dry samples (for example, hair or skin).In addition, a sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal tissue is of the same tissue type as tumor.A sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal sample is of a different tissue type from tumor.
[0159] The tumor may include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0160] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. Peptides can be acid-eluted from tumor cells or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.
[0161] IV. Neoantigens Neoantigens can comprise nucleotides or polynucleotides. For example, neoantigens can be RNA sequences that encode polypeptide sequences. Neoantigens useful in vaccines can therefore comprise nucleotide sequences or polypeptide sequences.
[0162] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the neoantigen comprises a nucleotide sequence (e.g., DNA or RNA) that encodes the associated polypeptide sequence.
[0163] The one or more polypeptides encoded by the neoantigen nucleotide sequences can comprise at least one of the following: a binding affinity to MHC with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides; the presence of a sequence motif within or near the peptide that promotes proteasomal cleavage; and the presence of a sequence motif within or near the peptide that promotes TAP transport; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II polypeptides; and the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.
[0164] One or more neoantigens can be present on the surface of a tumor.
[0165] The one or more neoantigens can be immunogenic in a tumor-bearing subject, for example, capable of eliciting a T cell or B cell response in the subject.
[0166] One or more neoantigens that induce an autoimmune response in a subject can be eliminated from consideration in the context of generating a vaccine for a tumor-bearing subject.
[0167] The size of at least one neoantigen peptide molecule is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35 , about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derivable therein. In a specific embodiment, the neoantigen peptide molecule is 50 or fewer amino acids.
[0168] Neoantigen peptides and polypeptides can be 15 residues or less in length for MHC class I, usually between about 8 and about 11 residues, particularly 9 or 10 residues; for MHC class II, they can be 6 to 30 residues.
[0169] If desired, longer peptides can be designed in several ways. In one example, if the likelihood of peptide presentation on HLA alleles is predicted or known, the longer peptides can consist of either (1) individual presented peptides with extensions of 2-5 amino acids toward the N- and C-termini of each corresponding gene product; or (2) a concatenation of some or all of the presented peptides, each with its extended sequence. In another example, if sequencing reveals long (more than 10 residues) neo-epitope sequences present in the tumor (e.g., due to frameshifts, readthrough, or intron inclusion resulting in novel peptide sequences), the longer peptides would (3) consist of the entire novel tumor-specific stretch of amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides that are presented to the strongest HLA alleles. In either example, the use of longer peptides may allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.
[0170] Neoantigen peptides and polypeptides can be presented on HLA proteins. In some embodiments, the neoantigen peptides and polypeptides are presented on HLA proteins with greater affinity than the wild-type peptide. In some embodiments, the neoantigen peptide or polypeptide can have an IC50 of at least 5000 nM or less, at least 1000 nM or less, at least 500 nM or less, at least 250 nM or less, at least 200 nM or less, at least 150 nM or less, at least 100 nM or less, at least 50 nM or less, or even less.
[0171] In some embodiments, the neoantigen peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.
[0172] Also provided are compositions comprising at least two or more neoantigen peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides can be derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which neoantigen peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some embodiments, the tumor-specific mutations are driver mutations for a particular cancer type.
[0173] Neoantigen peptides and polypeptides with desired activities or properties can be modified to confer certain desirable attributes, e.g., improved pharmacological characteristics, while enhancing or at least retaining substantially all of the biological activity of the unmodified peptide, which binds to desired MHC molecules and activates appropriate T cells. For example, neoantigen peptides and polypeptides can be further subjected to various modifications, such as conservative or non-conservative substitutions, which may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions refer to the replacement of an amino acid residue with another that is biologically and / or chemically similar, e.g., one hydrophobic residue with another hydrophobic residue, or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effects of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2nd Ed. (1984).
[0174] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful for increasing peptide and polypeptide stability in vivo. Stability can be assayed in a number of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally follows: Pooled human serum (type AB, non-heat-inactivated) is defatted by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.
[0175] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is typically composed of relatively small, neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that the optional spacer need not be composed of the same residues and can therefore be a hetero- or homo-oligomer. If present, the spacer will usually be at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.
[0176] The neoantigen peptide can be linked to the T helper peptide at either the amino or carboxy terminus of the peptide, either directly or via a spacer. The amino terminus of either the neoantigen peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxin 830-843, influenza 307-319, and malaria sporozoite peritoneal sites 382-398 and 378-389.
[0177] Proteins or peptides can be produced by any technique known to those of skill in the art, including expressing proteins, polypeptides, or peptides through standard molecular biology techniques, isolating proteins or peptides from natural sources, or chemically synthesizing proteins or peptides. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located on the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.
[0178] In a further embodiment, the neoantigen comprises a nucleic acid (e.g., a polynucleotide) encoding a neoantigen peptide or a portion thereof. The polynucleotide can be, for example, a single-stranded and / or double-stranded polynucleotide, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), or a polynucleotide having a phosphorothioate backbone, either in a natural or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further embodiment provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host; such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.
[0179] IV. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., a tumor-specific immune response. Vaccine compositions typically include multiple neoantigens selected, e.g., using the methods described herein. Vaccine compositions may also be referred to as vaccines.
[0180] The vaccine can contain 1 to 30 different peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine may contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, It may contain 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine contains 1-30 neoantigen sequences: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 It can contain 6, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.
[0181] In one embodiment, the different peptides and / or polypeptides, or the nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides are capable of binding to different MHC molecules, such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, a vaccine composition comprises coding sequences for peptides and / or polypeptides capable of binding to the most frequently occurring MHC class I molecules and / or MHC class II molecules. Thus, the vaccine composition can comprise different fragments capable of binding to at least two preferred, at least three preferred, or at least four preferred MHC class I molecules and / or MHC class II molecules.
[0182] The vaccine composition may generate a specific cytotoxic T cell response and / or a specific helper T cell response.
[0183] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are provided herein below. The composition can be combined with a carrier, such as a protein, or an antigen-presenting cell, such as a dendritic cell (DC), which can present peptides to T cells.
[0184] An adjuvant is any substance whose incorporation into a vaccine composition enhances or otherwise modifies the immune response to a neoantigen. The carrier can be a scaffold, such as a polypeptide or polysaccharide, to which the neoantigen can be bound. Optionally, the adjuvant is covalently or non-covalently conjugated.
[0185] The ability of adjuvant to increase the immune response to antigen is typically manifested by a significant or substantial increase in immune-mediated reaction or a reduction in disease symptoms.For example, the increase in humoral immunity is typically manifested by a significant increase in the titer of antibody produced against antigen, and the increase in T cell activity is typically manifested in an increase in cell proliferation, or cellular cytotoxicity, or cytokine secretion.Adjuvant can also change immune response, for example, by changing mainly humoral or Th response to mainly cellular or Th response.
[0186] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA206, Montanide ISA 50V, Montanide Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila's QS21 stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponins, mycobacterial extracts and synthetic bacterial cell wall mimics, and other proprietary adjuvants such as Ribi's Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are also useful. Several immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison AC; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Several cytokines have been directly linked to influencing dendritic cell migration to lymphoid tissues (e.g., TNF-α), accelerating dendritic cell maturation into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J ImmunotherEmphasis Tumor Immunol. 1996(6):414-418).
[0187] CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in vaccine settings. Other TLR-binding molecules, such as RNA that binds to TLR 7, TLR 8, and / or TLR 9, may also be used.
[0188] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies, such as cyclophosphamide, sunitinib, bevacizumab, Celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony-stimulating factors such as granulocyte-macrophage colony-stimulating factor (GM-CSF, sargramostim).
[0189] A vaccine composition can include more than one different adjuvant. Additionally, a therapeutic composition can include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.
[0190] The carrier (or excipient) can exist independently of the adjuvant. The function of the carrier can be, for example, to increase activity or immunogenicity, to provide stability, to increase biological activity, or to increase serum half-life, particularly to increase the molecular weight of the variant. Furthermore, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, a serum protein such as keyhole limpet hemocyanin, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is tolerated and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be a dextran, such as Sepharose.
[0191] Cytotoxic T cells (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. MHC molecules themselves are located on the cell surface of antigen-presenting cells. Therefore, CTL activation is possible when a trimeric complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when peptides are used to activate CTLs, but also when APCs bearing the respective MHC molecules are added, it can enhance the immune response. Therefore, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.
[0192] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenovirus (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880), etc. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0193] IV.A. Additional Considerations for Vaccine Design and Manufacturing IV.A.1. Determining a set of peptides covering all tumor subclones Truncal peptides, meaning those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine.53 Optionally, if there are no truncal peptides that are predicted to be highly likely to be presented and immunogenic, or if the number of truncal peptides that are predicted to be highly likely to be presented and immunogenic is small enough that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .
[0194] IV.A.2. Neoantigen Prioritization After applying all of the above neoantigen filters, it is possible that more candidate neoantigens remain available for vaccine inclusion than vaccine technology can accommodate. Additionally, uncertainty about various aspects of neoantigen analysis may remain, and trade-offs may exist between various attributes of candidate vaccine neoantigens. Therefore, instead of predetermined filters at each stage of the selection process, an integral multidimensional model can be considered, in which candidate neoantigens are placed in a space with at least the following axes, and selection is optimized using an integral approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmune risk is typically preferred) 2. Probability of sequencing artifacts (lower artifact probabilities are typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferable) 5. Gene Expression (higher expression is typically preferred) 6. HLA gene coverage (a greater number of HLA molecules involved in presenting a set of neoantigens may decrease the probability that tumors will evade immune attack through downregulation or mutation of HLA molecules) 7. HLA class coverage (covering both HLA-I and HLA-II may increase the chance of a therapeutic response and decrease the chance of tumor immune evasion).
[0195] Furthermore, in some cases, neoantigens can be deprioritized (e.g., excluded) from vaccination if they are predicted to be presented by HLA alleles that are lost or inactivated in all or part of the patient's tumor. Loss of HLA alleles can occur through somatic mutation, loss of heterozygosity, or homozygous deletion of the locus. Methods for detecting somatic mutations of HLA alleles are well known in the art (e.g., Shukla et al., 2015). Methods for detecting somatic loss of heterozygosity and homozygous deletion (including HLA loci) have also been described (Carter et al., 2012; McGranahan et al., 2017; Van Loo et al., 2010).
[0196] V. Methods of Treatment and Preparation Also provided are methods for inducing a tumor-specific immune response in a subject, vaccinating against the tumor, and treating and / or alleviating symptoms of cancer in a subject by administering to the subject one or more neoantigens, such as multiple neoantigens identified using the methods disclosed herein.
[0197] In some embodiments, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desired. The tumor can be any solid tumor, such as breast, ovarian, prostate, lung, kidney, stomach, colon, testicular, head and neck, pancreas, brain, melanoma, and other tissue organ tumors, as well as hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.
[0198] The neoantigen can be administered in an amount sufficient to induce a CTL response.
[0199] The neoantigen can be administered alone or in combination with other therapeutic agents, such as chemotherapeutic agents, radiation, or immunotherapy. Any suitable therapeutic treatment for the particular cancer can be administered.
[0200] In addition, the subject can be further administered with an anti-immunosuppressive / immunostimulatory substance, such as a checkpoint inhibitor.For example, the subject can be further administered with an anti-CTLA antibody, or anti-PD-1 or anti-PD-L1.Blocking CTLA-4 or PD-L1 with an antibody can enhance the immune response against cancerous cells in patients.In particular, blocking CTLA-4 has been shown to be effective when used in vaccination protocols.
[0201] The optimal amount of each neoantigen to be included in the vaccine composition and the optimal dosing regimen can be determined. For example, the neoantigen or its variants can be formulated for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Methods of injection include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administering vaccine compositions are known to those skilled in the art.
[0202] Vaccines can be edited so that the selection, number, and / or amount of neoantigens present in the composition are tissue-, cancer-, and / or patient-specific. For example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. Selection can depend on the specific type of cancer, the state of the disease, earlier treatment regimens, the patient's immune status, and, of course, the patient's HLA haplotype. Furthermore, vaccines can contain components that are personalized according to the individual needs of a particular patient. Examples include altering the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjusting for secondary treatments after a first round or scheme of treatment.
[0203] For compositions to be used as vaccines for cancer, neoantigens with similar normal self-peptides that are abundantly expressed in normal tissues can be avoided or present in low amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a high amount of a particular neoantigen, the respective pharmaceutical composition for treating that cancer can be present in high amount and / or can include more than one neoantigen specific to that particular neoantigen or pathway of that neoantigen.
[0204] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In therapeutic applications, the compositions are administered to patients in an amount sufficient to elicit an effective CTL response against the tumor antigen and cure or at least partially halt symptoms and / or complications. An amount adequate to achieve this is defined as a "therapeutically effective dose." An amount effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general health, and the judgment of the prescribing physician. It should be kept in mind that compositions can generally be used in severe disease states, i.e., life-threatening or potentially life-threatening situations, particularly when cancer has metastasized. In such instances, it is possible, and the treating physician may find it desirable, to administer substantial excesses of these compositions, taking into account the minimization of adventitious substances and the relatively non-toxic nature of the neoantigens.
[0205] For therapeutic use, administration can begin at the time of detection or surgical removal of a tumor, followed by boosting doses until at least symptoms are substantially abated, and for a period thereafter.
[0206] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against tumors. Disclosed herein are compositions for parenteral administration that contain a solution of a neoantigen, where the vaccine composition is dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. Various aqueous carriers can be used, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional, well-known sterilization techniques or sterile filtered. The resulting aqueous solutions can be packaged for use as is or lyophilized, with the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances required to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, and the like, for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.
[0207] Neoantigens can also be administered via liposomes, which target them to specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigen to be delivered is incorporated as part of the liposome, either alone or in combination with a molecule that binds to a receptor dominant among lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Liposomes filled with the desired neoantigen can thus be directed to the site of lymphoid cells, where they then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid lability, and stability of the liposomes in the bloodstream. Various methods are available for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.
[0208] For targeting to immune cells, the ligand to be incorporated into the liposome can include, for example, an antibody or fragment thereof specific for a cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at doses that vary according to, inter alia, the mode of administration, the peptide being delivered, and the stage of the disease being treated.
[0209] The peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient for therapeutic or immunization purposes. Numerous methods are conveniently used to deliver nucleic acids to a patient. For example, nucleic acids can be delivered directly as "naked DNA." This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), and U.S. Pat. Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery, as described, for example, in U.S. Pat. No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, DNA can be attached to particles, such as gold particles. Approaches for delivering nucleic acid sequences include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.
[0210] Nucleic acids can also be delivered by complexing them with cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO 96 / 18372; 9324640 WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833; 9106309 WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).
[0211] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter can also be included in a viral vector-based vaccine platform, such as the ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20 ( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0212] A means of administering nucleic acids uses minigene constructs encoding one or more epitopes. To generate DNA sequences (minigenes) encoding selected CTL epitopes for expression in human cells, the amino acid sequences of the epitopes are reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are then directly adjacent to generate a continuous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the minigene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.
[0213] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) compounds can also be complexed with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.
[0214] Also disclosed herein is a method of producing a tumor vaccine, comprising performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising multiple neoantigens or a subset of multiple neoantigens.
[0215] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector (e.g., a vector comprising at least one sequence encoding one or more neoantigens) disclosed herein can include culturing host cells under conditions suitable for expression of the neoantigen or vector, wherein the host cells comprise at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatographic, electrophoretic, immunological, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.
[0216] The host cell can comprise a Chinese hamster ovary (CHO) cell, an NS0 cell, yeast, or an HEK293 cell. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to the at least one nucleic acid sequence encoding the neoantigen or vector. In certain embodiments, the isolated polynucleotide can be a cDNA.
[0217] VI. Identification of neoantigens VI.A. Identification of Candidate Neoantigens A research method for NGS analysis of tumor and normal exomes and transcriptomes is described and applied in the specific space of neoantigens. 6,14,15The examples below consider certain optimizations for greater sensitivity and specificity for identifying neoantigens in a clinical setting. These optimizations can be grouped into two areas: those related to laboratory processes and those related to NGS data analysis.
[0218] VI.A.1. Laboratory Process Optimization The process improvements presented herein build on the concepts developed for reliable assessment of cancer driver genes in targeted cancer panels. 16 This addresses the challenges in high-precision neoantigen discovery from low tumor content and small volume clinical specimens by expanding the current approach to the whole-exome and whole-transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) unique average coverage across the tumor exome to detect mutations present at low mutant allele frequency due to either low tumor content or subclonal status. 2. Fewer than 5% of bases are covered at less than 100x to minimize missed potential neoantigens, e.g. a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for areas that are not sufficiently covered 3. Targeting uniform coverage across the normal exome, with less than 5% of bases covered below 20x, to minimize the chance of potential neoantigens remaining unclassified for somatic / germline status (and therefore unusable as TSNAs). 4. To minimize the total amount of sequencing required, sequence capture probes are designed only for the coding regions of genes, since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing 18 . b. Elimination of genes predicted to produce few or no candidate neoantigens due to factors such as poor expression, suboptimal digestion by the proteasome, or atypical sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples can be subjected to probe-based enrichment with the same or similar probes used to capture the exome in DNA. 19 It is extracted using
[0219] VI.A.2. Optimizing NGS Data Analysis Analytical method improvements address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically allow for customization relevant for identifying neoantigens in the clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment, as it contains multiple MHC region assemblies that better reflect population polymorphism, as opposed to earlier genome releases. 2. Various programs 5 Overcoming the limitations of single mutation callers 20 by merging results from a. Single nucleotide mutations and indels are detected in tumor DNA, tumor RNA, and normal DNA with a range of tools including: Strelka 21 and Mutec t22 and programs based on comparison of tumor and normal DNA, such as; and 23 , UNCeqR, and other programs that incorporate tumor DNA, tumor RNA, and normal DNA. b. Indels are found in Strelka and ABRA 24 This is determined by a program that performs local reassembly, such as c. Structural rearrangements are 25 or Breakseq 26It is determined using specialized tools such as 3. To detect and prevent sample swapping, mutation calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. Extensive filtering of artificial calls is performed, for example, by: a. Removal of mutations found in normal DNA, potentially with relaxed detection parameters in the case of low coverage and permissive proximity criteria in the case of indels. b. Removal of mutations due to poor mapping quality or poor base quality 27 . c. Elimination of mutations resulting from re-emerging sequencing artifacts, even if not observed in the corresponding normal 27 Examples include mutations that are detected primarily on one strand. d. Removal of mutations detected in a set of unrelated controls 27 . 5.seq2HLA 28 , ATHLATES 29 or Optitype, and also combine exome and RNA sequencing data 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing, such as long-read DNA sequencing. 30 or adaptation of methods for linking RNA fragments to maintain continuity. 31 Includes. 6. Robust Detection of Nascent ORFs Arising from Tumor-specific Splice Variants in CLASS 32 , Bayesembler 33 , StringTie 34 This is done by assembling transcripts from RNA-seq data using Cufflinks, or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate the entire transcripts from each experiment). 35Although commonly used for this purpose, it frequently produces an incredibly large number of splice variants, many of which are much shorter than the full-length gene, and may not be able to recover a simple positive control. The coding sequence and potential nonsense-mediated decay mechanisms reintroduced the variant sequence, SpliceR. 36 and MAMBA 37 Gene expression is determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels are determined using ASE. 38 or HTSeq 39 Potential filtering steps include: a. Removal of candidate nascent ORFs that are thought to be poorly expressed. b. Removal of candidate nascent ORFs predicted to trigger nonsense-mediated decay (NMD). 7. Candidate neoantigens observed only in RNA (e.g., neo-ORFs) that cannot be directly validated as tumor-specific are classified as likely to be tumor-specific according to additional parameters, for example, by considering the following: a. Presence of supporting cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmed trans-acting mutations in splicing factors in tumor DNA only. As an example, the gene that exhibited the most differential splicing in three independently published experiments with R625 mutant SF3B1 was 1. One experiment examined patients with uveal melanoma. 40 The second experiment examined uveal melanoma cell lines. 41 , and a third study looked at breast cancer patients. 42 Nevertheless, there was agreement. c. For novel splicing isoforms, the presence of confirmatory "novel" splice-junction reads in the RNASeq data. d. For de novo rearrangements, the presence of confirmatory exon-proximal reads in tumor DNA that are not present in normal DNA. e.GTEx 43 and absence from the gene expression compendium (i.e., making germline origin less likely). 8. Complementing reference genome alignment-based analyses by comparing tumor and normal reads (or k-mers derived from such reads) of assembled DNA to directly avoid alignment- and annotation-based errors and artifacts (e.g., for somatic mutations occurring near germline mutations or repeat-context indels).
[0220] In samples with polyadenylated RNA, the presence of viral and microbial RNA in the RNA-seq data will be assessed using RNA CoMPASS44 or similar methods to identify additional factors that may predict patient response.
[0221] VI.B. HLA Peptide Isolation and Detection Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) techniques after lysis and solubilization of tissue samples. 55~58 The clarified lysates were used for HLA-specific IP.
[0222] Immunoprecipitation was performed using antibodies coupled to beads, where the antibodies are specific for HLA molecules. For pan-class I HLA immunoprecipitation, a pan-class I CR antibody is used, and for class II HLA-DR, an HLA-DR antibody is used. The antibodies are covalently attached to NHS-Sepharose beads during overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60 Immunoprecipitation can also be performed using antibodies that are not covalently attached to beads. Typically, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to retain the antibody on the column. Some antibodies that can be used to selectively enrich MHC / peptide complexes are listed below. TIFF2025114587000002.tif42149
[0223] The clarified tissue lysate is added to antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is saved for further experiments, including additional IPs. The IP beads are washed to remove nonspecific binding, and the HLA / peptide complexes are eluted from the beads using standard techniques. Protein components are removed from the peptides using molecular weight spin columns or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and, in some cases, stored at -20°C prior to MS analysis.
[0224] The dried peptides were reconstituted in an HPLC buffer suitable for reversed-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution on a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of peptide mass / charge (m / z) were collected at high resolution on an Orbitrap detector, followed by MS2 low-resolution scans on an ion trap detector after HCD fragmentation of selected ions. Additionally, MS2 spectra can be acquired using either CID or ETD fragmentation methods, or any combination of the three techniques to obtain greater amino acid coverage of the peptide. MS2 spectra can also be measured with high-resolution mass accuracy on an Orbitrap detector.
[0225] The MS2 spectra from each analysis were analyzed using Comet 61、62 and peptide identifications were analyzed using Percolator 63~65 Further sequencing is performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or spectral matching and de novo sequencing are performed. 75 Sequencing methods including:
[0226] VI.B.1. Investigation of MS detection limits for comprehensive HLA peptide sequencing Using the peptide YVYVADVAAK (SEQ ID NO: 1), the limit of detection was determined using various amounts of peptide loaded onto the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol (Table 1). The results are shown in Figure 1F. These results indicate that the lowest limit of detection (LoD) was in the attomolar range (10 -18 ), a dynamic range spanning five orders of magnitude, and a signal-to-noise ratio in the low femtomole range (10 -15 ) appears to be sufficient for sequencing. TIFF2025114587000003.tif56128
[0227] VII. Presented Model VII.A. System Overview 2A is an overview of an environment 100 for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for implementing a presentation identification system 160, which itself includes a presentation information store 165.
[0228] The presentation identification system 160 is a computer model, embodied in a computational system such as that discussed below with respect to FIG. 30, that receives a peptide sequence associated with a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the set of associated MHC alleles. The presentation identification system 160 can be applied to both class I and class II MHC alleles, making it useful in a variety of contexts. One specific example application of the presentation identification system 160 is to receive the nucleotide sequence of a candidate neoantigen associated with a set of MHC alleles from tumor cells of a patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the tumor's associated MHC alleles and / or induce an immunogenic response in the patient's 110 immune system. Those candidate neoantigens with a high likelihood, as determined by the system 160, can be selected for inclusion in a vaccine 118, such that an anti-tumor immune response can be elicited from the immune system of the patient 110 that provided the tumor cells. Furthermore, T cells with TCRs that have reactivity to candidate neoantigens with a high likelihood of presentation can be generated for use in T cell therapy, which also elicits an anti-tumor immune response from the patient's 110 immune system.
[0229] The presentation identification system 160 determines presentation likelihoods through one or more presentation models. Specifically, the presentation models generate likelihoods of whether a given peptide sequence will be presented for a set of associated MHC alleles, and the likelihoods are generated based on the presentation information stored in the storage device 165. For example, the presentation model may generate likelihoods of whether the peptide sequence "YVYVADVAAK (SEQ ID NO: 1)" will be presented for the set of alleles HLA-A*02:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:03, and HLA-C*01:04 on the cell surface of a sample. The presentation information 165 contains information about whether these peptides bind to various types of MHC alleles so that the peptides are presented by the MHC alleles, which is determined in the model according to the position of the amino acid in the peptide sequence. Based on the presentation information 165, the presentation model can predict whether an unrecognized peptide sequence will be presented in association with the associated set of MHC alleles. As noted above, the presentation model can be applied to both class I and class II MHC alleles.
[0230] VII.B. Presentation information 2 illustrates a method for obtaining presentation information according to one embodiment. Presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. Allele interaction information includes information that affects presentation of peptide sequences that is dependent on the type of MHC allele. Allele non-interaction information includes information that affects presentation of peptide sequences that is independent of the type of MHC allele.
[0231] VII.B.1. Allelic Interaction Information Allele interaction information primarily includes identified peptide sequences known to be presented by one or more identified MHC molecules from humans, mice, etc. Of note, this may or may not include data obtained from tumor samples. Presented peptide sequences may be identified from cells expressing a single MHC allele. In this example, presented peptide sequences are generally collected from a single-allelic cell line engineered to express a predetermined MHC allele and then exposed to a synthetic protein. Peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows this example, in which an exemplary peptide, YEMFNDKSQRAPDDKMF (SEQ ID NO: 2), presented on the predetermined MHC allele HLA-DRB1*12:01, is isolated and identified by mass spectrometry. In this situation, because the peptide is identified through cells engineered to express a single predetermined MHC protein, the direct association between the presented peptide and the MHC protein to which it binds is definitively known.
[0232] Presented peptide sequences may also be collected from cells expressing multiple MHC alleles. Typically, in humans, six different types of MHC1 molecules and up to 12 different types of MHCII molecules are expressed by cells. Such presented peptide sequences may be identified from multi-allelic cell lines engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal or tumor tissue samples. In this particular example, MHC molecules can be immunoprecipitated from normal or tumor tissue. Peptides presented on multiple MHC alleles can also be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows this example in which six exemplary peptides, YEMFNDKSF (SEQ ID NO:3), HROEIFSHDFJ (SEQ ID NO:4), FJIEJFOESS (SEQ ID NO:5), NEIOREIREI (SEQ ID NO:6), JFKSIFEMMSJDSSUIFLKSJFIEIFJ (SEQ ID NO:7), and KNFLENFIESOFI (SEQ ID NO:8), are presented on identified class I MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, HLA-B*08:01, and class II MHC alleles HLA-DRB1*10:01, HLA-DRB1:11:01, isolated, and identified by mass spectrometry. In contrast to monoallelic cell lines, because bound peptides are isolated from MHC molecules prior to identification, the direct association between the presented peptide and the MHC protein to which it bound may be unknown.
[0233] Allele interaction information can also include mass spectrometry ion currents, which depend on both the concentration of peptide-MHC molecule complexes and the ionization efficiency of the peptides. Ionization efficiency varies from peptide to peptide in a sequence-dependent manner. Generally, ionization efficiency varies from peptide to peptide over approximately two orders of magnitude, while the concentration of peptide-MHC complexes varies over an even larger range.
[0234] Allele interaction information can also include measured or predicted binding affinities between a given MHC allele and a given peptide. One or more affinity models can generate such predictions (72, 73, 74). For example, returning to the example shown in Figure 1D, representation 165 may represent a binding affinity between the peptide YEMFNDKSF (SEQ ID NO: 3) and the class I allele HLA-A * The presentation information 165 may include a predicted binding affinity value of 1000 nM between the peptide KNFLENFIESOFI (SEQ ID NO: 8) and the class II allele HLA-DRB1:11:01. Few peptides with IC50 > 1000 nM are presented by the MHC, with lower IC50 values increasing the probability of presentation. The presentation information 165 may include a predicted binding affinity value between the peptide KNFLENFIESOFI (SEQ ID NO: 8) and the class II allele HLA-DRB1:11:01.
[0235] The allele interaction information can also include measured or predicted stability values for MHC complexes. One or more stability models can generate such predictions. More stable peptide-MHC complexes (i.e., complexes with longer half-lives) are more likely to be presented in high copy number on tumor cells and on antigen-presenting cells that encounter vaccine antigens. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include a predicted stability value for a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a predicted stability value for the half-life of the class II molecule HLA-DRB1:11:01.
[0236] Allele interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at faster rates are more likely to be presented at high concentrations on the cell surface.
[0237] Allele interaction information can also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of presented peptides have a length of 9 peptides. MHC class II molecules generally tend to present peptides with a length of 6-30 peptides.
[0238] Allele interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoded peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoded peptide. The presence of kinase motifs influences the probability of post-translational modifications that may enhance or interfere with MHC binding.
[0239] Allelic interaction information can also include expression or activity levels of proteins involved in post-translational modification processes, such as kinases (as measured or predicted by RNA-seq, mass spectrometry, or other methods).
[0240] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.
[0241] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual in question (e.g., as measured by RNA-seq or mass spectrometry): peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.
[0242] Allelic interaction information can also include the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals that express that particular MHC allele.
[0243] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and therefore, presentation of peptides by HLA-C is a priori less likely than presentation by HLA-A or HLA-B II. As another example, because HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, presentation of peptides by HLA-DP is predicted to be less likely than presentation by HLA-DR or HLA-DQ.
[0244] The allele interaction information can also include the protein sequence of a particular MHC allele.
[0245] Any of the MHC allele non-interacting information listed in the section below can also be modeled as MHC allele interacting information.
[0246] VII.B.2. Allelic Non-Interaction Information The allele-non-interacting information can include the C-terminal sequence adjacent to the neoantigen-encoded peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence (adjacent sequence) can affect the proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters an MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and therefore, the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in Figure 2C, the presentation information 165 can include the C-terminal flanking sequence FOEIFNDKSLDKFJI (SEQ ID NO: 9) of the presented peptide FJIEJFOESS (SEQ ID NO: 5), identified from the peptide's source protein.
[0247] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide mass spectrometry training data. As described later with respect to Figure 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, mRNA quantification measurements are determined from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).
[0248] The allele non-interacting information can also include N-terminal sequences adjacent to the peptide within its source protein sequence.
[0249] The allelic non-interaction information can also include a source gene for the peptide sequence. The source gene can be defined as an Ensembl protein family for the peptide sequence. In another example, the source gene can be defined as a source DNA or source RNA for the peptide sequence. The source gene can be represented, for example, as a string of nucleotides that encodes a protein, or alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode specific proteins. In another example, the allelic non-interaction information can also include a source transcript or isoform or a set of potential source transcripts or isoforms for the peptide sequence extracted from a database such as Ensembl or RefSeq.
[0250] The allelic non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.
[0251] The allele non-interaction information can also include the presence of protease cleavage motifs in peptides, optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and therefore less stable in cells, and therefore less likely to be presented.
[0252] Allelic non-interaction information can also include the turnover rate of the source protein when measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation, but this characteristic has low predictive power when measured in dissimilar cell types.
[0253] The allelic non-interaction information can also include the length of the source protein, optionally taking into account the specific splice variants ("isoforms") that are most highly expressed in tumor cells, as measured by RNA-seq or proteome mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.
[0254] Allele-free interaction information can also include the expression level of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Different proteasomes have different cleavage site preferences. More weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.
[0255] Allele-free interaction information can also include the expression of the peptide's source gene (e.g., as measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in the tumor sample. Peptides from genes with higher expression are more likely to be presented. Peptides from genes with undetectable levels of expression can be eliminated from consideration.
[0256] Allelic non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to nonsense-mediated decay as predicted by a model of nonsense-mediated decay, e.g., the model from Rivas et al., Science 2015.
[0257] Allelic non-interaction information can also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle. Genes that are expressed at low levels overall (as measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are displayed than genes that are stably expressed at very low levels.
[0258] Allelic non-interaction information can also include a comprehensive catalog of source protein properties, such as those provided in uniProt or the PDB (http: / / www.rcsb.org / pdb / home / home.do). These properties can include, among others, protein secondary and tertiary structure, subcellular localization, and Gene Ontology (GO) terms. Specifically, this information can include annotations operating at the protein level, e.g., 5'UTR length, and annotations operating at the level of specific residues, e.g., a helix motif between residues 300 and 310. These properties can also include turn motifs, sheet motifs, and disordered residues.
[0259] Allelic non-interacting information can also include features that describe the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (eg, alpha helix versus beta sheet); alternative splicing.
[0260] The allelic non-interaction information can also include properties that describe the presence or absence of presentation hotspots at the peptide's position in its source protein.
[0261] Allelic non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression level of the source protein in those individuals and the influence of the various HLA types of those individuals).
[0262] Allelic non-interaction information can also include the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias.
[0263] Expression of various gene modules / pathways (not necessarily containing the source protein of the peptides) as measured by gene expression assays such as RNASeq, microarrays, targeted panels such as Nanostring, or single / multiple genes representing gene modules measured by assays such as RT-PCR, that inform on the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs).
[0264] Allele non-interaction information can also include the copy number of the peptide's source gene in the tumor cell. For example, a peptide derived from a gene that is subject to homozygous deletion in the tumor cell can be assigned a presentation probability of zero.
[0265] The allele-non-interaction information can also include the probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP or that bind with higher affinity to TAP are more likely to be presented by MHC-I.
[0266] Allelic non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Higher TAP expression levels at MHC-I increase the probability of presentation of all peptides.
[0267] Allelic non-interaction information can also include the presence or absence of tumor mutations, including but not limited to: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3. ii. In genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery affected by loss-of-function mutations in the tumor have a reduced probability of presentation.
[0268] Presence or absence of functional germline polymorphisms, including but not limited to: i. In genes encoding proteins involved in the antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome).
[0269] The allelic non-interaction information can also include tumor type (eg, NSCLC, melanoma).
[0270] Allele non-interaction information can also include the known functionality of the HLA allele, e.g., as reflected by the HLA allele suffix. For example, the N suffix in the allele name HLA-A*24:09N indicates a null allele that is not expressed and therefore unlikely to present an epitope; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.
[0271] Allelic non-interaction information can also include clinical tumor subtype (eg, squamous cell lung cancer vs. non-squamous).
[0272] The allele non-interaction information can also include smoking history.
[0273] Allele non-interaction information can also include a history of sunburn, sun exposure, or exposure to other mutagens.
[0274] The allelic non-interaction information can also include regional expression of the peptide's source gene in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be represented.
[0275] The allelic non-interaction information can also include the frequency of the mutation in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.
[0276] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation can also include the mutation's annotation (e.g., missense, readthrough, frameshift, fusion, etc.) or whether the mutation is predicted to result in nonsense-mediated decay (NMD). For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in reduced mRNA translation, which reduces the probability of presentation.
[0277] VII.C. Presentation Identification System 3 is a high-level block diagram illustrating the computer logic components of presentation identification system 160, according to one embodiment. In this exemplary embodiment, presentation identification system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. Presentation identification system 160 also includes a training data store 170 and a presentation model store 175. Some embodiments of model management system 160 have different modules than those described herein. Likewise, functionality may be distributed among the modules in a manner different from that described herein.
[0278] VII.C.1. Data Management Module The data management module 312 generates sets of training data 170 from the representation information 165. Each training data set contains a number of data examples, each of which contains at least one of the represented or unrepresented peptide sequences p i and the peptide sequence p i one or more relevant MHC alleles combined with i and the dependent variable y, which represents information that the presentation identification system 160 is interested in predicting new values of the independent variables. i and the independent variable z i Contains a set of
[0279] In one particular implementation referred to throughout the remainder of this specification, the dependent variable y i is the peptide p i but one or more associated MHC alleles a i However, in other implementations, the dependent variable y i is the result of the presentation identification system 160 determining the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that one is interested in predicting. For example, in another implementation, the dependent variable y i σ may also be a numerical value indicating the mass analysis ion current determined for the example data.
[0280] Peptide sequence p for data example i i is k i is a sequence of k amino acids, i can vary within a range among data instances i. For example, the range can be 8 to 15 for MHC class I, or 6 to 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p in the training data set are i may have the same length, e.g., 9. The number of amino acids in a peptide sequence may vary depending on the type of MHC allele (e.g., MHC allele in humans). MHC allele a for data example i iis the peptide sequence p i indicates whether it existed in combination with
[0281] The data management module 312 also manages the peptide sequences p contained in the training data 170. i and bound MHC allele a i Together with the binding affinity b i and stability i For example, the training data 170 may include a predictor of the peptide p i and a i The predicted binding affinity b between each of the bound MHC molecules shown in i As another example, the training data 170 may contain a i The predicted stability value s for each of the MHC alleles shown in i may contain
[0282] The data management module 312 also receives the peptide sequence p i along with non-allele interacting variables such as C-terminal flanking sequences and mRNA quantification measurements. i It may also include.
[0283] The data management module 312 also identifies peptide sequences not presented by MHC alleles to generate the training data 170. Generally, this involves identifying a "longer" sequence of the source protein that contains the peptide sequence to be presented prior to presentation. If the presentation information contains an engineered cell line, the data management module 312 identifies a set of peptide sequences in the synthetic protein exposed to the cell that were not presented on the MHC alleles of the cell. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein from which the presented peptide sequences originated and identifies a set of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.
[0284] The data management module 312 also artificially generates peptides with random sequences of amino acids and identifies the generated sequences as peptides that are not presented on MHC alleles. This can be achieved by randomly generating peptide sequences, allowing the data management module 312 to easily generate large amounts of synthetic data for peptides that are not presented on MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, synthetically generated peptide sequences are very likely not presented by MHC alleles, even if they are included in proteins processed by cells.
[0285] 4 illustrates an exemplary set of training data 170A according to one embodiment. Specifically, the first three data examples in training data 170A show peptide presentation information from a monoallelic cell line containing the allele HLA-C*01:03 and three peptide sequences QCEIOWAREFLKEIGJ (SEQ ID NO: 10), FIEUHFWI (SEQ ID NO: 11), and FEWRHRJTRUJR (SEQ ID NO: 12). The fourth data example in training data 170A shows peptide presentation information from a multiallelic cell line containing the alleles HLA-B*07:02, HLA-C*01:03, and HLA-A*01:01 and the peptide sequence QIEJOEIJE (SEQ ID NO: 13). The first data example shows that the peptide sequence QCEIOWARE (SEQ ID NO: 14) was not presented by the allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, the negatively labeled peptide sequences may be randomly generated by the data management module 312 or may be identified from the source protein of the peptides being represented. The training data 170A also includes predicted binding affinities of 1000 nM and predicted stability values with half-lives of 1 hour for the peptide sequence-allele pairs. The training data 170A also includes the C-terminal flanking sequence of the peptide FJELFISBOSJFIE (SEQ ID NO: 15), and the C-terminal flanking sequence of the peptide FJELFISBOSJFIE (SEQ ID NO: 16). 2It also includes allele-non-interacting variables, such as mRNA quantification measurements of TPM. A fourth data example shows that the peptide sequence QIEJOEIJE (SEQ ID NO: 13) was presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. Training data 170A also includes predicted binding affinity and stability values for each of the alleles, as well as the C-terminal flanking sequence of the peptide and mRNA quantification measurements for the peptide.
[0286] VII.C.2. Coding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more representation models. In one implementation, the encoding module 314 one-hot encodes sequences (e.g., peptide sequences or C-terminal flanking sequences) for a predetermined 20-letter amino acid alphabet. Specifically, k i Peptide sequence p having amino acids i is 20·k i p, which is represented as a row vector of elements, corresponding to the alphabet of the amino acid at the j-th position of the peptide sequence. i 20·(j-1)+1 , p i 20·(j-1)+2 , …, p i 20·j A single element in has a value of 1. The remaining elements have a value of 0. As an example, for a given alphabet {A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y}, the three amino acid peptide sequence EAF of data instance i is represented as a 60-element row vector p i =[0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0]. The C-terminal flanking sequence c i , and the protein sequence for the MHC allele d h, and other sequence data in the presentation information can be similarly coded as above.
[0287] If the training data 170 contains sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equivalent length by adding PAD characters to extend the predetermined alphabet. For example, this may be done by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the peptide sequence with the longest length in the training data 170. Thus, if the peptide sequence with the longest length is k 最大 amino acids, the encoding module 314 encodes each sequence as (20+1) k 最大 It is represented numerically as a row vector of elements. For example, consider the extended alphabet {PAD,A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y} and k 最大 For a maximum amino acid length of 5, the same exemplary peptide sequence EAF of 3 amino acids is represented as a 105-element row vector p i =[1 0 ... i or other sequence data can be similarly encoded as above. Thus, the peptide sequence p i or c i Each argument or row in represents the occurrence of a particular amino acid at a particular position in the sequence.
[0288] Although the above method for encoding sequence data has been described with respect to sequences having amino acid sequences, the method can be similarly extended to other types of sequence data, such as, for example, DNA or RNA sequence data.
[0289] The encoding module 314 also encodes one or more MHC alleles a for data instance i. i is encoded into an m-element row vector, where each element h=1, 2, ..., m corresponds to a uniquely identified MHC allele. The element corresponding to the MHC allele identified for data instance i has a value of 1. The remaining elements have a value of 0. As an example, among the m=4 uniquely identified MHC allele types {HLA-A*01:01, HLA-C*01:08, HLA-B*07:02, HLA-DRB1*10:01}, the alleles HLA-B*07:02 and HLA-DRB1*10:01 for data instance i corresponding to a multi-allelic cell line are encoded into a four-element row vector a i = [0 0 1 1], and a3 i =1 and a4 i = 1. An example with four identified MHC allele types is described herein, but the number of MHC allele types can actually be hundreds or thousands. As noted above, each data instance i typically contains a peptide sequence p i Up to six different MHC allele class I types and / or peptide sequences p i Up to four different MHC class II DR allele types and / or peptide sequences p i Contains up to 12 different MHC class II allele types associated with
[0290] The encoding module 314 also generates a label y for each data instance i. i We code, as a binary variable with values from the set {0,1}, where a value of 1 indicates that the peptide x i However, the associated MHC allele a i a value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele ai The dependent variable y i If σ represents the mass analysis ion current, the encoding module 314 may additionally scale the value using various functions, such as a log function with a range of [-∞,∞] for ion current values between [0,∞].
[0291] The coding module 314 encodes the peptide p i and the allele interaction variable x for the associated MHC allele h h i pairs as row vectors in which the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], and b h i is the peptide p i and the predicted binding affinity for the associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).
[0292] In one example, the encoding module 314 encodes the measured or predicted values for binding affinity as a function of the allele interaction variable x h i The binding affinity information is expressed by incorporating the
[0293] In one example, the encoding module 314 encodes the measured or predicted values for binding stability as allele interaction variables x hi By incorporating the information into the
[0294] In one example, the encoding module 314 encodes the measured or predicted values for the binding on-rate as a function of the allele interaction variable x h i The combined on-rate information is expressed by incorporating
[0295] In one example, for peptides presented by class I MHC molecules, the encoding module 314 encodes the peptide length in the vector TIFF2025114587000004.tif4128 (However, TIFF2025114587000005.tif3128 is an index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i In another example, for peptides presented by class II MHC molecules, the encoding module 314 may include the peptide length in the vector TIFF2025114587000006.tif18146 (However, TIFF2025114587000007.tif3128 is an index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i can be included in
[0296] In one example, the encoding module 314 encodes the RNA-seq-based expression levels of the MHC alleles as a function of the allele interaction variable x h i By incorporating this information into the MHC alleles, RNA expression information is displayed.
[0297] Similarly, the encoding module 314 encodes the allele non-interacting variable w ican be represented as a row vector in which the numerical representations of the allele-non-interacting variables are concatenated one after the other. For example, w i is [c i ] or [c i m i w i ], and w i is the peptide p i mRNA quantitative measurements related to the C-terminal flanking sequence and peptide of m i Alternatively, one or more combinations of allele non-interacting variables may be stored individually (e.g., as individual vectors or matrices).
[0298] In one example, the encoding module 314 encodes the turnover rate or half-life as a function of the allele non-interacting variable w i represents the turnover rate of the source protein for the peptide sequence.
[0299] In one example, the encoding module 314 encodes the protein length as a function of the allele non-interacting variable w i represents the length of the source protein or isoform by incorporating
[0300] In one example, the encoding module 314 generates β1 i , β2 i , β5 i The mean expression of immunoproteasome-specific proteasome subunits, including the subunits, was calculated using the allele-noninteracting variable w i Incorporation into the IL-1 protein results in activation of the immunoproteasome.
[0301] In one example, the encoding module 314 encodes the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM) or a source protein of a gene or transcript of the peptide, by correlating the source protein abundance with an allele-non-interacting variable w i This is expressed by incorporating it into
[0302] In one example, the encoding module 314 calculates the probability that the transcript of the peptide's origin will undergo nonsense-mediated decay (NMD), for example, as estimated by the model in Rivas et al. Science, 2015, and calculates this probability as a function of the allele non-interaction variable w i This is expressed by incorporating it into
[0303] In one example, encoding module 314 represents the activation status of a gene module or pathway assessed via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM, for each of the genes in the pathway, and then computing a summary statistic, such as a mean, across the genes in the pathway. The mean is calculated using the allele-non-interaction variable w i can be incorporated into.
[0304] In one example, the encoding module 314 encodes the copy number of the source gene by dividing the copy number by the allele non-interacting variable w i This is expressed by incorporating it into
[0305] In one example, the encoding module 314 encodes the measured or predicted TAP binding affinity (e.g., in nanomolar units) relative to the allele-non-interacting variable w i The TAP binding affinity is expressed by including
[0306] In one example, the encoding module 314 encodes TAP expression levels measured by RNA-seq (and quantified, for example, by RSEM in units of TPM) as a function of the allele-non-interacting variable w i The expression level of TAP is represented by the inclusion of
[0307] In one example, the encoding module 314 encodes the tumor mutations as allele-non-interacting variables w i vector of indicator variables in (i.e., peptide pk is derived from a sample with a KRAS G12D mutation, d k = 1, otherwise 0).
[0308] In one example, the encoding module 314 encodes germline polymorphisms in antigen-presenting genes as a vector of indicator variables (i.e., peptide p k If is derived from a sample with a specific germline polymorphism in TAP, then d k = 1). These indicator variables are expressed as the allele non-interaction variables w i can be included in
[0309] In one example, the encoding module 314 represents tumor types as one-hot coded vectors of length 1 for an alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot coded variables are combined into an allelic non-interaction variable w i can be included in
[0310] In one example, the encoding module 314 represents MHC allele suffixes by processing four-digit HLA alleles with various suffixes. For example, HLA-A*24:09N is considered a different allele from HLA-A*24:09 for purposes of the model. Alternatively, because HLA alleles ending in an N suffix are not expressed, the probability of presentation by an MHC allele with an N suffix can be set to zero for all peptides.
[0311] In one example, the encoding module 314 represents tumor subtypes as one-hot coded vectors of length 1 for an alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot coded variables are combined into an allelic non-interaction variable w i can be included in
[0312] In one example, the encoding module 314 encodes smoking history as a function of the allele-non-interacting variable w iA binary indicator variable (d if the patient has a smoking history) can be included in k = 1 otherwise 0). Alternatively, smoking history can be coded as a one-hot coded variable of length 1 for the alphabet of smoking severity. For example, smoking status can be assessed on a scale of 1 to 5, with 1 indicating non-smoker and 5 indicating current heavy smoker. Because smoking history is primarily associated with lung tumors, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of smoking and the tumor type is lung tumor, and zero otherwise.
[0313] In one example, the encoding module 314 encodes sunburn history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a history of severe sunburn) can be included in k = 1 otherwise 0). Because severe sunburn is primarily associated with melanoma, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and zero otherwise.
[0314] In one example, the encoding module 314 represents the distribution of expression levels of a particular gene or transcript for each gene or transcript in the human genome as a summary statistic (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, the expression of peptide p in samples with the tumor type melanoma is expressed as a summary statistic (e.g., mean, median) of the distribution of expression levels. k Regarding peptide p k The measured gene or transcript expression levels of the genes or transcripts of origin are compared with the allele-non-interacting variable w i Not only can it be included in the peptide p in melanoma as measured by TCGA, k The mean and / or median gene or transcript expression of the genes or transcripts of a given source may also be included.
[0315] In one example, the encoding module 314 represents the variant types as one-hot coded variables of length 1 for an alphabet of variant types (e.g., missense, frameshift, NMD-induced, etc.). These one-hot coded variables are combined into an allele-non-interacting variable w i can be included in
[0316] In one example, the encoding module 314 encodes the protein-level characteristics of the protein as values of the source protein annotation (e.g., 5′ UTR length) and the allele-non-interacting variable w i In another example, the encoding module 314 encodes the peptide p i The residue-level annotation of the source protein for peptide p i is equal to 1 if overlaps with the helical motif, otherwise it is equal to 0, or i The indicator variable that is equal to 1 if is completely contained within the helical motif is the allele non-interaction variable w i In another example, the peptide p i The property that represents the proportion of residues in i can be included in
[0317] In one example, the encoding module 314 encodes the types of proteins or isoforms in the human proteome into an index vector o having a length equivalent to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is the peptide p k is 1 if comes from protein i, and 0 otherwise.
[0318] In one example, the encoding module 314 encodes the peptide p i Source gene G = gene(p i) as a categorical variable with L possible categories (where L denotes the upper bound 1, 2, ..., L on the number of subscripted source genes).
[0319] In one example, the encoding module 314 encodes the peptide p i T = tissue type, cell type, tumor type, or tumor histology type of T = tissue (p i ) as a categorical variable with M possible categories (where M denotes an upper limit on the number of subscripted types 1, 2, ..., M). Tissue types can include, for example, lung tissue, cardiac tissue, intestinal tissue, and neural tissue. Cell types can include, for example, dendritic cells, macrophages, and CD4 T cells. Cancers can include, for example, lung adenocarcinoma, lung squamous cell carcinoma, melanoma, and non-Hodgkin's lymphoma.
[0320] The encoding module 314 also encodes the peptide p i and the variable z for the associated MHC allele h i The entire set of alleles is expressed as the allele interaction variable x i and the allele non-interaction variable w i For example, the encoding module 314 may represent z h i [x h i w i ] or [w i x h i ] can be represented as a row vector equivalent to
[0321] VIII. Training Module The training module 316 constructs one or more presentation models that generate a likelihood of whether a peptide sequence will be presented by an MHC allele associated with the peptide sequence. k and peptide sequence p k MHC alleles associated with a k Given a set of peptide sequences, each proposed model k However, the associated MHC allele ak the estimate u, which indicates the likelihood that one or more of k Generate.
[0322] VIII.A. Overview The training module 316 constructs one or more representation models based on a training data set stored in storage 170, which is generated from the representation information stored in 165. Generally, regardless of the specific type of representation model, all representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function (y i∈S ,u i∈S ;θ) is the dependent variable y for one or more data examples S in the training data 170. i∈S and the estimated likelihood u for the data example S generated by the proposed model. i∈S In one particular implementation mentioned throughout the rest of this document, the loss function (y i∈S ,u i∈S ;θ) is the negative log likelihood function given by equation (1a) as follows: However, in practice, another loss function may be used. For example, if a prediction is made for mass spectrometry ion current, the loss function is the mean square loss given by Equation 1b as follows: TIFF2025114587000009.tif10128
[0323] The proposed model can be a parametric model, where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function (y i∈S ,u i∈SThe various parameters of the proposed parametric model that minimizes θ ( ; θ ) are determined through a gradient-based numerical optimization algorithm, such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the proposed model may be a non-parametric model, in which the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.
[0324] VIII.B. Allele-specific Model The training module 316 may build a presentation model to predict the presentation likelihood of a peptide on an allele-by-allele basis. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele.
[0325] In one implementation, the training module 316: TIFF2025114587000010.tif7128 identifies peptides for specific alleles. k The estimated presentation likelihood u k where the peptide sequence x h k is the peptide p k and the coded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, which for convenience of description will be referred to as a transformation function throughout this specification. h (·) is an arbitrary function, which for convenience of description will be referred to as the dependence function throughout this specification, and the parameter θ determined for the MHC allele h h Based on the set of allele interaction variables x h k Generate a dependency score for the parameter θ for each MHC allele h. h The set of values is θ h where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele h.
[0326] Dependence function g h (x h k ;θ h ) output is the MHC allele h with at least the allele interaction characteristic x h k and in particular the peptide p k The dependency score for MHC allele h indicates whether the corresponding neoantigen is presented based on the amino acid position of the peptide sequence of p. For example, the dependency score for MHC allele h is determined by the relationship between the MHC allele h and the peptide p. k The transformation function f(·) transforms the input, more specifically, g in this example. h (x h k ;θ h ) is used to calculate the dependency score for peptide p k is converted to an appropriate value indicating the likelihood that it will be presented by the MHC allele.
[0327] In one particular implementation referred to throughout the remainder of this specification, f(·) is a function with range in [0,1] for the appropriate domain range. In one example, f(·) is The expit function given by TIFF2025114587000011.tif10128. As another example, f(·) also has the following meaning: for values in the domain z greater than or equal to 0, It can also be the hyperbolic tangent function given by TIFF2025114587000012.tif6128. Alternatively, if the prediction is made for mass analysis ion currents with values outside the range [0,1], f(·) can be any function, for example, the identity function, the exponential function, the log function, etc.
[0328] Therefore, the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is given by the dependency function g for MHC allele h. h (·) is the peptide sequence pk to generate a corresponding dependency score. The dependency score can be generated by applying k may be transformed by a transformation function f(·) to generate the likelihood for each allele that will be presented by MHC allele h.
[0329] VIII.B.1 Dependence Functions for Allelic Interaction Variables In one particular implementation mentioned throughout this specification, the dependency function g h (·) is x h k Each allele interaction variable in is compared with the parameter θ determined for the relevant MHC allele h. h with the corresponding parameters in the set This is an affine function given by TIFF2025114587000013.tif5128.
[0330] In another specific implementation mentioned throughout this specification, the dependency function g h (·) is a network model NN with a set of nodes arranged in one or more layers. h (·), The network function is given by TIFF2025114587000014.tif5128. The nodes are connected by parameters θ h A node may be connected to other nodes through connections, each having an associated parameter in a set of . The value at one particular node may be represented as the sum of the values of the nodes connected to the particular node, weighted by the associated parameter mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and process data having amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how these interactions affect peptide presentation.
[0331] Generally speaking, the network model NN h (·) can be structured as feedforward networks such as artificial neural networks (ANNs), convolutional neural networks (CNNs), deep neural networks (DNNs), and / or recurrent neural networks (RNNs) such as long short-term memory networks (LSTMs), bidirectional LSTM networks, bidirectional recurrent networks, deep bidirectional recurrent networks, and multilayer neural networks (NLPs).
[0332] In one example, which will be mentioned throughout the remainder of this specification, each MHC allele in h=1, 2, ..., m is associated with a separate network model, NN h (·) denotes the output from the network model related to MHC allele h.
[0333] FIG. 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h=3. As shown in FIG. 5, the network model NN3(·) for MHC allele h=3 includes three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2), ..., θ3(10). The network model NN3(·) includes three allele interaction variables x3 for MHC allele h=3. k (1), x3 k (2) and x3 k (3) receives input values (individual data examples, including encoded polypeptide sequence data and any other training data used) and generates the value NN3(x3 k ) The network function may include one or more network models, each taking a different allele interaction variable as input.
[0334] In another example, the identified MHC alleles h=1, 2, ..., m are represented by a single network model NN H(·) and NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θ h may correspond to the set of parameters for a single network model, and thus the parameters θ h The set of can be shared by all MHC alleles.
[0335] FIG. 6A shows an exemplary network model NN shared by MHC alleles h=1, 2, ..., m. H (·). As shown in Figure 6A, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. The network model NN3(·) has an allele interaction variable x3 for MHC allele h=3. k , and the value NN3(x3 k ) and outputs m values.
[0336] In yet another example, a single network model NN H (·) is the allele interaction variable x for MHC allele h h k and the encoded protein sequence d h In such an example, the parameter θ h may again correspond to the set of parameters for a single network model, and thus the parameters θ h The set of MHC alleles can be shared by all MHC alleles. Therefore, in such an example, NNh(·) is a single network model with inputs [x h k d h ], a single network model NN HSuch a network model is advantageous because it can correctly predict peptide presentation probabilities for MHC alleles that were unknown in the training data simply by identifying their protein sequences.
[0337] Figure 6B shows an exemplary network model NN shared by MHC alleles. H (·). As shown in Figure 6B, the network model NN H (·) takes as input the allele interaction variables and protein sequence of MHC allele h=3 and calculates the dependency score NN3 (x3 k ) is output.
[0338] In yet another example, the dependency function g h (·)teeth, TIFF2025114587000015.tif5128, where g' h (x h k ;θ' h ) is an affine function, network function, etc., with a set of parameters θ'h, which represents the baseline probability of presentation for MHC allele h, and the bias parameter θ in the set of parameters for the allele interaction variables of the MHC alleles. h 0 accompanied by.
[0339] In another implementation, the bias parameter θ h 0 may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ for the MHC allele h h 0 is θ 遺伝子(h) 0where gene (h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family "HLA-A," and the bias parameter θ for each of these MHC alleles may be h 0 As another example, if the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 are assigned to the "HLA-DRB" gene family, and the bias parameters θ for each of these MHC alleles are h 0 can be shared.
[0340] As an example, returning to equation (2), the affine dependency function g h Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000016.tif5128, where x3k is the allele interaction variable identified for MHC allele h=3 and θ3 is the set of parameters determined for MHC allele h=3 through loss function minimization.
[0341] As another example, let us consider the MHC allele h = 3 for peptide p among m = 4 different identified MHC alleles using separate network transformation functions gh(·). k The likelihood that will be presented is TIFF2025114587000017.tif5128 can be generated by k is the allele interaction variable identified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.
[0342] Figure 7 shows the correlation coefficients of peptide p associated with MHC allele h=3 using the exemplary network model NN3(·). k As shown in Figure 7, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0343] VIII.B.2. Per allele with allele-noninteracting variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF2025114587000018.tif7128, peptide p k We model the estimated presentation likelihood uk, where w k is the peptide p k means the coded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interacting variable w Based on the set of allele-non-interacting variables w k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values of θ h and θ w where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele.
[0344] Dependence function g w (w k ;θ w ) output is a measure of the peptide p expression by one or more MHC alleles based on the influence of allele-non-interacting variables. k represents a dependency score for an allele-non-interacting variable, indicating whether peptide p kThe C-terminal flanking sequences and peptide p, which are known to positively influence the presentation of k If peptide p is bound, it may have a high value k The C-terminal flanking sequences and peptide p k If bound, it may have a low value.
[0345] According to equation (8), the peptide sequence p k The likelihood that is presented by MHC allele h is given by the function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. w (·) is also applied to the coded version of the allele-non-interacting variable to generate a dependency score for the allele-non-interacting variable. Both scores are combined, and the combined score is used to estimate the association of peptide sequence p with MHC allele h. k is transformed by a transformation function f(·) to produce the likelihood for each allele that will be presented.
[0346] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (2) k the allele interaction variable x h k One may include the allele non-interacting variable wk in the prediction by adding The image can be given by TIFF2025114587000019.tif7128.
[0347] VIII.B.3 Dependence Functions for Allelic Non-Interacting Variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g for allelic non-interacting variables w (·) is an affine function, or a separate network model for the allelic non-interaction variables w kIt can be a network function related to
[0348] Specifically, the dependency function g w (·) is w k The allele non-interaction variables in w with the corresponding parameters in the set This is an affine function given by TIFF2025114587000020.tif6128.
[0349] Dependence function g w (·) also corresponds to the parameter θ w the network model NN with relevant parameters in the set w (·), TIFF2025114587000021.tif6128. The network function may include one or more network models, each taking different allelic non-interacting variables as input.
[0350] In another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF2025114587000022.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of m k is the peptide p k is the mRNA quantitative measurement for , h(·) is a function that transforms the quantitative measurement, and θ w mis a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measurement to generate a dependency score for the mRNA quantification measurement. In one particular embodiment mentioned throughout the remainder of this specification, h(·) is a log function, although in practice h(·) can be any one of a variety of different functions.
[0351] In yet another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF2025114587000023.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of k is the peptide p k is the indicator vector described in Section VII.C.2, which represents proteins and isoforms in the human proteome, and θ w o is the set of parameters in the set of parameters for the allele non-interacting variables that are combined with the indicator vector. k and parameter θ w o If the dimension of the set is significantly higher, λ || θ w o A parameter regularization term such as || (where ||·|| represents the L1 norm, L2 norm, a combination, etc.) can be added to the loss function when determining the parameter value. The optimal value of the hyperparameter λ can be determined through an appropriate method.
[0352] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025114587000024.tif13128However, g' w (wk ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF2025114587000025.tif5128 is peptide p k is an indicator function equal to 1 if θ is derived from the source gene l as described above for allele-non-interacting variables, and θ w l is a parameter that indicates the "antigenicity" of the source gene l. In one variation, L is sufficiently large, and therefore the number of parameters θ w l=1, 2,...,L When is sufficiently large, λ ||θ w l A parameter regularization term such as || (where ||·|| is an L1 norm, L2 norm, or a combination thereof) can be added to the loss function when determining the parameter value. The optimal value of the hyperparameter λ can be determined by an appropriate method.
[0353] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025114587000026.tif13147However, g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF2025114587000027.tif5128 is a peptide p k is derived from source gene l, and peptide p k is an indicator function that is equal to 1 if originates from tissue type m, and θ w lmis a parameter indicating the antigenicity of the combination of source gene l and tissue type m. Specifically, the antigenicity of gene l in tissue type m may indicate the residual tendency of cells of tissue type m to present peptides derived from gene l after adjustment for RNA expression and peptide sequence context.
[0354] In one variation, L or M is sufficiently large, so that the number of parameters θ w lm=1, 2,...,LM When is sufficiently large, λ ||θ w lm A parameter regularization term such as || (where ||·|| is an L1 norm, an L2 norm, a combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the parameter values so that the coefficients for the same source gene do not differ significantly between tissue types. For example, a penalty term such as: TIFF2025114587000028.tif18128 (However, TIFF2025114587000029.tif5128 is the average antigenicity across tissue types for source gene l) can add a penalty to the standard deviation of antigenicity across different tissue types in the loss function.
[0355] In practice, the dependence function g on the allelic non-interacting variables can be calculated by combining any of the additional terms in equations (10), (11), (12a) and (12b). w For example, the term h(·) representing the mRNA quantification measurement in equation (10) and the term representing the antigenicity of the source gene in equation (12) can be added together along with any other affine or network functions to generate a dependency function for the allele-non-interacting variables.
[0356] As an example, returning to equation (8), the affine transformation function g h (·), gw Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000030.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0357] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000031.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0358] FIG. 8 shows exemplary network models NN3(·) and NN w Peptide p associated with MHC allele h=3 using (·) k As shown in Figure 8, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0359] VIII.C. Multi-Allele Models The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof.
[0360] VIII.C.1. Example 1: Maximum Values of Models Per Allele In one implementation, the training module 316 trains peptides p associated with a set of MHC alleles H. k The estimated presentation likelihood u k is the presentation likelihood u determined for each of the MHC alleles h in the set H determined based on cells expressing a single allele, as explained above in conjunction with equations (2)-(11). k h∈H Specifically, the presentation likelihood u k u k h∈H In one implementation, the function is a maximum function, as shown in equation (12), where the proposed likelihood u k can be determined as the maximum of the presentation likelihood for each MHC allele h in set H. TIFF2025114587000032.tif5128
[0361] VIII.C.2. Example 2.1: Sum Function Model In one implementation, the training module 316 trains peptides p k The estimated presentation likelihood u k of, TIFF2025114587000033.tif13128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles H associated with x h k is the peptide p kand the coded allele interaction variables for the corresponding MHC alleles. The parameter θ for each MHC allele h h The set of values is θ h The dependence function g can be determined by minimizing a loss function with respect to i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:
[0362] According to equation (13), the peptide sequence p k The likelihood that a given allele will be presented by one or more MHC alleles h is given by the dependency function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate a corresponding score for the allele interaction variable. The scores for each MHC allele h are combined to generate a corresponding score for the peptide sequence p k is transformed by a transformation function f(·) to produce the presentation likelihood that MHC allele H will be presented by the set of MHC alleles H.
[0363] The model presented in equation (13) is that for each peptide p k It differs from the allele-by-allele model of equation (2) in that the number of relevant alleles for a can be greater than 1. In other words, h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0364] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000034.tif5128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0365] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000035.tif5128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.
[0366] FIG. 9 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0367] VIII.C.3. Example 2.2: Sum Function Model with Allelic Non-Interacting Variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF2025114587000036.tif13128, peptide pk The estimated presentation likelihood u k where w k is the peptide p k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values of θ h and θ w The dependence function g can be determined by minimizing a loss function with respect to i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:
[0368] Therefore, according to equation (14), one or more MHC alleles H can bind to a peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores are combined and the combined score is used to estimate the association of the peptide sequence p with the MHC allele H. k is transformed by a transformation function f(·) to produce the presentation likelihood that
[0369] In the model presented in equation (14), each peptide p k The number of relevant alleles for a can be greater than 1. h k More than one element in the peptide sequence pk can have a value of 1 for multiple MHC alleles H associated with
[0370] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000037.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0371] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000038.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0372] FIG. 10 shows exemplary network models NN2(·), NN3(·), and NN w As shown in Figure 10, the network model NN2(·) uses the allele interaction variable x2 for MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. The network model NN3(·) generates the allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated.w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0373] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (15) k the allele interaction variable x h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF2025114587000039.tif13128.
[0374] VIII.C.4. Example 3.1: Model with Likelihood for Each Latent Allele In another implementation, the training module 316 k The estimated presentation likelihood u k of, TIFF2025114587000040.tif7128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles h∈H associated with u' k h is the likelihood of presentation for each potential allele for MHC allele h, and vector v has elements v h But, a h k ·u' k hwhere s(·) is a vector corresponding to v, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values of the input within a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, although it will be recognized that in other embodiments, s(·) may be any function, such as a maximum function. A set of values for the parameters θ for the likelihood of each potential allele can be determined by minimizing a loss function with respect to θ, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.
[0375] The presentation likelihood in the presentation model of equation (17) is the likelihood that each peptide p is presented by an individual MHC allele h. k The likelihood of presentation for each potential allele, u', corresponds to the likelihood that k h The likelihood of each potential allele differs from the allele-based presentation likelihood of Section VIII.B in that parameters for the likelihood of each potential allele can be learned from a multi-allelic setting, in which the direct association between the presented peptide and the corresponding MHC allele is unknown, in addition to a single-allelic setting. Thus, in a multi-allelic setting, the presentation model is based on the likelihood of the peptide p k Not only can we estimate whether peptide p is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are present in peptide p k The individual likelihood u' indicates which person is most likely to have presented k h∈H The advantage of this is that the presented model can generate potential likelihoods without training data for cells expressing a single MHC allele.
[0376] In one particular implementation that will be mentioned throughout the remainder of this specification, r(·) is a function with range [0,1]. For example, r(·) is a clip function: r(z)=min(max(z,0),1) may be the minimum value between z and 1, which represents the likelihood u k In another implementation, r(·) is chosen as r(z)=tanh(z) where the domain z is greater than or equal to 0.
[0377] VIII.C.5. Example 3.2: Sum of Functions Model In one particular implementation, s(·) is a summation function, and the presentation likelihood is given by summing the presentation likelihoods for each potential allele. TIFF2025114587000041.tif15128
[0378] In one implementation, the likelihood of presentation for each potential allele for MHC allele h is calculated as: Generated by TIFF2025114587000042.tif7128, the presented likelihood is Let it be estimated by TIFF2025114587000043.tif13128.
[0379] According to equation (19), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. Each dependency score is first calculated using the proposed likelihood u' for each potential allele. k h The likelihood of each allele is transformed by the function f(·) to generate the likelihood u' k h are combined and a clipping function is applied to the combined likelihood to clip the values into the range [0,1] to produce a peptide sequence p k A presentation likelihood can be generated that g will be presented by a set of MHC alleles H. The dependency function g h is the dependency function g introduced above in Section VIII.B.1.h It can be in any of the following forms:
[0380] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000044.tif7128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0381] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000045.tif7128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.
[0382] FIG. 11 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k) are then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.
[0383] In another implementation, if the prediction is made in terms of the log of the mass analysis ion current, then r(·) is the log function and f(·) is the exponential function.
[0384] VIII.C.6. Example 3.3: Sum of Functions Model with Allelic Non-Interacting Variables In one implementation, the likelihood of presentation for each potential allele for MHC allele h is calculated as: Generated by TIFF2025114587000046.tif7128, the presented likelihood is As generated by TIFF2025114587000047.tif13128, the effects of allelic non-interacting variables are incorporated into peptide presentation.
[0385] According to equation (21), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele non-interaction variables to generate dependency scores for the allele non-interaction variables. The scores of the allele non-interaction variables are combined into each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by a function f(·) to generate a presentation likelihood for each potential allele. The potential likelihoods are combined, and a clipping function is applied to the combined output to clip values into the range [0,1] to produce a representation of the peptide sequence p by the MHC allele H. k A likelihood of being presented can be generated. wis the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:
[0386] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000048.tif7128, where w k is the peptide p k are the allele-non-interacting variables identified for θw, and θw is the set of parameters determined for the allele-non-interacting variables.
[0387] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025114587000049.tif7128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0388] FIG. 12 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 12, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. w (·) indicates peptide p kAllele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·). The network model NN3(·) generates an allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated, which is also the same network model NN w (·) output NN w (w k ) and mapped by a function f(·). Both outputs are combined to produce the estimated presentation likelihood uk.
[0389] In another implementation, the likelihood of presentation for each potential allele for MHC allele h is calculated as: Generated by TIFF2025114587000050.tif7128, the presented likelihood is Generated by TIFF2025114587000051.tif13128.
[0390] VIII.C.7. Example 4: Quadratic Model In one implementation, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k teeth, TIFF2025114587000052.tif13128, where the element u' k h is the likelihood of presentation for each potential allele for MHC allele h. Values for a set of parameters θ for the likelihood for each potential allele can be determined by minimizing a loss function with respect to θ, where i is each example in subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The likelihood of presentation for each potential allele can be in any of the forms shown in equations (18), (20), and (22) above.
[0391] In one embodiment, the model of equation (23) is k However, there is a possibility that a given antigen may be simultaneously presented by two MHC alleles, which may imply that presentation by the two HLA alleles is statistically independent.
[0392] According to equation (23), one or more MHC alleles H bind to the peptide sequence p k The presentation likelihood is calculated by combining the presentation likelihoods for each potential allele and the likelihood of the peptide sequence p being presented by the MHC allele H. k Each pair of MHC alleles is assigned to a peptide p such that it generates a presentation likelihood that p will be presented. k can be generated by subtracting from the sum the likelihood that
[0393] For example, the affine transformation function g h The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF2025114587000053.tif5128, where x2 k , x3 k are the allele interaction variables identified for HLA alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for HLA alleles h=2, h=3.
[0394] As another example, the network transformation function g h (·), g w The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF2025114587000054.tif5128, where NN2(·) and NN3(·) are the network models specified for HLA alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for HLA alleles h=2 and h=3.
[0395] IX Example 5: Prediction Module The prediction module 320 receives sequence data and selects candidate neoantigens in the sequence data using the proposed model. Specifically, the sequence data may be DNA sequences, RNA sequences, and / or protein sequences extracted from tumor tissue cells of a patient. The prediction module 320 converts the sequence data into a plurality of peptide sequences p having 8-15 amino acids for MHC-I or 6-30 amino acids for MHC-II. k For example, the prediction module 320 can process a given sequence "IEFROEIFJEF (SEQ ID NO: 16)" into three nine amino acid peptide sequences "IEFROEIFJ (SEQ ID NO: 17)," "EFROEIFJE (SEQ ID NO: 18)," and "FROEIFJEF (SEQ ID NO: 19)." In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by comparing sequence data extracted from a patient's normal tissue cells with sequence data extracted from the patient's tumor tissue cells to identify segments containing one or more mutations.
[0396] The prediction module 320 applies one or more presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation models to the candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences with an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences with the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the selected candidate neoantigens for a given patient can be injected into the patient to induce an immune response.
[0397] X. Example 6: Patient Selection Module The patient selection module 324 selects a subset of patients for vaccine therapy and / or T cell therapy based on whether the patients meet the selection criteria. In one embodiment, the selection criteria are determined based on the patient's likelihood of presentation of neoantigen candidates generated by the presentation model. By adjusting the selection criteria, the patient selection module 324 can adjust the number of patients who receive vaccine administration and / or T cell therapy based on the patient's likelihood of presentation of neoantigen candidates. Specifically, strict selection criteria may result in a smaller number of patients being treated with the vaccine and / or T cell therapy, but a higher proportion of vaccine and / or T cell therapy-treated patients who receive effective treatment (e.g., one or more tumor-specific neoantigens (TSNAs) and / or one or more neoantigen-reactive T cells). In contrast, looser selection criteria may result in a larger number of patients being treated with the vaccine and / or T cell therapy, but a lower proportion of vaccine and / or T cell therapy-treated patients who receive effective treatment. The patient selection module 324 alters the selection criteria based on a desired balance between a target proportion of patients receiving treatment and the proportion of patients receiving effective treatment.
[0398] In some embodiments, the selection criteria for selecting patients to receive vaccine therapy are the same as the selection criteria for selecting patients to receive T cell therapy. However, in alternative embodiments, the selection criteria for selecting patients to receive vaccine therapy may differ from the selection criteria for selecting patients to receive T cell therapy. Sections XA and XB below discuss the selection criteria for selecting patients to receive vaccine therapy and T cell therapy, respectively.
[0399] Selection of patients for XA vaccine treatment In one embodiment, a patient is associated with a corresponding therapeutic subset of v neoantigen candidates that can potentially be included in a personalized vaccine for that patient, having a vaccine volume v. In one embodiment, the therapeutic subset for a patient is the neoantigen candidate with the highest likelihood of presentation as determined by the presentation model. For example, if a vaccine can include v=20 epitopes, the vaccine can include a therapeutic subset for each patient with the highest likelihood of presentation as determined by the presentation model. However, it will be recognized that in other embodiments, the therapeutic subset for a patient can be determined based on other methods. For example, the therapeutic subset for a patient can be randomly selected from the set of neoantigen candidates for that patient, or can be determined based in part on a combination of factors including prior art models that model the binding affinity or stability of peptide sequences, or presentation likelihood obtained from a presentation model and affinity or stability information for those peptide sequences.
[0400] In one embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the patient's tumor mutation burden is equal to or higher than a minimum mutation burden. A patient's tumor mutation burden (TMB) indicates the total number of nonsynonymous mutations in the tumor exome. In one embodiment, the patient selection module 324 selects a patient for vaccine treatment if the patient's absolute TMB number is equal to or higher than a predetermined threshold. In another implementation, the patient selection module 324 selects a patient for vaccine treatment if the patient's TMB is within a threshold percentile among the TMBs determined for the set of patients.
[0401] In another embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the patient's utility score based on the patient's therapeutic subset is equal to or greater than the minimum utility score. In one embodiment, the utility score is a measure of the estimated number of presented antigens from the therapeutic subset.
[0402] The estimated number of presented antigens can be predicted by modeling the presentation of neoantigens as random variables with one or more probability distributions. In one implementation, the utility score for patient i is the expected number of presented neoantigen candidates from the treatment subset, or a specific function thereof. As an example, the presentation of each neoantigen can be modeled as a Bernoulli random variable, where the probability of presentation (success) is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation (success) of each neoantigen candidate is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation of each neoantigen candidate is given by the probability of presentation of each neoantigen candidate with the highest presentation likelihood, u i1 , u i2 , …, u iv v neoantigen candidates p i1 , p i2 , …, p iv Treatment subset S i Regarding neoantigen candidate p ij The presentation of the random variable A ij where: The file is TIFF2025114587000055.tif6128. The expected number of neoantigens presented is given by the sum of the likelihoods of presentation of each neoantigen candidate. In other words, the utility score for patient i is expressed as: The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility score for the vaccine treatment.
[0403] In another implementation, the utility score for patient i is the probability that at least a threshold number of neoantigens k are presented. In one example, a therapeutic subset S of neoantigen candidates is i The number of presented antigens in is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the likelihood of presentation of each of the epitopes. In particular, the number of presented antigens for patient i is determined by the random variable N i can be given by TIFF2025114587000057.tif13128 where PBD(·) denotes the Poisson binomial distribution. The probability that at least a threshold number of neoantigens k are presented is given by the number of presented antigens N iis given by the operation of the probability that k is equal to or greater than k. In other words, the utility score for patient i is expressed as: The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility score for the vaccine treatment.
[0404] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of neoantigen candidates that have binding affinities or predicted binding affinities below a fixed threshold (e.g., 500 nM) for one or more of the patient's HLA alleles. i The threshold is the number of neoantigens in the target region. In one example, the fixed threshold ranges from 1000 nM to 10 nM. In some cases, the utility score may only count neoantigens detected as expressed by RNA-seq.
[0405] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of candidate neoantigens whose binding affinity to one or more HLA alleles for that patient is less than or equal to a threshold percentile of the binding affinity of a random peptide to that HLA allele. i The threshold percentile is the number of neoantigens in the target region. In one example, the threshold percentile ranges from the 10th percentile to the 0.1th percentile. In some cases, the utility score may only count neoantigens detected as expressed by RNA-seq.
[0406] It will be appreciated that the example utility scores described with respect to equations (25) and (27) are for illustrative purposes only, and the patient selection module 324 may use other statistics or probability distributions to generate utility scores.
[0407] Patient selection for XB T-cell therapy In another embodiment, instead of or in addition to receiving vaccine therapy, a patient can receive T cell therapy. Similar to vaccine therapy, in embodiments in which a patient receives T cell therapy, the patient can be associated with a corresponding therapeutic subset of the v candidate neoantigens as described above. This therapeutic subset of the v candidate neoantigens can be used to identify in vitro T cells from the patient that are reactive to one or more of the v candidate neoantigens. These identified T cells can then be expanded and infused into the patient in a personalized T cell therapy.
[0408] Patients can be selected for T cell therapy at two different time points: the first time point after the model predicts therapeutic subsets of v neoantigen candidates for the patient but before in vitro screening of T cells specific for the predicted therapeutic subsets of v neoantigen candidates; and the second time point after in vitro screening of T cells specific for the predicted therapeutic subsets of v neoantigen candidates.
[0409] First, patients can be selected for T cell therapy after a therapeutic subset of v neoantigen candidates for the patient has been predicted, but before in vitro identification of T cells from the patient that are specific for the predicted therapeutic subset of v neoantigen candidates has been performed. Specifically, because in vitro screening of neoantigen-specific T cells from patients can be costly, it may be desirable to select patients for screening for neoantigen-specific T cells only if the patient is likely to have neoantigen-specific T cells. Patient selection prior to the in vitro T cell screening step can use the same criteria as those used to select patients for vaccine therapy. Specifically, in some embodiments, the patient selection module 324 can select patients for T cell therapy if the patient's tumor mutational burden is equal to or higher than a minimum mutational burden, as described above. In another embodiment, the patient selection module 324 can select patients for T cell therapy if the patient's utility score based on the therapeutic subset of v neoantigen candidates for the patient is equal to or higher than a minimum utility score, as described above.
[0410] Second, in addition to or instead of selecting patients for T cell therapy before in vitro identification of T cells from the patient that are specific for a therapeutic subset of the predicted v neoantigen candidates, patients can also be selected for T cell therapy after in vitro identification of T cells that are specific for a therapeutic subset of the predicted v neoantigen candidates. Specifically, a patient can be selected for T cell therapy if at least a threshold amount of neoantigen-specific TCRs are identified for the patient in an in vitro screen of the patient's T cells for neoantigen recognition. For example, a patient can be selected for T cell therapy only if at least two neoantigen-specific TCRs have been identified for the patient, or only if neoantigen-specific TCRs have been identified for two different neoantigens.
[0411] In another embodiment, a patient can be selected to receive T cell therapy only if a threshold amount of neoantigens from a therapeutic subset of v neoantigen candidates for that patient is recognized by the patient's TCR. For example, a patient can be selected to receive T cell therapy only if at least one neoantigen from a therapeutic subset of v neoantigen candidates for that patient is recognized by the patient's TCR. In a further embodiment, a patient can be selected to receive T cell therapy only if at least a threshold amount of TCRs for that patient are identified as neoantigen-specific for a particular HLA-restricted class of neoantigen peptide. For example, a patient can be selected to receive T cell therapy only if at least one TCR for that patient is identified as a neoantigen-specific HLA class I-restricted neoantigen peptide.
[0412] In yet further embodiments, a patient can be selected to receive T cell therapy only if at least a threshold amount of a particular HLA-restricted class of neoantigen peptide is recognized by the patient's TCR. For example, a patient can be selected to receive T cell therapy only if at least one HLA class I-restricted neoantigen peptide is recognized by the patient's TCR. For example, a patient can be selected to receive T cell therapy only if at least two HLA class II-restricted neoantigen peptides are recognized by the patient's TCR. Any combination of the above criteria can also be used to select patients for T cell therapy after in vitro identification of T cells specific for the therapeutic subset of v predicted neoantigen candidates for the patient.
[0413] XI. Example 7: Experimental Results Demonstrating Exemplary Patient Selection Performance The validity of the patient selection described in Section X is validated by selecting patients from a set of simulated patients, each associated with a test set of simulated neoantigen candidates, for which a subset of the simulated neoantigens is known to be represented in the mass spectrometry data. Specifically, each simulated neoantigen candidate in the test set is associated with a label indicating whether the neoantigen is represented in the mass spectrometry data set of the multi-allelic JY cell line HLA-A*02:01 and HLA-B*07:02 from the Bassani-Sternberg dataset (dataset "D1") (data available at www.ebi.ac.uk / pride / archive / projects / PXD0000394). As described in more detail below in conjunction with Figure 13A, a number of neoantigen candidates for the simulated patients are sampled from the human proteome based on known frequency distributions of mutation burden in non-small cell lung cancer (NSCLC) patients.
[0414] Allele-specific presentation models for the same HLA allele are trained using a training set that is a subset of the mass spectrometry data for the single alleles HLA-A*02:01 and HLA-B*07:02 from the IEDB dataset (dataset "D2") (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Specifically, the presentation model for each allele is trained using the network dependency function g h (·) and g wThe presentation model for the HLA-A*02:01 allele generates the presentation likelihood of a particular peptide on the HLA-A*02:01 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables. The presentation model for the HLA-B*07:02 allele generates the presentation likelihood of a particular peptide on the HLA-B*07:02 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables.
[0415] As disclosed in the following examples with reference to Figures 13A-13E, different models, such as a presentation model trained for peptide binding prediction and a prior art model, are applied to a test set of neoantigen candidates for each simulated patient to identify different therapeutic subsets for the patient based on the predictions. Patients who meet selection criteria for vaccine treatment are selected and associated with personalized vaccines containing epitopes in the patient's therapeutic subset. The size of the therapeutic subsets varies depending on different vaccine doses. No overlap is introduced between the training set used to train the presentation model and the test set of simulated neoantigen candidates.
[0416] In the following example, the proportion of selected patients with at least a certain number of presented neoantigens among the epitopes included in the vaccine is analyzed. This statistic indicates the effectiveness of the simulated vaccine in delivering potential neoantigens that will elicit an immune response in patients. Specifically, simulated neoantigens in a test set are presented if the neoantigen is presented in mass spectrometry dataset D2. A high proportion of patients with presented neoantigens indicates the likelihood of successful treatment with the neoantigen vaccine by inducing an immune response.
[0417] XI.A. Example 7A: Frequency Distribution of Mutation Burden in NSCLC Cancer Patients Figure 13A shows the sample frequency distribution of mutation burden in NSCLC patients. Mutation burden and mutations in different tumor types, including NSCLC, can be found, for example, in the Cancer Genome Atlas (TCGA) (https: / / cancergenome.nih.gov). The X-axis represents the number of nonsynonymous mutations for each patient, and the Y-axis represents the proportion of sample patients with a specific number of nonsynonymous mutations. The sample frequency distribution in Figure 13A shows a range of 3 to 1786 mutations, with 30% of patients having fewer than 100 mutations. Although not shown in Figure 13A, studies have shown that mutation burden is higher in smokers compared to nonsmokers, and that mutation burden can be a strong indicator of neoantigen burden in patients.
[0418] As introduced at the beginning of Section XI above, each of the simulated patient populations is associated with a test set of neoantigen candidates. Each patient's test set is determined by selecting the mutation load m from the frequency distribution shown in Figure 13A for each patient. i The D1 dataset is generated by sampling the D1 sequence. For each mutation, a 21-mer peptide sequence from the human proteome is randomly selected to represent the mutant sequence to be simulated. A test set of candidate neoantigen sequences is generated for patient i by identifying each (8, 9, 10, 11)-mer peptide sequence across the mutations in the 21-mer. Each candidate neoantigen is associated with a label indicating whether the candidate neoantigen sequence is present in the mass spectrometry D1 dataset. For example, candidate neoantigen sequences present in dataset D1 can be associated with the label "1," and sequences not present in dataset D1 can be associated with the label "0." As described in more detail below, Figures 13B-13E show experimental results of patient selection based on the neoantigens presented by patients in the test set.
[0419] XI.B. Example 7B: Proportion of Selected Patients with Neoantigen Presentation Based on Selection Criteria for Mutational Burden Figure 13B shows the number of presented neoantigens in the simulated vaccine for patients selected based on the selection criteria of whether the patient met a minimum mutational load. The proportion of selected patients who had at least a certain number of presented neoantigens in the corresponding study was identified.
[0420] In Figure 13B, the x-axis shows the proportion of patients excluded from vaccine treatment based on tumor mutation burden, as indicated by the label "Minimum Number of Mutations." For example, the data point at "Minimum Number of Mutations" 200 indicates that the patient selection module 324 selected only a subset of simulated patients with a mutation burden of at least 200 mutations. As another example, the data point at "Minimum Number of Mutations" 300 indicates that the patient selection module 324 selected a lower proportion of simulated patients with at least 300 mutations. The y-axis shows the proportion of selected patients associated with at least a certain number of presented neoantigens in the test set without vaccine dose v. Specifically, the top plot shows the proportion of selected patients presenting at least one neoantigen, the middle plot shows the proportion of selected patients presenting at least two antigens, and the bottom plot shows the proportion of selected patients presenting at least three antigens.
[0421] As shown in Figure 13B, the proportion of patients with presented neoantigens significantly increased with increasing mutational burden, indicating that mutational burden as a selection criterion can be effective in selecting patients in whom neoantigen vaccines are likely to induce an effective immune response.
[0422] XI.C. Example 7C: Comparison of Neoantigen Presentation in Vaccines Identified by Presentation Models and Prior Art Models Figure 13C compares the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on the presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model. The plot on the left assumes a limiting vaccine volume of v=10, and the plot on the right assumes a limiting vaccine volume of v=20. Patients are selected based on a utility score indicating the expected number of presented neoantigens.
[0423] In Figure 13C, the solid lines indicate patients associated with a vaccine containing therapeutic subsets identified based on presentation models for the alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient is identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted lines indicate patients associated with a vaccine containing therapeutic subsets identified based on the prior art model NETMHCpan for the single allele HLA-A*02:01. Implementation details for NETMHCpan are provided at http: / / www.cbs.dtu.dk / services / NetMHCpan. The therapeutic subset for each patient is identified by applying the NETMHCpan model to the sequences in the test set and identifying the v neoantigen candidates with the highest estimated binding affinity. The x-axis of both graphs indicates the proportion of patients excluded from vaccine treatment based on the expected utility score, which indicates the expected number of presented neoantigens in the therapeutic subsets identified based on the presentation model. The expected utility score is determined as described in Section X in connection with equation (25). The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens) included in the vaccine.
[0424] As shown in Figure 13C, a significantly higher proportion of patients associated with vaccines containing therapeutic subsets based on the presentation model receive vaccines containing presented neoantigens than patients associated with vaccines containing therapeutic subsets based on the prior art model. For example, as shown in the graph on the right, 80% of selected patients associated with vaccines based on the presentation model receive at least one presented neoantigen in the vaccine, compared to only 40% of selected patients associated with vaccines based on the prior art model. These results demonstrate that the presentation model described herein is effective in selecting neoantigen candidates for vaccines that are likely to elicit an immune response to treat tumors.
[0425] XI.D. Example 7D: Effect of HLA Coverage on Neoantigen Presentation of Vaccines Identified by Presentation Models Figure 13D compares the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a single allele-specific presentation model for HLA-A*02:01 and selected patients associated with vaccines containing therapeutic subsets identified based on both an allele-specific presentation model for HLA-A*02:01 and an allele-specific presentation model for HLA-B*07:02. Vaccine volume is set to v = 20 epitopes. For each experiment, patients are selected based on expected utility scores determined based on different therapeutic subsets.
[0426] In Figure 13D, the solid line indicates patients associated with a vaccine containing therapeutic subsets based on both presentation models for HLA alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient was identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with a vaccine containing therapeutic subsets based on a single presentation model for HLA allele HLA-A*02:01. The therapeutic subset for each patient was identified by applying a presentation model for only a single HLA allele to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. In the solid plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility scores for the therapeutic subsets identified by both presentation models. In the dotted plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility scores for the therapeutic subsets identified by a single presentation model. The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens).
[0427] As shown in Figure 13D, patients associated with vaccines containing therapeutic subsets identified by presentation models for both HLA alleles presented neoantigens at a significantly higher rate than patients associated with vaccines containing therapeutic subsets identified by a single presentation model. These results demonstrate the importance of establishing presentation models with high HLA coverage.
[0428] XI.E. Example 7E: Comparison of Neoantigen Presentation in Patients Selected by Mutation Burden and Expected Number of Presented Neoantigens Figure 13E compares the number of presented neoantigens in simulated vaccines between patients selected based on mutation burden and patients selected by expected utility score, which is determined based on the therapeutic subset identified by the presentation model with a size of v = 20 epitopes.
[0429] In Figure 13E, the solid lines indicate patients selected based on the expected utility score associated with a vaccine containing the therapeutic subset identified by the proposed model. The therapeutic subset for each patient is identified by applying each of the proposed models to the sequences in the test set and identifying the v = 20 neoantigen candidates with the highest likelihood of presentation. The therapeutic utility score is determined based on the likelihood of presentation of the therapeutic subset identified in Section X based on Equation (25). The dotted lines indicate patients selected based on the mutational load associated with a vaccine containing the therapeutic subset identified by the proposed model. The x-axis indicates the proportion of patients excluded from vaccine treatment based on expected utility score in the solid plot and the proportion of patients excluded based on mutational load in the dotted plot. The y-axis indicates the proportion of selected patients who receive a vaccine containing at least a specific number of presented neoantigens (1, 2, or 3 neoantigens). As shown in Figure 13E, patients selected based on expected utility score receive a vaccine containing presented neoantigens at a higher rate than patients selected based on mutational load. However, patients selected based on mutational load receive a vaccine containing a higher proportion of presented neoantigens than unselected patients. Thus, while mutational load is an effective patient selection criterion for delayed delivery of effective neoantigen vaccines, expected utility scores are more effective.
[0430] XII. Example 8: Evaluation of Mass Spectrometry-Trained MHC Class II Presentation Models on Omitted MHC Class II Mass Spectrometry Data The effectiveness of the various proposed models described above was tested on test data T, which was a subset of training data 170 that was not used to train the proposed models, or a separate data set from training data 170 that had similar variables and data structure as training data 170.
[0431] As related indicators of the performance of the proposed model, TIFF2025114587000059.tif10128, which shows the ratio of the number of peptide instances correctly predicted to be presented on the relevant HLA allele to the number of peptide instances predicted to be presented on that HLA allele. i is the corresponding likelihood estimate u i is greater than or equal to a predetermined threshold t, the gene is predicted to be presented on one or more relevant HLA alleles. Another relevant indicator of the performance of the presentation model is TIFF2025114587000060.tif10128, which shows the ratio of the number of peptide instances correctly predicted to be presented on the relevant HLA allele to the number of peptide instances known to be presented on that HLA allele. Another related measure of the performance of a presentation model is the area under the curve (AUC) of the receiver operating characteristic (ROC). The ROC plots recall against the false positive rate (FPR), given by: TIFF2025114587000061.tif10128
[0432] XII.A. Performance of the proposed model on mass spectrometry data XII.A.1. Example 1 Figure 14A shows a histogram of peptide lengths eluted from class II MHC alleles on human tumor cells and tumor-infiltrating lymphocytes (TILs) using mass spectrometry. Specifically, mass spectrometry peptidomics was performed on the HLA-DRB1*12:01 homozygous allele ("Dataset 1") and a multi-allelic sample ("Dataset 2") of HLA-DRB1*12:01 and HLA-DRB1*10:01. The results show that the lengths of peptides eluted from class II MHC alleles range from 6 to 30 amino acids. The frequency distribution shown in Figure 14A is similar to the length distribution of peptides eluted from class II MHC alleles using a state-of-the-art mass spectrometry method, as shown in Figure 1C of reference 69.
[0433] Figure 14B shows the dependency of mRNA quantification and peptides presented per residue for Dataset 1 and Dataset 2. The results show that there is a strong dependency between mRNA expression and peptide presentation of class II MHC alleles.
[0434] Specifically, the horizontal axis of FIG. 14B is log 10 The vertical axis of Figure 14B shows mRNA expression in transcripts per million bins (TPM). -2 <log 10 TPM<10 -1 The plots are plotted as a multiple of the lowest bin corresponding to mRNA expression levels between 10 and 20. One solid line is a plot relating mRNA quantification and peptide presentation for dataset 1, and the other solid line is for dataset 2. As shown in Figure 14B, there is a strong correlation between mRNA expression levels and the amount of peptide presentation per residue within the corresponding gene. Specifically, when RNA expression levels are between 10 and 20, the correlation between mRNA expression levels and peptide presentation per residue within the corresponding gene is strong. 1 <log 10 TPM<10 2 Peptides from genes in the range of 0.01 are more than five times more likely to be represented compared to the lowest bin.
[0435] These results indicate that measures of mRNA quantification are strongly predictive of peptide presentation and that incorporating these measures can significantly improve the performance of presentation models.
[0436] Figure 14C compares the performance results of the exemplary proposed model trained and tested using Dataset 1 and Dataset 2. For each set of model features of the exemplary proposed model, Figure 14C shows the PPV values at 10% recall when features in that model feature set are classified as allele-interacting features or when features in that model feature set are classified as allele-non-interacting feature variables. As can be seen in Figure 14C, for each set of model features of the exemplary proposed model, the PPV values at 10% recall determined when features in that model feature set are classified as allele-interacting features are shown on the left, and the PPV values at 10% recall determined when features in that model feature set are classified as allele-non-interacting features are shown on the right. Note that peptide sequence features were always classified as allele-interacting features for the purposes of Figure 14C. The results show that the proposed model achieved PPV values at 10% recall ranging from 14% to 29%, which are significantly higher (approximately 500-fold) than the PPV for random prediction.
[0437] In this experiment, peptide sequences ranging in length from 9 to 20 residues were considered. The data were divided into training, validation, and test sets. Blocks of 50-residue peptides from both Dataset 1 and Dataset 2 were assigned to the training and test sets. Duplicate peptides anywhere within the proteome were removed to ensure that no peptide sequences appeared in both the training and test sets. The incidence of peptide presentation in the training and test sets increased 50-fold by removing non-presented peptides. This is because Dataset 1 and Dataset 2 were derived from human tumor samples in which only a portion of the cells possessed class II HLA alleles, resulting in approximately 10-fold lower peptide yield than a pure sample of class II HLA alleles (although this is still an underestimate due to the imperfect sensitivity of mass spectrometry). The training set contained 1,064 presented peptides and 3,810,070 non-presented peptides. The test set contained 314 presented peptides and 807,400 non-presented peptides.
[0438] Exemplary model 1 is a network dependency function g h The network dependency function g is the sum of the functions in Eq. (22) using the expit function f(·), and the identity function r(·). h (·) was structured as a multilayer perceptron (MLP) with 256 hidden nodes and rectified linear unit (ReLU) activation. Besides the peptide sequence, the allele interaction variable w includes one-hot coded C-terminal and N-terminal flanking sequences, peptide p i The index of the source gene G = gene(p i ) and a variable indicating mRNA quantitative measurements. Exemplary Model 2 was the same as Exemplary Model 1 except that C-terminal and N-terminal flanking sequences were omitted from the allele interaction variables. Exemplary Model 3 was the same as Exemplary Model 1 except that source gene subscripts were omitted from the allele interaction variables. Exemplary Model 4 was the same as Exemplary Model 1 except that mRNA measurements were omitted from the allele interaction variables.
[0439] Exemplary model 5 is a network dependency function g h (·), expit function f(·), identity function r(·), and dependency function g in Eq. w The dependence function g w (·) also included a network model with mRNA quantification measurements as input and structured as an MLP with 16 hidden nodes and rectified linear unit (ReLU) activations, and a network model with C-terminal flanking sequences as input and structured as an MLP with 32 hidden nodes and rectified linear unit (ReLU) activations. The network dependency function g h(·) was structured as a multilayer perceptron with 256 hidden nodes and rectified linear unit (ReLU) activations. Example Model 6 was the same as Example Model 5 except that the network models for the C- and N-terminal flanking sequences were omitted. Example Model 7 was the same as Example Model 5 except that the source gene subscripts were omitted from the allele non-interaction variables. Example Model 8 was the same as Example Model 5 except that the network model for the mRNA quantification measurements was omitted.
[0440] The occurrence rate of the displayed peptides in the test set was approximately 1 / 2400, and therefore the PPV of random prediction would also be approximately 1 / 2400 = 0.00042. As shown in Figure 14C, the most accurate display model achieved a PPV value of approximately 29%, which is approximately 500 times better than the PPV value of random prediction.
[0441] XII.A.2. Example 2 Figure 14D is a histogram showing the amount of peptides sequenced using mass spectrometry for each of 73 samples consisting of human tumors (NSCLC, lymphoma, and ovarian cancer) and cell lines (EBV) containing HLA class II molecules. As shown in Figure 14D, an average of 900 peptides were sequenced for each sample. Furthermore, for each sample of multiple samples, the histogram shown in Figure 14D shows the amount of peptides sequenced using mass spectrometry at different q-value thresholds. Specifically, for each sample of multiple samples, Figure 14D shows the amount of peptides sequenced using mass spectrometry at q-values less than 0.01, q-values less than 0.05, and q-values less than 0.2.
[0442] As noted above, each of the 73 samples in Figure 14D contained an HLA class II molecule. More specifically, each of the 73 samples in Figure 14D contained an HLA-DR molecule. An HLA-DR molecule is a type of HLA class II molecule. Even more specifically, each of the 73 samples in Figure 14D contained an HLA-DRB1 molecule, an HLA-DRB3 molecule, an HLA-DRB4 molecule, and / or an HLA-DRB5 molecule. HLA-DRB1 molecule, an HLA-DRB3 molecule, an HLA-DRB4 molecule, and an HLA-DRB5 molecule are types of HLA-DR molecules.
[0443] While this particular experiment was performed using a sample containing HLA-DR molecules, specifically HLA-DRB1, HLA-DRB3, HLA-DRB4, and HLA-DRB5 molecules, in alternative embodiments, the experiment can be performed using a sample containing one or more of any type(s) of HLA class II molecules. For example, in alternative embodiments, the same experiment can be performed using a sample containing HLA-DP and / or HLA-DQ molecules. Those skilled in the art will appreciate that the same method can be used to model any type(s) of MHC class II molecules and still obtain reliable results. See, e.g., Jensen, Kamilla Kjaergaard et al. 76 is an example of a recent scientific paper that uses the same method to model binding affinity for HLA-DR molecules, as well as for HLA-DP and HLA-DQ molecules. Thus, one skilled in the art will understand that the experiments and models described herein can be used to model not only HLA-DR molecules, but also any other MHC class II molecules, separately or simultaneously, and still obtain reliable results.
[0444] Mass spectrometry was performed on each of the 73 samples to sequence the peptides in each sample. The mass spectra obtained for the samples were searched using Comet and scored using Percolator to sequence the peptides. The amount of peptides sequenced in the sample was then determined for several different Percolator q-value thresholds. Specifically, the amount of peptides sequenced for the sample was determined using a Percolator q-value of less than 0.01, a Percolator q-value of less than 0.05, and a Percolator q-value of less than 0.2.
[0445] The amount of peptides sequenced at each of the different Percolator q-value thresholds for each of the 73 samples is shown in Figure 14D. For example, as can be seen in Figure 14D, for the first sample, approximately 4700 peptides were sequenced using mass spectrometry at a q-value of less than 0.2, approximately 3600 peptides were sequenced using mass spectrometry at a q-value of less than 0.05, and approximately 3200 peptides were sequenced using mass spectrometry at a q-value of less than 0.01.
[0446] Overall, Figure 14D shows that mass spectrometry can be used to sequence a large number of peptides from samples containing MHC class II molecules at low q values. In other words, the data shown in Figure 14D demonstrates that mass spectrometry can be used to sequence with high confidence peptides that can be presented by MHC class II molecules.
[0447] Figure 14E is a histogram showing the amount of samples in which specific MHC class II molecule alleles were identified. More specifically, Figure 14E shows the amount of samples in which specific MHC class II molecules were identified for a total of 73 samples containing HLA class II molecules.
[0448] As described above with respect to Figure 14D, each of the 73 samples in Figure 14D contained HLA-DRB1, HLA-DRB3, HLA-DRB4, and / or HLA-DRB5 molecules. Accordingly, Figure 14E shows the amount of samples in which a particular allele was identified for HLA-DRB1, HLA-DRB3, HLA-DRB4, and HLA-DRB5 molecules. To identify the HLA alleles present in a sample, the sample is subjected to HLA class II DR typing. The amount of samples in which a particular HLA allele was identified is then determined by simply adding up the number of samples in which the HLA allele was identified using HLA class II DR typing. For example, as shown in Figure 14E, 17 of the 73 samples contained the HLA class II molecule allele HLA-DRB3*01:01. In other words, 17 of the 73 samples contained the HLA-DRB3 allele HLA-DRB3*01:01. Overall, Figure 14E shows that a wide range of HLA class II molecule alleles can be identified from the 73 samples containing HLA class II molecules.
[0449] Figure 14F is a histogram showing the proportion of peptides presented by MHC class II molecules for each of a range of peptide lengths across a total of 73 samples. To determine the length of each peptide in each of the 73 samples, each peptide was sequenced using mass spectrometry as described above with respect to Figure 14D, and then the number of residues in the sequenced peptide was simply quantified.
[0450] As discussed above, MHC class II molecules typically present peptides between 9 and 20 amino acids in length. Accordingly, Figure 14F shows the percentage of peptides presented by MHC class II molecules in 73 samples for each peptide length between 9 and 20 amino acids. For example, as shown in Figure 14F, approximately 23% of peptides presented by MHC class II molecules in 73 samples were 14 amino acids in length.
[0451] Based on the data shown in Figure 14F, the most frequent lengths of peptides presented by MHC class II molecules in the 73 samples were identified as 14 and 15 amino acids. These frequent lengths identified for peptides presented by MHC class II molecules in the 73 samples are consistent with previous reports on the most frequent lengths of peptides presented by MHC class II molecules. Furthermore, and also consistent with previous reports, the data in Figure 14F indicate that over 60% of the peptides presented by MHC class II molecules from the 73 samples had lengths other than 14 and 15 amino acids. In other words, Figure 14F indicates that while peptides presented by MHC class II molecules are most frequently 14 or 15 amino acids in length, a large proportion of peptides presented by MHC class II molecules are lengths other than 14 or 15 amino acids. Therefore, it is an incorrect assumption to assume that peptides of all lengths have an equal probability of being presented by MHC class II molecules, or that only peptides 14 or 15 amino acids in length are presented by MHC class II molecules. As discussed in more detail below with respect to Figure 14L, these incorrect assumptions are currently used in many state-of-the-art models for predicting presentation by MHC class II molecules, and therefore the presentation likelihoods predicted by these models are often unreliable.
[0452] Figure 14G is a line graph showing the relationship between gene expression and the incidence of presentation of gene expression products by MHC class II molecules for genes present in 73 samples. More specifically, Figure 14G shows the relationship between gene expression and the proportion of residues resulting from that gene expression that form the N-terminus of peptides presented by MHC class II molecules. To quantify gene expression in each of the 73 total samples, RNA sequencing was performed on the RNA contained in each sample. In Figure 14G, gene expression is measured by RNA sequencing in units of transcripts per million (TPM). To determine the incidence of presentation of gene expression for each of the 73 samples, HLA class II DR peptidome data was identified for each sample.
[0453] As shown in Figure 14G, for 73 samples, a strong correlation is observed between gene expression levels and the presentation of residues of expressed gene products by MHC class II molecules. Specifically, as shown in Figure 14G, peptides resulting from expression of the least expressed genes are over 100 times less likely to be presented by MHC class II molecules than peptides resulting from expression of the most highly expressed genes. In simpler terms, products of more highly expressed genes are presented more frequently by MHC class II molecules.
[0454] Figures 14H-I and 14K-L are line graphs comparing the performance of different presentation models in predicting the likelihood that a peptide in a test dataset of peptides will be presented by at least one MHC class II molecule present in the test dataset. As shown in Figures 14H-I and 14K-L, the performance of a model in predicting the likelihood that a peptide will be presented by at least one MHC class II molecule present in the test dataset is determined by determining the ratio of the true positive rate to the false positive rate for each prediction generated by the model. These ratios determined for a given model can be visualized in a line graph as a receiver operator characteristic (ROC) curve, where the x-axis quantifies the false positive rate and the y-axis quantifies the true positive rate. The area under the curve (AUC) is used to quantify model performance. Specifically, models with a larger AUC have better performance (i.e., higher accuracy) compared to models with a smaller AUC. In Figures 14H, 14I, and 14L, the dashed black line with a slope of 1 (the ratio of true positive rate to false positive rate is 1) represents the predicted curve of the likelihood of randomly estimated peptide presentation. The AUC of the dashed line is 0.5. ROC curves and AUC measurements are discussed in detail in the first half of Section XII above.
[0455] Figure 14H is a line graph comparing the performance of five exemplary models in predicting the likelihood that a peptide in a test dataset of peptides will be presented by an MHC class II molecule, given different sets of allele-interacting and allele-non-interacting variables. In other words, Figure 14H quantifies the relative importance of different allele-interacting and allele-non-interacting variables in predicting the likelihood that a peptide will be presented by an MHC class II molecule.
[0456] The model architecture of each of the five exemplary models used to generate the ROC curves in the line graphs of Figure 14H included an ensemble of five sigmoid summation models. Each sigmoid summation model in the ensemble was configured to model peptide presentation for up to four unique HLA-DR alleles per sample. Furthermore, each sigmoid summation model in the ensemble was configured to predict the likelihood of peptide presentation based on the following allele-interacting and allele-non-interacting variables: peptide sequence, flanking sequence, RNA expression in TPM, gene identifier, and sample identifier. The allele-interacting component of each sigmoid summation model in the ensemble was a one-hidden-layer MLP with ReLU activation as the 256 hidden units.
[0457] Prior to using the exemplary model to predict the likelihood that peptides in a test dataset of peptides will be presented by MHC class II molecules, the exemplary model was trained and validated. To train, validate, and ultimately test the exemplary model, the data described above for the 73 samples was divided into training, validation, and test datasets.
[0458] To prevent peptides from appearing in multiple datasets among the training, validation, and test datasets, the following procedure was performed. First, all peptides from a total of 73 samples that appeared at multiple positions in the proteome were removed. Next, the peptides from a total of 73 samples were divided into blocks of 10 adjacent peptides. Each block of peptides from a total of 73 samples was individually assigned to the training dataset, the validation dataset, or the test dataset. This resulted in no peptides appearing in multiple datasets among the training, validation, and test datasets.
[0459] Of the 38,035,453 peptides in 73 total samples, the training dataset contained 33,570 peptides presented by MHC class II molecules from 69 of the 73 total samples. The 33,570 peptides included in the training dataset ranged from 9 to 20 amino acids in length. The example model used to generate the ROC curve in Figure 14H was trained on the training dataset using the ADAM optimization algorithm and early stopping.
[0460] The validation dataset consisted of 3,925 peptides presented by MHC class II molecules from the same 69 samples used in the training dataset. The validation set was used for early stopping only.
[0461] The test dataset included peptides presented by MHC class II molecules identified from tumor samples using mass spectrometry. Specifically, the test dataset included 232 peptides identified from four tumor samples. The peptides included in the test dataset were excluded from the training dataset described above.
[0462] As noted above, Figure 14H quantifies the relative importance of different allele-interacting and allele-non-interacting variables in predicting the likelihood that a peptide will be presented by an MHC class II molecule. Also as noted above, the exemplary model used to generate the ROC curve for the line graph in Figure 14H was configured to predict peptide presentation likelihood based on the following allele-interacting and allele-non-interacting variables: peptide sequence, flanking sequence, RNA expression in TPM, gene identifier, and sample identifier. To quantify the relative importance of four of these five variables (peptide sequence, flanking sequence, RNA expression, and gene identifier) for predicting the likelihood that a peptide will be presented by an MHC class II molecule, each of the five exemplary models described above was tested using data from a test dataset with different combinations of the four variables. Specifically, for each peptide in the test dataset, exemplary model 1 generated a prediction of peptide presentation likelihood based on peptide sequence, flanking sequence, gene identifier, and sample identifier, excluding RNA expression. Similarly, for each peptide in the test dataset, exemplary model 2 generated a prediction of peptide presentation likelihood based on the peptide sequence, RNA expression, gene identifier, and sample identifier, excluding the flanking sequence. Similarly, for each peptide in the test dataset, exemplary model 3 generated a prediction of peptide presentation likelihood based on the flanking sequence, RNA expression, gene identifier, and sample identifier, excluding the peptide sequence. Similarly, for each peptide in the test dataset, exemplary model 4 generated a prediction of peptide presentation likelihood based on the flanking sequence, RNA expression, peptide sequence, and sample identifier, excluding the gene identifier. Finally, for each peptide in the test dataset, exemplary model 5 generated a prediction of peptide presentation likelihood based on all five variables: flanking sequence, RNA expression, peptide sequence, gene identifier, and sample identifier.
[0463] The performance of each of these five exemplary models is shown in the line graph of Figure 14H. Specifically, each of the five exemplary models is associated with an ROC curve that shows the ratio of the true positive rate to the false positive rate for each prediction generated by the model. For example, Figure 14H shows a curve for exemplary Model 1, which generated a prediction of peptide display likelihood based on peptide sequence, flanking sequence, gene identifier, and sample identifier, excluding RNA expression. Figure 14H shows a curve for exemplary Model 2, which generated a prediction of peptide display likelihood based on peptide sequence, RNA expression, gene identifier, and sample identifier, excluding flanking sequence. Figure 14H also shows a curve for exemplary Model 3, which generated a prediction of peptide display likelihood based on flanking sequence, RNA expression, gene identifier, and sample identifier, excluding peptide sequence. Figure 14H also shows a curve for exemplary Model 4, which generated a prediction of peptide display likelihood based on flanking sequence, RNA expression, peptide sequence, and sample identifier, excluding gene identifier. And finally, Figure 14H shows the curves for exemplary Model 5, which generated predictions of peptide presentation likelihood based on all five variables: flanking sequence, RNA expression, peptide sequence, sample identifier, and gene identifier.
[0464] As described above, the performance of a model in predicting the likelihood of a peptide being presented by an MHC class II molecule is quantified by determining the AUC of the ROC curve, which represents the ratio of the true positive rate to the false positive rate, for each prediction generated by the model. Models with a higher AUC have better performance (i.e., higher accuracy) compared to models with a lower AUC. As shown in Figure 14H, the curve for exemplary Model 5, which generated a prediction of peptide presentation likelihood based on all five variables (flanking sequence, RNA expression, peptide sequence, sample identifier, and gene identifier), achieved the highest AUC of 0.98. Thus, exemplary Model 5, which used all five variables to generate a prediction of peptide presentation, achieved the best performance. The curve for exemplary Model 2, which generated a prediction of peptide presentation likelihood based on peptide sequence, RNA expression, gene identifier, and sample identifier, excluding flanking sequence, achieved the second highest AUC of 0.97. Thus, flanking sequence can be identified as the least important variable in predicting the likelihood of a peptide being presented by an MHC class II molecule. The curve for exemplary Model 4, which generated a prediction of peptide presentation likelihood based on flanking sequence, RNA expression, peptide sequence, and sample identifier, excluding gene identifier, achieved the third-highest AUC of 0.96. Therefore, gene identifier can be identified as the second-least important variable in predicting the likelihood of a peptide being presented by an MHC class II molecule. The curve for exemplary Model 3, which generated a prediction of peptide presentation likelihood based on flanking sequence, RNA expression, gene identifier, and sample identifier, excluding peptide sequence, achieved the lowest AUC of 0.88. Therefore, peptide sequence can be identified as the most important variable in predicting the likelihood of a peptide being presented by an MHC class II molecule. The curve for exemplary Model 1, which generated a prediction of peptide presentation likelihood based on peptide sequence, flanking sequence, gene identifier, and sample identifier, excluding RNA expression, achieved the second-lowest AUC of 0.95. Therefore, RNA expression can be identified as the second-most important variable in predicting the likelihood of a peptide being presented by an MHC class II molecule.
[0465] FIG. 14I is a line graph comparing the performance of four different presentation models in predicting the likelihood that peptides in a test dataset of peptides will be presented by MHC class II molecules.
[0466] The first model tested in Figure 14I is referred to herein as the "binding affinity model." The binding affinity model in Figure 14I is the NetMHCII2.3 model, a best-in-class model that uses the minimum NetMHCII2.3 predicted binding affinity as the basis for generating predictions. Specifically, the NetMHCII2.3 model generates predictions of peptide presentation likelihood based on MHC class II molecule type and peptide sequence. The NetMHCII2.3 model is available on the NetMHCII2.3 website (www.cbs.dtu.dk / services / NetMHCII / , PMID 29315598). 76 The test was carried out using
[0467] The second model tested in Figure 14I is referred to herein as "MLP." The MLP (Multilayer Perceptron) model uses the allele-non-interacting variable w k and allele interaction variable x h k is an embodiment of the proposed model described above in which the alleles are input to separate dependency functions, such as neural networks, and then the outputs of these separate dependency functions are added together. Specifically, the fully non-interacting model is k is the dependency function g w and the allele interaction variable x h k is another dependency function g h and the dependency function g w and the dependency function g h is one embodiment of the presentation model described above, in which the outputs of alleles and alleles are added together. Thus, in some embodiments, the fully non-interacting model determines the likelihood of peptide presentation using Equation 8 shown above. Additionally, the allele non-interacting variable w k is the dependency function g wand the allele interaction variable x h k is another dependency function g h and the dependency function g w and the dependency function g h An embodiment of the fully non-interacting model in which the outputs of are summed is described in detail above with respect to the first half of Section VIII.B.2, the second half of Section VIII.B.3, the first half of Section VIII.C.3, and the first half of Section VIII.C.6.
[0468] The third model tested in Figure 14I is referred to herein as "RNN." RNN models include recurrent neural networks and are similar to the fully non-interactive model described above. However, the layers of the recurrent neural network in the RNN model differ from the layers of the neural network in the MLP model. Specifically, the input layer of the recurrent neural network in the RNN model accepts a variable-length peptide string, which models one peptide at a time. This peptide is fed to a neural network node, one amino acid at a time, and the output is piped to the node's input along with the next amino acid in the sequence until the entire sequence is modeled. Recurrent layers are particularly applicable to modeling MHC class II peptides for two reasons: (1) the continuity of the data is captured in the model, and (2) the length of the peptide can be varied without artificial padding. The next layer in the recurrent neural network is a dropout layer with p = 0.2 and finally a fully connected 64-node layer using ReLu activation.
[0469] The fourth model tested in Figure 14I is referred to herein as "Bi-LSTM." The Bi-LSTM model consists of a bidirectional long short-term memory neural network. The Bi-LSTM model is identical to the non-interactive model except for the peptide input layer. The input layer of the Bi-LSTM model accepts a 20-mer peptide string and then embeds this 20-mer peptide string as an (n, 20, 21) tensor. Each subsequent layer of the Bi-LSTM model's bidirectional long short-term memory neural network consists of a 128-node recurrent long short-term memory layer, a dropout layer with p = 0.2, and finally a fully connected 64-node layer with ReLu activation. In traditional LSTM models, the order of sequential data is assumed to be directional (e.g., read from left to right or right to left). In bidirectional LSTM, sequential data is processed in both directions, proceeding from left to right and right to left. Because peptide binding is essentially a non-directional task, modeling the sequence in both directions ensures that information from both ends of the sequence carries equal weight in the model's predictions.
[0470] Briefly referring to FIG. 14J , FIG. 14J illustrates an exemplary embodiment of the Bi-LSTM model of FIG. 14I configured to predict peptide presentation by HLA-DRB (an MHC class II gene). As shown in FIG. 14J , the Bi-LSTM model includes a shared neural network that accepts allele-noninteracting features (e.g., RNA sequence, sample number, protein number, and flanking sequence) and a set of separate neural networks, each associated with a different HLA-DRB allele and configured to accept an encoded peptide sequence (allele-interacting feature). Each separate neural network in the set of neural networks includes a Bi-LSTM neural network. In the exemplary embodiment of the Bi-LSTM model of FIG. 14J , the set of separate neural networks associated with different alleles includes four separate neural networks because up to four different alleles are associated with the HLA-DRB gene per patient sample. However, in another embodiment in which the Bi-LSTM model is configured to predict peptide presentation by different HLA genes, the set of separate neural networks includes a number of separate neural networks equal to the maximum possible amount of each allele in a patient sample for a given HLA gene. Each separate neural network in the set of neural networks determines the likelihood that a peptide input to the model will be presented by the HLA-DRB allele associated with the given neural network. Each of these likelihoods is then combined with the output from the shared neural network. Finally, the combined likelihoods are summed to generate an overall likelihood that the peptide will be presented by the HLA-DRB gene.
[0471] Referring again to FIG. 14I, before using each of the four models in FIG. 14I to predict the likelihood that a peptide in the peptide test dataset would be presented by an MHC class II molecule, each model was trained and validated. The binding affinity model was trained and validated using its own training and validation dataset based on HLA peptide binding affinity assays deposited in the Immune Epitope Database (IEDB, www.iedb.org). The other three models were trained using the training dataset of 69 samples described above and validated using the validation dataset described above. Following this training and validation of each model, each of the four models was tested using four omitted tumor samples from the test dataset described above. Specifically, for each of the four models, each peptide from the four omitted tumor samples from the test dataset was input into the model, and the model then output the presentation likelihood of that peptide.
[0472] The performance of each of these four models is shown in the line graphs of Figure 14I. Specifically, each of the four models is associated with an ROC curve that shows the ratio of the true positive rate to the false positive rate for each prediction generated by the model. For example, Figure 14I shows the ROC curves for the binding affinity model, the ROC curve for the RNN model, the ROC curve for the MLP model, and the ROC curve for the Bi-LSTM model.
[0473] As mentioned above, the performance of a model in predicting the likelihood that a peptide will be presented by an MHC class II molecule is quantified by determining the AUC of the ROC curve, which indicates the ratio of true positives to false positives for each prediction generated by the model. Models with a higher AUC have better performance (i.e., higher accuracy) compared to models with a lower AUC. As shown in Figure 14I, the Bi-LSTM model curve achieved the highest AUC of 0.98. Therefore, the Bi-LSTM model achieved the best performance. This peak performance of the Bi-LSTM model is due in part to Bi-LSTM's greatest ability to accurately predict peptides of variable length, relatively long length, and peptides with repeating amino acids. The MLP and RNN model curves achieved the second highest AUC of 0.97. Therefore, the MLP and RNN models achieved the second best performance. The binding affinity model curve achieved the lowest AUC of 0.79. Therefore, the binding affinity model performed the worst. Note that each of the Bi-LSTM, MLP, and RNN models tested in Figure 14I has an AUC greater than 0.9. Thus, despite differences in the architecture of the peptide input layer between these models, these models are able to obtain relatively accurate predictions of peptide presentation, unlike the binding affinity model, which has a significantly lower AUC.
[0474] Figure 14K is a line graph showing the perfect precision-recall curves for the "Bi-LSTM," "MLP," "RNN," and "binding affinity" models described above with respect to Figure 14I. As shown in Figure 14K, and as expected based on Figure 14I, the "Bi-LSTM" model achieved the best performance with AUC = 0.23, the "RNN" model achieved the second-best performance with AUC = 0.16, the "MLP" model achieved the third-best performance with AUC = 0.11, and the "binding affinity" model achieved the worst performance with AUC = 0.01. Specifically, the Bi-LSTM model trained on mass spectrometry data significantly outperformed the binding affinity model, increasing AUC by more than 20-fold.
[0475] Figure 14L is a line graph comparing the performance of two exemplary best-in-class conventional models given two different criteria and two exemplary models given two different sets of allele interaction and allele non-interaction variables in predicting the likelihood that a peptide in a test dataset of peptides will be presented by an MHC class II molecule. Specifically, Figure 14L is a line graph comparing the performance of an exemplary best-in-class conventional model (exemplary Model 1) that uses the binding affinity predicted by min NetMHCII2.3 as the criterion to generate predictions, an exemplary best-in-class conventional model (exemplary Model 2) that uses the binding rank predicted by min NetMHCII2.3 as the criterion to generate predictions, an exemplary model (exemplary Model 4) that generates predictions of peptide presentation likelihood based on MHC class II molecule type and peptide sequence, and an exemplary model (exemplary Model 3) that generates predictions of peptide presentation likelihood based on MHC class II molecule type, peptide sequence, RNA expression, gene identifier, and flanking sequence.
[0476] The best-in-class conventional model used as Exemplary Model 1 and Exemplary Model 2 in Figure 14L is the NetMHCII2.3 model. The NetMHCII2.3 model generates predictions of peptide presentation likelihood based on the type of MHC class II molecule and the peptide sequence. The NetMHCII2.3 model can be found on the NetMHCII2.3 website (www.cbs.dtu.dk / services / NetMHCII / , PMID 29315598). 76 The test was carried out using
[0477] As mentioned above, the NetMHCII2.3 models were tested according to two criteria: specifically, Exemplary Model 1 generated predictions of peptide presentation likelihood according to the binding affinity predicted by minimum NetMHCII2.3, and Exemplary Model 2 generated predictions of peptide presentation likelihood according to the binding rank predicted by minimum NetMHCII2.3.
[0478] The display models used as Exemplary Model 3 and Exemplary Model 4 are embodiments of the display models disclosed herein that are trained using data obtained by mass spectrometry. As noted above, the display models generated predictions of peptide presentation likelihood based on two different sets of allele-interacting and allele-non-interacting variables. Specifically, Exemplary Model 4 generated predictions of peptide presentation likelihood based on MHC class II molecule type and peptide sequence (the same variables used in the NetMHCII2.3 model), and Exemplary Model 3 generated predictions of peptide presentation likelihood based on MHC class II molecule type, peptide sequence, RNA expression, gene identifier, and flanking sequence.
[0479] Prior to using the exemplary models in Figure 14L to predict the likelihood that peptides in a test dataset of peptides will be presented by MHC class II molecules, each model was trained and validated. The NetMHCII2.3 models (Exemplary Model 1 and Exemplary Model 2) were trained and validated using their own training and validation datasets based on HLA peptide binding affinity assays deposited in the Immune Epitope Database (IEDB, www.iedb.org). The training dataset used to train the NetMHCII2.3 models is known to contain almost exclusively 15-mer peptides. In contrast, Exemplary Models 3 and 4 were trained using the training dataset described above with respect to Figure 14H and validated using the validation dataset described above with respect to Figure 14H.
[0480] Following training and validation of each model, each model was tested using a test dataset. As mentioned above, the NetMHCII2.3 model was trained on a dataset containing almost exclusively 15-mer peptides. This meant that NetMHCII2.3 could not prioritize peptides of different weights, which reduced its predictive performance on mass spectrometry data of HLA class II presentation, which contained peptides of all lengths. Therefore, to provide a fair comparison between models unaffected by varying peptide lengths, the test dataset contained only 15-mer peptides. Specifically, the test dataset contained 933 15-mer peptides. Of the 933 peptides in the test dataset, 40 were presented by MHC class II molecules, specifically, HLA-DRB1*07:01, HLA-DRB1*15:01, HLA-DRB4*01:03, and HLA-DRB5*01:01 molecules. The peptides included in the test dataset were excluded from the training dataset described above.
[0481] To test each exemplary model using the test dataset, for each of the exemplary models, a prediction of the likelihood of presentation of the peptide was generated for each of the 933 peptides in the test dataset. Specifically, for each peptide in the test dataset, the model in Example 1 generated a presentation score for that peptide by an MHC class II molecule by using the type of MHC class II molecule and the peptide sequence, and ranking the peptide by the binding affinity predicted by minimum NetMHCII2.3 across the four HLA class II DR alleles in the test dataset. Similarly, for each peptide in the test dataset, the model in Example 2 generated a presentation score for that peptide by an MHC class II molecule by using the type of MHC class II molecule and the peptide sequence, and ranking the peptide by the binding rank (i.e., quantile-normalized binding affinity) predicted by minimum NetMHCII2.3 across the four HLA class II DR alleles in the test dataset. For each peptide in the test dataset, the model in Example 4 generated a likelihood of presentation of that peptide by an MHC class II molecule based on the type of MHC class II molecule and the peptide sequence. Similarly, for each peptide in the test dataset, the model of Example 3 generated a likelihood of presentation of that peptide by an MHC class II molecule based on the type of MHC class II molecule, peptide sequence, RNA expression, gene identifier, and flanking sequence.
[0482] The performance of each of these four exemplary models is shown in the line graph of Figure 14L. Specifically, each of the four exemplary models is associated with an ROC curve that shows the ratio of the true positive rate to the false positive rate for each prediction generated by the model. For example, Figure 14L shows the ROC curve for the model of Example 1, which used binding affinities predicted by Minimum NetMHCII2.3 to generate predictions; the ROC curve for the model of Example 2, which used binding ranks predicted by Minimum NetMHCII2.3 to generate predictions; the ROC curve for the model of Example 4, which generated peptide presentation likelihoods based on MHC class II molecule type and peptide sequence; and the ROC curve for the model of Example 3, which generated peptide presentation likelihoods based on MHC class II molecule type, peptide sequence, RNA expression, gene identifier, and flanking sequence.
[0483] As discussed above, the performance of a model in predicting the likelihood that a peptide will be presented by an MHC class II molecule is quantified by determining the AUC of the ROC curve, which indicates the ratio of the true positive rate to the false positive rate for each prediction generated by the model. Models with a higher AUC have better performance (i.e., higher accuracy) compared to models with a lower AUC. As shown in Figure 14L, the curve for the model in Example 3, which generated the peptide presentation likelihood based on the MHC class II molecule type, peptide sequence, RNA expression, gene identifier, and flanking sequence, achieved the highest AUC of 0.95. Therefore, the model in Example 3, which generated the peptide presentation likelihood based on the MHC class II molecule type, peptide sequence, RNA expression, gene identifier, and flanking sequence, achieved the best performance. The curve for the model in Example 4, which generated the peptide presentation likelihood based on the MHC class II molecule type and peptide sequence, achieved the second highest AUC of 0.91. Therefore, the model in Example 4, which generated the peptide presentation likelihood based on the MHC class II molecule type and peptide sequence, achieved the second best performance. The curve for the model in Example 1, which used binding affinity predicted by minimum NetMHCII2.3 to generate the predictions, had the lowest AUC of 0.75. Therefore, the curve for the model in Example 1, which used binding affinity predicted by minimum NetMHCII2.3 to generate the predictions, was the worst performing. The curve for the model in Example 2, which used binding rank predicted by minimum NetMHCII2.3 to generate the predictions, had the second lowest AUC of 0.76. Therefore, the curve for the model in Example 2, which used binding rank predicted by minimum NetMHCII2.3 to generate the predictions, was the second worst performing.
[0484] As shown in Figure 14L, the performance gap between exemplary models 1 and 2 and exemplary models 3 and 4 is large. Specifically, the performance of the NetMHCII2.3 model (which uses either the criteria of binding affinity predicted by minimum NetMHCII2.3 or binding rank predicted by minimum NetMHCII2.3) is approximately 25% lower than the performance of the presentation model disclosed herein (which generates peptide presentation likelihoods based on MHC class II molecule type and peptide sequence, or MHC class II molecule type, peptide sequence, RNA expression, gene identifier, and flanking sequence). Thus, Figure 14L demonstrates that the presentation model disclosed herein can achieve significantly more accurate presentation predictions than the NetMHCII2.3 model, the current best-in-class conventional model.
[0485] Furthermore, as mentioned above, the NetMHCII2.3 model is trained on a training dataset containing almost exclusively 15-mer peptides. As a result, the NetMHCII2.3 model is not trained to learn which peptide lengths are more likely to be presented by MHC class II molecules. Therefore, the NetMHCII2.3 model does not weight its prediction of the likelihood of peptide presentation by MHC class II molecules according to peptide length. In other words, the NetMHCII2.3 model does not modify its prediction of the likelihood of peptide presentation by MHC class II molecules for peptides with lengths outside the most common peptide length of 15 amino acids. As a result, the NetMHCII2.3 model over-predicts the likelihood of presentation of peptides with lengths longer or shorter than 15 amino acids.
[0486] In contrast, the display models disclosed herein are trained using peptide data obtained by mass spectrometry, and therefore can be trained on a training dataset containing peptides of all different lengths. As a result, the models disclosed herein can learn which peptide lengths are more likely to be presented by MHC class II molecules. Therefore, the display models disclosed herein can weight their predictions of the likelihood of peptide presentation by MHC class II molecules according to peptide length. In other words, the display models disclosed herein can modify their predictions of the likelihood of peptide presentation by MHC class II molecules for peptides with lengths outside the most common peptide length of 15 amino acids. As a result, the display models disclosed herein can achieve significantly more accurate presentation predictions for peptides longer or shorter than 15 amino acids compared to the NetMHCII2.3 model, the current best-in-class conventional model. This is one advantage of using the display models disclosed herein to predict the likelihood of peptide presentation by MHC class II molecules.
[0487] XII.A.3. Example 3 Figure 14M is a histogram showing the amount of peptides sequenced using mass spectrometry at a q value of less than 0.1 for each of 230 samples, including human tumors (NSCLC, lymphoma, and ovarian cancer) and cell lines (EBV) containing HLA class II molecules. As shown in Figure 14M, an average of 1300 peptides were sequenced for each sample at a q value of less than 0.1.
[0488] As described above with respect to Figure 14D, each of the 230 samples in Figure 14M contained an HLA class II molecule. More specifically, each of the 230 samples in Figure 14M contained an HLA-DR molecule. An HLA-DR molecule is a type of HLA class II molecule. Even more specifically, each of the 230 samples in Figure 14M contained an HLA-DRB1 molecule, an HLA-DRB3 molecule, an HLA-DRB4 molecule, and / or an HLA-DRB5 molecule. HLA-DRB1 molecule, an HLA-DRB3 molecule, an HLA-DRB4 molecule, and an HLA-DRB5 molecule are types of HLA-DR molecules.
[0489] While this particular experiment was performed using a sample containing HLA-DR molecules, specifically HLA-DRB1, HLA-DRB3, HLA-DRB4, and HLA-DRB5 molecules, in alternative embodiments, the experiment can be performed using a sample containing one or more of any type(s) of HLA class II molecules. For example, in alternative embodiments, the same experiment can be performed using a sample containing HLA-DP and / or HLA-DQ molecules. Those skilled in the art will appreciate that the same method can be used to model any type(s) of MHC class II molecules and still obtain reliable results. See, e.g., Jensen, Kamilla Kjaergaard et al. 76 is an example of a recent scientific paper that uses the same method to model binding affinity for HLA-DR molecules, as well as for HLA-DP and HLA-DQ molecules. Thus, one skilled in the art will understand that the experiments and models described herein can be used to model not only HLA-DR molecules, but also any other MHC class II molecules, separately or simultaneously, and still obtain reliable results.
[0490] Mass spectrometry was performed on each sample to sequence the peptides in each of the 230 samples. The mass spectra obtained for the samples were searched using Comet and scored using Percolator to sequence the peptides. The amount of sequenced peptides in the sample was then determined for several different Percolator q-value thresholds. Specifically, the amount of sequenced peptides for the sample was determined using a Percolator q-value of less than 0.01, a Percolator q-value of less than 0.05, and a Percolator q-value of less than 0.2.
[0491] The amount of peptides sequenced at each of the different Percolator q-value thresholds for each of the 203 samples is shown in Figure 14M. For example, as can be seen in Figure 14M, for the first sample, approximately 8000 peptides with q-values below 0.1 were sequenced using mass spectrometry.
[0492] Overall, Figure 14M shows that mass spectrometry can be used to sequence a large number of peptides from samples containing MHC class II molecules at low q values. In other words, the data shown in Figure 14M demonstrates that mass spectrometry can be used to sequence with high confidence peptides that can be presented by MHC class II molecules.
[0493] Figure 14N is a histogram showing the amount of samples in which a particular MHC class II molecule allele was identified. More specifically, Figure 14N shows the amount of samples in which a particular MHC class II molecule was identified for a total of 230 samples containing HLA class II molecules.
[0494] As noted above with respect to Figure 14M, each of the 230 samples in Figure 14M contained an HLA-DRB1 molecule, an HLA-DRB3 molecule, an HLA-DRB4 molecule, and / or an HLA-DRB5 molecule. Accordingly, Figure 14N shows the amount of a sample in which a particular allele was identified for HLA-DRB1, HLA-DRB3, HLA-DRB4, and HLA-DRB5 molecules.
[0495] To identify which HLA-DRB1, HLA-DRB3, HLA-DRB4, and HLA-DRB5 alleles were present in the samples, HLA class II DR typing was performed on the samples. To determine the amount of samples in which a particular HLA allele was identified, the number of samples in which the HLA allele was identified using HLA class II DR typing was simply added up. For example, as shown in Figure 14N, 28 samples out of a total of 230 samples contained the HLA class II molecule allele HLA-DRB3*03:01. In other words, 28 samples out of a total of 230 samples contained the HLA-DRB3*03:01 allele for the HLA-DRB3 molecule. Overall, Figure 14N demonstrates that a wide range of HLA class II molecule alleles can be identified from 230 samples containing HLA class II molecules. Regarding human populations, the allele frequencies of HLA-DRB1 alleles in the Caucasian population can be found in U.S. Patent Application No. Maiers, M, et al. 161 .
[0496] FIG. 14O shows peptides bound to MHC class I molecules and peptides bound to MHC class II molecules. 162 As shown in Figure 14O, each peptide comprises a peptide backbone and multiple amino acids. Each MHC molecule has a binding groove. However, as described below, the manner in which peptides bind within the binding groove of MHC class I and MHC class II molecules differs.
[0497] As described throughout this disclosure, the length of peptides presented by MHC molecules can vary. Specifically, peptides presented by MHC molecules can be 9 to 20 amino acids long. When a peptide binds to an MHC molecule and is presented by the MHC molecule, the "binding core" of the peptide is located within the binding groove of the MHC molecule. Specifically, the binding core of a peptide is the amino acid sequence of the peptide that is located within the binding groove of the MHC molecule when the peptide binds to the MHC molecule and is presented by the MHC molecule. Furthermore, when a peptide binds to an MHC molecule and is presented by the MHC molecule, the "binding anchor" of the binding core of the peptide physically binds within the binding groove of the MHC molecule. Specifically, the binding anchor of the binding core of the peptide is a specific amino acid of the binding core that binds to the binding groove of the MHC molecule when the peptide binds to the MHC molecule and is presented by the MHC molecule.
[0498] As shown in Figure 14O, the binding core of a peptide presented by an MHC class I molecule comprises the entire peptide. Specifically, as shown in Figure 14O, the entire peptide presented by an MHC class I molecule is located within the binding groove of the MHC class I molecule. In contrast, in a molecule presented by an MHC class II molecule, only a partial sequence of amino acids of the peptide may be included in the binding core of the peptide. Specifically, as shown in Figure 14O, both ends of the peptide presented by an MHC class II molecule are not located within the binding groove of the MHC class II molecule. The partial sequence of am...
Claims
[Claim 1] The invention described herein.
Citation Information
Patent Citations
Prediction of Immunogenicity of T-Cell Epitopes
JP2016521128A
Neoantigen identification, manufacture, and use
WO2018195357A1