Immunogenic mutant peptide screening platform

JP2024156702A5Pending Publication Date: 2025-06-17GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024112292
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2014-09-10
Filing Date
2024-07-12
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Current methods for identifying immunogenic epitopes, particularly for cancer therapy, are time-consuming, costly, and inefficient due to the patient-specific nature of somatic mutations in cancer cells, and existing predictive algorithms are not effective for high-throughput screening of personalized immunogenic epitopes.

Method used

A method combining sequence-based variant identification with immunogenicity prediction and mass spectrometry to identify disease-specific immunogenic variant peptides by analyzing variant coding sequences, predicting their binding to MHC molecules, and validating their immunogenicity through functional analysis.

Benefits of technology

Enables efficient and personalized identification of immunogenic peptides for cancer therapy, facilitating high-throughput screening and validation of peptides that can stimulate targeted immune responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000043_0000
    Figure 00000043_0000
  • Figure 00000043_0001
    Figure 00000043_0001
  • Figure 00000043_0002
    Figure 00000043_0002
Patent Text Reader

Abstract

To provide methods of identifying a disease-specific immunogenic peptide which may be performed in a high-throughput manner.SOLUTION: A method of identifying a disease-specific immunogenic mutant peptide from a disease tissue in an individual is provided, comprising (a) providing a set of variant-coding sequences of the disease tissue in the individual, each variant-coding sequence having a sequence variation compared to a reference sample; and (b) selecting immunogenic variant-coding sequences from the set of variant-coding sequences, comprising predicting the immunogenicity of the peptide comprising variant amino acids encoded by variant-coding sequences. This process enables the identification of the disease-specific immunogenic mutant peptide.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 048,742, filed Sep. 10, 2014, the entire contents of which are incorporated herein by reference for all purposes.

[0002] Submitting a sequence listing as an ASCII text file The contents of the following ASCII text file submission are incorporated by reference in their entirety into this specification: Sequence Listing in Computer Readable Form (CRF) (Filename: 146392027600SeqList.txt, Date Recorded: September 9, 2015, Size: 6KB).

[0003] The present disclosure relates to methods for identifying variant peptides useful for developing immunotherapies. [Background technology]

[0004] Cytotoxic T lymphocytes (cytotoxic T cells or CD8 T cells), involved in cell-mediated immunity, monitor changes in cellular health by scanning for peptide epitopes or antigens on the cell surface. Peptide epitopes are derived from cellular proteins and serve as a display mechanism that allows cells to present evidence of current cellular processes. Both native and non-native proteins (often referred to as self and non-self, respectively) are processed for peptide epitope presentation. Most self-peptides are derived from natural protein turnover and defective ribosomal products. Non-self peptides can be derived from proteins generated during events such as viral and bacterial infections, diseases, and cancer.

[0005] Human tumors are characterized by carrying a significant number of somatic mutations. Thus, expression of peptides containing mutations can be recognized as non-self neoepitopes by the adaptive immune system. Upon recognition of non-self antigens, cytotoxic T cells trigger an immune response, resulting in apoptosis of cells displaying non-self neoepitopes. Cytotoxic T cell adaptive immunity is a highly specific mechanism and an efficient means to eliminate infected, diseased, and cancerous cells. Identifying immunogenic epitopes is of great therapeutic value, since exposure to immunogenic epitopes via vaccination can be used to trigger the desired cytotoxic T cell immune response. Although the role of immunogenic epitopes has been known to the scientific and medical communities for decades, the identification of antigens that drive effective anti-tumor CD8 T cell responses remains largely unknown. The complexities involved in epitope presentation and cytotoxic T cell activation have made their discovery and therapeutic use difficult.

[0006] Major histocompatibility complex (MHC) class I molecules are responsible for peptide epitope presentation to cytotoxic T cells. In humans, the human leukocyte antigen (HLA) system is the genetic locus that codes for MHC class I and class II molecules. HLA-A, -B, and -C genes code for MHC class I (MHCI) proteins. Peptides, typically 8-11 amino acids in length, bind to MHCI molecules through interactions with a groove formed by two alpha helices located on top of an antiparallel beta sheet. Processing and presentation of peptide-MHC class I (pMHCI) molecules involves a series of sequential steps including a) protease-mediated digestion of the protein, b) peptide transport into the endoplasmic reticulum (ER) mediated by transporters associated with antigen processing (TAPs), c) formation of pMHCI using newly synthesized MHCI molecules, and d) transport of pMHCI to the cell surface.

[0007] On the cell surface, pMHCI interacts with cytotoxic T cells via the T cell receptor (TCR). Following complex pMHCI-TCR interactions, identification of non-self antigens can lead to cytotoxic T cell activation through a series of biochemical events mediated by associated enzymes, co-receptors, adaptor molecules, and transcription factors. The activated cytotoxic T cells proliferate and generate an effector T cell population expressing a TCR specific for the identified immunogenic peptide epitope. Expansion of T cells with TCR specificity for the identified non-self epitope leads to immune-mediated apoptosis of cells displaying the activating non-self epitope.

[0008] The use of immunogenic epitopes to activate the immune system is currently being investigated for use in cancer therapy. While cancer cells originate from normal tissues, somatic mutations result in numerous alterations in the cancer proteasome. The resulting MHCI-presented peptide epitopes, termed tumor-associated antigens (TAA) or neoepitopes, enable cytotoxic T cell differentiation between normal and cancer tissues. Recent studies have established that mutant peptides can serve as epitopes recognized as non-self by CD4 or CD8 T cells, but few mutant neoepitopes have been described since.

[0009] The use of peptide-based immunotherapy depends on the selection of peptide epitopes that stimulate the desired cytotoxic T cell response. Specifically, tumor antigens can be divided into two categories: tumor-associated self-antigens (e.g., cancer-testis antigens, differentiation antigens) and antigens derived from shared or patient-specific mutant proteins. Mutant neoantigens are likely to be more immunogenic, since the presentation of self-antigens in the thymus can result in the elimination of high-avidity T cells. The development of such epitopes for immunotherapy is a challenging pursuit, and efficient methods useful for the identification of effective epitopes have yet to be developed.

[0010] The time-consuming and costly nature of identifying and validating immunogenic peptide epitopes has hindered the development of effective peptide-based cancer vaccinations. Further complicating the problem of identifying immunogenic epitopes is that the permutations of mutations in cancer cells are often patient-specific. Discovery of mutant neoepitopes requires laborious screening of a patient's tumor-infiltrating lymphocytes for their ability to recognize antigens from libraries constructed based on information from the patient's tumor exome sequence. Alternatively, mutant neoepitopes can be detected by mass spectrometry. However, mutant sequences have escaped detection because their identification cannot be achieved using public proteome databases that do not contain patient-specific variants. The use of predictive algorithms such as peptide-MHCI binding or peptide immunogenicity may be applied to identify personalized immunogenic epitopes. However, the vast number of somatic mutations and expression level changes contained within cancer cells results in a scale of predicted immunogenic epitopes that is too large for high-throughput immunogenicity screening. Furthermore, evidence of poor immunogenicity of predicted epitopes has raised doubts about the utility of current methods.

[0011] There is a need in the art to identify immunogenic epitopes suitable for use in peptide-based immunotherapy.Specifically, there is a need in the art to identify immunogenic epitopes for use in peptide-based cancer therapy.Furthermore, there is a need in the art for high-throughput methods for predicting immunogenic epitopes based on personalized genetic and / or proteomic analysis.

[0012] All references cited herein are hereby specifically incorporated by reference. Summary of the Invention

[0013] In one aspect, the application provides a method for identifying disease specific immunogenic variant peptides comprising: a) providing a set of variant coding sequences of disease tissue in an individual, each variant coding sequence having a sequence difference compared to a reference sample; and b) selecting immunogenic variant coding sequences from the set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the variant coding sequences, thereby identifying disease specific immunogenic variant peptides. In some embodiments, the method comprises: a) obtaining a first set of variant coding sequences based on a genomic sequence of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of diseased tissue in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the expressed variant coding sequences to bind to MHC class I molecules (MHCI); and d) selecting immunogenic variant coding sequences from the third set, which comprise predicting the immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides.

[0014] In some embodiments, according to any of the above-mentioned methods, the method further comprises: i) obtaining a plurality of peptides bound to MHCI molecules from the diseased tissue; ii) subjecting the MHCI-binding peptides to mass spectrometry-based sequencing; and iii) correlating the mass spectrometry-derived sequence information of the MHCI-binding peptides with immunogenic variant coding sequences.

[0015] In another aspect, the application provides a method of identifying disease specific immunogenic variant peptides comprising: a) obtaining a plurality of peptides bound to MHC molecules from diseased tissue of an individual; b) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and c) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with a set of variant coding sequences of diseased tissue in the individual, where each variant coding sequence has a sequence difference compared to a reference sample, thereby identifying the disease specific immunogenic variant peptides.

[0016] In some embodiments, the method includes: a) obtaining a first set of variant coding sequences based on the genomic sequence of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on the transcriptome sequence of diseased tissue in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on the predicted ability of peptides encoded by the expressed variant coding sequences to bind to MHC class I molecules (MHCI); d) obtaining a plurality of peptides bound to MHCI molecules from the diseased tissue; e) subjecting the MHCI-binding peptides to mass spectrometry-based sequencing; and f) correlating the mass spectrometry-derived sequence information of the MHCI-binding peptides with the third set of epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes predicting the immunogenicity of the disease-specific immunogenic variant peptides, wherein the disease-specific immunogenic variant peptides comprise variant amino acids.

[0017] According to any of the above-mentioned methods comprising a step of predicting immunogenicity, predicting immunogenicity is based on one or more of the following parameters: i) binding affinity of the peptide to MHCI molecule, ii) protein level of peptide precursor containing the peptide, iii) expression level of transcript encoding the peptide precursor, iv) processing efficiency of the peptide precursor by the immunoproteasome, v) timing of expression of transcript encoding the peptide precursor, vi) binding affinity of the peptide to TCR molecule, vii) position of mutated amino acid in the peptide, viii) solvent exposure of the peptide when bound to MHCI molecule, ix) solvent exposure of mutated amino acid when bound to MHCI molecule, x) content of aromatic residues in the peptide, xi) properties of mutated amino acid compared to wild type residue, and xii) properties of peptide precursor. In some embodiments, predicting immunogenicity is further based on HLA typing analysis.

[0018] In some embodiments, according to any one of the above-mentioned methods comprising obtaining a plurality of peptides bound to MHCI molecules from diseased tissue, the peptides bound to MHCI are obtained by isolating MHCI / peptide complexes from diseased tissue and eluting the peptides from MHCI. In some embodiments, the isolation of MHCI / peptide complexes is carried out by immunoprecipitation. In some embodiments, the immunoprecipitation is carried out using an antibody specific to MHCI. In some embodiments, the isolated peptides are further separated by chromatography before being subjected to mass spectrometry.

[0019] In some embodiments, according to any one of the above-mentioned methods comprising the step of obtaining a first set of mutant coding sequences, obtaining the first set of mutant coding sequences comprises: i) obtaining a first set of mutant sequences based on a genomic sequence of diseased tissue in the individual, each mutant sequence having a sequence difference compared to a reference sample; and ii) identifying mutant coding sequences from the first set of mutant sequences.

[0020] In some embodiments, according to any of the above methods, the method further comprises synthesizing a peptide based on the sequence of the identified disease-specific immunogenic variant peptide. In some embodiments, according to any of the above methods, the method further comprises synthesizing a nucleic acid encoding a peptide based on the sequence of the identified disease-specific immunogenic variant peptide. In some embodiments, the method further comprises testing the synthesized peptide for immunogenicity in vivo.

[0021] In some embodiments, according to any of the methods described above, the disease is cancer. In some embodiments, according to any of the methods described above, the individual is a human.

[0022] In another aspect, the present application also provides a disease-specific variant peptide or a composition of disease-specific variant peptides identified by any of the methods described herein. In some embodiments, the composition comprises two or more disease-specific immunogenic variant peptides described herein. In some embodiments, the composition further comprises an adjuvant.

[0023] In yet another aspect, the present application also provides a method of treating a disease in an individual comprising administering to the individual an effective amount of a composition comprising a disease-specific variant peptide identified using any of the methods for identifying a disease-specific immunogenic variant peptide disclosed herein. In some embodiments, the individual is the same individual from whom the disease-specific immunogenic variant peptide was identified.

[0024] The present application also provides an immunogenic composition comprising at least one disease-specific peptide or a precursor of such a disease-specific peptide, said disease-specific peptide being identified by any of the methods described herein. In some embodiments, the immunogenic composition comprises a plurality of disease-specific peptides.

[0025] The present application also provides an immunogenic composition comprising at least one nucleic acid encoding a disease-specific peptide, said disease-specific peptide being identified by any of the methods described herein. In some embodiments, the immunogenic composition comprises a plurality of nucleic acids, each encoding at least one disease-specific peptide. In some embodiments, the immunogenic composition comprises nucleic acids encoding two or more (e.g., any number of 3, 4, 5, 6, 7, 8, 9, or more) disease-specific peptides.

[0026] The present application in yet another aspect also provides a method of stimulating an immune response in an individual having a disease, comprising administering any of the immunogenic compositions described herein. In some embodiments, the method further comprises administering another agent. In some embodiments, the other agent is an immunomodulatory agent. In some embodiments, the other agent is a checkpoint protein. In some embodiments, the other agent is an antagonist of PD-1 (e.g., an anti-PD1 antibody). In some embodiments, the other agent is an antagonist of PD-L1 (e.g., an anti-PD-L1 antibody).

[0027] In yet another aspect, the present application also provides a method of stimulating an immune response in an individual having a disease, comprising: a) identifying a disease-specific immunogenic variant peptide from diseased tissue in the individual by any one of the identification methods described above; b) producing a composition comprising a peptide or a nucleic acid encoding the peptide based on the sequence of the identified disease-specific immunogenic variant peptide; and c) administering the composition to the individual. In some embodiments, the method further comprises administering a PD-1 antagonist (e.g., an anti-PD1 antibody) to the individual. In some embodiments, the method further comprises administering a PD-L1 antagonist (e.g., an anti-PD-L1 antibody) to the individual. [Brief description of the drawings]

[0028] [Figure 1]FIG. 1 shows an exemplary method for immunogenic peptide identification. [Diagram 2] A) Distribution of identified genes identified as epitopes presented on MHC molecules of MC-38 cell line against measured reads per kilobase of exon model per million mapped reads (RPKM). B) Distribution of identified genes identified as epitopes presented on MHC molecules of TRAMP-C1 cell line against measured RPKM. [Diagram 3] FIG. 1 shows structural modelling of a peptide bound to an MHC molecule. [Figure 4A-B] A) Percentage of peptide-specific CD8 T cells in wild-type C57BL / 6 mice immunized with selected peptides, and B) Percentage of dextramer-positive CD8 T cells in the spleen and tumor. [Figure 4C-D] (C) Measurement of CD8 T cells and CD45 T cells relative to tumor volume. (D) Percentage of tumor-specific CD8 TILs co-expressing PD-1 and TIM-3 in the total and Adpgk-positive CD8 TIL populations. [Figure 5A-B] A shows the tumor volumes of control and immunogenic vaccine-treated mice after tumor challenge with MC-38 tumor cells, and the percentage of Adpgk-positive CD8 T cells after vaccination. Arrows indicate measurements from a single animal. B shows the percentage of peptide-specific CD8 T cells in the spleen and tumor. [Figure 5C-D] (C) Percentage of live cells in the tumor measured as CD45-expressing and CD8-expressing T cells. (D) Percentage of Adgpk-specific CD8 TILs co-expressing PD-1 and TIM-3 in the total CD8 T cell population after vaccination. [Figure 5E] FIG. 1 shows levels of PD-1 and TIM-3 surface expression after vaccination. [Figure 5F]FIG. 1 shows the percentage of IFN-γ expressing CD8 and CD4 TILs in tumors and spleens after vaccination. [Figure 5G] FIG. 1 shows tumor volume measurements after vaccination. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0029] The present application provides a highly efficient screening platform for identifying disease-specific immunogenic variant peptides. By combining sequence-based variant identification methods with immunogenicity prediction and / or mass spectrometry, the methods described herein allow for the robust and efficient identification of disease-specific immunogenic variant peptides from diseased tissues (e.g., tumor cells) of individuals. These peptides or nucleotide-based precursors (e.g., DNA or RNA) can be useful for a variety of different applications, such as developing vaccines, developing variant peptide-specific therapeutics (e.g., antibody therapeutics or T cell receptor ("TCR")-based therapeutics), and developing tools for monitoring the dynamics and distribution of T cell responses. For example, individual peptides or collections of peptides can be utilized to perform comparative binding affinity measurements, or can be multimerized to measure antigen-specific T cell responses by MHC multimer flow cytometry. The methods described herein are particularly useful in the context of personalized medicine, where variant peptides identified from diseased individuals can be used to develop therapeutics (e.g., peptide, DNA, or RNA-based vaccines) to treat the individuals.

[0030] Thus, the present application in one aspect provides a method for identifying disease-specific immunogenic variant peptides from diseased tissue of an individual by combining sequence-based variant identification methods with immunogenicity prediction.

[0031] In another aspect, the application provides a method for identifying disease-specific immunogenic variant peptides from diseased tissue of an individual by combining sequence-based variant identification methods with mass spectrometry.

[0032] Also provided are kits and systems useful for the methods described herein. Further included are immunogenic compositions comprising peptides, cells that present such peptides, and nucleic acids that encode such identified peptides.

[0033] definition As used in this disclosure, the singular forms "a," "an," and "the" also encompass the plurals of the terms they refer to, unless the context clearly dictates otherwise. Reference herein to "about" a value or parameter includes (and describes) a variation directed to that value or parameter itself. For example, a description that refers to "about X" includes the description of "X."

[0034] It is understood that aspects and embodiments of the invention described herein include "consisting of" and / or "consisting essentially of" aspects and embodiments.

[0035] As used herein, a "disease-specific variant peptide" refers to a peptide that contains at least one variant amino acid that is present in diseased tissue but not in normal tissue. A "disease-specific immunogenic variant peptide" refers to a disease-specific variant peptide that is capable of eliciting an immune response in an individual. Disease-specific variant peptides can result from, for example, non-synonymous mutations (e.g., point mutations) that result in different amino acids in the protein; read-through mutations, in which a stop codon is modified or removed, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; splice site mutations, which result in the inclusion of an intron in the mature mRNA, thus resulting in a unique tumor-specific protein sequence; chromosomal rearrangements (i.e., gene fusions), which result in a chimeric protein with a tumor-specific sequence at the junction of two proteins; and frameshift mutations or deletions, which result in a new open reading frame with a new tumor-specific protein sequence. See, for example, Sensi and Anichini, Clin Cancer Res, 2006, v.12, 5023-5032.

[0036] As used herein, a "variant coding sequence" refers to a sequence that has differences compared to a sequence in a reference sample, where the sequence differences result in a change in the amino acid sequence contained in or encoded by the variant coding sequence. A variant coding sequence can be a nucleic acid sequence that has a mutation that results in an amino acid change in the encoded amino acid sequence. Alternatively, a variant coding sequence can be an amino acid sequence that contains an amino acid mutation.

[0037] An "expressed variant coding sequence" refers to a variant coding sequence that is expressed in diseased tissue of an individual.

[0038] A nucleic acid sequence that "encodes" a peptide refers to a nucleic acid that contains the coding sequence of the peptide. An amino acid sequence that "encodes" a peptide refers to an amino acid sequence that contains the sequence of the peptide.

[0039] "Epitope variant coding sequence" refers to a variant coding sequence that encodes a peptide that binds, or is predicted to bind, to an MHC molecule (eg, an MHC class I molecule, or MHCI).

[0040] "Immunogenic variant coding sequence" refers to a variant coding sequence that encodes a peptide that is predicted to be immunogenic.

[0041] As used herein, the term "disease tissue" refers to tissue associated with a disease in an individual and comprises a plurality of cells. "Disease tissue sample" refers to a sample of diseased tissue.

[0042] As used herein, "peptide precursor" refers to a polypeptide present in a diseased tissue of an individual that contains a peptide of interest. For example, a peptide precursor may be a polypeptide present in a diseased tissue that can be processed by the immunoproteasome to yield the peptide of interest.

[0043] Methods for identifying immunogenic variant peptides The method of the present application in an embodiment combines a sequence-specific variant identification method with an immunogenicity prediction method. For example, in some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissues in an individual, comprising: a) providing a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having a sequence difference compared to a reference sample; and b) selecting immunogenic variant coding sequences from the first set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissues in an individual that contribute as neo-epitopes in disease tissues. In some embodiments, the set of variant coding sequences comprises 1, 10, 100, 1,000, or more than 10,000 different variant coding sequences. In some embodiments, methods are provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual comprising: a) obtaining a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; and b) selecting immunogenic variant coding sequences from the first set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the variant coding sequences, thereby identifying the disease specific immunogenic variant peptides.In some embodiments, the selecting step comprises predicting the immunogenicity of the peptide based on one or more (e.g., any number of 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11) of the following parameters: i) binding affinity of the peptide to an MHCI molecule; ii) protein levels of the peptide precursor containing the peptide; iii) expression levels of transcripts encoding the peptide precursor; iv) efficiency of processing of the peptide precursor by the immunoproteasome; v) timing of expression of transcripts encoding the peptide precursor; vi) binding affinity of the peptide to a TCR molecule; vii) position of the mutated amino acid within the peptide; viii) solvent exposure of the peptide when bound to an MHCI molecule; ix) solvent exposure of the mutated amino acid when bound to an MHCI molecule; x) content of aromatic residues in the peptide; xi) properties of the mutated amino acid compared to the wild-type residue (e.g., change from charged to hydrophobic, or vice versa); and xii) properties of the peptide precursor.

[0044] In some embodiments, the first set of variant coding sequences can be first filtered to obtain a smaller set of variant coding sequences (referred to as "epitope variant coding sequences") that encode peptides predicted to bind to MHC molecules, and the smaller set of variant coding sequences is then subjected to selection based on predicted immunogenicity. In such embodiments, the method may include: a) providing a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of epitope variant coding sequences from the first set based on predicted ability of peptides encoded by the first set of variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); and c) selecting immunogenic variant coding sequences from the second set of epitope variant coding sequences, the selecting step including predicting immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method includes: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of epitope variant coding sequences from the first set based on the predicted ability of peptides encoded by the first set of variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); and c) selecting immunogenic variant coding sequences from the second set of epitope variant coding sequences, wherein the selecting step includes predicting the immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes validating the disease-specific immunogenic variant peptides by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0045] In some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissues in an individual, comprising: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having sequence differences compared to a reference sample based on the genomic sequence of the disease tissues in the individual; and b) selecting immunogenic variant coding sequences from the first set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the genomic sequence is obtained by whole genome sequencing. In some embodiments, the genomic sequence is obtained by whole exome sequencing. In some embodiments, the genomic sequence is obtained by targeted genome or exome sequencing. For example, the genomic sequence in the disease tissue and / or reference sample can be first enriched with a set of probes (e.g., probes specific to disease-associated genes) and then processed for variant identification. In some embodiments, a first set of variant coding sequences can be first filtered to obtain a smaller set of epitope variant coding sequences, and then the smaller set of variant coding sequences is subjected to selection based on predicted immunogenicity.For example, in some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissue in an individual, comprising: a) obtaining a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having sequence differences compared to a reference sample based on the genomic sequence of the disease tissue in the individual; b) selecting a second set of epitope variant coding sequences from the first set based on the predicted ability of peptides encoded by the first set of variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); and c) selecting immunogenic variant coding sequences from the second set of epitope variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further comprises validating the disease-specific immunogenic variant peptides by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0046] In some embodiments, a method for identifying disease-specific immunogenic variant peptides from disease tissues in an individual is provided, comprising: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having sequence differences compared to a reference sample based on the transcriptome sequences of disease tissues in the individual; and b) selecting immunogenic variant coding sequences from the first set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the transcriptome sequences are obtained by whole transcriptome RNA-Seq sequencing. In some embodiments, the transcriptome sequences are obtained by targeted transcriptome sequencing. For example, RNA or cDNA sequences in disease tissues and / or reference samples can be first enriched with a set of probes (e.g., probes specific to disease-related genes), and then processed for variant identification. In some embodiments, the first set of variant coding sequences can be filtered to obtain a smaller set of epitope variant coding sequences, and then the smaller set of variant coding sequences is subjected to immunogenicity prediction. For example, in some embodiments, a method for identifying disease-specific immunogenic variant peptides from disease tissues in an individual is provided, comprising: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having sequence differences compared to a reference sample based on the transcriptome sequence of the disease tissues in the individual; b) selecting a second set of epitope variant coding sequences from the first set based on the predicted ability of peptides encoded by the first set of epitope variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); and c) selecting immunogenic variant coding sequences from the second set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides.

[0047] In some embodiments, a method is provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual, comprising: a) providing a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the disease tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the disease tissue in the individual; and c) selecting immunogenic variant coding sequences from the second set of expressed variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of a peptide comprising a variant amino acid encoded by the expressed variant coding sequence, thereby identifying the disease specific immunogenic variant peptide. In some embodiments, the method comprises: a) obtaining a first set of variant coding sequences of diseased tissue in the individual based on a genomic sequence of the diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the diseased tissue in the individual; and c) selecting immunogenic variant coding sequences from the second set of expressed variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the expressed variant coding sequences, thereby identifying disease-specific immunogenic variant peptides.

[0048] In some embodiments, the second set of expressed variant coding sequences can be filtered to obtain a smaller set of epitope variant coding sequences, and the smaller set of variant coding sequences is then subjected to prediction of immunogenicity. Thus, for example, in some embodiments, methods are provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual comprising: a) providing a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the disease tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the disease tissue in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHCI); and d) selecting immunogenic variant coding sequences from the third set of epitope variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying the disease specific immunogenic variant peptides. In some embodiments, the method comprises: a) obtaining a first set of variant coding sequences of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the diseased tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the diseased tissue in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHCI); and d) selecting immunogenic variant coding sequences from the third set of epitope variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides.In some embodiments, the method further comprises validating the disease-specific immunogenic variant peptide by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0049] In some embodiments, the disease-specific immunogenic variant peptides identified by the methods described herein are further verified by correlating variant coding sequence information with information of peptides physically bound to MHC molecules.The method may further include, for example, obtaining a plurality of peptides bound to MHC molecules from diseased tissue, subjecting the MHC-binding peptides to mass spectrometry-based sequencing, and correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with peptides predicted to be immunogenic variant coding sequences.Mass spectrometry and correlation methods are further described in the following sections.

[0050] In another aspect, a method is provided that combines sequence-specific variant identification with mass spectrometry. For example, in some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissue in an individual, comprising: a) obtaining a plurality of peptides bound to MHC molecules from disease tissue in an individual; b) subjecting the MHC-bound peptides to mass spectrometry-based sequencing; and c) correlating the mass spectrometry-derived sequence information of the MHC-bound peptides with a set of variant coding sequences of disease tissue in the individual, each variant coding sequence having sequence differences compared to a reference sample, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the plurality of peptides bound to MHC are obtained by isolating MHC / peptide complexes from disease tissue (e.g., by immunoprecipitation) and eluting peptides from MHC. In some embodiments, the peptides are subjected to tandem mass spectrometry. In some embodiments, the mass spectrometry-based sequencing comprises subjecting the peptides to mass spectrometry and comparing the mass spectrometry spectrum to a reference spectrum (e.g., a hypothetical mass spectrometry spectrum of a putative protein encoded by a sequence in the reference sample). In some embodiments, the mass spectrometry sequence information is filtered by peptide length and / or the presence of anchor motifs before the correlation step. In some embodiments, the method further comprises validating the disease-specific immunogenic variant peptides by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0051] In some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissues in an individual, comprising: a) obtaining a first set of variant coding sequences from disease tissues in an individual, each variant coding sequence having sequence differences compared to a reference sample; b) obtaining a plurality of peptides bound to MHC molecules from disease tissues in an individual; c) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and d) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the first set of variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the first set of variant coding sequences can be filtered to obtain a smaller set of variant coding sequences (hereinafter referred to as "epitope variant coding sequences") that code for peptides predicted to bind to MHC molecules, and then subjecting the smaller set of variant coding sequences to correlation analysis. In such embodiments, the method may comprise: a) providing a first set of variant coding sequences of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of epitope variant coding sequences from the first set based on the predicted ability of peptides encoded by the first set of variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHC I); c) obtaining a plurality of peptides bound to MHC molecules from the diseased tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides.In some embodiments, the method includes: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of epitope variant coding sequences from the first set based on the predicted ability of peptides encoded by the first set of variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); c) obtaining a plurality of peptides bound to MHC molecules from disease tissues of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes validating the disease-specific immunogenic variant peptides by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0052] In some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissue in an individual, comprising: a) obtaining a first set of variant coding sequences of disease tissue in an individual, each variant coding sequence having sequence differences compared to a reference sample based on the genomic sequence of the disease tissue in the individual; b) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; c) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and d) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the first set of variant coding sequences, thereby identifying the disease-specific immunogenic variant peptides. In some embodiments, the genomic sequence is obtained by whole genome sequencing. In some embodiments, the genomic sequence is obtained by whole exome sequencing. In some embodiments, the genomic sequence is obtained by targeted genome or exome sequencing. For example, the genomic sequence in the disease tissue and / or reference sample can be first enriched with a set of probes (e.g., probes specific to disease-associated genes) and then processed for variant identification.

[0053] In some embodiments, the first set of variant coding sequences can be filtered to obtain a smaller set of variant coding sequences that encode peptides predicted to bind to MHC molecules (hereinafter referred to as "epitope variant coding sequences"), and the smaller set of variant coding sequences is then subjected to prediction of immunogenicity. For example, in some embodiments, a method is provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual comprising: a) obtaining a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the disease tissue in the individual; b) selecting a second set of epitope variant coding sequences from the first set based on a predicted ability of peptides encoded by the first set of variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHC I); c) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of epitope variant coding sequences, thereby identifying the disease specific immunogenic variant peptides.

[0054] In some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissues in an individual, comprising: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having sequence differences compared to a reference sample based on the transcriptome sequence of the disease tissues in the individual; b) obtaining a plurality of peptides bound to MHC molecules from the disease tissues in the individual; c) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and d) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the first set of variant coding sequences, thereby identifying the disease-specific immunogenic variant peptides. In some embodiments, the transcriptome sequence is obtained by whole transcriptome RNA-Seq sequencing. In some embodiments, the transcriptome sequence is obtained by targeted transcriptome sequencing. For example, RNA sequences or cDNA sequences in disease tissues and / or reference samples can be first enriched with a set of probes (e.g., probes specific to disease-related genes) and then processed for variant identification. In some embodiments, the first set of variant coding sequences can be filtered to obtain a smaller set of epitope variant coding sequences, and the smaller set of variant coding sequences is then subjected to prediction of immunogenicity.For example, in some embodiments, a method is provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual comprising: a) obtaining a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a transcriptome sequence of the disease tissue in the individual; b) selecting a second set of epitope variant coding sequences from the first set based on a predicted ability of peptides encoded by the first set of variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHC I); c) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of epitope variant coding sequences, thereby identifying the disease specific immunogenic variant peptides.

[0055] In some embodiments, a method is provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual, comprising: a) providing a first set of variant coding sequences of disease tissue in the individual based on a genomic sequence of the disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the disease tissue in the individual; c) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of expressed variant coding sequences, thereby identifying the disease specific immunogenic variant peptides. In some embodiments, the method comprises: a) obtaining a first set of variant coding sequences of diseased tissue in the individual based on a genomic sequence of the diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the diseased tissue in the individual; c) obtaining a plurality of peptides bound to MHC molecules from the diseased tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of expressed variant coding sequences, thereby identifying disease-specific immunogenic variant peptides.

[0056] In some embodiments, the second set of expression variant coding sequences can be filtered to obtain a smaller set of epitope variant coding sequences, and the smaller set of variant coding sequences is then subjected to prediction of immunogenicity. Thus, for example, in some embodiments, methods are provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual comprising: a) providing a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the disease tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the disease tissue in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHCI); d) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; e) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and f) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the third set of epitope variant coding sequences, thereby identifying the disease specific immunogenic variant peptides.In some embodiments, the method includes: a) obtaining a first set of variant coding sequences of diseased tissues in an individual based on the genomic sequence of the diseased tissues in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on the transcriptomic sequence of the diseased tissues in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on the predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); d) obtaining a plurality of peptides bound to MHC molecules from the diseased tissues of the individual; e) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and f) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the third set of epitope variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes validating the disease-specific immunogenic variant peptides by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (eg, a human individual with cancer).

[0057] In some embodiments, the disease-specific immunogenic variant peptides identified by the mass spectrometry-based methods described herein are further selected by predicting the immunogenicity of the peptide. In some embodiments, the selecting step comprises predicting the immunogenicity of the peptide based on one or more of the following parameters (e.g., any number of 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11): i) the binding affinity of the peptide to the MHCI molecule, ii) the protein level of the peptide precursor containing the peptide, iii) the expression level of the transcript encoding the peptide precursor, iv) the processing efficiency of the peptide precursor by the immunoproteasome, v) the timing of the expression of the transcript encoding the peptide precursor, vi) the binding affinity of the peptide to the TCR molecule, vii) the position of the mutated amino acid in the peptide, viii) the solvent exposure of the peptide when bound to the MHCI molecule, ix) the solvent exposure of the mutated amino acid when bound to the MHCI molecule, x) the content of aromatic residues in the peptide, and xi) the nature of the peptide precursor.

[0058] Thus, for example, in some embodiments, there is provided a method of identifying disease specific immunogenic variant peptides from disease tissue in an individual comprising: a) obtaining a plurality of peptides bound to MHC molecules from disease tissue of the individual; b) subjecting the MHC binding peptides to mass spectrometry based sequencing; and c) correlating the mass spectrometry derived sequence information of the MHC binding peptides with a set of variant coding sequences of the disease tissue in the individual, each variant coding sequence having a mutation in sequence compared to a reference sample to obtain a second set of variant coding sequences; and d) selecting immunogenic variant coding sequences from the second set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the second set of variant coding sequences, thereby identifying the disease specific immunogenic variant peptides. In some embodiments, a method is provided for identifying disease specific immunogenic variant peptides from disease tissue in an individual, comprising: a) obtaining a set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; c) subjecting the MHC binding peptides to mass spectrometry-based sequencing; d) correlating the mass spectrometry-derived sequence information of the MHC binding peptides with the first set of variant coding sequences to obtain a second set of variant coding sequences; and e) selecting immunogenic variant coding sequences from the second set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the second set of variant coding sequences, thereby identifying the disease specific immunogenic variant peptides.In some embodiments, the method comprises: a) providing a first set of variant coding sequences of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of epitope variant coding sequences from the first set based on a predicted ability of peptides encoded by the first set of variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or an MHC class I molecule); c) obtaining a plurality of peptides bound to MHC molecules from the diseased tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of epitope variant coding sequences to obtain a third set of variant coding sequences; and f) selecting immunogenic variant coding sequences from the third set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the third set of variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method includes: a) obtaining a first set of variant coding sequences of disease tissues in an individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of epitope variant coding sequences from the first set based on the predicted ability of peptides encoded by the first set of variant coding sequences to bind to MHC molecules (e.g., MHC class I molecules, or MHCI); c) obtaining a plurality of peptides bound to MHC molecules from disease tissues of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the set of epitope variant coding sequences to obtain a third set of variant coding sequences; and f) selecting immunogenic variant coding sequences from the third set of variant coding sequences, wherein the selecting step includes predicting the immunogenicity of peptides comprising variant amino acids encoded by the third set of variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes validating the disease-specific immunogenic variant peptides by functional analysis.In some embodiments, the method further comprises validating the disease-specific immunogenic variant peptide by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0059] In some embodiments, a method is provided for identifying disease-specific immunogenic variant peptides from disease tissue in an individual, comprising: a) providing a first set of variant coding sequences of disease tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample, based on a genomic sequence of the disease tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set, based on a transcriptomic sequence of the disease tissue in the individual; c) obtaining a plurality of peptides bound to MHC molecules from the disease tissue of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of expressed variant coding sequences to obtain a third set of variant coding sequences; and f) selecting immunogenic variant coding sequences from the third set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the third set of variant coding sequences, thereby identifying the disease-specific immunogenic variant peptides. In some embodiments, the method includes: a) obtaining a first set of variant coding sequences of diseased tissues in an individual based on the genomic sequence of the diseased tissues in the individual, each variant coding sequence having a sequence difference compared to a reference sample; b) selecting a second set of expressed variant coding sequences from the first set based on the transcriptomic sequence of the diseased tissues in the individual; c) obtaining a plurality of peptides bound to MHC molecules from the diseased tissues of the individual; d) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and e) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with the second set of expressed variant coding sequences to obtain a third set of variant coding sequences; and f) selecting immunogenic variant coding sequences from the third set of variant coding sequences, wherein the selecting step includes predicting the immunogenicity of peptides comprising variant amino acids encoded by the third set of variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes validating the disease-specific immunogenic variant peptides by functional analysis.In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0060] In some embodiments, the method includes: a) providing a first set of variant coding sequences of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the diseased tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the diseased tissue in the individual; c) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHC I molecule); d) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHC I molecule) from the diseased tissue of the individual. e) subjecting the MHC binding peptides to mass spectrometry based sequencing, and f) correlating the mass spectrometry derived sequence information of the MHC binding peptides with a third set of epitope variant coding sequences to obtain a fourth set of variant coding sequences, and g) selecting immunogenic variant coding sequences from the fourth set of variant coding sequences, wherein the selecting step comprises predicting the immunogenicity of peptides comprising variant amino acids encoded by the fourth set of variant coding sequences, thereby identifying the disease specific immunogenic variant peptides.In some embodiments, the method includes: a) obtaining a first set of variant coding sequences of diseased tissue in the individual, each variant coding sequence having a sequence difference compared to a reference sample based on a genomic sequence of the diseased tissue in the individual; b) selecting a second set of expressed variant coding sequences from the first set based on a transcriptomic sequence of the diseased tissue in the individual; and c) selecting a third set of epitope variant coding sequences from the second set based on a predicted ability of peptides encoded by the second set of expressed variant coding sequences to bind to an MHC molecule (e.g., an MHC class I molecule, or MHC I). and d) obtaining a plurality of peptides bound to MHC molecules from diseased tissue of the individual; e) subjecting the MHC-binding peptides to mass spectrometry-based sequencing; and f) correlating the mass spectrometry-derived sequence information of the MHC-binding peptides with a third set of epitope variant coding sequences to obtain a fourth set of variant coding sequences; and g) selecting immunogenic variant coding sequences from the fourth set of variant coding sequences, wherein the selecting step includes predicting the immunogenicity of peptides comprising variant amino acids encoded by the fourth set of variant coding sequences, thereby identifying disease-specific immunogenic variant peptides. In some embodiments, the method further includes validating the disease-specific immunogenic variant peptides by functional analysis. In some embodiments, the disease is cancer. In some embodiments, the individual is a human individual (e.g., a human individual with cancer).

[0061] Also provided herein is a disease-specific immunogenic variant peptide obtained by any one of the methods described herein.The disease-specific immunogenic variant peptide can be used, for example, to produce a composition (e.g., a vaccine composition) for treating a disease.Alternatively, the disease-specific immunogenic variant peptide can be used to produce a variant peptide-specific therapeutic agent, for example, a therapeutic antibody.

[0062] The methods described herein are particularly useful in the context of personalized medicine, where a disease-specific immunogenic variant peptide obtained by any one of the methods described herein is used to develop a therapeutic agent (e.g., a vaccine or a therapeutic antibody) for the same individual. Thus, for example, in some embodiments, a method is provided for treating a disease (e.g., cancer) in an individual, comprising a) identifying a disease-specific immunogenic variant peptide in the individual, and b) synthesizing the peptide, and c) administering the peptide to the individual. In some embodiments, a method is provided for treating a disease (e.g., cancer) in an individual, comprising a) obtaining a disease tissue sample from the individual, b) identifying a disease-specific immunogenic variant peptide in the individual, and c) synthesizing the peptide, and d) administering the peptide to the individual. In some embodiments, a method is provided for treating a disease (e.g., cancer) in an individual, comprising a) identifying a disease-specific immunogenic variant peptide in the individual, b) producing an antibody (or TCR analog, e.g., a soluble TCR) that specifically recognizes the variant peptide, and c) administering the peptide to the individual. In some embodiments, a method for treating a disease (e.g., cancer) in an individual is provided, comprising: a) obtaining a disease tissue sample from the individual; b) identifying a disease-specific immunogenic variant peptide in the individual; c) producing an antibody (or TCR analog, e.g., a soluble TCR) that specifically recognizes the variant peptide; and d) administering the peptide to the individual. In some embodiments, the identifying step combines a sequence-specific variant identification method with an immunogenicity prediction method. In some embodiments, the identifying step combines a sequence-specific variant identification method with mass spectrometry. Any method for identifying a disease-specific immunogenic variant peptide described herein can be used for the treatment method described herein.

[0063] Obtaining mutant coding sequences The methods described herein in various embodiments include providing and / or obtaining a mutant coding sequence, which can generally be obtained, for example, by sequencing the genome or RNA sequence in a diseased tissue sample from an individual and comparing the sequence to that obtained from a reference sample.

[0064] In some embodiments, the diseased tissue is blood. In some embodiments, the diseased tissue is solid tissue (e.g., a solid tumor). In some embodiments, the diseased tissue is a population of cells (e.g., circulating cancer cells in the blood). In some embodiments, the diseased tissue is a population of lymphocytes. In some embodiments, the diseased tissue is a population of leukocytes. In some embodiments, the diseased tissue is a population of epithelial cells. In some embodiments, the diseased tissue is connective tissue. In some embodiments, the diseased tissue is a population of germ cells and / or pluripotent cells. In some embodiments, the diseased tissue is a population of blast cells.

[0065] Suitable disease tissue samples include, but are not limited to, tumor tissue, tumor adjacent normal tissue, tumor distant normal tissue, or peripheral blood lymphocytes. In some embodiments, the disease tissue sample is tumor tissue. In some embodiments, the disease tissue sample is a biopsy containing cancer cells, such as cancer cells (e.g., pancreatic cancer cells) obtained by fine needle aspiration or laparoscopy of cancer cells (e.g., pancreatic cancer cells). In some embodiments, the biopsy cells are centrifuged into a pellet, fixed, and embedded in paraffin before analysis. In some embodiments, the biopsy cells are flash frozen before analysis.

[0066] In some embodiments, the diseased tissue sample comprises circulating metastatic cancer cells. In some embodiments, the diseased tissue sample can be obtained by separating circulating tumor cells (CTCs) from blood. In further embodiments, the CTCs are separated from primary tumors and circulate in bodily fluids. In further embodiments, the CTCs are separated from primary tumors and circulate in bloodstream. In further embodiments, the CTCs are indicators of metastasis. In some embodiments, the CTCs are pancreatic cancer cells. In some embodiments, the CTCs are colorectal cancer cells. In some embodiments, the CTCs are non-small cell lung cancer cells.

[0067] The difference can be identified based on the genomic sequence of diseased tissue in an individual. For example, genomic DNA can be obtained from diseased tissue in an individual and subjected to sequence analysis. The sequence thus obtained can then be compared with that obtained from a reference sample. In some embodiments, the diseased sample is subjected to whole genome sequencing. In some embodiments, the diseased sample is subjected to whole exome sequencing, i.e., only the exons in the genomic sequence are sequenced. In some embodiments, the genomic sequence is "enriched" for specific sequences before being compared with the reference sample. For example, specific probes can be designed to enrich for specific desired sequences (e.g., disease-specific sequences) before being subjected to sequence analysis. Methods for whole genome sequencing, whole exome sequencing, and targeted sequencing are known in the art and are reported, for example, in Bentley, DR et al., Accurate whole human genome sequencing using reversible terminator chemistry, Nature, 2008, v.456, 53-59; Choi, M. et al., Genetic diagnosis by whole exome capture and massively parallel DNA sequencing, Proceedings of the National Academy of Sciences, 2009, v.106(45), 19096-19101; and Ng, SB et al., Targeted capture and massively parallel sequencing of 12 human exomes, Nature, 2009, v.461, 272-276, which are incorporated herein by reference.

[0068] In some embodiments, the difference is identified based on the transcriptome sequence of diseased tissue in an individual. For example, a whole or partial transcriptome sequence (e.g., by a method such as RNA-Seq) can be obtained from diseased tissue in an individual and subjected to sequence analysis. The sequence thus obtained can then be compared to that obtained from a reference sample. In some embodiments, the diseased sample is subjected to whole transcriptome RNA-Seq sequencing. In some embodiments, the transcriptome sequence is "enriched" for specific sequences before comparison with the reference sample. For example, specific probes can be designed to enrich for specific desired sequences (e.g., disease-specific sequences) before being subjected to sequence analysis. Methods for whole-transcriptome sequencing and targeted sequencing are known in the art and have been reported, for example, in Tang, F. et al., mRNA-Seq whole-transcriptome analysis of a single cell, Nature Methods, 2009, v.6, 377-382; Ozsolak, F., RNA sequencing: advances, challenges and opportunities, Nature Reviews, 2011, v.12, 87-98; German, M. A. et al., Global identification of microRNA-target RNA pairs by parallel analysis of RNA ends, Nature Biotechnology, 2008, v.26, 941-946; and Wang, Z. et al., RNA-Seq: a revolutionary tool for transcriptomics, Nature Reviews, 2009, v.10, p.57-63. In some embodiments, the transcriptome sequencing method includes, but is not limited to, RNA poly(A) library, microarray analysis, parallel sequencing, massively parallel sequencing, PCR, and RNA-Seq.RNA-Seq is a high-throughput method for sequencing a portion or substantially all of the transcriptome.Briefly, the isolated transcriptome sequence population is converted into a library of cDNA fragments with adaptors attached to one or both ends. With or without amplification, each cDNA molecule is then analyzed to obtain short stretches of sequence information, typically 30-400 base pairs. These pieces of sequence information are then aligned to a reference genome, reference transcript, or combined de novo to reveal the structure (i.e., transcription boundaries) and / or levels of expression of the transcript.

[0069] Once the sequence in diseased tissue is obtained, it can be compared with the corresponding sequence in reference sample. Sequence comparison can be performed at the nucleic acid level by aligning the nucleic acid sequence in diseased tissue with the corresponding sequence in reference sample. Sequence differences that lead to one or more changes in encoded amino acids are then identified. Alternatively, sequence comparison can be performed at the amino acid level, i.e., the nucleic acid sequence is first converted in silico into amino acid sequence, and then comparison is performed.

[0070] In some embodiments, the comparison of sequences from disease tissues with those from standards can be completed by methods known in the art, such as manual alignment, FAST-All (FASTA), and Basic Local Alignment Search Tool (BLAST). The sequence comparison completed by BLAST requires input of disease sequence and input of standard sequence. BLAST first compares disease sequence with standard database by identifying short sequence matches between two sequences, a process called seeding. Once sequence matches are found, an extension of sequence alignment is performed using a scoring matrix.

[0071] In some embodiments, the reference sample is a matched disease-free tissue sample. As used herein, a "matched" disease-free tissue sample is one selected from the same or similar tissue type as the diseased tissue. In some embodiments, the matched disease-free and diseased tissues may be from the same individual. In some embodiments, the reference sample described herein is a disease-free sample from the same individual. In some embodiments, the reference sample is a disease-free sample from a different individual (e.g., a disease-free individual). In some embodiments, the reference sample is obtained from a population of different individuals. In some embodiments, the reference sample is a database of known genes associated with an organism. In some embodiments, the reference sample may be a combination of known genes associated with an organism and genomic information from a matched disease-free tissue sample. In some embodiments, the variant coding sequence may code for or include a point mutation in the amino acid sequence. In some embodiments, the variant coding sequence may code for or include an amino acid deletion or insertion.

[0072] In some embodiments, a set of variant coding sequences is first identified based on genome sequences.Then, this initial set is further filtered to obtain a narrower set of expressed variant coding sequences based on the presence of variant coding sequences in a transcriptome sequencing database (and thus considered to be "expressed").In some embodiments, the set of variant coding sequences is reduced by at least about 10, 20, 30, 40, 50 or more times by filtering through a transcriptome sequencing database.

[0073] In some embodiments, the mutant coding sequence is a sequence resulting from a non-synonymous mutation (e.g., a point mutation) resulting in a different amino acid in the protein. In some embodiments, the mutant coding sequence is a sequence resulting from a read-through mutation in which a stop codon is modified or removed resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus. In some embodiments, the mutant coding sequence is a sequence resulting from a splice site mutation resulting in the inclusion of an intron in the mature mRNA, thus resulting in a unique tumor-specific protein sequence. In some embodiments, the mutant coding sequence is a sequence resulting from a chromosomal rearrangement (i.e., gene fusion) resulting in a chimeric protein with a tumor-specific sequence at the junction of two proteins. In some embodiments, the mutant coding sequence is a sequence resulting from a frameshift mutation or deletion resulting in a new open reading frame with a new tumor-specific protein sequence. In some embodiments, the mutant coding sequence is a sequence resulting from more than one mutation. In some embodiments, the mutant coding sequence is a sequence resulting from more than one mutation mechanism.

[0074] Obtaining epitope variant coding sequences In some embodiments, the variant coding sequences described herein are filtered to obtain a smaller set of variant coding sequences that encode peptides predicted to bind to MHC molecules ("epitope variant coding sequences"). In some embodiments, the set of variant coding sequences is reduced by at least about 10, 20, 30, 40, 50, 60, 80, 100, 150, 200, 250, 300, or more fold by filtering through the MHC binding prediction process.

[0075] The MHC (e.g., MHCI) binding ability of peptides encoded by variant coding sequences can be evaluated by prediction algorithms, such as NETMHC. Briefly, NETMHC is an algorithm trained on quantitative peptide data using both affinity data from the Immune Epitope Database and Analysis Resource (IEDB) and elution data from SYFPEITHI. It can predict MHCI binding for peptides between 8 and 14 amino acids in length. NETMHC uses predictors trained on 9 amino acid sequences from 55 MHC alleles (43 human and 12 non-human). To allow prediction of shorter input sequences, NETMHC virtually extends 8 amino acid sequences. For input sequences longer than 9 amino acids in length, NETMHC generates every 9 amino acid sequence contained within the input sequence. NETMHC then predicts MHC binding using a trained artificial neural network and a position-specific scoring matrix.

[0076] Selecting immunogenic variant coding sequences The methods provided herein in some embodiments further comprise selecting an immunogenic variant coding sequence, comprising predicting the immunogenicity of a peptide comprising a variant amino acid encoded by the variant coding sequence. Predicting immunogenicity can be performed, for example, by a process (e.g., an in silico process) that considers one or more parameters of the peptide and the corresponding peptide precursor and predicts the likelihood that the peptide is immunogenic. These parameters include, but are not limited to, i) the binding affinity of the peptide to the MHCI molecule, ii) the protein level of the peptide precursor containing the peptide, iii) the expression level of the transcript encoding the peptide precursor, iv) the processing efficiency of the peptide precursor by the immunoproteasome, v) the timing of expression of the transcript encoding the peptide precursor, vi) the binding affinity of the peptide to the TCR molecule, vii) the position of the variant amino acid within the peptide, viii) the solvent exposure of the peptide when bound to the MHCI molecule, ix) the solvent exposure of the variant amino acid when bound to the MHCI molecule, x) the content of aromatic residues in the peptide, and xi) the nature of the peptide precursor. In some embodiments, immunogenicity is based on at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 parameters described herein.

[0077] In some embodiments, the binding affinity of peptides to MHC molecules is used to predict immunogenicity. The binding affinity of peptides to MHC molecules can predict the stability of pMHC, which in turn can allow for sustained presentation of pMHC, thus increasing cell surface exposure for potential interaction with immune cells. Binding affinity can be predicted using techniques known in the art, such as RankPep, MHCBench, nHLAPred, SVMHC, NETMHCpan, and POPI, which are based on methods such as artificial neural networks, average relative binding matrix, quantitative matrix, and stabilized matrix methods. In some embodiments, the binding affinity of peptides to MHC is based on the presence of specific amino acid residues (amino acids involved in MHC binding) located at known anchor positions. In some embodiments, each residue of the peptide is evaluated for its contribution to binding. In some embodiments, the analysis system is trained with peptides known to bind to MHC. In some embodiments, the binding energy of peptide-MHC molecules is calculated. In some embodiments, a binding threshold, such as an IC50 value <500 nM, is used to evaluate the predicted affinity of peptides that bind to MHC.

[0078] In some embodiments, expression levels of peptide precursors in diseased tissues are used to predict immunogenicity. In some embodiments, protein levels are measured biochemically (e.g., Western blot and ELISA). In some embodiments, protein levels can be measured by known quantitative mass spectrometry. As demonstrated by La Gruta et al., high expression levels of peptide precursors in diseased tissues can be correlated with predicted immunogenicity. As described, the availability of higher amounts of peptide precursors fed into the epitope processing pathway positively correlated with increased epitope presentation and immunogenic responses (La Gruta, NL et al., A virus-specific CD8+ T cell immunodominance hierarchy determined by antigen dose and precursor frequencies, Proceedings of the National Academy of Sciences, 2006, v.103, 994-999, incorporated herein by reference).

[0079] In some embodiments, the expression level of transcripts encoding peptide precursors in disease tissues can be used to predict immunogenicity. In some embodiments, RNA expression levels are measured by RT-PCR. In some embodiments, RNA expression levels are measured by sequence analysis. As discussed above, due to epitope processing pathways, increasing the availability of peptide precursors positively correlates with the immunogenic response associated with the epitope obtained from the peptide precursor. Furthermore, a positive correlation between mRNA levels and protein abundance has been observed (Ghaemmaghami, S. et al., Global analysis of protein expression in yeast, Nature, 2003, v.425, 737-741, incorporated herein by reference). Thus, the expression level of transcripts encoding peptide precursors in disease tissues can be used to predict immunogenicity.

[0080] In some embodiments, the efficiency of processing peptide precursors by immunoproteasomes is used to predict immunogenicity. As used herein, "processing efficiency" of peptide precursors refers to the efficiency with which the starting amino acid sequence (i.e., a larger peptide or protein) undergoes expression, translation, transcription, digestion, transport, and any further processing before binding to MHCI molecules. "Immunoproteasomes" are a collection of proteases that enzymatically digest peptides and / or protein precursors into small amino acid sequences with the ultimate goal of epitope formation. For example, as demonstrated by Chen et al., the ability of immunoproteasomes to generate epitope precursors directly correlates with immunogenicity (Chen, W. et al., Immunoproteasomes shape immunodominance hierarchies of antiviral CD8+ T cell repertoire and presentation of viral antigens, The Journal of Experimental Medicine, 2001, v.193, 1319-1326, incorporated herein by reference). Knowledge of the steps involved in epitope processing can be used to predict the immunogenicity of the resulting epitopes. In particular, studies on the immunoproteasome suggest that it is efficient in producing MHC epitopes, and understanding of the immunoproteasome machinery will facilitate the prediction of immunogenic peptides.

[0081] In some embodiments, the timing of expression of peptide precursors can be used to predict immunogenicity. Proteins expressed relatively early in disease progression are more likely to be presented as MHC epitopes (Moutaftsi, M. et al., A consensus epitope prediction approach identifies the breadth of murine T CD8+-cell responses to vaccina virus, Nature Biotechnology, 2006, v.24, 817-819, incorporated herein by reference). In some embodiments, comparative analysis of disease tissues can be used to determine the temporal expression patterns of gene products. In some embodiments, estimation of temporal expression patterns can allow identification of expressed gene products that are abundantly presented at early time points compared to other identified expressed gene products.

[0082] In some embodiments, the binding affinity of peptide epitopes with T cell receptor (TCR) can be used to predict immunogenicity.The method of predicting the binding affinity of peptide epitopes with TCR is known in the art, and is reported in, for example, Tung, C.-W. et al., POPISK: T-cell reactivity prediction using support vector machines and string kernels, BMC Bioinformatics, 2011, v.12, 446, which is incorporated herein by reference.

[0083] In some embodiments, the position of the mutated amino acid in the peptide is used to predict immunogenicity. The peptide epitope binds to the MHCI molecule at two different anchor positions. The spacing between the anchor positions is separated by about 6-7 amino acids, not including the amino acid occupying the anchor position, as measured by the epitope peptide sequence. It has been reported that the amino acid spacing between the two MHCI anchor positions, i.e., mutations at amino acids 4-6, are more likely to be positively correlated with immunogenic responses (Calis, JJA et al., Properties of MHC class I presented peptides that enhance immunogenicity, PLOS Computational Biology, 2013, v.9, 1-13, incorporated herein by reference). The sequence position of the amino acid is determined by starting the sequence position count at number 1 with the terminal amino acid.

[0084] In some embodiments, the structural characteristics of peptides presented on MHC-presented epitopes can predict immunogenicity. Structural evaluation of MHC-binding peptides can be performed by in silico 3D analysis and / or protein docking programs. Methods for predicting the structure of pMHC molecules are known in the art and are reported, for example, in Marti-Renom, MA et al., Comparative protein structure modeling of genes and genomes, Annual Review of Biophysics and Biomolecular Structure, 2000, v.29, 291-325; Chivian, D. et al., Homology modeling using parametric alignment ensemble generation with consensus and energy-based model selection, Nucleic Acids Research, 2006, v.34, e112; and McRobb, FM et al., Homology modeling and docketing evaluation of aminergic G protein-coupled receptors, Journal of Chemical Information and Modeling, 2010, v.50, 626-637, which are incorporated herein by reference. The use of predicted epitope structures when bound to MHC molecules, e.g., obtained from the Rosetta algorithm, can be used to assess the solvent accessibility of amino acid residues of said epitopes when bound to MHC molecules. This information can be substantially correlated with the immunogenicity of the peptide.For example, as described by Park et al., mutant peptides in which mutated amino acid residues exhibit additional solvent exposure compared to wild-type sequences are positively correlated with increased immunogenicity (Park, M.-S. et al., Accurate structure prediction of peptide-MHC complexes for identifying highly immunogenic antigens, Molecular Immunology, 2013, v.56, 81-90, incorporated herein by reference). In some embodiments, the solvent exposure of the entire peptide when presented on an MHC complex is used to predict immunogenicity. In some embodiments, the solvent exposure of mutated amino acids on a peptide is used to predict immunogenicity.

[0085] In some embodiments, epitope content of bulky and / or aromatic residues in a peptide can be used to predict immunogenicity. Calis et al. observed a link between the presence of bulky and / or aromatic amino acid residues in an epitope amino acid sequence and immunogenicity. Specifically, it was reported that phenylalanine and isoleucine content are positively correlated with epitope immunogenicity. As discussed above, both position and structural evaluation of variant amino acids can be used to predict epitope immunogenicity. Further to this discussion, it was hypothesized that bulky and / or aromatic residues in positions 4-6 of an epitope can be predicted to have a high degree of solvent exposure.

[0086] In some embodiments, the properties of the peptide precursor can be used to predict the immunogenicity of the variant-encoded amino acid sequence. For example, when the diseased tissue is a tumor, peptide precursor sequences known to be associated with cancer can be useful in predicting immunogenicity. However, peptide precursor sequences from proteins that are not directly associated with cancer can also be useful within the methods of the present invention, for example, when mutated.

[0087] In some embodiments, at least two (e.g., at least any one of 3, 4, 5, 6, 7, 8, 9, or 10) parameters described herein can be used to predict the immunogenicity of a peptide. In some embodiments, a first predictive evaluation can be used to select a set of variant-encoding amino acid sequences, which are then processed by a second predictive evaluation to provide a cumulative prediction of epitope immunogenicity. It is contemplated as part of this disclosure that multiple evaluations can be used to predict epitope immunogenicity. Alternatively, multiple parameters can be evaluated in parallel, resulting in a composite score based on the evaluation of the various parameters. For example, a score can be calculated for each of the parameters, and a percentage weight can be assigned to each parameter. A composite score can then be calculated based on the scores and percentage weights of each of the evaluated parameters. A combination of sequential evaluation combined with parallel parameter evaluation can also be used.

[0088] For a given difference in the variant coding sequence, multiple overlapping putative variant peptides of various lengths can be generated. In one embodiment, these multiple overlapping putative variant peptides are ranked based on immunogenicity, which may include one or more analyses described herein. In addition, ranking of putative variant peptides can be achieved by any one or more means known in the art that can be included in the immunogenicity determination or considered as a separate analysis, depending on the preference of the person performing the ranking process. Some non-limiting examples of such means include the abundance of the precursor protein, the abundance of the peptide in the processed proteome, including the efficiency of processing, and the abundance of the peptide in the peptide-MHCI complex. Further examples of possible ranking analyses include, but are not limited to, the binding affinity of the peptide, the stability of the peptide-MHCI complex, and the similarity or difference of the peptide with the self-peptide. Alternatively, peptides can be ranked based on correlation with data generated from mass spectrometry that measures these or other properties that are well known to those skilled in the art or described herein. Some representative properties include the presence or absence of large or aromatic amino acids that increase immunogenicity, and the specific positions at which these amino acids are found within the peptide, with the preference for the immunogenic impact given to those amino acids found in intermediate positions of the peptide, e.g., peptides 4-6 (see, e.g., Calis et al. PLOS, 9(10):e1003266 (2013)).

[0089] In some embodiments, the prediction of immunogenicity further includes HLA (human leukocyte antigen) typing analysis. Due to the polygenic nature of MHC, every person expresses at least three different antigen-presenting MHC class I molecules and three (or sometimes four) MHC class II molecules on their cells. In fact, the number of different MHC molecules expressed on the cells of most people is much higher due to the extreme polymorphism of MHC and the codominant expression of MHC gene products. There are more than 200 alleles of some human MHC class I and class II genes, and each allele is present at a relatively high frequency in the population. Therefore, there is only a small probability that the corresponding MHC loci on both homologous chromosomes of an individual have the same allele, and most individuals are heterozygous for MHC loci. The specific combination of MHC alleles found on a single chromosome is known as MHC haplotype. The expression of MHC alleles is codominant, and the protein products of both alleles of a locus are expressed in cells, and both gene products can present antigens to T cells. Thus, extensive polymorphisms on each locus have the potential to double the number of different MHC molecules expressed in an individual, thereby increasing the diversity already present through polygenicity (see, for example, Janeway's Immunobiology, edited by Murphy, Kenneth, Garland Science, New York, NY (2011) for an overview). HLA typing can be achieved using any one of several methods known in the prior art, such as DNA-based histocompatibility assays. Specific examples of methods in the art include those in which the polymerase chain reaction (PCR) product is further analyzed, such as PCR-RFLP (restriction fragment length polymorphism), PCR-SSO (sequence-specific oligonucleotides), PCR-SSP (sequence-specific primers), and PCR-SBT (sequence-based typing) methods. Thus, determining the specific gene polymorphism type involved in the presentation of a peptide of interest can provide further information about the immunogenicity of the variant peptide.

[0090] Obtaining peptides bound to MHC molecules The method provided herein in some embodiments includes obtaining peptides bound to MHC molecules from diseased tissue of an individual. In some embodiments, the MHC-binding peptides are isolated by immunoaffinity methods. In some embodiments, the MHC-binding peptides are isolated by affinity chromatography. In some embodiments, the MHC-binding peptides are isolated by immunoaffinity chromatography. In some embodiments, the MHC-binding peptides are isolated by immunoprecipitation.

[0091] In some embodiments, an anti-MHC antibody is used to capture the MHC / peptide molecules. In some embodiments, multiple anti-MHC antibodies, optionally with different affinities and / or binding characteristics, can be used to capture the MHC / peptide complexes. Suitable antibodies include, but are not limited to, monoclonal antibody W6 / 32 specific for HLA class I, and monoclonal antibody BB7.2 specific for HLA-A2.

[0092] In some embodiments, the MHC / peptide complexes can be isolated first and the MHC-binding peptides subsequently separated from the MHC molecules. In some embodiments, the MHC-binding peptides are separated from the MHC molecules by acid elution. In some embodiments, the acid-mediated separation of the MHC-binding peptides from the MHC molecules can be performed on whole intact cells, optionally in the presence of lysed cells and / or cell debris. In some embodiments, the MHC-binding peptides can be separated from the MHC molecules after exposure of the pMHC to a buffer with an acidic pH. In some embodiments, the MHC-binding peptides can be separated from the MHC molecules by mild acid elution (MAE). In some embodiments, the MHC-binding peptides can be separated from the MHC molecules by mild acid elution (MAE) of the extracellular surface. In some embodiments, the MHC-binding peptides can be separated from the MHC molecules by denaturation of the pMHC molecules.

[0093] In some embodiments, the MHC binding peptides can be further processed before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be enriched before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be fractionated before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be enriched before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be further enzymatically digested before mass spectrometry-based sequencing. In some embodiments, the buffer in which the MHC binding peptides may be contained can be exchanged before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be labeled before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be covalently labeled before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be enzymatically labeled before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides can be chemically labeled before mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides may be labeled to allow for enhanced ionization during mass spectrometry-based sequencing. In some embodiments, the MHC binding peptides may be labeled to allow for quantification during mass spectrometry-based sequencing. In some embodiments, the isolated MHC binding peptides may be derived from multiple sources and / or enrichment procedures, and may optionally be pooled collectively prior to mass spectrometry-based sequencing.

[0094] Mass spectrometry-based peptide sequencing The peptides bound to the MHC molecules in the methods described herein are subjected to mass spectrometry sequencing. As used herein, "mass spectrometry-based sequencing" refers to the use of mass spectrometry to identify the amino acid sequence of peptides and / or proteins. A mass spectrometer is an instrument that can measure the mass-to-charge (m / z) ratio of individual ionized molecules, allowing researchers to identify unknown compounds, quantify known compounds, and elucidate the structure and chemical properties of molecules. The methods provided herein can be used to obtain sequence information of peptide epitopes bound to MHCI molecules. In some embodiments, the entire sequence of the peptide epitope can be determined. In some embodiments, a partial sequence of the peptide epitope can be determined. In some embodiments, the peptides are subjected to tandem mass spectrometry, such as tandem chromatography mass spectrometry (e.g., LC-MS or LC-MS-MS).

[0095] In some embodiments, mass spectrometry begins with isolating a sample and loading it onto the instrument. In some embodiments, the MHC-binding peptides can be chromatographed prior to mass spectrometry. In some embodiments, the chromatography is liquid chromatography. In some embodiments, the chromatography is reversed-phase chromatography. In some embodiments, the MHC-binding peptides can be chromatographed and simultaneously concentrated prior to introduction into the mass spectrometer. In some embodiments, the chromatographic separation can be online, where peptides eluting from the chromatographic source enter directly into the mass spectrometer. In some embodiments, the chromatographic separation can be offline. In some embodiments, offline chromatographic separation can be used to fractionate the isolated MHC-binding peptides. Offline chromatographic separation typically involves separation and / or fractionation of a mass spectrometry sample, where the resulting separated and / or fractionated sample is not directly introduced into the mass spectrometer upon exiting the chromatographic system.

[0096] In some embodiments, MHC-binding peptides can be sequenced using known mass spectrometric ionization methods (e.g., matrix-assisted laser desorption / ionization, electrospray ionization, and / or nanoelectrospray ionization, atmospheric pressure chemical ionization). In some embodiments, MHC-binding peptides can be ionized outside, inside, and / or when entering a mass spectrometer. In some embodiments, positive ions of MHC-binding peptides can be analyzed in a mass spectrometer. The ions are then separated according to their mass-to-charge ratio via exposure to a magnetic field. In some embodiments, a sector-type instrument is used, and the ions are quantified according to the strength of the deflection of the ions' trajectory as they pass through the magnetic field of the instrument, which directly correlates with the mass-to-charge ratio of the ions. In other embodiments, the ion mass-to-charge ratio is measured based on their motion as they pass through a quadrupole, or in a three-dimensional or linear ion trap or orbitrap, or in the magnetic field of a Fourier transform ion cyclotron resonance mass spectrometer. The instrument records the relative abundance of each ion, which is used to determine the chemical, molecular and / or isotopic composition of the original sample. In some embodiments, a time-of-flight instrument is used, which uses an electric field to accelerate ions through the potential and measures the time it takes each ion to reach the detector. This technique relies on the charge of each ion being uniform so that the kinetic energy of each ion is the same. The only variable that affects velocity in this scenario is mass, with lighter ions traveling faster and therefore reaching the detector sooner. The resulting data is represented as a mass spectrum or histogram of intensity vs. mass-to-charge ratio, with peaks representing ionized compounds or fragments.

[0097] Mass spectral data can be obtained by tandem mass spectrometry. In some embodiments, the mass spectrometric acquisition method for obtaining information for peptide sequencing of MHC binding peptides can be data-dependent. In some embodiments, the mass spectrometric acquisition method for obtaining information for peptide sequencing of MHC binding peptides can be data-independent. In some embodiments, the mass spectrometric acquisition method for obtaining information for peptide sequencing of MHC binding peptides can be based on measured accurate mass analysis. In some embodiments, the mass spectrometric acquisition method for obtaining information for peptide sequencing of MHC binding peptides can be peptide mass fingerprinting. Mass spectral data useful in the present invention can be obtained by peptide mass fingerprinting. Peptide mass fingerprinting involves inputting observed masses from a spectrum of a mixture of peptides generated by proteolytic digestion into a database and correlating the observed masses in silico with predicted masses of fragments resulting from digestion of known proteins. A known mass corresponding to a sample mass provides evidence that the known protein is present in the sample tested.

[0098] In some embodiments, tandem mass spectrometry involves a process of colliding peptide ions with gas and fragmenting them (e.g., due to vibrational energy imparted by the collision). The fragmentation process results in cleavage products that break peptide bonds at various sites along the protein. The masses of the observed fragments can be matched with a database of predicted masses for one of many given peptide sequences, and the presence of the protein can be predicted. In some embodiments, the mass spectrometry acquisition method can utilize a fragmentation method (e.g., collision-induced dissociation, pulsed Q dissociation, high-energy collision dissociation, electron transfer dissociation, and electron transfer dissociation, infrared multiphoton dissociation).

[0099] In some embodiments, data acquired from the mass spectrometer can be used to identify peptide sequences. In some embodiments, a search algorithm (e.g., SEQUEST and Mascot) can be used to assign peptide sequences to acquired mass spectra. In some embodiments, the assigned peptide sequences can have a false discovery rate of less than about 5%. In some embodiments, the assigned peptide sequences can have a false discovery rate of less than about 1%. In some embodiments, the assigned peptide sequences can have a false discovery rate of less than about 0.5%. In some embodiments, a database can be used to assign peptide sequences to acquired spectra by the search algorithm. In some embodiments, the database used to assign peptide sequences to acquired spectra by the search algorithm can be a database of known sequences of the organism. In some embodiments, the database used to assign peptide sequences to acquired spectra by the search algorithm can be a database of known proteins of the organism. In some embodiments, the database used to assign peptide sequences to acquired spectra by the search algorithm can be a database of known genomic sequences of the organism. In some embodiments, the database used to assign peptide sequences to acquired spectra by the search algorithm can include sequence information obtained from diseased tissue. In some embodiments, the database used to assign peptide sequences to acquired spectra by the search algorithm can include sequence information obtained from non-diseased tissue.

[0100] In some embodiments, the sequence assigned to the spectrum can be manually verified to confirm the correct fragment ion assignment by the algorithm. In some embodiments, synthetic peptide standards can be used to confirm the algorithm-assigned sequence. In some embodiments, the spectrum generated from the MHC-binding peptide can be compared to the spectrum generated from the peptide standard. For example, the comparison can involve matching the fragment ion pattern, and optionally fragment abundance or intensity, based on the m / z value of the spectrum obtained from the diseased tissue source against the standard. In some embodiments, manual verification can confirm the sequence assignment of the complete peptide sequence. In some embodiments, manual verification can confirm the sequence assignment of a partial segment of the peptide sequence.

[0101] Correlation of mass spectrometry and genomic data In some embodiments, the method provided herein includes correlating mass spectrometry-derived sequence information of MHC-binding peptides with a set of variant coding sequences to identify disease-specific immunogenic variant peptides. For example, mass spectrometry sequences of MHC-binding peptides can be used to further select a population of predicted disease-specific immunogenic variant peptides. In some embodiments, mass spectrometry-based epitope identification can complement genomic-based immunogenic epitope identification and / or prediction. In some embodiments, mass spectrometry-based epitope identification can confirm genomic-based immunogenic epitope identification and / or prediction.

[0102] Data obtained from mass spectrometry may be correlated with immunogenic peptide predictions based on genomic and / or transcriptomic sequence analysis of diseased tissue. In general, the amino acid sequence identified from mass spectrometry-based sequencing is compared to the amino acid sequence of the predicted immunogenic peptide to find regions that include partial sequence alignment. In some embodiments, the peptide identified via mass spectrometry is a sequence that exactly matches the sequence of the predicted immunogenic peptide. In some embodiments, the length of the amino acid sequence may vary between the peptide identified via mass spectrometry and that of the predicted immunogenic peptide. For example, the peptide identified via mass spectrometry may contain additional amino acids amended at the C- and / or N-terminus of the peptide compared to the predicted immunogenic peptide. Alternatively, the peptide identified via mass spectrometry containing the mutant amino acid may have fewer amino acids on the C- and / or N-terminus compared to the predicted immunogenic peptide. In these exemplary embodiments, the mutant amino acid and the sequence surrounding the mutant amino acid must be the same in both the peptide identified via mass spectrometry and the immunogenic and predicted peptide. In some embodiments, the results obtained from correlating mass spectrometry-based sequences with immunogenic peptide predictions match predicted immunogenic sequences with sequences confirmed to be physically presented by MHC molecules via mass spectrometry-based sequencing.

[0103] In some embodiments, the mass spectrometry sequence identifications obtained can be further filtered by peptide length. For example, in some embodiments, the population of MHC binding peptides identified by mass spectrometry can be further filtered to only include identified peptide sequences that are 8 or 9 amino acids in length.

[0104] Functional validation of immunogenic mutant peptides Disease-specific immunogenic variant peptides identified by the methods described herein can be further validated by functional studies. For example, peptides can be synthesized and tested based on their ability to activate targeted immune responses (e.g., those mediated by cytotoxic T cells). In some embodiments, peptides are chemically synthesized. In some embodiments, peptides are synthesized by recombinant methods. In some embodiments, peptides are synthesized by first expressing a peptide precursor molecule, which is then processed (e.g., by the immune proteasome) to generate the peptide of interest. The synthesized peptides can be subjected to further purification before being subjected to functional analysis.

[0105] In some embodiments, synthetic predicted disease-specific immunogenic peptides are used in vitro to test for cytotoxic T cell responses. In some embodiments, synthetic predicted disease-specific immunogenic peptides are used in vivo to test for cytotoxic T cell responses.

[0106] In some embodiments, the immunogenicity of disease-specific peptides can be tested by immunizing mice. In some embodiments, the immunogenicity of disease-specific peptides can be tested by measuring CD8 T cell response after immunization. In some embodiments, the CD8 T cell response can be measured using MHCI / peptide-specific dextramer. In some embodiments, the immunogenicity of disease-specific peptides can be tested by analyzing tumor infiltrating cells (TIL).

[0107] In some embodiments, the presence of specific epitopes and / or cell surface proteins can be measured. In some embodiments, the presence of epitopes derived from vaccination can be measured. In some embodiments, interferon gamma (IFN-γ) can be measured. In some embodiments, programmed cell death 1 (PD-1) can be measured. In some embodiments, T cell immunoglobulin mucin-3 (TIM-3) can be measured. In some embodiments, cytotoxic T cells expressing specific proteins and / or epitopes can be measured. In some embodiments, cytotoxic T cells displaying specific epitopes can be measured.

[0108] In some embodiments, the immunogenicity of disease-specific peptides can be tested by first expressing the peptide in dendritic cells, and then testing the ability of presenting antigens recognized by T cells.In some embodiments, dendritic cells are obtained from patients, and disease-specific peptides are identified in said patients.In some embodiments, the immunogenicity of disease-specific peptides can be tested by first expressing the peptide in B lymphocytes, and then testing the ability of presenting antigens recognized by T cells.In some embodiments, B lymphocytes are obtained from patients, and disease-specific peptides are identified in said patients.See, for example, US Patent No. 8,349,558.

[0109] Composition of immunogenic peptides The present disclosure provides a method for identifying disease-specific immunogenic peptides. Immunogenic peptides can be identified based on their ability to activate targeted immune responses (e.g., those mediated by cytotoxic T cells). In some embodiments, the amino acid sequence of the identified disease-specific immunogenic peptide can be used to develop a pharma- ceutically acceptable composition. In some embodiments, the composition can include a synthetic disease-specific immunogenic peptide. In some embodiments, the composition can include a synthetic disease-specific immunogenic variant peptide. In some embodiments, the composition can include two or more disease-specific immunogenic peptides. In some embodiments, the composition can include two or more disease-specific immunogenic variant peptides. In some embodiments, the two or more disease-specific immunogenic peptides can activate a cytotoxic T cell response against two or more unique epitopes.

[0110] In some embodiments, the composition may include precursors (e.g., proteins, peptides, DNA and RNA) of the disease-specific immunogenic peptide. In some embodiments, the precursors of the disease-specific immunogenic peptide may generate or be generated into the identified disease-specific immunogenic peptide. In some embodiments, the precursors of the disease-specific immunogenic peptide may be prodrugs.

[0111] In some embodiments, the composition comprising the disease-specific immunogenic peptide may be pharma- ceutically acceptable. In some embodiments, the composition comprising the disease-specific immunogenic peptide may further comprise an adjuvant. For example, mutant peptides may be used as vaccines (see Sahin et al., Int. J. Cancer, 78:387-9 (1998); Stumiolo et al., Nature Biotechnol, 17:555-61 (1999); Rammensee et al., Immunol Rev 188:164-76 (2002); and Hannani et al. Cancer J 17:351-358 (2011)). In addition, the vaccine may contain individualized components according to the personal needs of a particular patient. In some embodiments, the vaccine may be specific to a predicted immunogenic peptide in a particular patient. In some embodiments, the vaccine contains more than one immunogenic peptide or peptide precursor. In some embodiments, the length of the peptides used in the vaccine may vary in length. In some embodiments, the peptide is about 7-50 amino acids in length (e.g., about any of 8, 9, 10, 11, 12, 13, 14, 15, 17, 20, 22, 25, 30, 35, 40, 45, or 50 amino acids in length). In some embodiments, the peptide is about 8-12 amino acids in length. In some embodiments, the peptide is about 8-10 amino acids in length. The peptides can be utilized in their isolated form, or alternatively, the peptides can be added to the ends of MHC isolated peptides to generate "longer peptides" that may prove to be more immunogenic (see, e.g., Castle et al., Cancer Res 72:1081-1091 (2012)). In some embodiments, the peptides can also be tagged or can be fusion proteins or hybrid molecules. In some embodiments, the peptides are in the form of a pharma- ceutically acceptable salt.

[0112] In some embodiments, the vaccine is a nucleic acid vaccine. In some embodiments, the nucleic acid encodes an immunogenic peptide or peptide precursor. In some embodiments, the nucleic acid vaccine comprises a sequence adjacent to the sequence encoding the immunogenic peptide or peptide precursor. In some embodiments, the nucleic acid vaccine comprises more than one immunogenic epitope. In some embodiments, the nucleic acid vaccine is a DNA-based vaccine. In some embodiments, the nucleic acid vaccine is an RNA-based vaccine. In some embodiments, the RNA-based vaccine comprises mRNA. In some embodiments, the RNA-based vaccine comprises naked mRNA. In some embodiments, the RNA-based vaccine comprises modified mRNA (e.g., mRNA protected from degradation using protamine, mRNA containing a modified 5'CAP structure, or mRNA containing modified nucleotides). In some embodiments, the RNA-based vaccine comprises a single-stranded mRNA.

[0113] The polynucleotide may be substantially pure or may be contained in a suitable vector or delivery system.Suitable vectors and delivery systems include systems based on viruses, such as adenovirus, vaccinia virus, retrovirus, herpes virus, adeno-associated virus, or hybrids that contain elements of more than one virus.Non-viral delivery systems include cationic lipids and cationic polymers (e.g., cationic liposomes).In some embodiments, physical delivery can be used, such as using a "gene gun".

[0114] In some embodiments, the peptides described herein can be used to produce variant peptide-specific therapeutics, such as antibody therapeutics. For example, variant peptides can be used to generate and / or identify antibodies that specifically recognize variant peptides. These antibodies can be used as therapeutics. Synthetic short peptides have been used to generate protein-reactive antibodies. The advantage of immunizing with synthetic peptides is that unlimited amounts of pure stable antigens can be used. This approach involves synthesizing short peptide sequences, linking them to large carrier molecules, and immunizing selected animals with the peptide-carrier molecules. The properties of the antibodies depend on the primary sequence information. Careful selection of the sequence and linking method can usually generate a good response to the desired peptide. Most peptides can elicit a good response. The advantage of anti-peptide antibodies is that they can be prepared immediately after determining the amino acid sequence of the variant peptide, and specific regions of the protein can be specifically targeted for antibody production. Because variant peptides have been screened for high immunogenicity, there is a high probability that the resulting antibodies will recognize the native protein in a tumor setting. As in the vaccine situation, the length of the peptide is another important factor to consider. Approximately, peptides of 10-15 residues are optimal for anti-peptide antibody production, and longer peptides are better, because the number of possible epitopes increases with peptide length. However, longer peptides increase the difficulty in synthesis, purification, and linking to carrier proteins. The quality of the antibody depends on the quality of the peptide. By-products contained in the peptide product may result in poor quality antibodies.

[0115] Peptide-carrier protein linkage is another factor involved in the production of high titer antibodies. Most linkage methods rely on reactive functional groups of amino acids, such as -NH2, -COOH, -SH and phenolic -OH. Site-specific linkage is the best method. Any suitable method used in anti-peptide antibody production can be utilized with the peptides identified by the method of the present invention. Two such known methods are the multiple antigen peptide system (MAPs) and lipid core peptide (LCP method). The advantage of MAPs is that it does not require a conjugation method. No carrier protein or linkage bond is introduced into the immunized host. One of the disadvantages is that the purity of the peptide is more difficult to control. In addition, MAPs can bypass the immune response system of some hosts. The LCP method is known to provide higher titers than other anti-peptide vaccine systems and therefore may be advantageous.

[0116] Also provided herein are isolated MHC / peptide complexes comprising disease-specific immunogenic variant peptides disclosed herein. Such MHC / peptide complexes can be used, for example, to identify antibodies, soluble TCRs, or TCR analogs. Certain types of these antibodies have been called TCR mimics because they are antibodies that bind peptides from tumor-associated antigens in the context of a particular HLA environment. This type of antibody has been shown to mediate lysis of cells expressing the complex on their surface, as well as to protect mice from transplanted cancer cell lines expressing the complex (see, for example, Wittman et al., J. of Immunol. 177:4187-4195 (2006)). One advantage of TCR mimics as IgG mAbs is that affinity maturation can be performed and the molecule is linked to immune effector functions through the presented Fc region. These antibodies can also be used to target therapeutic molecules, such as toxins, cytokines, or drug products, to tumors. Other types of molecules have been developed using peptides, such as those selected using the methods of the invention using non-hybridoma-based antibody production or antibody fragments with binding capacity, such as anti-peptide Fab molecules produced on bacteriophage. These fragments can also be conjugated to other therapeutic molecules for tumor delivery, such as anti-peptide MHC Fab-immunotoxin conjugates, anti-peptide MHC Fab-cytokine conjugates and anti-peptide MHC Fab-drug conjugates.

[0117] Methods of Treatment Including Immunogenic Vaccines The present disclosure provides a method of treatment comprising an immunogenic vaccine. In some embodiments, a method of treatment for a disease (e.g., cancer) is provided, which may include administering to an individual an effective amount of a composition comprising an immunogenic peptide. In some embodiments, a method of treatment for a disease (e.g., cancer) is provided, which may include administering to an individual an effective amount of a composition comprising a precursor of an immunogenic peptide. In some embodiments, the immunogenic vaccine may comprise a pharma- ceutically acceptable disease-specific immunogenic peptide. In some embodiments, the immunogenic vaccine may comprise a pharma- ceutically acceptable precursor (e.g., protein, peptide, DNA, and RNA) of a disease-specific immunogenic peptide. In some embodiments, a method of treatment for a disease (e.g., cancer) is provided, which may include administering to an individual an effective amount of an antibody that specifically recognizes a disease-specific immunogenic variant peptide. In some embodiments, a method of treatment for a disease (e.g., cancer) is provided, which may include administering to an individual an effective amount of a soluble TCR or TCR analog that specifically recognizes a disease-specific immunogenic variant peptide.

[0118] In some embodiments, the cancer is carcinoma, lymphoma, blastoma, sarcoma, leukemia, squamous cell carcinoma, lung cancer (including small cell lung cancer, non-small cell lung cancer, adenocarcinoma of the lung, and squamous cell carcinoma of the lung), cancer of the peritoneum, hepatocellular carcinoma, gastric cancer, cancer) (including gastrointestinal cancer), pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, liver cancer, breast cancer, colon cancer, melanoma, endometrial or uterine cancer, salivary gland cancer, kidney or renal cancer, liver cancer, prostate cancer, vulvar cancer, thyroid cancer, hepatocellular carcinoma, head and neck cancer, colorectal cancer, rectal cancer, soft tissue sarcoma, Kaposi's sarcoma, B-cell lymphoma (low-grade / follicular non-Hodgkin's lymphoma (NHL), small lymphocytic (SL) NHL, intermediate-grade / follicular NHL, intermediate-grade diffuse NHL, high-grade immunoblastic and any one of the following: chronic lymphocytic leukemia (NCL), high-grade lymphoblastic NHL, high-grade small noncleaved cell NHL, bulky disease NHL, mantle cell lymphoma, AIDS-related lymphoma, and Waldenstrom's macroglobulinemia; chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), myeloma, hairy cell leukemia, chronic myeloblastic leukemia, and post-transplant lymphoproliferative disorder (PTLD); and abnormal blood vessel proliferation associated with nevus syndrome, edema (e.g., associated with a brain tumor), and Meigs' syndrome.

[0119] The methods described herein are particularly useful in the context of personalized medicine, where a disease-specific immunogenic variant peptide obtained by any one of the methods described herein is used to develop a therapeutic agent (e.g., a vaccine or a therapeutic antibody) for the same individual. Thus, for example, in some embodiments, a method for treating a disease (e.g., cancer) in an individual is provided, comprising a) identifying a disease-specific immunogenic variant peptide in the individual, and b) synthesizing a peptide or peptide precursor, and c) administering the peptide to the individual. In some embodiments, a method for treating a disease (e.g., cancer) in an individual is provided, comprising a) identifying a disease-specific immunogenic variant peptide in the individual, b) producing an antibody that specifically recognizes the variant peptide, and c) administering the peptide to the individual. In some embodiments, the identifying step combines a sequence-specific variant identification method with an immunogenicity prediction method. In some embodiments, the identifying step combines a sequence-specific variant identification method with mass spectrometry. Any method for identifying a disease-specific immunogenic variant peptide described herein can be used for the treatment method described herein. In some embodiments, the method further comprises obtaining a sample of diseased tissue from the individual.

[0120] The methods provided herein can be used to treat an individual (e.g., a human) who has been diagnosed with or is suspected of having cancer. In some embodiments, the individual may be a human. In some embodiments, the individual may be at least about any of the following ages: 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or 85 years of age. In some embodiments, the individual may be male. In some embodiments, the individual may be female. In some embodiments, the individual may have refused surgery. In some embodiments, the individual may be medically inoperable. In some embodiments, the individual may be at clinical stage Ta, Tis, Tl, T2, T3a, T3b, or T4. In some embodiments, the cancer may be recurrent. In some embodiments, the individual may be a human who exhibits one or more symptoms associated with cancer. In some embodiments, the individual may be genetically or otherwise predisposed (e.g., have risk factors) to developing cancer.

[0121] The methods provided herein can be performed in an adjuvant setting. In some embodiments, the methods are performed in a neoadjuvant setting, i.e., the methods can be performed before primary / definitive therapy. In some embodiments, the methods are used to treat previously treated individuals. Any of the methods of treatment provided herein can be used to treat previously untreated individuals. In some embodiments, the methods are used as first-line therapy. In some embodiments, the methods are used as second-line therapy.

[0122] In some embodiments, a method is provided for reducing the incidence or burden of existing cancer tumor metastasis (e.g., lung metastasis or lymph node metastasis) in an individual, comprising administering to the individual an effective amount of a composition comprising an immunogenic vaccine.

[0123] In some embodiments, a method of increasing the time to disease progression of cancer in an individual is provided, comprising administering to the individual an effective amount of a composition comprising an immunogenic vaccine.

[0124] In some embodiments, a method of extending survival of an individual with cancer is provided comprising administering to the individual an effective amount of a composition comprising an immunogenic vaccine.

[0125] In some embodiments, at least one or more chemotherapeutic agents can be administered in addition to the composition comprising the immunogenic vaccine. In some embodiments, the one or more chemotherapeutic agents can (but do not necessarily) belong to different classes of chemotherapeutic agents.

[0126] In some embodiments, a method of treating a disease (e.g., cancer) in an individual is provided, comprising administering a) an immunogenic vaccine, and b) an immunomodulatory agent. In some embodiments, a method of treating a disease (e.g., cancer) in an individual is provided, comprising a) an immunogenic vaccine, and b) an antagonist of a checkpoint protein. In some embodiments, a method of treating a disease (e.g., cancer) in an individual is provided, comprising administering a) an immunogenic vaccine, and b) an antagonist of programmed cell death 1 (PD-1), e.g., anti-PD-1. In some embodiments, a method of treating a disease (e.g., cancer) in an individual is provided, comprising administering a) an immunogenic vaccine, and b) an antagonist of programmed death-ligand 1 (PD-L1), e.g., anti-PD-L1. In some embodiments, a method of treating a disease (e.g., cancer) in an individual is provided, comprising administering a) an immunogenic vaccine, and b) an antagonist of cytotoxic T-lymphocyte-associated protein 4 (CTLA-4), e.g., anti-CTLA-4. EXAMPLES

[0127] Example 1 This example demonstrates an exemplary method for the prediction of immunogenic peptide epitopes.

[0128] Whole-exome sequencing was performed on MC-38 and TRAMP-C1 mouse tumor cell lines to identify tumor-specific point mutations. Coding variants were called against the reference mouse genome, identifying 4285 and 949 nonsynonymous variants in MC-38 and TRAMP-C1, respectively. Data were subsequently filtered for gene expression by RNA-Seq analysis, revealing that 1290 and 67 mutated genes were expressed in MC-38 and TRAMP-C1, respectively. 170 predicted neoepitopes in MC-38 and 6 predicted neoepitopes in TRAMP-C1 tumors were identified using the NETMHC-3.4 algorithm.

[0129] Next, mass spectrometry analysis using the transcriptome-generated FASTA database revealed 797 unique H-2Kb epitopes and 725 unique H-2Db epitopes presented by the MC-38 cell line, and 477 unique H-2Kb epitopes and 332 unique H-2Db epitopes presented by the TRAMP-C1 cell line. We observed that peptides derived from abundant transcripts were more likely to be presented by MHC1 in MC-38 (Figure 2A) and TRAMP-C1 (Figure 2B) cells.

[0130] Of the 1290 and 67 amino acid changes in MC-38 and TRAMP-C1, respectively, only 7 (7 in MC-38 and 0 in TRAMP-C1) were found by mass spectrometry to be presented on MHCI (Table 1). One epitope derived from the cancer-testis autoantigen MAGE-D1 was also detected by mass spectrometry in MC-38 cells. These peptides were manually verified for accuracy and compared to synthetically generated versions of the peptides. All but one of these neoepitopes was predicted to bind MHCI (IC50<500nm, Table 1). Although both wild-type (WT) and mutant transcripts were expressed by tumor cells and most of the corresponding WT peptides were also predicted to bind MHCI, only three of them were detected by mass spectrometry.

[0131] Although there is a correlation between peptide binding affinity to MHCI and immunogenicity, other factors also contribute. For example, interaction of mutant amino acids with TCR is likely essential to recognize mutant peptides as "non-self". This is especially true when the corresponding WT peptide is also presented on MHCI. Five of the seven neoepitopes showed high binding prediction scores (IC50<50nM by NETMHC-3.4, Table 1). Other neoepitopes showed lower binding prediction scores, suggesting that they may be less immunogenic. Utilizing published crystal structures of H-2Db and H-2Kb and a Rosetta-based algorithm, we modeled each of the mutant peptides in complex with MHCI and analyzed the possibility of mutant residues in each neoepitope interacting with the T cell receptor. In general, TCR recognition of displayed peptides is mediated by interactions with peptide residues 3 to 7. Among peptides with high binding scores, only Reps1 and Adpgk have peptides with mutations within this range. Structural modeling also predicted that the mutated residues were oriented toward the solvent accessible surface and therefore judged to have a good chance of being immunogenic (Table 1 and Figure 3). On the other hand, mutations in the Irgq, Aatf, and Dpagt1 neoepitopes were found near the C-terminus of the peptide, which is likely outside the TCR binding region, suggesting that these neoepitopes are unlikely to be immunogenic (Table 1 and Figure 3).

[0132] Next, the immunogenicity of the mutant tumor antigens was evaluated by immunizing wild-type C57BL / 6 mice with long peptides encoding the mutant epitopes in combination with adjuvants, and CD8 T cell responses were measured using MHCI / peptide-specific dextramers. As shown in Figure 4A, three of the six peptides induced CD8 T cell responses compared to the adjuvant-only group. Reps1 and Adpgk were predicted to be immunogenic based on structure and binding affinity predictions, and both induced strong CD8 T cell responses. Of the four peptides predicted to be non-immunogenic, only Dpagt1 induced a weak CD8 T cell response.

[0133] The immunogenicity of these mutant peptides was confirmed by analyzing tumor-infiltrating cells (TILs) in the tumor context. Reps1, Adpgk, and Dpagt1-specific T cells were observed to be enriched in the tumor bed (Figure 4B). Although there was heterogeneity, Adpgk-specific CD8 T cells were the most abundant of the three, which was specific for MC-38 tumors, since Adpgk-specific CD8 T cells were not detected in syngeneic TRAMP-C1 tumors. Interestingly, peptides derived from a single cancer-testis autoantigen (MAGE-D1) identified by mass spectrometry showed poor immunogenicity, and MAGE-D1-specific CD8 T cells could not be detected in the tumor bed (data not shown).

[0134] Bulk TILs are usually analyzed to monitor antitumor responses, which may not provide a true assessment, since only a small proportion of TILs are tumor-specific. The frequency and phenotype of antitumor TILs compared to bulk TILs were examined using MHCI / peptide-specific dextramers for three immunogenic peptides. The frequency of tumor-specific CD8 T cells infiltrating the tumor initially increased but decreased as the tumor further grew, suggesting an inverse correlation of tumor growth with the frequency of tumor-specific CD8 T cells in the tumor (Figure 4C). Interestingly, the majority (76.9±7.1%) of tumor-specific CD8 TILs co-expressed PD-1 and TIM-3, markers of T cell exhaustion, compared to bulk TILs (52.6±3.6%) (Figure 4D). Tumor-specific CD8 TILs also expressed high levels of surface PD-1.

[0135] To determine whether CD8 T cells induced against the neoepitopes could provide protective antitumor immunity, healthy mice were immunized with the mutant peptide vaccine and subsequently challenged with MC-38 tumor cells. Compared to adjuvant alone, tumor growth was completely inhibited in most animals in the vaccine group (Figure 5A). The single animal that grew a tumor in this experiment did not actually respond to the vaccine, strongly supporting the possibility that CD8 T cell responses specific to the mutant peptide could confer protection (Figure 5A).

[0136] Next, neoepitope-specific CD8 T cell responses were evaluated to see whether they could be further amplified upon immunization in tumor-bearing mice. After a single immunization, the frequency of Adpgk-reactive CD8 T cells was significantly increased in the spleens of tumor-bearing mice compared to naive healthy animals (Figure 4B). A nearly three-fold increase in the accumulation of Adgpk-specific CD8 T cells among total CD8 TILs in tumors was also observed (Figure 5B). Peptide vaccination also increased the CD45 + Cells and CD8 +Increased overall T cell infiltration, resulting in an almost 20-fold increase in the frequency of neoepitope-specific CD8 T cells among all live cells in the tumor (Figure 5C).

[0137] Furthermore, we analyzed the phenotype of peptide-specific cells induced by vaccination. + PD-1 + We found that the frequency of Adgpk-specific CD8 TILs was reduced after vaccination, and the surface expression of PD-1 and TIM-3 on these cells was also reduced (Figure 5D and Figure 5E). This may be an adjuvant effect, since it was also seen in the adjuvant-only group. This result suggests that tumor-specific T cells exhibit a less exhausted phenotype after vaccination, which was further confirmed by the higher percentage of IFN-γ-expressing CD8 and CD4 TILs in vaccinated tumors (Figure 5F).

[0138] Finally, it was assessed whether these vaccine-induced qualitative and quantitative changes in tumor-specific CD8 T cells could be translated into regression of established tumors. Even in this more challenging therapeutic setting, vaccinated mice showed a significant inhibition of tumor growth compared to untreated controls or adjuvant-only groups (Figure 5G). Thus, simple peptide vaccination with predicted neoepitopes generated sufficient T cell immunity to reject previously established tumors.

[0139] method MHCI peptide profiling was performed on the H-2Kb and H-2Db ligandomes of two mouse cell lines in an H-2b background, TRAMP-C1 (ATCC) and MC-38 (Academisch Ziekenhuis Leiden). Cells derived from C57BL / 6 mice were prepared as previously described. For a complete description of the methods for preparing the cell lines, see US Patent Application No. 13 / 087948 and US Patent Application No. 11 / 00474. MHCI molecules from each sample were immunoprecipitated using two different antibodies and H-2Kb-specific and H-2Db-specific peptides were extracted, respectively. Peptides were separated by reversed-phase chromatography (nanoAcquity UPLC system, Waters, Milford, MA) using a 180-minute gradient. Eluted peptides were analyzed by data-dependent acquisition (DDA) on an LTQ-Orbitrap Velos hybrid mass spectrometer (Thermo Fisher Scientific, Bremen, Germany) equipped with an electrospray ionization (ESI) source. Mass spectral data were acquired using a method that included a full scan (survey scan) with high mass accuracy in the Orbitrap (R = 30,000 for TOP3, R = 60,000 for TOP5), followed by an MS / MS (profile) scan either in the Orbitrap (R = 7500) for the five most abundant precursor ions (TOP5) or in the LTQ for the three most abundant precursor ions (TOP3). Seven replicate injections and analyses were performed for each set of samples.

[0140] Synthetic peptides corresponding to the identification of mutant MC-38 and TRAMP-C1 antigenic peptides were analyzed on an LTQ-Oritrap Elite mass spectrometer (ThermoFisher, Bremen, Germany) and ionized at a spray voltage of 1.2 kV using an ADVANCE source (Michrom-Bruker, Fremont, CA). Mass spectral data were acquired using a method consisting of one full MS scan (375-1600 m / z) at a resolution of 60,000 M / ΔM at m / z 400 in the Orbitrap, followed by an MS / MS (centroid) scan in the LTQ of peptide fragment ions.

[0141] One microgram of total RNA from MC-38 and TRAMP-C1 cancer cell lines was used to generate RNA-Seq libraries using the TruSeq RNA Sample Preparation Kit (Illumina, CA). Total RNA was purified from the cell lines and fragmented to 200-300 base pairs (bp) with an average length of 260 bp. RNA-Seq libraries were multiplexed (two per lane) and sequenced on a HiSeq 2000 according to the manufacturer's recommendations (Illumina, CA).

[0142] Approximately >50 million paired-end (2 x 100 bp) sequencing reads were generated per sample. Exome capture was performed using the SureSelect Human All Exome Kit (50 Mb) (Aglient, CA). Exome capture libraries were then sequenced (200 cycles) using the HiSeq sequencing kit on a HiSeq 2000 (Illumina, CA).

[0143] 92.9 million RNA fragments were sequenced from MC-38 and 65.3 million from TRAMP-C1. For exome sequencing, 60 million reads were sequenced from each cell line. Reads were mapped to the mouse genome (NCBI build 37 or mm9) using GSNAP (Wu and Nacu, Bioinformatics, 2010, v.26, 873-881). Only uniquely mapped reads were kept for further analysis. 80.6 million RNA fragments were uniquely mapped in MC-38 samples and 57.6 million in TRAMP-C1 samples. 50.9 million exome fragments were uniquely mapped in MC-38 and 52 million fragments in TRAMP-C1. To obtain mouse gene models, Refseq mouse genes were mapped to the mm9 genome using GMAP, and the genome sequences were then used to generate gene models.

[0144] Exome-seq based variants were called using GATK1. Variants with allele frequency ≥ 10% were retained. Variants were annotated for their effect on the transcript using the variant effect predictor tool 2. Only variants with interpretable amino acid changes were retained. To obtain variants with evidence of expression, exome-based variant locations were checked for evidence of difference with RNA-Seq read alignment. Variants that were documented by more than two RNA-Seq reads and expressed at ≥ 10% allele frequency based on RNA-Seq were retained.

[0145] For each amino acid difference, mutant full protein sequences were generated to form a set of putative proteins that served as a reference database for searching LC-MS spectra. Without haplotype information, multiple differences in the same protein are characterized as separate mutant proteins in the database.

[0146] Tandem mass spectral results were submitted for protein database searching using the Mascot algorithm version 2.3.02 (MatrixScience, London, UK) against the concatenated target-decoy database Uniprot version 2011_12 or a transcriptome-generated FASTA database containing mouse proteins and common laboratory contaminants such as trypsin. Data were searched with no enzyme specificity, methionine oxidation (+15.995 Da), and 20 ppm precursor ion mass tolerance.

[0147] Fragment ion mass tolerance was specified at 0.8 Da or 0.05 Da for MS / MS data acquired on LTQ or Orbitrap, respectively. Search results were filtered using a linear discriminant algorithm (LDA) to an estimated peptide false discovery rate (FDR) of 5%. For higher confidence in variant peptide identification, data were further filtered by either a peptide length of 8 for H-2Kb data and 9 for H-2Db, or by using regular expressions to isolate peptides with the following well-characterized anchor motifs: H-2Kb: XXXX[FY]XX[MILV] (SEQ ID NO: 22) and H-2Db: XXXX[N]XXX[MIL] (SEQ ID NO: 23). Synthetic peptides were generated to verify the sequences.

[0148] For the generation of the first model, peptide-MHC complex structures were selected from the PDB based on sequence similarity between the mutant peptides and the peptides in the model structure. For each mutant peptide model, the following PDB codes were used: Reps1, 2ZOL 4; Adpgk, 1HOC 5; Dpagt1, 3P9L 6; Cpne1, 1JUF 7; Irgq, 1FFN 8; Aatf, 1BZ9. The Med12 peptide was not modeled due to the lack of a published H-2Kb crystal structure in complex with a 10 amino acid long peptide that could be used as a reasonable starting model. The peptide was then modified into the mutant form using COOT 10. These first models were then optimized using the Rosetta FlexPepDock web server 11, and the top scoring model was selected for display.

[0149] The top-scoring FlexPepDock model for each peptide was also examined, and backbone arrangements were found to be similar for the top 10 models generated. Peptide-MHCI images were generated using Pymol (Schrodinger, LLC).

[0150] Age-matched 6-8 week old C57BL / 6 mice (The Jackson Laboratory) were intraperitoneally injected with 50 mg of each long peptide combined with adjuvant (50 μg of anti-CD40 Ab clone FJK45 plus 100 mg of poly(I:C) (Invivogen)) in PBS. Mice were immunized on days 0 and 14, and one week after the final injection, either blood or splenocytes were used for detection of Ag-specific CD8 T cells. To identify peptide-specific T cells, cells were stained with PE-conjugated dextramer (MHCI / peptide complex; Immudex, Denmark) for 20 min, followed by staining with cell surface markers CD3, CD4, CD8 and B220 (BD Biosciences). The peptide sequences are as follows: Reps1: GRVLELFRAAQLANDVVLQIMELCGATR (SEQ ID NO: 1), Adpgk: GIPVHLELASMTNMELMSSIVHQQVFPT (SEQ ID NO: 2), Dpagt1: EAGQSLVISASIIVFNLLELEGDYR (SEQ ID NO: 3), Aatf: SKLLSFMAPIDHTTMSDDARTELFRS (SEQ ID NO: 4), Irgq: KARDETAALLNSAVLGAAPLFVPPAD (SEQ ID NO: 5), Cpne1: DFTGSNGDPSSPYSLHYLSPTGVNEY (SEQ ID NO: 6), Med12: GPQEKQQRVELSSISNFQAVSELLTFE (SEQ ID NO: 7).

[0151] C57BL / 6 mice were injected with 1 × 10 5 of MC-38 tumor cells were implanted subcutaneously. Whole tumors were isolated, digested with collagenase and DNAase, and TILs were isolated. TILs were stained with dextramer (as described above) followed by antibodies against CD45, Thy1.2, CD4, CD8 (BD Biosciences), PD-1 (eBiosciences), and TIM-3 (R&D Systems). Live / dead staining was used to gate on live cells.

[0152] All animals received 1 × 10 5 of MC-38 cells were inoculated subcutaneously (right hind flank). For prophylaxis studies, mice were immunized with adjuvant (50 mg anti-CD40 plus 100 mg poly(I:C)) or adjuvant with 50 μg each of Reps1, Adpgk, and Dpagt1 peptides 3 weeks before tumor inoculation. Induction of peptide-specific CD8 T cells was measured in blood 1 day before tumor cell inoculation. For vaccination in tumor-bearing mice, 1 × 10 5 Ten days after inoculation with MC-38 tumor cells (approximately 100-150 mm on day 10), 3 Mice were injected with adjuvant or adjuvant with 50 μg of Reps1, Adpgk, and Dpagt1 peptides, respectively, at 10 min (only tumors with a volume of 0.01 μg / mL were included in the study). Measurements and body weights were collected twice a week. Animals showing weight loss of more than 15% of their initial body weight were weighed daily and euthanized if they lost more than 20% of their initial body weight.

[0153] Animals that showed adverse clinical problems were observed more frequently, up to daily, depending on severity, and were euthanized if moribund. Mice were cultured until tumor volumes reached 3,000 mm 3 Mice were euthanized after 3 months if their body weight exceeded 100 mg / kg or if no tumors had formed. Clinical observations of all mice were performed twice weekly throughout the study.

[0154] TIFF2024156702000001.tif209170

Claims

1. 1. A method for identifying a disease specific immunogenic variant peptide from diseased tissue in an individual, comprising: (a) obtaining a nucleic acid sequence from a diseased tissue that includes a set of variant coding sequences, each variant coding sequence encoding a polypeptide having a variation in sequence compared to a reference sample; and (b) selecting an immunogenic mutant coding sequence from the set of mutant coding sequences, comprising predicting the immunogenicity of a peptide containing the mutant amino acid encoded by the mutant coding sequence. Including, The predicting step comprises: (i) the stability of the peptide bound to an MHC molecule; (ii) orientation of the mutated amino acid such that the mutated amino acid faces the solvent accessible surface so as to interact with the TCR; (iii) the position of the mutated amino acid within the peptide; (iv) the properties of the mutated amino acid compared to the wild-type residue; or (v) the level of a peptide precursor comprising said peptide. based on one or more of the above, thereby identifying disease specific immunogenic variant peptides.

2. Predicting immunogenicity i) the binding affinity of the peptide to an MHC molecule; ii) the expression level of transcripts encoding peptide precursors; iii) the unique tumor-specific sequence of said peptide; iv) efficiency of processing of peptide precursors by the immunoproteasome; v) timing of expression of peptide precursors; vi) the binding affinity of the peptide to a TCR molecule; vii) solvent exposure of the peptide when bound to the MHCI molecule; viii) the solvent exposure of the mutated amino acid when bound to an MHCI molecule; or ix) Nature of the peptide precursor The method of claim 1 further based on one or more of:

3. The method of claim 1 or 2, wherein obtaining the set of mutant coding sequences of the diseased tissue in the individual is based on the genomic sequence of the diseased tissue in the individual.

4. 4. The method of claim 3, wherein obtaining a set of variant coding sequences of diseased tissue in the individual comprises selecting genomic variant coding sequences based on a transcriptome sequence of diseased tissue in the individual.

5. 5. The method of claim 4, wherein obtaining a set of mutant coding sequences of diseased tissue in an individual comprises selecting transcriptome sequences based on the predicted binding ability of the peptide to an MHC molecule.

6. The method of claim 5 , wherein the MHC molecule is an MHC class I molecule.

7. 7. The method of claim 1, wherein the property of the mutated amino acid compared to the wild-type residue is based on the mutated amino acid being an aromatic residue.

8. The method according to any one of claims 1 to 7, wherein the prediction of immunogenicity further comprises HLA typing analysis.

9. 9. The method of claim 1, further comprising synthesizing a peptide based on the sequence of the identified disease-specific immunogenic variant peptide.

10. The method of claim 1 , further comprising synthesizing a nucleic acid encoding the peptide based on the sequence of the identified disease-specific immunogenic variant peptide.

11. 11. The method of claim 1, further comprising testing the peptide in vivo for immunogenicity.

12. 11. The method of claim 1, further comprising testing the peptide in vitro for immunogenicity.

13. 13. The method of any one of claims 1 to 12, wherein the disease is cancer.

14. 14. The method of any one of claims 1 to 13, wherein the individual is a human.

15. 15. A method for producing a disease specific immunogenic variant peptide, comprising identifying a disease specific immunogenic variant peptide by the method of any one of claims 1 to 14.

16. 16. A method for producing a composition comprising a disease-specific immunogenic variant peptide, comprising the method of claim 15.

17. The method of claim 16, wherein the composition comprises two or more disease-specific immunogenic variant peptides.

18. The method of claim 16 or 17, wherein the composition further comprises an adjuvant.

19. 19. The method of any one of claims 16 to 18, wherein the composition is used to treat a disease in an individual.

20. 20. The method of claim 19, wherein the individual is the same individual in which the disease specific immunogenic variant peptide was identified.

21. A method for producing an immunogenic composition comprising at least one disease-specific peptide, comprising identifying a disease-specific peptide by the method of any one of claims 1 to 14.

22. A method for producing an immunogenic composition comprising at least one nucleic acid encoding at least one disease-specific peptide, comprising identifying a disease-specific peptide by the method of any one of claims 1 to 14.

23. 15. A method for producing an immunogenic composition comprising a plurality of disease-specific peptides, comprising identifying disease-specific peptides by the method of any one of claims 1 to 14.

24. 15. A method for producing an immunogenic composition comprising nucleic acids encoding a plurality of disease-specific peptides, comprising identifying disease-specific peptides by the method of any one of claims 1 to 14.

25. 20. A method for producing a composition for stimulating an immune response in an individual with a disease, comprising producing the composition by a method according to any one of claims 16 to 18.

26. 26. The method of claim 25, wherein the composition is administered in combination with another agent.

27. a) identifying disease-specific immunogenic variant peptides from diseased tissue in an individual by the method of any one of claims 1 to 14; and b) producing a composition comprising a peptide or a nucleic acid encoding the peptide based on the sequence of the identified disease-specific immunogenic variant peptide; A method for producing a composition for stimulating an immune response in an individual with a disease, comprising:

28. A system for identifying disease-specific immunogenic variant peptides from diseased tissue in an individual, comprising a memory storing one or more algorithms comprising instructions for carrying out the method of any one of claims 1 to 14, configured to be executed by one or more processors.