Neoantigen identification, manufacture, and use

An optimized approach using next-generation sequencing and machine-learning models improves the identification of tumor-specific neoantigens, addressing low PPV issues in current methods and enhancing the effectiveness of personalized cancer vaccines by increasing the likelihood of inducing anti-tumor immunity and reducing autoimmunity.

JP2025094018APending Publication Date: 2025-06-24GRITSTONE BIO INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025039865
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-06-09
Filing Date
2025-03-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Current methods for identifying tumor-specific neoantigens have low positive predictive value (PPV), leading to ineffective neoantigen vaccination in a significant number of patients due to inaccurate prediction of peptide presentation on tumor surfaces, and may induce autoimmunity or inefficiencies in vaccine design.

Method used

An optimized approach using next-generation sequencing and machine-learning models to identify neoantigens, considering peptide-MHC binding, stability, and tumor-specific presentation likelihoods, to select neoantigens for personalized cancer vaccines.

Benefits of technology

Improves the positive predictive value of neoantigen selection, increasing the likelihood of inducing anti-tumor immunity and reducing autoimmunity, thereby enhancing the effectiveness of personalized cancer vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094018000001_ABST
    Figure 2025094018000001_ABST
Patent Text Reader

Abstract

To provide an optimized approach for identifying and selecting neoantigens for personalized cancer vaccines.SOLUTION: Provided is a system and methods for determining the alleles, neoantigens, and vaccine composition as determined on the basis of an individual's tumor mutations. Also provided are systems and methods for obtaining high quality sequencing data from a tumor. Further, provided are systems and methods for identifying somatic changes in polymorphic genome data. Also further provided are systems and methods for selecting a subset of patients for treatment. A utility score indicating an estimated number of neoantigens presented on the surface of tumor cells is determined for each patient based on one or more neoantigen candidates identified for the patient. The subset of patients are selected based on the determined utility scores. The selected subset of patients can receive treatment, such as neoantigen vaccines or checkpoint inhibitor therapy. Still further provided are unique cancer vaccines.SELECTED DRAWING: Figure 13B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 517,786, filed Jun. 9, 2017, which is hereby incorporated by reference in its entirety.

Background Art

[0002] Background Therapeutic vaccines based on tumor - specific neoantigens are extremely promising as the next generation of personalized cancer immunotherapies. 1~3 Cancers with a high amount of genetic mutations, such as non - small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies because they are relatively likely to generate neoantigens. 4,5 Initial evidence shows that vaccination with neoantigen - based vaccines induces T - cell responses 6 and that cell therapies targeting neoantigens can cause tumor regression in selected patients. 7 Both MHC class I and MHC class II affect T - cell responses. 70~71 .

[0003] One question regarding the design of neoantigen vaccines is which of the numerous coding mutations present in the target tumor can give rise to the "best" therapeutic neoantigens (e.g., antigens capable of inducing anti - tumor immunity and causing tumor regression).

[0004] Initial methods incorporating mutation - based analysis using next - generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides have been proposed. 8However, in these proposed methods, many steps other than gene expression and MHC binding (e.g., TAP transport, proteasome cleavage, MHC binding, transport of the peptide-MHC complex to the cell surface, and / or recognition of MHC-I by TCR; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsin), competition with CLIP peptides for HLA binding catalyzed by HLA-DM, transport of the peptide-MHC complex to the cell surface, and / or recognition of MHC-II by TCR) are included 9 The entire epitope generation process cannot be modeled. Therefore, existing methods tend to have the problem of low positive predictive value (PPV) (Figure 1A).

[0005] In fact, analysis of peptides presented by tumor cells performed by multiple groups has shown that less than 5% of the peptides predicted to be presented using gene expression and MHC binding affinity are found on MHC on the tumor surface 10,11 (Figure 1B). Such a low correlation between binding prediction and MHC presentation is further indicated by the lack of improvement in the predictive accuracy of neoantigens restricted to binding for checkpoint inhibitor response with respect to the number of mutations alone. 12 。

[0006] Such a low positive predictive value (PPV) of existing methods for predicting presentation presents a problem in the design of neoantigen-based vaccines. If a vaccine is designed using a low-PPV prediction, the likelihood of administering therapeutic neoantigens to most patients will be low, and the number of patients receiving multiple neoantigens will be even lower (even assuming that all presented peptides are immunogenic). Therefore, neoantigen vaccination by current methods is likely to be ineffective in a significant number of subjects with tumors (Figure 1C).

[0007] Furthermore, previous approaches have generated candidate neoantigens using only cis-acting mutations, mutations in splicing factors that occur in multiple tumor types and lead to aberrant splicing in many genes 13 , and in most cases have not considered additional sources of neo-ORFs, including mutations that create or remove protease cleavage sites.

[0008] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard approaches to tumor analysis may erroneously promote sequence artifacts or germline polymorphisms as neoantigens, leading to inefficient use of vaccine capacity or risk of autoimmunity, respectively. SUMMARY OF THE INVENTION

[0009] Summary Disclosed herein is an optimized approach for identifying and selecting neoantigens for personalized cancer vaccines. First, efforts are made towards an optimized tumor exome and transcriptome analysis approach for identifying neoantigen candidates using next-generation sequencing (NGS). These methods are based on the standard approach for tumor analysis by NGS such that the most sensitive and specific neoantigen candidates are developed across all classes of genomic changes. Second, a novel approach to high PPV neoantigen selection is provided to overcome the problem of specificity and to make the neoantigens developed for vaccine addition more likely to induce anti-tumor immunity. These approaches include, depending on the embodiment, a trained statistical regression or non-linear deep learning model that co-models peptide-allergen mapping, and an allele-wise motif for peptides of multiple lengths that share statistical power across peptides of different lengths. In particular, the non-linear deep learning model can be designed and trained to treat different MHC alleles within the same cell as independent, thus solving the problems associated with linear models where linear models interfere with each other. Finally, further concerns regarding the design and manufacture of neoantigen-based personalized vaccines are addressed.

[0010] Also disclosed herein is a method for identifying a subset of patients suitable for treatment. From the tumor cells and normal cells of each patient, at least one of exosome, transcriptome, or whole-genome tumor nucleotide sequencing data is obtained. Using the tumor nucleotide sequencing data, the peptide sequence of each set of neoantigens identified by comparing the nucleotide sequencing data from tumor cells with the nucleotide sequencing data from normal cells is obtained. The peptide sequence of each neoantigen of a patient includes at least one change that makes it different from the corresponding wild-type parental peptide sequence identified from the patient's normal cells. By inputting the peptide sequence of each set of neoantigens into a machine-learned presentation model, a set of numerical presentation likelihoods of the set of neoantigens is generated for each patient. Each presentation likelihood represents the likelihood that the corresponding neoantigen is presented by one or more MHC alleles on the surface of the patient's tumor cells. The set of presentation likelihoods is determined based at least on mass spectrometry data. One or more neoantigens are identified from the set of neoantigens of the patient. A utility score is determined for each patient, indicating the estimated number of neoantigens presented on the surface of the patient's tumor cells, determined by the corresponding presentation likelihoods of the one or more neoantigens for the patient. A subset of patients is selected for treatment. Each patient within this subset of patients is associated with a utility score that meets a predetermined inclusion criterion. Treatments such as neoantigen vaccines or checkpoint inhibitor therapies can be administered to the selected subset of patients. [The present invention 1001] A method for identifying a subset of patients suitable for treatment, comprising Obtaining for each patient at least one of exosome, transcriptome, or whole-genome tumor nucleotide sequencing data from the patient's tumor cells and normal cells, wherein the tumor nucleotide sequencing data is used to obtain the respective peptide sequences of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, and the peptide sequence of each neoantigen for the patient comprises at least one change that makes it different from the corresponding wild-type parental peptide sequence identified from the patient's normal cells, said obtaining, Generating for each patient a set of numerical presentation likelihoods for the set of neoantigens for the patient by inputting the respective peptide sequences of the set of neoantigens into a presentation model trained by machine learning, wherein each presentation likelihood represents the likelihood that the corresponding neoantigen is presented by one or more MHC alleles on the surface of the patient's tumor cells, and the set of presentation likelihoods is identified based at least on mass spectrometry data, said generating, Identifying for each patient one or more neoantigens from the set of neoantigens of the patient, Determining for each patient a utility score indicating the estimated number of neoantigens presented on the surface of the patient's tumor cells, determined by the corresponding presentation likelihoods for the one or more neoantigens for the patient, Selecting a subset of patients suitable for treatment, wherein each patient within the subset of patients is associated with a utility score that meets a predetermined inclusion criterion, said selecting The method comprising. [Invention 1002] The method of Invention 1001, wherein identifying the one or more neoantigens for the patient comprises selecting a subset of neoantigens from the set of neoantigens for the patient. [Invention 1003] The method of the present invention 1002, wherein the subset of neoantigens is a neoantigen having the highest presentation likelihood among the set of presentation likelihoods for the patient. [The present invention 1004] The method of the present invention 1001, further comprising treating each patient within the selected subset of the patients with a corresponding neoantigen vaccine comprising at least one of the one or more neoantigens identified for the patient. [The present invention 1005] The method of the present invention 1001, further comprising identifying, for each patient within the selected subset of the patients, one or more T cells or T cell receptors that are antigen-specific for at least one of the one or more neoantigens identified for the patient. [The present invention 1006] The method of the present invention 1001, wherein identifying the one or more neoantigens for the patient comprises selecting the entire set of neoantigens identified for the patient. [The present invention 1007] The method of the present invention 1006, further comprising administering checkpoint inhibitor therapy to each patient within the selected subset of the patients. [The present invention 1008] The method of the present invention 1001, wherein selecting a subset of patients suitable for treatment comprises selecting a subset of patients having a tumor mutation burden (TMB) higher than a minimum threshold, and the TMB of a patient indicates the number of neoantigens within the set of neoantigens associated with that patient. [The present invention 1009] Selecting a subset of patients suitable for treatment comprises selecting a subset of patients having a utility score higher than a minimum threshold The method of the present invention 1001. [The present invention 1010] The method of the present invention 1001, wherein the utility score is the sum of the presentation likelihoods for each neoantigen within the identified subset of neoantigens of the patient. [The present invention 1011] The method of the present invention 1001, wherein the usefulness score is the probability that the number of presented neoantigens among the one or more identified neoantigens for the patient exceeds a minimum threshold. [The present invention 1012] The machine-learned presentation model a label obtained by mass spectrometry that measures the presence of a peptide bound to at least one MHC allele identified as being present in at least one of a plurality of samples, the training peptide sequence comprising a plurality of amino acids constituting the training peptide sequence and information regarding the set of positions of the amino acids within the training peptide sequence, at least one MHC allele associated with the training peptide sequence and a plurality of parameters identified based at least on a training dataset comprising; a function representing the relationship between the peptide sequence and the presentation likelihood based on the plurality of parameters The method of the present invention 1001. [The present invention 1013] The training dataset (a) data related to a measured value of peptide-MHC binding affinity for at least one of the isolated peptides, and (b) data related to a measured value of peptide-MHC binding stability for at least one of the isolated peptides The method of the present invention 1012, further comprising at least one of [The present invention 1014] The set of numerical likelihoods (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within its source protein sequence, and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within its source protein sequence The method of the present invention 1001, further specified by a characteristic comprising at least one of [The present invention 1015] The method of the present invention 1001, wherein the set of presentation likelihoods is further specified by at least the expression levels of the one or more MHC alleles of the subject measured by RNA-seq or mass spectrometry. [The present invention 1016] The set of presentation likelihoods is (a) the predicted affinity between the neoantigens within the set of neoantigens and the one or more MHC alleles, and (b) the predicted stability of the neoantigen-encoding peptide-MHC complex The method of the present invention 1001, which is further specified by a property including at least one of them. [The present invention 1017] Inputting the peptide sequence into the machine-learned presentation model is Applying the machine-learned presentation model to the peptide sequence of each neoantigen to generate, for each of the one or more MHC alleles, a dependency score indicating whether the MHC allele presents the neoantigen based on specific amino acids at specific positions of the peptide sequence The method of the present invention 1001, which includes this. [The present invention 1018] Inputting the peptide sequence into the machine-learned presentation model is Converting the dependency score to generate, for each MHC allele, an allele-specific likelihood indicating the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and Combining the allele-specific likelihoods to generate the presentation likelihood of the neoantigen The method of the present invention 1017, which includes this. [The present invention 1019] The method of the present invention 1018, wherein converting the dependency score models the presentation of the neoantigen as mutually exclusive across the one or more class MHC alleles. [The present invention 1020] Inputting the peptide sequence into the machine-learned presentation model is Converting the combination of the dependency scores to generate the presentation likelihood The method of the present invention 1017, which includes and models the conversion of the combination of the dependency scores such that the presentation of the neoantigen interferes between the one or more MHC alleles. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] These features, aspects, and facets of the present invention, as well as other features, aspects, and facets, will be better understood with reference to the following description and the accompanying drawings.

[0012]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 1F

Figure 1G

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0013] DETAILED DESCRIPTION I. DEFINITIONS In general, the terms used in the claims and the specification are to be construed as having the ordinary meaning understood by those of ordinary skill in the art. Specific terms are defined below for further clarity. In the event of a conflict between the ordinary meaning and the given definition, the given definition shall be used.

[0014] As used herein, the term "antigen" refers to a substance that induces an immune response.

[0015] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes the antigen different from the corresponding wild-type parental antigen, for example, by a mutation in a tumor cell or a post-translational modification specific to a tumor cell. A neoantigen may comprise a polypeptide sequence or a nucleotide sequence. The mutation can include a frameshift or non-frameshift insertion / deletion (indel), a missense or nonsense substitution, a splice site change, a genomic rearrangement or gene fusion, or any genomic or expression change that results in a novel ORF. The mutation can also include splice variants. Post-translational modifications specific to tumor cells can include abnormal phosphorylation. Post-translational modifications specific to tumor cells can also include splice antigens generated by the proteasome. See Liepe et al., A large fraction of HLA class I ligands are proteasome-generated spliced peptides; Science. 2016 Oct 21;354(6310):354-358.

[0016] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in a subject's tumor cells or tissue but not in the subject's corresponding normal cells or tissue.

[0017] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct based on one or more neoantigens, for example, multiple neoantigens.

[0018] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.

[0019] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.

[0020] As used herein, the term "coding mutation" refers to a mutation that occurs in a coding region.

[0021] As used herein, the term "ORF" means open reading frame.

[0022] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that results from other abnormalities such as mutations or splicing.

[0023] As used herein, the term "missense mutation" refers to a mutation that causes a substitution of one amino acid for another.

[0024] As used herein, the term "nonsense mutation" refers to a mutation that causes a substitution of an amino acid for a stop codon.

[0025] As used herein, the term "frameshift mutation" refers to a mutation that causes a change in the protein frame.

[0026] As used herein, the term "indel" refers to an insertion or deletion of one or more nucleic acids.

[0027] As used herein, the term "identity" (%) in the context of the sequences of two or more nucleic acids or polypeptides refers to the percentage of specific nucleotides or amino acid residues that are the same when compared and aligned using one of the following sequence comparison algorithms (e.g., BLASTP and BLASTN, or other algorithms available to those skilled in the art), or by visual inspection, for the greatest match. Depending on the application, "identity" (%) can exist over the region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.

[0028] In array comparison, generally, one array functions as a reference array against which a test array is compared. When using an array comparison algorithm, the test array and the reference array are input into a computer, and if necessary, partial array coordinates are specified and the parameters of the array algorithm program are specified. Then, the array comparison algorithm calculates the percent array identity of the test array relative to the reference array based on the specified program parameters. Alternatively, the similarity or dissimilarity of arrays can also be established by the presence or absence of specific nucleotides at selected array positions (e.g., array motifs) or, in the translated array, by the combination of amino acids.

[0029] Optimal alignment of the arrays for comparison can be carried out, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson & Lipman, Proc. Nat’l. Acad. Sci. USA 85:2444 (1988), by computer-implemented execution of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (generally, see Ausubel et al. below).

[0030] One example of an algorithm suitable for determining percent array identity and percent array similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information.

[0031] As used herein, the term "non-stop or read-through" refers to a mutation that causes the removal of a natural stop codon.

[0032] As used herein, the term "epitope" refers to the specific part of an antigen to which an antibody or T cell receptor generally binds.

[0033] As used herein, the term "immunogenicity" refers to the ability to induce an immune response, for example, via T cells, B cells, or both.

[0034] As used herein, the terms "HLA binding affinity", "MHC binding affinity" refer to the affinity of the binding of a specific antigen to a specific MHC allele.

[0035] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.

[0036] As used herein, the term "mutation" refers to the difference between a subject nucleic acid and a reference human genome used as a control.

[0037] As used herein, the term "mutation call" refers to an algorithmic determination of the presence of a mutation, typically from sequencing.

[0038] As used herein, the term "polymorphism" refers to a germline mutation, that is, a mutation found in all DNA-bearing cells of an individual.

[0039] As used herein, the term "somatic mutation" refers to a mutation that occurs in non-germline cells of an individual.

[0040] As used herein, the term "allele" refers to one version of a gene or one version of a gene sequence or one version of a protein.

[0041] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.

[0042] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the cellular degradation of mRNA due to premature stop codons.

[0043] As used herein, the term "truncal mutation" refers to a mutation that occurs early in tumor development and is present in most of the tumor cells.

[0044] As used herein, the term "subclonal mutation" refers to a mutation that occurs late in tumor development and is present in only a subset of the tumor cells.

[0045] As used herein, the term "exome" refers to a subset of the genome that encodes proteins. The exome can be the collective exons of the genome.

[0046] As used herein, the term "logistic regression" refers to a regression model for binary data from statistics in which the logit of the probability that the dependent variable equals 1 is modeled as a linear function of the dependent variable.

[0047] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise non-linear transformations typically trained by stochastic gradient descent and backpropagation.

[0048] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.

[0049] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. The peptidome may also refer to the properties of a cell or a collection of cells (e.g., the tumor peptidome means the union of the peptidomes of all cells containing the tumor).

[0050] As used herein, the term "ELISPOT" means enzyme-linked immunosorbent spot assay, a common method for observing immune responses in humans and animals.

[0051] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.

[0052] As used herein, the term "tolerance or immunotolerance" refers to a state of immune non-responsiveness to one or more antigens, such as self-antigens.

[0053] As used herein, the term "central tolerance" refers to the tolerance imparted in the thymus by either deleting autoreactive T cell clones or promoting the differentiation of autoreactive T cell clones into immunosuppressive regulatory T cells (Tregs).

[0054] As used herein, the term "peripheral tolerance" refers to the tolerance imparted in the peripheral system by downregulating or anergizing autoreactive T cells that have survived central tolerance, or by promoting the differentiation of these T cells into Tregs.

[0055] The term "sample" can include a single cell, or a plurality of cells, or a fragment of a cell, or an aliquot of a body fluid, collected from a subject by means including, but not limited to, venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, wash sample, swabbing, surgical incision, or intervention, or other means known in the art.

[0056] The term "subject" includes any cell, tissue, or organism, human or non-human, male or female, in vivo, ex vivo, or in vitro. The term "subject" includes mammals including humans.

[0057] The term "mammal" includes both human and non-human mammals, including but not limited to humans, non-human primates, dogs, cats, mice, cows, horses, and pigs.

[0058] The term "clinical factor" refers to a measure of the state of a subject, such as the activity or severity of a disease. "Clinical factor" includes all markers of the health state of the subject, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and gender. A clinical factor can be a score, value, or set of values obtained from the assessment of a subject or a sample (or population of samples) from a subject under defined conditions. A clinical factor can also be predicted by other parameters such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.

[0059] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.

[0060] As used in this specification and the appended claims, it should be noted that the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise.

[0061] Terms not directly defined herein should be understood to have the meanings commonly associated with them as would be understood within the scope of the technical field of the present invention. Specific terms are considered herein for the purpose of providing further guidance to the practitioner in describing the compositions, devices, methods, etc. of aspects of the present invention, as well as the methods of making or using them. It will be recognized that there may be multiple ways of referring to the same thing. Thus, alternative words and synonyms may be used for any one or more of the terms considered herein. It should not be weighted whether a term is detailed or considered herein. Some synonyms or alternative ways, materials, etc. are provided. The recitation of one or several synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the aspects of the invention herein.

[0062] All references, issued patents, and patent applications cited in the text of this specification are hereby incorporated by reference in their entirety for all purposes.

[0063] II. Methods for Identifying Neoantigens Disclosed herein are methods for identifying neoantigens derived from a subject's tumor that are likely to be presented on the cell surface of the tumor and / or are likely to be immunogenic. As an example, one such method includes obtaining from the subject's tumor cells at least one of exosome, transcriptome, or whole-genome tumor nucleotide sequencing data, wherein data representing the peptide sequence of each of a set of neoantigens is obtained using the tumor nucleotide sequencing data, and the peptide sequence of each neoantigen includes at least one change that makes the peptide sequence different from the corresponding wild-type parental peptide sequence; inputting the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each neoantigen is presented by one or more MHC alleles on the tumor cell surface of the subject's tumor cells or by cells present within the tumor, wherein the set of numerical likelihoods is identified based at least on received mass spectrometry data; and selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a selected set of neoantigens.

[0064] The presentation model can include a statistical regression or machine learning (e.g., deep learning) model trained on a set of reference data (also referred to as a training data set) that includes a set of corresponding labels, the set of reference data being obtained from each of a plurality of distinct subjects, some of whom may have a tumor, and the set of reference data including at least one of data representing an exome nucleotide sequence from tumor tissue, data representing an exome nucleotide sequence from normal tissue, data representing a transcriptome nucleotide sequence from tumor tissue, data representing a proteome sequence from tumor tissue, data representing an MHC peptidome sequence from tumor tissue, and data representing an MHC peptidome sequence from normal tissue. The reference data can further include mass spectrometry data, sequencing data, RNA sequencing data, and proteomics data of single allele cell lines engineered to express a given MHC allele, as well as T cell assays (e.g., ELISPOT), that are subsequently exposed to synthetic proteins, normal and tumor human cell lines, and fresh and frozen primary samples. In certain embodiments, the set of reference data includes each form of reference data.

[0065] The presentation model can include a set of characteristics that are at least partially derived from the set of reference data, the set of characteristics including at least one of allele-dependent characteristics and allele-independent characteristics. In certain embodiments, each characteristic is included.

[0066] The characteristics of dendritic cell presentation to naive T cells can include at least one of the above characteristics. The dose and type of antigen in the vaccine (e.g., peptide, mRNA, virus, etc.): (1) the pathway by which dendritic cells (DCs) take up the antigen type (e.g., endocytosis, micropinocytosis); and / or (2) the efficiency by which the antigen is taken up by DCs. The dose and type of adjuvant in the vaccine. The length of the vaccine antigen sequence. The number and site of vaccine administrations. The baseline immune function of the patient (measured, for example, by the history of recent infections, blood cell counts, etc.). For RNA vaccines, (1) the metabolic turnover rate of the mRNA protein product in dendritic cells, (2) the translation rate of the mRNA after uptake by dendritic cells, as measured by in vitro or in vivo experiments, and / or (3) the number or rounds of translation of the mRNA after uptake by dendritic cells, as measured by in vivo or in vitro experiments. Optionally, the presence of protease cleavage motifs in the peptide, which gives additional weight to proteases typically expressed in dendritic cells (measured, for example, by RNA-seq or mass spectrometry). The levels of expression of the proteasome and immunoproteasome in typical activated dendritic cells (which can be measured by RNA-seq, mass spectrometry, immunohistochemistry, or other standard techniques). Optionally, the expression level of a specific MHC allele in the individual being targeted, specifically measured in activated dendritic cells or other immune cells (measured, for example, by RNA-seq or mass spectrometry). Optionally, the probability of peptide presentation by a specific MHC allele in other individuals expressing the specific MHC allele, specifically measured in activated dendritic cells or other immune cells. Optionally, the probability of peptide presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals, specifically measured in activated dendritic cells or other immune cells.

[0067] The immune tolerance escape characteristics can include at least one of the following: direct measurement of the self-peptidome by protein mass spectrometry performed on one or several cell types. Estimation of the self-peptidome by taking the union of all k-mers (e.g., 5-25) substrings of self-proteins. Estimation of the self-peptidome using a presentation model similar to the above-presented model applied to all non-mutated self-proteins, optionally explaining germline mutations.

[0068] Ranking can be performed using a plurality of neoantigens provided by at least one model based at least in part on numerical likelihood. After ranking, selection can be made to select a subset of the ranked neoantigens according to selection criteria. After selection, a subset of the ranked peptides can be given as output.

[0069] The number of the selected set of neoantigens can be 20.

[0070] The presentation model can represent the dependence between the presence of a pair of a specific one of the MHC alleles and a specific amino acid at a specific position of the peptide sequence, and the likelihood of presentation on the tumor cell surface of such a peptide sequence containing the specific amino acid at the specific position by a specific one of the pair of MHC alleles.

[0071] The method disclosed herein may also include applying one or more presentation models to the peptide sequences of the corresponding neoantigens to generate a dependence score indicating whether the corresponding neoantigen is presented by an MHC allele for each of the one or more MHC alleles, based at least on the amino acids at least at the positions of the peptide sequences of the corresponding neoantigens.

[0072] The methods disclosed herein may also include converting a dependency score to generate an allele - specific likelihood for each MHC allele, which indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining the allele - specific likelihoods to generate a numerical likelihood.

[0073] The step of converting the dependency score can model the presentation of the peptide sequence of the corresponding neoantigen as mutually exclusive.

[0074] The methods disclosed herein may further include converting a combination of dependency scores to generate a numerical likelihood.

[0075] The step of converting the combination of dependency scores can model the presentation of the peptide sequence of the corresponding neoantigen as interference between MHC alleles.

[0076] The set of numerical likelihoods can be further specified at least by allele - non - interacting properties. The methods disclosed herein may also include applying an allele - non - interacting model among one or more presentation models to the allele - non - interacting properties to generate a dependency score for the allele - non - interacting properties, which indicates whether the peptide sequence of the corresponding neoantigen is presented based on the allele - non - interacting properties.

[0077] The methods disclosed herein may also include combining the dependency score for each MHC allele at one or more MHC alleles with the dependency score for the allele - non - interacting properties; converting the combined dependency scores for each MHC allele to generate a corresponding allele - specific likelihood for the MHC allele, which indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining the allele - specific likelihoods to generate a numerical likelihood.

[0078] The methods disclosed herein may also include generating a numerical likelihood by transforming a combination of a dependency score for each of the MHC alleles and a dependency score for the allele non-interaction characteristics.

[0079] A set of numerical parameters for the presentation model can be trained based on a training dataset that includes at least a set of training peptide sequences identified as present in a plurality of samples and one or more MHC alleles associated with each training peptide sequence, where the training peptide sequences are identified by mass spectrometry of isolated peptides eluted from MHC alleles derived from a plurality of samples.

[0080] The sample may also include a cell line engineered to express a single MHC class I or class II allele.

[0081] The sample may also include a cell line engineered to express multiple MHC class I or class II alleles.

[0082] The sample may also include human cell lines obtained from or derived from multiple patients.

[0083] The sample may also include fresh or frozen tumor samples obtained from multiple patients.

[0084] The sample may also include fresh or frozen tissue samples obtained from multiple patients.

[0085] The sample may also include peptides identified using a T cell assay.

[0086] The training dataset can further include the peptide abundance of the set of training peptides present in the sample; data related to the peptide length of the set of training peptides in the sample.

[0087] The training dataset can be generated by comparing a set of training peptide sequences by alignment with a database containing a set of known protein sequences, where the set of training protein sequences is longer than the training peptide sequences and contains the training peptide sequences.

[0088] The training dataset may be generated by performing nucleotide sequencing on a cell line to obtain at least one of exosome, transcriptome, or whole genome sequencing data from the cell line, or may be generated based on nucleotide sequencing having been performed previously, where the sequencing data contains at least one nucleotide sequence containing variations.

[0089] The training dataset may be generated based on obtaining at least one of exosome, transcriptome, or whole genome normal nucleotide sequencing data from a normal tissue sample.

[0090] The training dataset may further contain data related to proteome sequences related to the sample.

[0091] The training dataset may further contain data related to MHC peptidome sequences related to the sample.

[0092] The training dataset may further contain data related to measured values of peptide-MHC binding affinity for at least one of the isolated peptides.

[0093] The training dataset may further contain data related to measured values of peptide-MHC binding stability for at least one of the isolated peptides.

[0094] The training dataset may further contain data related to the transcriptome related to the sample.

[0095] The training dataset may further include data related to the genome associated with the sample.

[0096] The training peptide sequence can have a length within the range of k-mers (where k is 8 or more and 15 or less for MHC class I, or 6 or more and 30 or less for MHC class II).

[0097] The method disclosed herein may also include encoding the peptide sequence using a one-hot encoding scheme.

[0098] The method disclosed herein may also include encoding the training peptide sequence using a left-padded one-hot encoding scheme.

[0099] A method of treating a subject having a tumor, further comprising obtaining a tumor vaccine comprising a selected set of neoantigens, the method including performing the step of claim 1, and administering the tumor vaccine to the subject.

[0100] Also disclosed herein is a method for producing a tumor vaccine, comprising: obtaining from a subject's tumor cells at least one of exosome, transcriptome, or whole-genome tumor nucleotide sequencing data, wherein data representing the peptide sequence of each of a set of neoantigens is obtained using the tumor nucleotide sequencing data, and the peptide sequence of each neoantigen comprises at least one mutation that makes the peptide sequence different from the corresponding wild-type parental peptide sequence; generating a set of numerical likelihoods that each of the neoantigens is presented by one or more MHC alleles on the tumor cell surface of the subject's tumor cells by inputting the peptide sequence of each neoantigen into one or more presentation models, wherein the set of numerical likelihoods is determined based at least on received mass spectrometry data; generating a set of selected neoantigens by selecting a subset of the set of neoantigens based on the set of numerical likelihoods; and producing or having previously produced a tumor vaccine comprising the set of selected neoantigens.

[0101] Also provided herein is a method of obtaining at least one of exosome, transcriptome, or whole-genome tumor nucleotide sequencing data from a target tumor cell, wherein data representing each peptide sequence of a set of neoantigens is obtained using the tumor nucleotide sequencing data, and each peptide sequence of each neoantigen contains at least one mutation that makes the peptide sequence different from the corresponding wild-type parental peptide sequence; a step of generating a set of numerical likelihoods that each of the neoantigens is presented by one or more MHC alleles on the tumor cell surface of the target tumor cell by inputting the peptide sequence of each neoantigen into one or more presentation models, wherein the set of numerical likelihoods is determined based at least on the received mass spectrometry data; a step of generating a selected set of neoantigens by selecting a subset of the set of neoantigens based on the set of numerical likelihoods; and a step of producing or having previously produced a tumor vaccine comprising the selected set of neoantigens. Also provided is a tumor vaccine comprising the selected set of neoantigens selected by performing a method comprising these steps.

[0102] The tumor vaccine may comprise one or more of a nucleotide sequence, a polypeptide sequence, RNA, DNA, a cell, a plasmid, or a vector.

[0103] The tumor vaccine may comprise one or more neoantigens presented on the tumor cell surface.

[0104] The tumor vaccine may comprise one or more neoantigens that are immunogenic in a subject.

[0105] The tumor vaccine may not comprise one or more neoantigens that induce an autoimmune response against normal tissue in a subject.

[0106] The tumor vaccine may comprise an adjuvant.

[0107] The tumor vaccine may comprise an excipient.

[0108] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of being presented on the surface of tumor cells for neoantigens not selected based on the presentation model.

[0109] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of being able to induce a tumor-specific immune response in a subject for neoantigens not selected based on the presentation model.

[0110] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of being able to be presented to naive T cells by professional antigen-presenting cells (APCs) for neoantigens not selected based on the presentation model, optionally, the APC is a dendritic cell (DC).

[0111] The methods disclosed herein may also include selecting neoantigens that have a decreased likelihood of being inhibited by central or peripheral tolerance for neoantigens not selected based on the presentation model.

[0112] The methods disclosed herein may also include selecting neoantigens that have a decreased likelihood of being able to induce an autoimmune response against normal tissue in a subject for neoantigens not selected based on the presentation model.

[0113] Nucleotide sequencing data of exosomes or transcriptomes can be obtained by sequencing tumor tissue.

[0114] The sequencing may be next-generation sequencing (NGS) or any massively parallel processing sequencing approach.

[0115] The set of numerical likelihoods can be further specified by at least one of the following at least MHC allele interaction characteristics. That is, the predicted affinity of the MHC allele for the neoantigen-encoding peptide; the predicted stability of the neoantigen-encoding peptide-MHC complex; the sequence and length of the neoantigen-encoding peptide; the probability of presentation of a neoantigen-encoding peptide having a similar sequence to that of cells from other individuals expressing a particular MHC allele, as evaluated by mass spectrometry proteomics or other means; the expression level of a particular MHC allele of the subject of interest (e.g., measured by RNA-seq or mass spectrometry); the probability of presentation by a particular MHC allele in another distinct individual expressing the particular MHC allele, independent of the overall neoantigen-encoding peptide sequence; the probability of presentation by an MHC allele of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in another distinct subject, independent of the overall neoantigen-encoding peptide sequence.

[0116] The set of numerical likelihoods is further specified by at least one MHC allele non-interaction characteristic among the following. That is, the C-terminal and N-terminal sequences adjacent to the neoantigen-encoding peptide within its source protein sequence; the presence of protease cleavage motifs within the neoantigen-encoding peptide, optionally weighted according to the expression of the corresponding protease in tumor cells (measured by RNA-seq or mass spectrometry); the metabolic turnover rate of the source protein measured in appropriate cell types; the length of the source protein, optionally considering the specific splice variant ("isoform") most highly expressed in tumor cells, predicted from annotation of germline or somatic splicing mutations measured by RNA-seq or proteomic mass spectrometry, or detected in DNA or RNA sequence data; the level of expression of the proteasome, immunoproteasome, thymoproteasome, or other protease in tumor cells (measurable by RNA-seq, proteomic mass spectrometry, or immunohistochemistry); the expression of the source gene of the neoantigen-encoding peptide (e.g., measured by RNA-seq or mass spectrometry); the typical tissue-specific expression of the source gene of the neoantigen-encoding peptide at different stages of the cell cycle; a comprehensive catalog of the properties of the source protein and / or its domains, such as can be found in, for example, uniProt or PDB http: / / www.rcsb.org / pdb / home / home.do; properties that describe the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (e.g., alpha helix versus beta sheet); alternative splicing; the probability of presentation of peptides derived from the source protein of the neoantigen-encoding peptide of interest in other distinct subjects; the probability that the peptide is not detected or overrepresented by mass spectrometry due to technical bias; the expression of various gene modules / pathways measured by RNASeq that provide information about the state of tumor cells, stroma, or tumor-infiltrating lymphocytes (TIL) (not necessarily including the source protein of the peptide); the copy number of the source gene of the neoantigen-encoding peptide in tumor cells;The probability that a peptide binds to TAP, or the measured or predicted binding affinity of a peptide for TAP; the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, immunohistochemistry); the presence or absence of tumor mutations, including but not limited to: driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, NTRK3, and mutations in genes encoding proteins involved in the antigen presentation machinery (e.g., any of B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation depends on components of the antigen presentation machinery that result in loss-of-function mutations in the tumor have a low probability of presentation; the presence or absence of functional germline polymorphisms, including but not limited to: polymorphisms in genes encoding proteins involved in the antigen presentation machinery (e.g., any of B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or genes encoding components of the proteasome or immunoproteasome); tumor type (e.g., NSCLC, melanoma); clinical tumor subtype (e.g., squamous cell lung cancer vs. non-squamous cell); smoking history;Typical expression of the source gene of the peptide in related tumor types or clinical subtypes that are sometimes stratified by driver mutations;

[0117] At least one mutation may be a frameshift or non-frameshift insertion deletion, missense or nonsense substitution, splice site change, genomic rearrangement or gene fusion, or any genomic or expression change that results in a neo-ORF.

[0118] The tumor cells can be selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0119] The methods disclosed herein may also include obtaining a tumor vaccine comprising a selected set of neoantigens or a subset thereof, and optionally further comprising administering the tumor vaccine to a subject.

[0120] When at least one of the neoantigens within the selected set of neoantigens is in polypeptide form, it may include at least one of the following: binding affinity to MHC with an IC50 value of less than 1000 nM, 8 to 15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids in length for MHC class I polypeptides, 6 to 30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids in length for MHC class II polypeptides, the presence of a sequence motif within or near the polypeptide in the parent protein sequence that promotes proteasome cleavage, and the presence of a sequence motif that promotes TAP transport. For MHC class II, the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.

[0121] Also disclosed herein is a method for generating a model for identifying one or more neoantigens likely to be presented on the tumor cell surface of tumor cells, the method comprising receiving mass spectrometry data comprising data related to a plurality of isolated peptides eluted from major histocompatibility complex (MHC) derived from a plurality of samples, and obtaining a training dataset by at least identifying a set of training peptide sequences present in the samples and one or more MHCs associated with each training peptide sequence, and training a set of numerical parameters of a presentation model using the training dataset comprising the training peptide sequences, wherein the presentation model gives a plurality of numerical likelihoods of a peptide sequence derived from a tumor cell being presented by one or more MHC alleles on the tumor cell surface.

[0122] The presentation model can represent the dependence between the presence of a particular amino acid at a particular position in a peptide sequence and the likelihood of presentation of a peptide sequence having the particular amino acid at the particular position by one of the MHC alleles on the tumor cell.

[0123] The sample may also include a cell line engineered to express a single MHC class I or class II allele.

[0124] The sample may also include a cell line engineered to express multiple MHC class I or class II alleles.

[0125] The sample may also include human cell lines obtained from or derived from multiple patients.

[0126] The sample may also include fresh or frozen tumor samples obtained from multiple patients.

[0127] The sample may also include peptides identified using a T cell assay.

[0128] The training dataset can further include data related to the peptide abundance of the set of training peptides present in the sample; data related to the peptide length of the set of training peptides in the sample.

[0129] The method disclosed herein may also include obtaining a set of training protein sequences that are longer than the training peptide sequences and include the training peptide sequences, based on the training peptide sequences, by comparing the set of training peptide sequences by alignment with a database containing a set of known protein sequences.

[0130] The method disclosed herein may also include performing mass spectrometry on a cell line or mass spectrometry having been performed heretofore to obtain at least one of exosome, transcriptome, or genomic nucleotide sequencing data from the cell line, wherein the nucleotide sequencing data includes at least one protein sequence containing a mutation.

[0131] The method disclosed herein may also include encoding the training peptide sequences using a one-hot encoding scheme.

[0132] The method disclosed herein may also include obtaining at least one of exosome, transcriptome, and genomic normal nucleotide sequencing data from a normal tissue sample, and training a set of parameters of a presentation model using the normal nucleotide sequencing data.

[0133] The training dataset may further include data related to the proteome sequence related to the sample.

[0134] The training dataset may further include data related to the MHC peptidome sequence related to the sample.

[0135] The training dataset may further include data related to measurements of peptide-MHC binding affinity for at least one of the isolated peptides.

[0136] The training dataset may further include data related to measurements of peptide-MHC binding stability for at least one of the isolated peptides.

[0137] The training dataset may further include data related to the transcriptome related to the sample.

[0138] The training dataset may further include data related to the genome related to the sample.

[0139] The method disclosed herein may also include performing logistic regression on a set of parameters.

[0140] The training peptide sequences can have lengths within the range of k-mers (where k is 8 or more and 15 or less for MHC class I, or 6 or more and 30 or less for MHC class II).

[0141] The method disclosed herein may also include encoding the training peptide sequences using a left-padded one-hot encoding scheme.

[0142] The method disclosed herein may also include using a deep learning algorithm to determine values for a set of parameters.

[0143] A method for identifying one or more neoantigens likely to be presented on the surface of tumor cells in this specification, comprising the steps of receiving mass spectrometry data including data related to a plurality of isolated peptides eluted from a major histocompatibility complex (MHC) derived from a plurality of fresh or frozen tumor samples; obtaining a training dataset by at least identifying a set of training peptide sequences present in the tumor sample and presented on one or more MHC alleles associated with each training peptide sequence; obtaining a set of training protein sequences based on the training peptide sequences; and training a set of numerical parameters of a presentation model using the training protein sequences and the training peptide sequences, wherein the presentation model gives a plurality of numerical likelihoods of peptide sequences derived from tumor cells being presented by one or more MHC alleles on the surface of tumor cells, is disclosed.

[0144] The presentation model can represent the dependence between the presence of a pair of a specific one of the MHC alleles and a specific amino acid at a specific position of the peptide sequence, and the likelihood of presentation on the surface of tumor cells of such a peptide sequence containing the specific amino acid at the specific position by the specific one of the MHC alleles.

[0145] The method disclosed in this specification may also include selecting a subset of neoantigens, wherein the subset of neoantigens is selected because each has an increased likelihood of being presented on the cell surface of the tumor relative to one or more distinct tumor neoantigens.

[0146] The method disclosed in this specification may also include selecting a subset of neoantigens, wherein the subset of neoantigens is selected because each has an increased likelihood of being able to induce a tumor-specific immune response in a subject relative to one or more distinct tumor neoantigens.

[0147] The methods disclosed herein may also include selecting a subset of neoantigens, wherein the subset of neoantigens is selected because each has an increased likelihood of being presented by professional antigen-presenting cells (APCs) to naive T cells against one or more distinct tumor neoantigens, and optionally, the APC is a dendritic cell (DC).

[0148] The methods disclosed herein may also include selecting a subset of neoantigens, wherein the subset of neoantigens is selected because each has a decreased likelihood of being inhibited by central or peripheral tolerance against one or more distinct tumor neoantigens.

[0149] The methods disclosed herein may also include selecting a subset of neoantigens, wherein the subset of neoantigens is selected because each has a decreased likelihood of inducing an autoimmune response against normal tissue in a subject against one or more distinct tumor neoantigens.

[0150] The methods disclosed herein may also include selecting a subset of neoantigens, wherein the subset of neoantigens is selected because each has a decreased likelihood of being differentially post-translationally modified in tumor cells relative to APCs, and optionally, the APC is a dendritic cell (DC).

[0151] In carrying out the methods described herein, unless otherwise indicated, conventional methods of protein chemistry, biochemistry, recombinant DNA technology and pharmacology within the skill in the art are used. Such techniques are well described in the literature. See, for example, T.E. Creighton, Proteins: Structures and Molecular Properties (W.H. Freeman and Company, 1993); A.L. Lehninger, Biochemistry (Worth Publishers, Inc., current addition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (2nd Edition, 1989); Methods In Enzymology (S. Colowick and N. Kaplan eds., Academic Press, Inc.); Remington’s Pharmaceutical Sciences, 18th Edition (Easton, Pennsylvania: Mack Publishing Company, 1990); Carey and Sundberg Advanced Organic Chemistry 3rd Ed. (Plenum Press) Vols A and B (1992).

[0152] The set of presentation likelihoods can also be generated based on the source genes of the set of neoantigens.

[0153] The set of presentation likelihoods can also be generated based on the source genes and source tissue types of the set of neoantigens.

[0154] The methods disclosed herein may include identifying a subset of patients suitable for treatment with a neoantigen vaccine, the steps including: Obtaining, for each patient, at least one of exosome, transcriptome, or whole-genome tumor nucleotide sequencing data from the patient's tumor cells, wherein the tumor nucleotide sequencing data is used to obtain each peptide sequence of a set of neoantigens, and each peptide sequence of the neoantigens contains at least one change that makes it different from the corresponding wild-type parental peptide sequence; said obtaining; Generating, for each patient, a set of numerical presentation likelihoods for the set of neoantigens for the patient by inputting each peptide sequence of the set of neoantigens into one or more presentation models, wherein the set of presentation likelihoods represents the likelihood that each of the set of neoantigens is presented by one or more MHC alleles on the surface of the patient's tumor cells, and the set of presentation likelihoods is determined based at least on the received mass spectrometry data; said generating; Identifying, for each patient, a therapeutic subset of neoantigens from the set of neoantigens of the patient, wherein the therapeutic subset corresponds to a predetermined number of neoantigens having the highest presentation likelihood within the set of presentation likelihoods generated for that patient; said identifying; Selecting a subset of patients suitable for treatment with a neoantigen vaccine, wherein the selected subset of patients meets inclusion criteria based on the set of neoantigens obtained for each patient within the selected subset or based on the tumor nucleotide sequencing data; said selecting.

[0155] The method disclosed herein may include treating each patient within the selected subset of patients with the corresponding neoantigen vaccine, and the neoantigen vaccine for the patient includes the therapeutic subset identified by the set of presentation likelihoods for the patient.

[0156] The methods disclosed herein may include selecting a subset of patients having a tumor mutation burden (TMB) higher than a minimum threshold, where the TMB of a patient indicates the number of neoantigens within the set of neoantigens associated with that patient.

[0157] The methods disclosed herein may include determining, for each patient, a utility score indicative of a measure of the estimated number of neoantigens presented from a treatment subset of the patient; and selecting a subset of patients having a utility score higher than a minimum threshold.

[0158] Presentation of neoantigens can be modeled as a Bernoulli random variable, the utility score can represent the expected number of presented neoantigens in a treatment subset for a patient, and the utility score can be given by the sum of the presentation likelihoods for each neoantigen in the treatment subset of the patient.

[0159] Presentation of neoantigens can also be modeled as a Poisson binomial random variable, and the utility score can be the probability that the number of presented neoantigens in a treatment subset for a patient exceeds a minimum threshold.

[0160] III. Identification of Tumor-Specific Mutations in Neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells), in particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject having cancer, but not in normal tissue from the subject.

[0161] Genetic mutations in tumors can be considered useful for tumor immunological targeting when they exclusively result in changes in the amino acid sequence of proteins in the tumor. Useful mutations include the following: (1) Nonsynonymous mutations that result in different amino acids in the protein; (2) Read-through mutations where the stop codon is modified or deleted, resulting in the translation of a longer protein with a novel tumor-specific sequence at the C-terminus; (3) Splice site mutations that result in the inclusion of an intron in the mature mRNA and thus a unique tumor-specific protein sequence; (4) Chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) Frameshift mutations or deletions that result in a new open reading frame with a novel tumor-specific protein sequence. Mutations can also include one or more of non-frameshift insertions / deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in a de novo ORF.

[0162] For example, peptides or mutant polypeptides with mutations resulting from splice site, frameshift, read-through, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.

[0163] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.

[0164] A variety of methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advancements in this field provide accurate, easy, and inexpensive large-scale SNP genotyping. For example, several techniques have been described, including various DNA "chip" technologies such as dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, TaqMan systems, and Affymetrix SNP chips. These methods typically utilize amplification of the target gene region by PCR. Still other methods are based on the generation of small signal molecules by invasive cleavage and subsequent mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.

[0165] PCR-based detection means can simultaneously include multiplex amplification of a number of markers. For example, it is well known in the art to select PCR primers to generate PCR products that do not overlap in size and can be analyzed simultaneously. Alternatively, it is possible to amplify different markers with primers that are differentially labeled and thus can be differentially detected. Of course, hybridization-based detection means enable differential detection of multiple PCR products in a sample. Other techniques that enable multiplex analysis of multiple markers are known in the art.

[0166] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA. For example, single nucleotide polymorphisms can be detected by using specialized exonuclease-resistant nucleotides, as disclosed, for example, in Mundy, C.R. (U.S. Patent No. 4,656,127). According to this method, a primer complementary to the allelic sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a particular animal or human. If the polymorphic site on the target molecule contains a nucleotide that is complementary to a particular exonuclease-resistant nucleotide derivative present, the derivative is incorporated onto the end of the hybridized primer. Because of such incorporation, the primer becomes resistant to exonuclease, thereby allowing its detection. Since the identity of the exonuclease-resistant derivative in the sample is known, the finding that the primer has become resistant to exonuclease reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage of not requiring the determination of large amounts of exogenous sequence data.

[0167] To determine the identity of the nucleotides at the polymorphic site, solution-based methods can be used (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087). As in the method of Mundy of U.S. Patent No. 4,656,127, a primer that is complementary to the allelic sequence immediately 3' of the polymorphic site is used. This method determines the identity of the nucleotide at that site using a labeled dideoxynucleotide derivative that will be incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site. An alternative method, known as Genetic Bit Analysis or GBA, is described by Goelet, P. et al. (PCT Application No. 92 / 15712). The method of Goelet, P. et al. uses a mixture of a labeled terminator and a primer that is complementary to the 3' sequence of the polymorphic site. The method of Goelet, P. et al. uses a mixture of a labeled terminator and a primer that is complementary to the 3' sequence of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which the primer or the target molecule is immobilized on a solid phase.

[0168] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J.S. et al., Nucl. Acids Res. 17:7779-7784 (1989); Sokolov, B.P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M.N. et al., Proc. Natl. Acad. Sci. (U.S.A.) 88:1143-1147 (1991); Prezant, T.R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, since the signal is proportional to the number of incorporated deoxynucleotides, polymorphisms occurring in runs of the same nucleotide can result in signals proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).

[0169] Numerous initiatives directly obtain sequence information in parallel from millions of individual molecules of DNA or RNA. Sequencing technologies by real-time single molecule synthesis rely on the detection of fluorescent nucleotides when incorporated into the nascent DNA strand that is complementary to the template being sequenced. In one method, oligonucleotides 30 to 50 bases in length are covalently attached at the 5' end to a glass coverslip. These attached strands serve two functions. First, they act as capture sites for target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis of sequence readout. The capture primer functions as a fixed position site for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and cleavage of the dye. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. When the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing technologies by synthesis also exist.

[0170] Any suitable sequencing platform by synthesis can be used to identify mutations. As noted above, four major sequencing platforms by synthesis are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing platforms by synthesis are also described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the nucleic acid molecules to be sequenced are attached to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' end and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. The capture sequence (also referred to as a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to the support that can act double as a universal primer.

[0171] As an alternative to the capture sequence, members of a coupling pair (e.g., antibody / antigen, receptor / ligand, or an avidin-biotin pair such as described in, for example, U.S. Patent Application No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.

[0172] Following capture, the array can be analyzed by single molecule detection / sequencing, such as that described in the Examples and U.S. Patent No. 7,283,337, including, for example, sequencing by template-dependent synthesis. In sequencing by synthesis, molecules bound to a surface are exposed to a large number of labeled nucleotide triphosphates in the presence of polymerase. The sequence of the template is determined by the order of the labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time and can be done in a step-and-repeat mode. For real-time analysis, different optical labels can be incorporated for each nucleotide and multiple lasers can be utilized for excitation of the incorporated nucleotides.

[0173] Sequencing can also include other massively parallel sequencing, or next generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, Thermo PGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel sequencing technologies, and future generations of these technologies, can be used.

[0174] Any cell type or tissue can be utilized to obtain a nucleic acid sample for use in the methods described herein. For example, a DNA or RNA sample can be obtained from a tumor or a body fluid, such as blood or saliva obtained by known techniques (e.g., venipuncture). Alternatively, a nucleic acid test can be performed on a dried sample (e.g., hair or skin). In addition, a sample can be obtained from a tumor for sequencing, and another sample can be obtained from normal tissue for sequencing if the normal tissue is of the same tissue type as the tumor. A sample can be obtained from a tumor for sequencing, and another sample can be obtained from normal tissue for sequencing if the normal sample is of a tissue type distinct from the tumor.

[0175] The tumor can include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0176] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. The peptides can be acid eluted from tumor cells or from HLA molecules immunoprecipitated from tumors and then identified using mass spectrometry.

[0177] IV. Neoantigens Neoantigens can include nucleotides or polynucleotides. For example, a neoantigen can be an RNA sequence encoding a polypeptide sequence. Neoantigens useful in a vaccine can thus include a nucleotide sequence or a polypeptide sequence.

[0178] Isolated peptides containing tumor-specific mutations identified by the methods disclosed herein, peptides containing known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein are disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the nucleotide sequences (e.g., DNA or RNA) encoding the polypeptide sequences to which the neoantigens are related are included.

[0179] One or more polypeptides encoded by a neoantigen nucleotide sequence can include at least one of the following: binding affinity to MHC with an IC50 value of less than 1000 nM, for MHC class I peptides, 8 to 15 amino acids in length, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids in length, the presence of sequence motifs within or near the peptide that facilitate proteasome cleavage, and the presence of sequence motifs that facilitate TAP transport. For MHC class II peptides, 6 to 30 amino acids in length, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids in length, the presence of sequence motifs within or near the peptide that facilitate cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.

[0180] One or more neoantigens can be present on the surface of a tumor.

[0181] One or more neoantigens can be immunogenic in a subject having a tumor, e.g., can elicit a T cell response or a B cell response in the subject.

[0182] One or more neoantigens that induce an autoimmune response in a subject can be excluded from consideration in the context of vaccine generation for a subject having a tumor.

[0183] The size of at least one neoantigenic peptide molecule can include, but is not limited to, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derived from these ranges. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.

[0184] Neoantigenic peptides and polypeptides can be 15 residues or less in length for MHC class I, usually consisting of between about 8 and about 11 residues, and can in particular be 9 or 10 residues; for MHC class II, they can be 6 to 30 residues.

[0185] If desired, longer peptides can be designed in several ways. In one example, if the likelihood or the knowledge of peptide presentation on an HLA allele is predicted, the longer peptides can consist of (1) individual presented peptides having an extension of 2-5 amino acids towards the N-terminus and C-terminus of each corresponding gene product; (2) any of some or all of the chains of the presented peptides, each having an extended sequence. In another example, if sequencing reveals the presence of long (longer than 10 residues) neoepitope sequences in a tumor (e.g., due to frameshift, read-through, or inclusion of an intron resulting in a novel peptide sequence), the longer peptides can consist of (3) the entire stretch of novel tumor-specific amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides presented by the strongest HLA. In any of the examples, the use of longer peptides allows for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.

[0186] Neoantigenic peptides and polypeptides can be presented on HLA proteins. In some embodiments, the neoantigenic peptides and polypeptides are presented on HLA proteins with a stronger affinity than wild-type peptides. In some embodiments, the neoantigenic peptide or polypeptide can have an IC50 of at least less than 5000 nM, at least less than 1000 nM, at least less than 500 nM, at least less than 250 nM, at least less than 200 nM, at least less than 150 nM, at least less than 100 nM, at least less than 50 nM, or less than that.

[0187] In some embodiments, the neoantigenic peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.

[0188] Also provided is a composition comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides are derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which the neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some aspects, the tumor-specific mutations are driver mutations for a particular cancer type.

[0189] Neoantigenic peptides and polypeptides having desirable activities or properties can be modified to confer certain desirable attributes, such as improved pharmacological characteristics, while increasing or at least substantially retaining all of the biological activities of unmodified peptides that bind to desirable MHC molecules and activate appropriate T cells. By way of example, neoantigenic peptides and polypeptides can be further subjected to various modifications, such as either conservative or non-conservative substitutions, which can provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions mean replacing an amino acid residue with another that is biologically and / or chemically similar, e.g., replacing one hydrophobic residue with another, or one polar residue with another. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effect of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (N.Y., Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2d Ed. (1984).

[0190] Modification of peptides and polypeptides with various amino acid mimics or unnatural amino acids can be particularly useful for increasing the stability of peptides and polypeptides in vivo. Stability can be assayed in many ways. By way of example, proteases, as well as various biological media such as human plasma and serum, have been used to test stability. See, for example, Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). The half-life of a peptide can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally is as follows. Pooled human serum (type AB, non-heat inactivated) is defatted by centrifugation prior to use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The turbid reaction samples are cooled (4° C.) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse phase HPLC using stability-specific chromatography conditions.

[0191] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. By way of example, the ability of a peptide to induce CTL activity can be enhanced by ligation to a sequence containing at least one epitope capable of inducing a T helper cell response. Immunogenic peptide / T helper conjugates can be linked by a spacer molecule. The spacer typically consists of relatively small neutral molecules such as amino acids or amino acid mimics that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that any optionally present spacer need not be composed of the same residues and can thus be a hetero-oligomer or a homo-oligomer. When present, the spacer will usually be at least 1 or 2 residues, more usually 3 - 6 residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.

[0192] Neoantigenic peptides can be linked to T helper peptides either directly or via a spacer at either the amino or carboxy terminus of the peptide. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include 830 - 843 of tetanus toxoid, 307 - 319 of influenza, around 382 - 398 and 378 - 389 of malaria sporozoite.

[0193] A protein or peptide can be made by any technique known to those of skill in the art, including expression of a protein, polypeptide, or peptide through standard molecular biological techniques, isolation of a protein or peptide from a natural source, or chemical synthesis of a protein or peptide. Nucleotide as well as protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computer-processed databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located on the website of the National Institutes of Health. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.

[0194] In a further aspect, the neoantigen comprises a nucleic acid (e.g., a polynucleotide) encoding the neoantigenic peptide or a portion thereof. The polynucleotide can be, for example, single-stranded and / or double-stranded polynucleotides, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), polynucleotides having a phosphorothioate backbone, etc., either in their native form or a stabilized form, or a combination thereof, and may or may not contain introns. A further aspect provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be ligated to appropriate transcriptional and translational regulatory nucleotide sequences recognized by the desired host, but such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Protocols can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.

[0195] IV. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can elicit a specific immune response, e.g., a tumor-specific immune response. The vaccine composition typically comprises a plurality of neoantigens selected, for example, using the methods described herein. The vaccine composition can also be referred to as a vaccine.

[0196] The vaccine can contain 1 to 30 types of peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine can contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine can contain 1 to 30 types of neoantigen sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 types or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.

[0197] In one embodiment, the different peptides and / or polypeptides, or nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides can bind to different MHC molecules such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, one vaccine composition comprises a coding sequence for a peptide and / or polypeptide that can bind to the most frequently occurring MHC class I molecule and / or MHC class II molecule. Thus, the vaccine composition can comprise different fragments that can bind to at least 2 preferred, at least 3 preferred, or at least 4 preferred MHC class I molecules and / or MHC class II molecules.

[0198] The vaccine composition can generate a specific cytotoxic T cell response and / or a specific helper T cell response.

[0199] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are shown below in the present specification. The composition can bind to a carrier such as a protein, for example, or an antigen-presenting cell such as a dendritic cell (DC) that can present a peptide to T cells.

[0200] An adjuvant is any substance whose mixing into the vaccine composition increases or otherwise modifies the immune response to neoantigens. A carrier can be a scaffold structure to which a neoantigen can bind, such as a polypeptide or a polysaccharide. Optionally, the adjuvant is conjugated covalently or non-covalently.

[0201] The ability of an adjuvant to increase the immune response to an antigen is typically manifested by a significant or substantial increase in an immune-mediated reaction or a reduction in disease symptoms. For example, an increase in humoral immunity is typically manifested by a significant increase in the titer of antibodies generated against the antigen, and an increase in T cell activity is typically manifested in an increase in cell proliferation, or cytotoxicity, or cytokine secretion. An adjuvant can also change the immune response, for example, by changing a primarily humoral or Th response to a primarily cellular or Th response.

[0202] Suitable adjuvants include, but are not limited to, 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, MontanideISA206, Montanide ISA 50V, Montanide ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila’s QS21 stimulon derived from saponin (Aquila Biotech, Worcester, Mass., USA), bacterial extracts and synthetic bacterial cell wall mimetics, and other proprietary adjuvants such as Ribi’s Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are useful. Some immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison A C; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Some cytokines are directly linked to effects on the migration of dendritic cells to lymphoid tissues (e.g., TNF-α), acceleration of the maturation of dendritic cells into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Patent No. 5,849,589, which is specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich D I, et al., J ImmunotherEmphasis Tumor Immunol. 1996(6):414-418).

[0203] CpG immunostimulatory oligonucleotides have also been reported to enhance the effect of adjuvants in vaccine settings. Other TLR-binding molecules such as RNA that bind to TLR 7, TLR 8, and / or TLR 9 may also be used.

[0204] Other examples of useful adjuvants include chemically modified CpGs (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies such as cyclophosphamide, sunitinib, bevacizumab, celecoxib, NCX-4016, sildenafil, tadalafil, vardenafil, sorafenib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175 that can act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony-stimulating factors such as granulocyte macrophage colony-stimulating factor (GM-CSF, sargramostim).

[0205] Vaccine compositions can contain more than one different adjuvant. Further, therapeutic compositions can contain any adjuvant substance, including any of the foregoing or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.

[0206] The carrier (or excipient) can exist independently from the adjuvant. The function of the carrier can be, for example, to increase the activity or immunogenicity, to confer stability, to increase the biological activity, or to increase the serum half-life, particularly by increasing the molecular weight of the variant. Further, the carrier can help to present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, keyhole limpet hemocyanin, serum proteins such as transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, immunoglobulins, or hormones such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is human-acceptable and safe. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be dextran, such as sepharose.

[0207] Cytotoxic T lymphocytes (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. The MHC molecules themselves are located on the cell surface of antigen-presenting cells. Thus, activation of CTLs is possible when a trimolecular complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when the peptide is used for activation of CTLs, but additionally when APCs having the respective MHC molecules are added, it can enhance the immune response. Thus, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.

[0208] Neoantigens can also be included in virus vector-based vaccine platforms, including but not limited to vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses of the second, third, or hybrid second / third generations, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61, Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18, Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690, Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880). Depending on the packaging capacity of the virus vector-based vaccine platforms described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The arrays may be adjacent non-mutated arrays, may be separated by linkers, or may be preceded by one or more arrays targeting intracellular compartments (see, for example, Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8, Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41, Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express neoantigens, thereby eliciting a host immune (e.g., CTL) response to the peptides. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guerin). The BCG vector is described in Stover et al. (Nature 351:456-460 (1991)). A variety of other vaccine vectors useful for the therapeutic administration or immunization with neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.

[0209] IV.A. Further Considerations in Vaccine Design and Manufacture IV.A.1. Determination of a Set of Peptides That Covers All Tumor Subclones Truncal peptides, which mean those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides predicted to be presented with high probability and be immunogenic, or if the number of truncal peptides predicted to be presented with high probability and be immunogenic is so small that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 。

[0210] IV.A.2. Neoantigen Prioritization After applying all of the above neoantigen filters, there may still be more candidate neoantigens available for vaccine inclusion than the vaccine technology can handle. Additionally, there may be uncertainties remaining about various aspects of neoantigen analysis, and there may be trade-offs between the various properties of candidate vaccine neoantigens. Thus, instead of pre-determined filters at each stage of the selection process, an integrated multidimensional model can be considered that places candidate neoantigens in a space with at least the following axes and optimizes the selection using an integrated approach. 1. Risk of autoimmunity or tolerance (germline risk) (lower risk of autoimmunity is typically preferred) 2. Probability of sequencing artifacts (lower probability of artifacts is typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferred) 5. Gene expression (higher expression is typically preferred) 6. Coverage of HLA genes (A larger number of HLA molecules involved in the presentation of the set of neoantigens may reduce the probability that the tumor evades immune attack through downregulation or mutation of HLA molecules.) 7. Coverage of HLA classes (Covering both HLA-I and HLA-II may increase the probability of treatment response and decrease the probability of tumor immune escape.)

[0211] V. Treatment and manufacturing methods Also provided is a method of inducing a tumor-specific immune response in a subject, vaccinating against the tumor, and treating and / or alleviating the symptoms of cancer in the subject by administering to the subject one or more neoantigens such as a plurality of neoantigens identified using the methods disclosed herein.

[0212] In some embodiments, the subject is diagnosed with cancer or at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desired. The tumor can be any solid tumor, such as a tumor of the breast, ovary, prostate, lung, kidney, stomach, colon, testis, head and neck, pancreas, brain, melanoma, and tumors of other tissues and organs, as well as hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.

[0213] The neoantigen can be administered in an amount sufficient to induce a CTL response.

[0214] The neoantigen can be administered alone or in combination with other therapeutic substances. The therapeutic substance can be, for example, a chemotherapeutic agent, radiation, or immunotherapy. Any suitable therapeutic treatment for a particular cancer can be administered.

[0215] In addition, an anti-immunosuppressive / immunostimulatory substance such as a checkpoint inhibitor can be further administered to the subject. For example, an anti-CTLA antibody or anti-PD-1 or anti-PD-L1 can be further administered to the subject. Blockade of CTLA-4 or PD-L1 by an antibody can enhance the immune response against cancer cells in a patient. In particular, CTLA-4 blockade has been shown to be effective when a vaccination protocol is employed.

[0216] The optimal amount of each neoantigen to be included in the vaccine composition, and the optimal dosing regimen, can be determined. For example, the neoantigen or its variant can be prepared for intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal (i.p.) injection, intramuscular (i.m.) injection. Methods of injection include s.c., i.d., i.p., i.m., and i.v. Methods of DNA or RNA injection include i.d., i.m., s.c., i.p., and i.v. Other methods of administration of the vaccine composition are known to those skilled in the art.

[0217] The vaccine can be engineered such that the selection, number, and / or amount of neoantigens present in the composition are specific to the tissue, cancer, and / or patient. As an example, the stringent selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. The selection can depend on the cancer-specific type, disease state, earlier treatment regimen, the patient's immune status, and of course, the patient's HLA haplotype. Further, the vaccine can contain individualized components according to the specific needs of a particular patient. Examples include changing the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjustments for secondary treatment after the first round or scheme of treatment.

[0218] For a composition to be used as a vaccine for cancer, neoantigens having similar normal self - peptides that are highly expressed in normal tissues can be avoided or present in small amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a certain neoantigen in large amounts, each pharmaceutical composition for the treatment of this cancer can be present in large amounts and / or can include this specific neoantigen or more than one neoantigen specific to the pathway of this neoantigen.

[0219] A composition containing neoantigens can be administered to an individual already suffering from cancer. In a therapeutic application, the composition is administered to the patient in an amount sufficient to elicit an effective CTL response against tumor antigens and to cure or at least partially arrest the symptoms and / or complications. The amount considered appropriate to achieve this is defined as the "therapeutically effective dose". The effective amount for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general state of health, as well as the judgment of the prescribing physician. It should be borne in mind that the composition can generally be used in severe disease states, i.e., life - threatening or potentially life - threatening situations, especially when the cancer has metastasized. In such cases, taking into account the minimization of foreign substances and the relatively non - toxic nature of the neoantigens, it is possible to administer a substantially excessive amount of these compositions, and the treating physician can feel it is desirable.

[0220] For therapeutic use, administration can be initiated at the time of tumor detection or surgical removal. This is followed by boost doses until at least the symptoms are substantially reduced and then for a certain period of time.

[0221] Pharmaceutical compositions (e.g., vaccine compositions) for therapeutic treatment are intended for parenteral, topical, nasal, oral, or local administration. The pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against the tumor. Parenteral administration compositions containing a solution of neoantigens are disclosed herein, and the vaccine compositions are dissolved or suspended in an acceptable carrier, for example, an aqueous carrier. Various aqueous carriers, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, etc. can be used. These compositions can be sterilized by conventional well-known sterilization techniques or can be sterile filtered. The resulting aqueous solution can be packaged as is for use or can be lyophilized, and the lyophilized preparation is combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances necessary to approximate physiological conditions, such as pH adjusters and buffers, tonicity agents, wetting agents, etc., for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, etc.

[0222] Neoantigens can also be administered via liposomes that target them to specific cell tissues such as lymphoid tissues. Liposomes are also useful for increasing the half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, etc. In these preparations, the neoantigens to be delivered are incorporated as part of the liposomes, alone or together with molecules that bind to dominant receptors among, for example, lymphoid cells, such as monoclonal antibodies that bind to the CD45 antigen, or other therapeutic or immunogenic compositions. Thus, liposomes filled with the desired neoantigens can be directed to the site of lymphoid cells, where the liposomes then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including phospholipids having neutral and negative charges, and sterols such as cholesterol. The choice of lipids is generally guided by considerations such as, for example, liposome size, acid lability, and stability of the liposomes in the bloodstream. For example, various methods are available for preparing liposomes, as described in Szoka et al., Ann.Rev.Biophys.Bioeng.9;467 (1980), U.S. Patent Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.

[0223] For targeting to immune cells, the ligand to be incorporated into the liposomes can include, for example, an antibody or a fragment thereof specific for the cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at a dose that varies, inter alia, according to the mode of administration, the peptide being delivered, and the stage of the disease being treated.

[0224] For therapeutic or immunization purposes, the peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient. Numerous methods are conveniently used to deliver nucleic acids to a patient. By way of example, nucleic acids can be delivered directly as "naked DNA". This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), as well as in U.S. Patent Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery as described, for example, in U.S. Patent No. 5,204,253. Particles consisting of just DNA can be administered. Alternatively, DNA can be attached to particles such as gold particles. Approaches for delivering nucleic acid sequences can include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.

[0225] Nucleic acids can also be delivered by complexing them with cationic compounds such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372WOAWO 96 / 18372;9324640WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833;9106309WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).

[0226] Neoantigens can also be included in virus vector-based vaccine platforms, including but not limited to vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses of the second, third, or hybrid second / third generations, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61, Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18, Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690, Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880). Depending on the packaging capacity of the virus vector-based vaccine platforms described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The arrays may be adjacent non-mutated arrays, may be separated by linkers, or may be preceded by one or more arrays targeting intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8, Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41, Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20(13):3401-10). Upon introduction into the host, the infected cells express neoantigens, thereby eliciting a host immune (e.g., CTL) response to the peptides. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guerin). BCG vectors are described in Stover et al. (Nature 351:456-460 (1991)). A variety of other vaccine vectors useful for the therapeutic administration or immunization with neoantigens, such as typhoid vectors, will be apparent to those skilled in the art from the description herein.

[0227] The means for administering nucleic acids uses a mini-gene construct encoding one or more epitopes. To create a DNA sequence (mini-gene) encoding a selected CTL epitope for expression in human cells, the amino acid sequence of the epitope is reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are placed directly adjacent to each other to create a continuous polypeptide sequence. Additional elements can be incorporated during mini-gene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the mini-gene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of the CTL epitope can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences proximal to the CTL epitope. The mini-gene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the mini-gene. Overlapping oligonucleotides (30 - 100 bases in length) are synthesized, phosphorylated, purified, and annealed using well-known techniques under appropriate conditions. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic mini-gene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.

[0228] Purified plasmid DNA can be prepared for injection using various formulations. The simplest of these is the reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) can also complex with purified plasmid DNA to affect variables such as stability, intramuscular dispersion, or transport to specific organs or cell types.

[0229] Also disclosed herein is a method for manufacturing a tumor vaccine, which includes performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising a plurality of neoantigens or a subset of a plurality of neoantigens.

[0230] The neoantigens disclosed herein can be manufactured using methods known in the art. For example, a method for producing a neoantigen or a vector (e.g., a vector comprising at least one sequence encoding one or more neoantigens) disclosed herein can include culturing a host cell under conditions suitable for expressing the neoantigen or vector, wherein the host cell comprises at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods can include chromatographic techniques, electrophoretic techniques, immunological techniques, precipitation techniques, dialysis techniques, filtration techniques, concentration techniques, and chromatofocusing techniques.

[0231] The host cell can include Chinese hamster ovary (CHO) cells, NS0 cells, yeast, or HEK293 cells. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to at least one nucleic acid sequence encoding a neoantigen or vector. In certain embodiments, the isolated polynucleotide can be cDNA.

[0232] VI. Identification of Neoantigens VI.A. Identification of Neoantigen Candidates Research methods for NGS analysis of tumor and normal exomes and transcriptomes are described and applied in the space of neoantigen identification. 6,14,15. The following examples consider certain optimizations for greater sensitivity and specificity in the identification of neoantigens in a clinical setting. These optimizations can be grouped into two areas, those related to laboratory processes and those related to NGS data analysis.

[0233] VI.A.1. Optimization of Laboratory Processes The process improvements presented herein expand the concept developed for the evaluation of reliable cancer driver genes in a targeted cancer panel 16 to address the challenges in high-precision neoantigen discovery from low tumor content and small volume clinical specimens by expanding it to the whole exome and whole transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) native mean coverage across the tumor exome to detect mutations present at low variant allele frequencies due to either low tumor content or subclonal status. 2. As an example, less than 5% of bases covered at less than 100x, so that the chance of missing potential neoantigens is minimized, a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for regions that are not sufficiently covered 3. Targeting uniform coverage across the normal exome such that less than 5% of bases are covered at less than 20x, so that potential neoantigens are least likely to remain unclassified with respect to somatic / germline status (and thus not usable as TSNA). 4. To minimize the total amount of sequencing required, the sequence capture probes are designed for only the coding regions of genes, as non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing18 。 b. Exclusion of genes predicted to generate few or no candidate neoantigens due to factors such as insufficient expression, suboptimal digestion by the proteasome, or aberrant sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant (“isoform”) expression, and fusion detection. RNA from FFPE samples is extracted using probe-based enrichment with the same or similar probes used to capture exomes in DNA. 19 using.

[0234] VI.A.2. Optimization of NGS Data Analysis Improvements in analysis methods address the suboptimal sensitivity and specificity of common research variant calling approaches, specifically considering customizations relevant for the identification of neoantigens in a clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment (as it contains multiple MHC region assemblies that better reflect population polymorphisms, in contrast to previous genome releases). 2. Overcoming the limitations of single variant caller 20 by merging results from various programs 5 from. a. Single nucleotide variants and indels are detected from tumor DNA, tumor RNA, and normal DNA using a series of tools including: Strelka 21 and Mutec t22 such as programs based on the comparison of tumor and normal DNA; and, particularly advantageous in low purity samples 23 , programs that incorporate tumor DNA, tumor RNA, and normal DNA such as UNCeqR. b. Indels are determined using programs that perform local reassembly such as Strelka and ABRA 24 such as. c. Structural rearrangements are detected using Pindel 25 or Breakseq26 It is determined using dedicated tools such as 3. To detect and prevent sample swaps, variant calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. As an example, extensive filtering of artificial calls is performed as follows: a. Removal of mutations found in normal DNA with lenient detection parameters in the case of potentially low coverage examples and with acceptable proximity criteria in the case of indels. b. Removal of mutations due to low mapping quality or low base quality 27 。 c. Removal of mutations resulting from recurring sequencing artifacts even if not observed in the corresponding normal. 27 。Examples mainly include mutations detected on mainly one strand. d. Removal of mutations detected in a set of unrelated controls 27 。 5. seq2HLA 28 、ATHLATES 29 、or one of Optitype is used, and also exome and RNA sequencing data are combined 28 、accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing such as long-read DNA sequencing 30 、or the adaptation of methods for ligating RNA fragments to maintain continuity 31 is included. 6. Robust detection of de novo ORFs arising from tumor-specific splice variants is performed by assembling transcripts from RNA-seq data using CLASS 32 、Bayesembler 33 、StringTie 34 、or similar programs in their reference-guided mode (i.e., using known transcript structures rather than attempting to recreate their entire transcripts from each experiment). Cufflinks 35is commonly used for this purpose, but it frequently produces a prohibitively large number of splice variants, many of which are much shorter than the full-length gene and may not be able to recover a simple positive control. The coding sequence and potential nonsense mutation-dependent decay mechanisms were determined by tools such as SpliceR 36 and MAMBA 37 . Gene expression was determined by tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels were determined by tools developed for these purposes such as ASE 38 or HTSeq 39 . Potential filtering steps include: a. Removal of candidate de novo ORFs that are considered to be insufficiently expressed. b. Removal of candidate de novo ORFs predicted to cause nonsense mutation-dependent decay (NMD). 7. Candidate neoantigens (e.g., de novo ORFs) observed only in RNA that cannot be directly verified as tumor-specific are classified as likely to be tumor-specific according to additional parameters by considering, for example: a. The presence of support for cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmation of trans-acting mutations in tumor DNA only in splicing factors. As an example, in three independently published experiments with the R625 mutant SF3B1, the genes that exhibited the most differential splicing were consistent 40 even though one experiment examined uveal melanoma patients 41 , the second experiment examined uveal melanoma cell lines 42 , and the third experiment examined breast cancer patients. c. For novel splicing isoforms, the presence of confirmation of "novel" splice-junction reads in RNASeq data. d. For novel rearrangements, the presence of validation of exon - proximal reads in tumor DNA that is not present in normal DNA. e. GTEx 43 Absence from gene expression summaries such as (i.e., making the likelihood of germline origin lower). 8. To directly avoid alignment - and annotation - based errors and artifacts, complementation of reference - genome - alignment - based analysis (e.g., for somatic mutations occurring near germline mutations or repeat - context indels) by comparing tumor and normal reads of the assembled DNA (or k - mers derived from such reads).

[0235] In samples with polyadenylated RNA, the presence of viral and microbial RNA in RNA - seq data is evaluated using RNA CoMPASS44 or similar methods towards the identification of additional factors that can predict patient response.

[0236] VI.B. Isolation and Detection of HLA Peptides Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) methods after lysis and solubilization of tissue samples. 55~58 The clarified lysate was used for HLA - specific IP.

[0237] Immunoprecipitation was performed using an antibody coupled to beads where the antibody is specific for the HLA molecule. For pan - class I HLA immunoprecipitation, a pan - class I CR antibody was used and for class II HLA - DR, an HLA - DR antibody was used. The antibody was covalently attached to NHS - sepharose beads during an overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60Immunoprecipitation can also be performed using antibodies that are not covalently bound to beads. Generally, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to hold the antibodies. Some antibodies that can be used to selectively enrich MHC / peptide complexes are shown below. TIFF2025094018000002.tif46160

[0238] The clarified tissue lysate is added to the antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is stored for additional experiments including additional IP. Using standard techniques, the IP beads are washed to remove non-specific binding and the HLA / peptide complex is eluted from the beads. The protein components are removed from the peptides using a molecular weight spin column or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and in some cases stored at -20 °C prior to MS analysis.

[0239] The dried peptides are reconstituted in HPLC buffer suitable for reverse-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution in a Fusion Lumos mass spectrometer (Thermo). The MS1 spectrum of the peptide mass / charge (m / z) is collected at high resolution in an Orbitrap detector, and then the MS2 low-resolution scan is collected in an ion trap detector after HCD fragmentation of the selected ions. Additionally, the MS2 spectrum can be obtained using either the CID or ETD fragmentation method, or any combination of the three techniques to obtain greater amino acid coverage of the peptide. The MS2 spectrum can also be measured at high-resolution mass accuracy in an Orbitrap detector.

[0240] The MS2 spectrum from each analysis is searched against a protein database using Comet 61、62 and peptide identification is performed using Percolator63~65 Score using it. Further sequencing can be performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or sequencing methods including spectral matching and de novo sequencing 75 can be used.

[0241] VI.B.1. Study of the MS detection limit for comprehensive HLA peptide sequencing Using the peptide YVYVADVAAK (SEQ ID NO:1), what is the detection limit was determined using various amounts of peptide loaded on the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol. (Table 1) The results are shown in Figure 1F. These results show that the lowest limit of detection (LoD) is in the attomolar range (10 -18 ), the dynamic range extends over 5 digits, and the signal-to-noise appears to be sufficient for sequencing in the low femtomolar range (10 -15 ).

[0242] TIFF2025094018000003.tif61128

[0243] VII. Presentation model VII.A. Overview of the system Figure 2A is an overview of an environment 100 for determining the likelihood of peptide presentation in a patient according to one embodiment. The environment 100 provides a context for introducing a presentation specific system 160 that itself includes a presentation information storage device 165.

[0244] The presentation prediction system 160 is one or more computer models embodied in a computer computing system as discussed below with respect to FIG. 14, and receives a peptide sequence related to a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the related set of MHC alleles. The presentation prediction system 160 can be applied to both class I and class II MHC alleles. This is useful in a variety of contexts. One example of a specific use of the presentation prediction system 160 is to receive a nucleotide sequence of a candidate neoantigen related to a set of MHC alleles from the tumor cells of patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the tumor-related MHC alleles and / or induce an immunogenic response in the immune system of patient 110. Those candidate neoantigens having a high likelihood when determined by system 160 can be selected for inclusion in vaccine 118, and such an anti-tumor immune response can be elicited from the immune system of patient 110 presenting the tumor cells.

[0245] The presented specific system 160 determines the presentation likelihood through one or more presentation models. Specifically, the presentation model generates the likelihood of whether a given peptide sequence will be presented for a set of related MHC alleles, and the likelihood is generated based on the presentation information stored in the storage device 165. For example, the presentation model can generate the likelihood of whether the peptide sequence "YVYVADVAAK (SEQ ID NO:1)" will be presented for the set of alleles HLA-A*02:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:03, HLA-C*01:04 on the cell surface of the sample. The presentation information 165 contains information about whether these peptides bind to various types of MHC alleles so that the peptides are presented by the MHC alleles, which is determined according to the positions of the amino acids in the peptide sequence in the model. The presentation model can predict whether an unrecognized peptide sequence will bind and be presented with a related set of MHC alleles based on the presentation information 165. As described above, the presentation model can be applied to both class I and class II MHC alleles.

[0246] VII.B. Presentation Information Figure 2 illustrates a method for obtaining presentation information according to one embodiment. The presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. The allele interaction information includes information that affects the presentation of peptide sequences depending on the type of MHC allele. The allele non-interaction information includes information that affects the presentation of peptide sequences independent of the type of MHC allele.

[0247] VII.B.1. Allele Interaction Information The allergenic interaction information includes a specified peptide sequence, which is known to be presented mainly by one or more specified MHC molecules derived from humans, mice, etc. Notably, this may or may not include data obtained from tumor samples. The presented peptide sequence may be identified from cells expressing a single MHC allele. In this example, the presented peptide sequence is generally collected from a single allele cell line that has been engineered to express a predetermined MHC allele and then exposed to a synthetic protein. The peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows an exemplary peptide presented on the predetermined MHC allele HLA-DRB1*12:01 This example shows that TIFF2025094018000004.tif5128 is isolated and identified by mass spectrometry. In this context, since the peptide is identified through cells engineered to express a single predetermined MHC protein, the direct relationship between the presented peptide and the MHC protein to which it binds is deterministically known.

[0248] The presented peptide sequence may also be collected from cells expressing multiple MHC alleles. Typically in humans, six different types of MHC-I molecules and up to twelve different types of MHC-II molecules are expressed in cells. Such presented peptide sequences may be identified from multiple allele cell lines that have been engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal tissue samples or tumor tissue samples. In this example in particular, the MHC molecules can be immunoprecipitated from normal or tumor tissue. The peptides presented on multiple MHC alleles can likewise be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows six exemplary peptides TIFF2025094018000005.tif17166 is presented on the identified class I MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, HLA-B*08:01, and the class II MHC alleles HLA-DRB1*10:01, HLA-DRB1:11:01, isolated, and identified by mass spectrometry, as shown in this example. In contrast to single-allele cell lines, since the bound peptides are isolated from the MHC molecules before being identified, the direct relationship between the presented peptides and the MHC proteins to which they are bound may be unknown.

[0249] Allele interaction information can also include mass spectrometry ion current, which depends on both the concentration of the peptide-MHC molecule complex and the ionization efficiency of the peptide. Ionization efficiency varies for each peptide in a sequence-dependent manner. Generally, the ion efficiency varies for each peptide over approximately two orders of magnitude, while the concentration of the peptide-MHC complex varies over a much larger range.

[0250] Allele interaction information can also include a measured or predicted value of the binding affinity between a given MHC allele and a given peptide. One or more affinity models can generate such predicted values (72,73,74). For example, returning to the example shown in FIG. 1D, the presentation information 165 can include a predicted binding affinity value of 1000 nM between the peptide YEMFNDKSF (SEQ ID NO:3) and the class I allele HLA-A * 01:01. Peptides with an IC50>1000 nM are presented by the MHC only rarely, and lower IC50 values increase the probability of presentation. The presentation information 165 can include the binding affinity prediction value between the peptide TIFF2025094018000006.tif4128 and the class II allele HLA-DRB1:11:01.

[0251] The allelic interaction information can also include a measured or predicted value of the stability of the MHC complex. One or more stability models can generate such predicted values. More stable peptide-MHC complexes (i.e., complexes with a longer half-life) are more likely to be presented at high copy numbers on tumor cells and on antigen-presenting cells that encounter the vaccine antigen. For example, returning to the example shown in Figure 2C, the presentation information 165 can include a stability prediction value with a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a stability prediction value for the half-life of the class II molecule HLA-DRB1:11:01.

[0252] The allelic interaction information can also include the measured or predicted rate of the peptide-MHC complex formation reaction. Complexes that form at a faster rate are more likely to be presented on the cell surface at high concentrations.

[0253] The allelic interaction information can also include the sequence and length of the peptide. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of the presented peptides have a length of 9. MHC class II molecules generally tend to present peptides with a length of 6-30 peptides.

[0254] The allelic interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoding peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoding peptide. The presence of the kinase motif affects the probability of post-translational modifications that can enhance or interfere with MHC binding.

[0255] The allelic interaction information can also include the expression or activity level of proteins involved in the process of post-translational modification (when measured or predicted by RNA seq, mass spectrometry, or other methods), such as kinases.

[0256] Allele interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing a particular MHC allele, as evaluated by mass spectrometry proteomics or other means.

[0257] Allele interaction information can also include the expression level of a particular MHC allele in the individual in question (e.g., as measured by RNA-seq or mass spectrometry). Peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.

[0258] Allele interaction information can also include the overall de novo antigen-encoding peptide sequence-independent probability of presentation by a particular MHC allele in other individuals expressing that particular MHC allele.

[0259] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of molecules of the same family (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and thus the presentation of peptides by HLA-C is a priori less probable than presentation by HLA-A or HLA-B II. As another example, since HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, it is presumed that the presentation of peptides by HLA-DP is less probable than presentation by HLA-DR or HLA-DQ.

[0260] Allele interaction information can also include the protein sequence of a particular MHC allele.

[0261] Any MHC allele non-interaction information listed in the following sections can also be modeled as MHC allele interaction information.

[0262] VII.B.2. Allele non-interaction information Allele non-interaction information can include the C-terminal sequence adjacent to the neoantigen-encoding peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence can affect the proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters the MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and thus the effect of the C-terminal flanking sequence cannot vary according to the MHC allele type. For example, returning to the example shown in Figure 2C, the presentation information 165 is the C-terminal flanking sequence of the presented peptide FJIEJFOESS (SEQ ID NO:5) identified from the source protein of the peptide may include TIFF2025094018000007.tif5128.

[0263] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide the mass spectrometry training data. As described later with respect to Figure 13H, RNA expression was identified as a strong predictor of peptide presentation. In one embodiment, the mRNA quantification measurements are identified from the software tool RSEM. Details of the execution of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).

[0264] The allergen non-interaction information can also include the N-terminal sequence adjacent to the peptide within the source protein sequence.

[0265] The allergen non-interaction information can also include the source gene of the peptide sequence. The source gene can be defined as the Ensembl protein family of the peptide sequence. In other examples, the source gene can be defined as the source DNA or source RNA of the peptide sequence. The source gene can be represented, for example, as a string of nucleotides encoding a protein or, alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode a particular protein. In another example, the allergen non-interaction information can also include the source transcript or isoform of the peptide sequence or a set of potential source transcripts or isoforms extracted from a database such as Ensembl or RefSeq.

[0266] The allergen non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.

[0267] The allergen non-interaction information can also include, optionally (when measured by RNA-seq or mass spectrometry), the presence of protease cleavage motifs in the peptide, weighted according to the expression of the corresponding protease in tumor cells. Peptides containing protease cleavage motifs are more easily degraded by proteases and thus have lower stability within cells and are therefore less likely to be presented.

[0268] The allergen non-interaction information can also include the metabolic turnover rate of the source protein when measured in an appropriate cell type. A faster metabolic turnover rate (i.e., a lower half-life) increases the probability of presentation, but when measured in dissimilar cell types, the predictive power of this property is low.

[0269] Allele non-interaction information can also optionally include the length of the source protein, taking into account the specific splice variant ("isoform") that is most highly expressed in tumor cells, when measured by RNA-seq or proteomic mass spectrometry, or when predicted from the annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.

[0270] Allele non-interaction information can also include the level of expression of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry). Different proteasomes have different preferences for cleavage sites. A greater weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.

[0271] Allele non-interaction information can also include the expression of the source gene of the peptide (when measured, for example, by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes within the tumor sample. Peptides derived from more highly expressed genes are more likely to be presented. Peptides derived from genes with undetectable levels of expression can be excluded from consideration.

[0272] Allele non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to a nonsense-mediated decay mechanism, such as predicted by a model from Rivas et al, Science 2015.

[0273] Allergen non-interaction information can also include the typical tissue-specific expression of the peptide's source gene during various stages of the cell cycle. Genes that are expressed at generally low levels (when measured by RNA-seq or sample analysis proteomics), but are known to be expressed at high levels during specific stages of the cell cycle, are more likely to produce peptides that are presented than genes that are stably expressed at very low levels.

[0274] Allergen non-interaction information can also include a comprehensive catalog of the properties of the source protein, such as those provided in, for example, uniProt or PDB http: / / www.rcsb.org / pdb / home / home.do. These properties can include, among other things, the secondary and tertiary structure of the protein, intracellular localization 11, gene ontology (GO) terms. Specifically, this information can contain annotations that act at the protein level, such as the 5'UTR length, and annotations that act at the level of specific residues, such as the helix motif of residues 300-310. These properties can also include turn motifs, sheet motifs, and disordered residues.

[0275] Allergen non-interaction information can also include properties that describe the nature of the domain of the source protein that contains the peptide, such as secondary or tertiary structure (e.g., alpha helix vs. beta sheet); alternative splicing.

[0276] Allergen non-interaction information can also include properties that describe the presence or absence of presentation hotspots at the position of the peptide in the source protein of the peptide.

[0277] Allergen non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression levels of the source protein in those individuals and the influence of the various HLA types of those individuals).

[0278] Allele non-interaction information can also include the probability that a peptide will not be detected by mass spectrometry or will be overrepresented due to technical bias.

[0279] Expression of various gene modules / pathways (need not contain the source protein of the peptide), as measured by gene expression assays such as RNASeq, microarrays, targeted panels such as Nanostring, or by assays such as RT-PCR, that provide information about the state of tumor cells, stroma, or tumor-infiltrating lymphocytes (TILs), or single / multiple genes representative of gene modules measured by such assays.

[0280] Allele non-interaction information can also include the copy number of the source gene of the peptide in tumor cells. For example, a peptide derived from a gene that undergoes homozygous deletion in tumor cells can be assigned a presentation probability of zero.

[0281] Allele non-interaction information can also include the probability that a peptide binds to TAP, or the measured or predicted binding affinity of the peptide for TAP. Peptides that are more likely to bind to TAP, or that bind to TAP with higher affinity, are more likely to be presented by MHC-I.

[0282] Allele non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, immunohistochemistry). In MHC-I, higher TAP expression levels increase the probability of presentation of all peptides.

[0283] Allele non-interaction information can also include, but is not limited to, the presence or absence of tumor mutations: i.Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, NTRK3. ii. in a gene encoding a protein involved in antigen presentation machinery (for example, any of B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or a gene encoding a component of the proteasome or immunoproteasome). Peptides whose presentation depends on components of the antigen presentation machinery that are under the influence of loss-of-function mutations in the tumor have a reduced probability of presentation.

[0284] The presence or absence of functional germline polymorphisms, including but not limited to: i. in a gene encoding a protein involved in antigen presentation machinery (for example, any of B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or a gene encoding a component of the proteasome or immunoproteasome).

[0285] Allele non-interaction information can also include tumor type (for example, NSCLC, melanoma).

[0286] Allele non-interaction information can also include known functionality of HLA alleles, such as as reflected by HLA allele suffixes as an example. For example, the suffix N in the allele name HLA-A*24:09N indicates a null allele that does not express and thus has a low likelihood of presenting epitopes; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.

[0287] Allele non-interaction information can also include clinical tumor subtypes (e.g., squamous cell lung cancer vs. non-squamous).

[0288] Allele non-interaction information can also include smoking history.

[0289] Allele non-interaction information can also include history of sunburn, sunlight exposure, or exposure to other mutagens.

[0290] Allele non-interaction information can also include local expression of the source gene of the peptide in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be presented.

[0291] Allele non-interaction information can also include the frequency of mutations in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.

[0292] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation may also include the annotation of the mutation (e.g., missense, read-through, frameshift, fusion, etc.), or whether the mutation is predicted to result in a nonsense-mediated decay (NMD) mechanism. For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in a decrease in mRNA translation, which decreases the probability of presentation.

[0293] VII.C. Presentation Specific System FIG. 3 is a high-level block diagram illustrating the computer logic components of a presentation specific system 160 according to one embodiment. In this exemplary embodiment, the presentation specific system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. The presentation specific system 160 also consists of a training data storage device 170 and a presentation model storage device 175. Some embodiments of the model management system 160 have modules different from those described herein. Similarly, the functions can be distributed among the modules in a manner different from that described herein.

[0294] VII.C.1. Data Management Module The data management module 312 generates a set of training data 170 from the presentation information 165. Each set of training data contains a number of data examples, and each data example i contains at least a peptide sequence p that is presented or not presented i and one or more associated MHC alleles a i combined with the peptide sequence p i and a dependent variable y representing information that the presentation specific system 160 is interested in predicting a new value of the independent variable i and contains a set of independent variables z i .

[0295] In one particular implementation referred to throughout the remainder of this specification, the dependent variable y i is a binary label indicating whether the peptide p i is presented by one or more associated MHC alleles a i . However, in other implementations, it is recognized that the dependent variable y i may represent any other kind of information that the presentation prediction system 160 is interested in predicting depending on the independent variable z i . For example, in another implementation, the dependent variable y i may also be a numerical value indicating the mass spectrometry ion current specified for the data example.

[0296] The peptide sequence p for data example i i is a sequence of k i amino acids, where k i can vary within a certain range among data examples i. For example, the range can be 8 - 15 for MHC class I or 6 - 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p in the training dataset i can have the same length, e.g., 9. The number of amino acids in the peptide sequence can vary depending on the type of MHC allele (such as MHC alleles in humans, etc.). The MHC allele a for data example i i indicates which MHC allele was present in combination with the corresponding peptide sequence p i .

[0297] The data management module 312 may also include additional allele interaction variables such as predicted values of the binding affinity b i and stability s i along with the peptide sequence p i and the bound MHC allele a i contained in the training data 170. For example, the training data 170 includes the predicted values b i of the binding affinity between the peptide p i and each of the bound MHC molecules indicated by a imay contain. As another example, the training data 170 may contain a i stability prediction value s for each of the MHC alleles shown in i .

[0298] The data management module 312 may also include the peptide sequence p i along with allelic non-interaction variables w such as the C-terminal flanking sequence and mRNA quantification measurements i as well.

[0299] The data management module 312 also identifies peptide sequences not presented by the MHC alleles and generates the training data 170. Generally, this involves identifying the "longer" sequence of the source protein containing the peptide sequence to be presented prior to presentation. If the presentation information contains the engineered cell line, the data management module 312 identifies a series of peptide sequences in the synthetic protein to which the cells were exposed that were not presented on the MHC alleles of the cells. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein that is the origin of the presented peptide sequence and identifies a series of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.

[0300] The data management module 312 also artificially generates peptides having random sequences of amino acids and identifies the generated sequences as peptides not presented on the MHC alleles. This can be achieved by randomly generating peptide sequences, enabling the data management module 312 to easily generate large amounts of synthetic data for peptides not presented on the MHC alleles. In practice, since a small percentage of peptide sequences are presented by the MHC alleles, the synthetically generated peptide sequences are very likely not to be presented by the MHC alleles even if they are contained in proteins processed by the cells.

[0301] Figure 4 illustrates an exemplary set of training data 170A according to one embodiment. Specifically, the first three data examples in the training data 170A are a single allele cell line containing allele HLA-C*01:03, and three peptide sequences show peptide presentation information from TIFF2025094018000008.tif9128. The fourth data example in the training data 170A shows a multi-allele cell line containing alleles HLA-B*07:02, HLA-C*01:03, HLA-A*01:01, and peptide information from the peptide sequence QIEJOEIJE (SEQ ID NO:13). The first data example shows that the peptide sequence QCEIOWARE (SEQ ID NO:14) was not presented by allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, the negatively labeled peptide sequences may be randomly generated by the data management module 312, or may be identified from the source protein of the presented peptides. The training data 170A also includes, for the peptide sequence-allele pairs, a predicted binding affinity value of 1000 nM and a predicted stability value of a half-life of 1 hour. The training data 170A also includes the C-terminal flanking sequence of the peptide FJELFISBOSJFIE (SEQ ID NO:15), and 10 2 also includes allele non-interaction variables such as the mRNA quantification measurement value of 10 TPM. The fourth data example shows that the peptide sequence QIEJOEIJE (SEQ ID NO:13) was presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. The training data 170A also includes the predicted binding affinity value and stability value for each of the alleles, as well as the C-terminal flanking sequence of the peptide and the mRNA quantification measurement value for the peptide.

[0302] VII.C.2. Encoding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more presentation models. In one implementation, the encoding module 314 encodes an array (e.g., a peptide array or a C-terminal flanking array) in one-hot for a predefined 20-letter amino acid alphabet. Specifically, a peptide sequence p i having k i amino acids is represented as a row vector of 20·k i elements, and the single element in p i 20·(j-1)+1 , p i 20·(j-1)+2 ,..., p i 20·j corresponding to the amino acid at the j-th position of the peptide sequence has a value of 1. All other remaining elements have a value of 0. As an example, for a given alphabet {A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y}, the peptide sequence EAF of 3 amino acids for data example i can be represented by a row vector of 60 elements represented by TIFF2025094018000009.tif12138. The C-terminal flanking sequence c i , as well as the protein sequence d h for the MHC allele, and other array data in the presentation information can be encoded in the same way as described above.

[0303] If the training data 170 contains amino acid sequences of different lengths, the encoding module 314 can further encode the peptides into vectors of equal length by adding PAD characters to extend the predefined alphabet. For example, this can be done by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the length of the peptide sequence with the maximum length in the training data 170. Thus, if the peptide sequence with the maximum length has k 最大 amino acids, the encoding module 314 encodes each sequence into a vector of (20 + 1)·k 最大Numerically represent as a row vector of elements. As an example, for the extended alphabet {PAD, A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y} and k 最大 = 5 for the maximum amino acid length, the same exemplary peptide sequence EAF of 3 amino acids is a 105 - element row vector TIFF2025094018000010.tif19158. The C - terminal adjacent sequence c i or other sequence data can similarly be encoded as described above. Thus, for the peptide sequence p i or c i each independent variable or column in represents the presence of a specific amino acid at a specific position in the array.

[0304] The above - described method of encoding array data has been described for arrays with amino acid sequences, but the method can be extended, in the same way, to other types of array data, such as, for example, DNA or RNA sequence data.

[0305] The encoding module 314 also encodes one or more MHC alleles a i for data example i into an m - element row vector, where each element h = 1, 2,..., m corresponds to a specific MHC allele. The element corresponding to the MHC allele specified for data example i has a value of 1. The remaining elements, otherwise, have a value of 0. As an example, for data example i corresponding to a multi - allele cell line among the specific MHC allele types {HLA - A*01:01, HLA - C*01:08, HLA - B*07:02, HLA - DRB1*10:01} with m = 4, the alleles HLA - B*07:02 and HLA - DRB1*10:01 can be represented by the 4 - element row vector a i = [0 0 1 1], where a3 i = 1 and a4 i = 1. Examples for four specific MHC allele types are described herein, but the number of MHC allele types can actually be in the hundreds or thousands. As described above, each data example i typically has a peptide sequence pi includes up to six different MHC allele types in relation to

[0306] The encoding module 314 also encodes the label y for each data example i i as a binary variable having a value from the set {0,1}, where a value of 1 indicates that the peptide x i is presented by one of the associated MHC alleles a i and a value of 0 indicates that the peptide x i is not presented by any of the associated MHC alleles a i . When the dependent variable y i represents a mass spectrometry ion current, the encoding module 314 can additionally scale the values using various functions such as a log function having a range of [-∞,∞] for ion current values between [0,∞].

[0307] The encoding module 314 can represent the pair of the peptide p i and the allele interaction variable x for the associated MHC allele h h i as a row vector in which numerical representations of the allele interaction variables are concatenated one after another. For example, the encoding module 314 can represent x h i as a row vector equivalent to [p i , [p i b h i , [p i s h i , or [p i b h i s h i , where b h i is the predicted binding affinity value for the peptide pi and the associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of the allele interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0308] In one example, the encoding module 314 represents the binding affinity information by incorporating the value measured or predicted for the binding affinity into the allelic interaction variable x h i

[0309] In one example, the encoding module 314 represents the binding stability information by incorporating the value measured or predicted for the binding stability into the allelic interaction variable x h i

[0310] In one example, the encoding module 314 represents the binding on-rate information by incorporating the value measured or predicted for the binding on-rate into the allelic interaction variable x h i

[0311] In one example, for a peptide presented by a class I MHC molecule, the encoding module 314 represents the peptide length as a vector TIFF2025094018000011.tif5132 (where TIFF2025094018000012.tif3128 is an indicator function and L k means the length of the peptide p k ). The vector T k can be included in the allelic interaction variable x h i In another example, for a peptide presented by a class II MHC molecule, the encoding module 314 represents the peptide length as a vector TIFF2025094018000013.tif19161 (where TIFF2025094018000014.tif3128 is an indicator function and L k means the length of the peptide p k ). The vector T k can be included in the allelic interaction variable x h ​​​i can be included in.

[0312] In one example, the encoding module 314 represents the RNA expression information of the MHC allele by incorporating the RNA-seq based expression level of the MHC allele into the allelic interaction variable xhi.

[0313] Similarly, the encoding module 314 can represent the allelic non-interaction variable w i as a row vector in which numerical representations of allelic non-interaction variables are chained one after another. For example, w i can be a row vector equivalent to [c i or [c i m i w i , and w i is a row vector representing the C-terminal flanking sequence of the peptide pi and any other allelic non-interaction variables in addition to the mRNA quantification measurement value m i related to the peptide. Alternatively, one or more combinations of allelic non-interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0314] In one example, the encoding module 314 represents the turnover rate or half-life of the source protein for the peptide sequence by incorporating the turnover rate or half-life into the allelic non-interaction variable w i .

[0315] In one example, the encoding module 314 represents the length of the source protein or isoform by incorporating the protein length into the allelic non-interaction variable w i .

[0316] In one example, the encoding module 314 includes the average expression of immunoproteasome-specific proteasome subunits including β1 i , β2 i , β5 i subunits into the allelic non-interaction variable w iBy incorporating it, it represents the activation of the immunoproteasome.

[0317] In one example, the encoding module 314 represents the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM), or the source protein of the gene or transcript of the peptide, by incorporating the abundance of the source protein into the allelic non-interaction variable w i By incorporating it.

[0318] In one example, the encoding module 314 represents the probability that the transcript of origin of a peptide will undergo nonsense-mediated decay (NMD), as estimated by a model in Rivas et.al. Science, 2015, by incorporating this probability into the allelic non-interaction variable w i By incorporating it.

[0319] In one example, the encoding module 314 represents the activation status of a gene module or pathway evaluated via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM for each gene in the pathway, and then computationally calculating summary statistics, such as the mean, across the genes in the pathway. The mean can be incorporated into the allelic non-interaction variable w i Can be incorporated.

[0320] In one example, the encoding module 314 represents the copy number of the source gene by incorporating the copy number into the allelic non-interaction variable w i By incorporating it.

[0321] In one example, the encoding module 314 represents the measured or predicted TAP binding affinity (e.g., in nanomolar units) by including it in the allelic non-interaction variable w i To represent the TAP binding affinity.

[0322] In one example, the encoding module 314 represents the TAP expression level by including the TAP expression level measured by RNA-seq (and quantified, e.g., in units of TPM by RSEM) as the allelic non-interaction variable w i in it.

[0323] In one example, the encoding module 314 represents the tumor mutation as a vector of indicator variables in the allelic non-interaction variable w i (i.e., if the peptide p k is derived from a sample having the KRAS G12D mutation, then d k = 1, otherwise 0).

[0324] In one example, the encoding module 314 represents the germline polymorphism in the antigen-presenting gene as a vector of indicator variables (i.e., if the peptide p k is derived from a sample having a germline polymorphism specific to TAP, then d k = 1). These indicator variables can be included in the allelic non-interaction variable w i .

[0325] In one example, the encoding module 314 represents the tumor type as a length-1 one-hot encoded vector for the alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot encoded variables can be included in the allelic non-interaction variable w i .

[0326] In one example, the encoding module 314 represents the MHC allele suffix by processing the 4-digit HLA allele with various suffixes. For example, HLA-A*24:09N is considered a different allele than HLA-A*24:09 for the purposes of the model. Alternatively, since HLA alleles ending with the N suffix are not expressed, the probability of presentation by MHC alleles with the N suffix can be set to zero for all peptides.

[0327] In one example, the encoding module 314 represents the tumor subtype as a length-1 one-hot encoded vector for the alphabet of the tumor subtype (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot encoded variables can be included in the allelic non-interaction variable w i can be included in.

[0328] In one example, the encoding module 314 can include the smoking history in the allelic non-interaction variable w i as a binary indicator variable (d k = 1 if the patient has a smoking history, 0 otherwise). Alternatively, the smoking history can be encoded as a length-1 one-hot encoded variable for the alphabet of the severity of smoking. For example, the smoking status can be rated on a scale of 1 to 5, where 1 indicates a non-smoker and 5 indicates a current heavy smoker. Since the smoking history is mainly associated with lung tumors, when training a model for multiple tumor types, this variable can also be defined as 1 if the patient has a smoking history and the tumor type is a lung tumor, and 0 otherwise.

[0329] In one example, the encoding module 314 can include the sunburn history in the allelic non-interaction variable w i as a binary indicator variable (d k = 1 if the patient has a history of severe sunburn, 0 otherwise). Since severe sunburn is mainly associated with melanoma, when training a model for multiple tumor types, this variable can also be defined as 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and 0 otherwise.

[0330] In one example, the encoding module 314 represents the distribution of the expression levels of a particular gene or transcript for each gene or transcript in the human genome as summary statistics (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, for peptide p in a sample having the tumor type melanoma, the measured expression level of the gene or transcript from which peptide p k is derived can be included in the allelic non-interaction variable w k , and the mean and / or median gene or transcript expression of the gene or transcript from which peptide p i is derived in melanoma as measured by TCGA can also be included. k

[0331] In one example, the encoding module 314 represents the mutation type as a one-hot encoded variable of length 1 for the alphabet of mutation types (e.g., missense, frameshift, NMD-inducing, etc.). These one-hot encoded variables can be included in the allelic non-interaction variable w i .

[0332] In one example, the encoding module 314 represents the protein-level characteristics of a protein as values of annotations of the source protein (e.g., 5’UTR length) in the allelic non-interaction variable w i . In another example, the encoding module 314 represents the residue-level annotation of the source protein for peptide p i as an indicator variable that is equivalent to 1 if peptide p i overlaps with a helix motif and 0 otherwise, or as equivalent to 1 if peptide p i is completely contained within a helix motif, by including it in the allelic non-interaction variable wi. In another example, the encoding module 314 represents the characteristic representing the proportion of residues in peptide p i contained within a helix motif annotation as the allelic non-interaction variable w ican be included.

[0333] In one example, the encoding module 314 represents the type of protein or isoform in the human proteome as an index vector o having a length equal to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is 1 if the peptide p k is derived from protein i and 0 otherwise.

[0334] In one example, the encoding module 314 represents the source gene G = gene(p i ) of the peptide p as a categorical variable having L possible categories (where L indicates the upper limit of the number of indexed source genes 1, 2,..., L). i

[0335] In one example, the encoding module 314 represents the tissue type, cell type, tumor type, or tumor histology type T = tissue(p i ) of the peptide p as a categorical variable having M possible categories (where M indicates the upper limit of the number of indexed types 1, 2,..., M). Examples of tissue types include, for example, lung tissue, heart tissue, intestinal tissue, nerve tissue, etc. Examples of cell types include dendritic cells, macrophages, CD4 T cells, etc. Examples include lung adenocarcinoma, squamous cell carcinoma of the lung, melanoma, non-Hodgkin lymphoma, etc. i

[0336] The encoding module 314 can also represent the entire set of variables z i for the peptide p i and the associated MHC allele h as a row vector in which the numerical representations of the allele interaction variable x i and the allele non-interaction variable w i are chained one after another. For example, the encoding module 314 represents z h i as [xh i w i or [w i x h i can be represented as a row vector equivalent to

[0337] VIII. Training Module The training module 316 constructs one or more presentation models that generate the likelihood that a peptide sequence will be presented by an MHC allele related to the peptide sequence. Specifically, for the peptide sequence p k and the peptide sequence p k and the MHC allele a k associated therewith, each presentation model generates an estimated value u k indicating the likelihood that the peptide sequence p k will be presented by one or more of the associated MHC alleles a k .

[0338] VIII.A. Overview The training module 316 constructs one or more presentation models based on a training data set stored in the storage device 170, which is generated from the presentation information stored in 165. Generally, regardless of the specific type of the presentation model, all of the presentation models capture the dependency between the independent variable and the dependent variable in the training data 170 such that the loss function is minimized. Specifically, the loss function TIFF2025094018000015.tif5128 represents the conflict between the value of the dependent variable y i∈S for one or more data examples S in the training data 170 and the estimated likelihood u i∈S for the data example S generated by the presentation model. In one particular implementation referred to throughout the remainder of this specification, the loss function TIFF2025094018000016.tif4128 is the negative log-likelihood function given by the following equation (1a). TIFF2025094018000017.tif11132However, in practice, other loss functions may be used. For example, when making predictions about mass spectrometry ion currents, the loss function is the mean squared loss given by Equation 1b below. TIFF2025094018000018.tif11128

[0339] The presentation model can be a parametric model where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function TIFF2025094018000019.tif4128The various parameters of the parametric type presentation model that minimize the loss function are determined through gradient-based numerical optimization algorithms such as, for example, batch gradient algorithms, stochastic gradient algorithms, etc. Alternatively, the presentation model can be a non-parametric model where the model structure is determined from the training data 170 and is not strictly based on a fixed set of parameters.

[0340] VIII.B. Per-allele Model The training module 316 can construct a presentation model for predicting the presentation likelihood of peptides on a per-allele basis. In this example, the training module 316 can train the presentation model based on the data example S in the training data 170 generated from cells expressing a single MHC allele.

[0341] In one implementation, the training module 316 TIFF2025094018000020.tif7128models the estimated presentation likelihood u of the peptide pk for a particular allele h, where the peptide sequence x k represents the peptide p h k and the encoded allele interaction variable for the corresponding MHC allele h, and f(·) is an arbitrary function, which for convenience of description is referred to as the transformation function throughout this specification. Further, g k h ​(·) is an arbitrary function, which for convenience of description is referred to as a dependency function throughout this specification, and is a parameter θ determined for MHC allele h. h Based on the set of h k generate a dependency score for the allelic interaction variable x for each MHC allele h. The values of the set of parameters θ for each MHC allele h h can be determined by minimizing a loss function with respect to θ h , where i is each example in a subset S of the training data 170 generated from cells expressing a single MHC allele h.

[0342] The dependency function g h (x h k ; θ h ) outputs a dependency score for MHC allele h, indicating whether MHC allele h presents the corresponding neoantigen based on at least the allelic interaction characteristic x h k and in particular based on the position of the amino acid in the peptide sequence of peptide p k . For example, the dependency score for MHC allele h can have a high value if MHC allele h is likely to present peptide p k , and can have a low value if the likelihood of presentation is low. The conversion function f(·) converts the input, and more specifically, in this example, converts the dependency score generated by g h (x h k ; θ h ) into an appropriate value indicating the likelihood that peptide p k will be presented by the MHC allele.

[0343] In one particular implementation referred to throughout the remainder of this specification, f(·) is a function having a range within [0,1] for an appropriate domain range. In one example, f(·) is the expit function given by TIFF2025094018000021.tif10128. As another example, f(·) can also be the hyperbolic tangent function given by TIFF2025094018000022.tif5128. Alternatively, if the prediction is made for a mass spectrometry ion current having a value outside the range [0,1], f(·) can be any function such as, for example, an identity function, an exponential function, a log function, etc.

[0344] Thus, the per-allele likelihood that the peptide sequence p k will be presented by the MHC allele h can be generated by applying the dependency function g h (·) to the encoded version of the peptide sequence p k to generate the corresponding dependency score. The dependency score may be transformed by the transformation function f(·) to generate the per-allele likelihood that the peptide sequence p k will be presented by the MHC allele h.

[0345] VIII.B.1 Dependency Function for Allele Interaction Variables In one particular implementation referred to throughout this specification, the dependency function g h (·) linearly combines each allele interaction variable in x h k with the corresponding parameter in the set of parameters θ h determined for the associated MHC allele h, which is an affine function given by TIFF2025094018000023.tif6128.

[0346] In another particular implementation referred to throughout this specification, the dependency function g h (·) is a network function given by TIFF2025094018000024.tif6128 represented by a network model NN h (·) having a series of nodes arranged in one or more layers, where the nodes have parameters θh It can be connected to other nodes through connections each having related parameters in the set. The value at one particular node can be represented as the sum of the values of the nodes connected to the particular node, weighted by the related parameters mapped by the activation function related to the particular node. In contrast to an affine function, the network model is advantageous because the presentation model can incorporate non-linearity and process data having amino acid sequences of different lengths. Specifically, through non-linear modeling, the network model can capture the interactions between amino acids at different positions in the peptide sequence and how this interaction affects peptide presentation.

[0347] Generally, the network model NN h (·) can be structured as a feed-forward network such as an artificial neural network (ANN), a convolutional neural network (CNN), a deep neural network (DNN), etc., and / or a recurrent network such as a long short-term memory network (LSTM), a bidirectional recurrent network, a deep bidirectional recurrent network, etc.

[0348] In one example referred to throughout the remainder of this specification, each MHC allele at h = 1, 2,..., m is associated with a separate network model, and NN h (·) means the output from the network model associated with MHC allele h.

[0349] Figure 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h = 3. As shown in Figure 5, the network model NN3(·) for MHC allele h = 3 includes three types of input nodes at layer l = 1, four types of nodes at layer l = 2, two types of nodes at layer l = 3, and one type of output node at layer l = 4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2),..., θ3(10). The network model NN3(·) receives input values (individual data examples including encoded polypeptide sequence data and any other training data used) for the three allelic interaction variables x3 k (1), x3 k (2), and x3 k (3), and outputs the value NN3(x3 k ). The network function may include one or more network models that each take a different allelic interaction variable as input.

[0350] In another example, the specified MHC alleles h = 1, 2,..., m are associated with a single network model NN H (·), where NN h (·) represents one or more outputs of a single network model associated with MHC allele h. In such an example, the set of parameters θ h may correspond to the set of parameters for a single network model, and thus the set of parameters θ h may be shared by all MHC alleles.

[0351] Figure 6A illustrates an exemplary network model NN H (·) shared by MHC alleles h = 1, 2,..., m. As shown in Figure 6A, the network model NN H (·) includes m output nodes, each corresponding to an MHC allele. The network model NN3(·) receives the allelic interaction variable x3 k for MHC allele h = 3 and outputs the value NN3(x3k Outputs m values including

[0352] In yet another example, a single network model NN H (·) is a network model that takes the allele interaction variable x of MHC allele h h k and the encoded protein sequence d h and outputs a dependency score. In such an example, the set of parameters θ h can again correspond to the set of parameters for a single network model, and thus the set of parameters θ h can be shared by all MHC alleles. Thus, in such an example, NNh(·) takes as input [x h k d h and gives the output of a single network model NN H (·). Such a network model is advantageous because it can correctly predict the peptide presentation probability for MHC alleles that were unknown in the training data simply by identifying their protein sequences.

[0353] Figure 6B illustrates an exemplary network model NN H (·) shared by MHC alleles. As shown in Figure 6B, the network model NN H (·) receives as input the allele interaction variable and protein sequence of MHC allele h = 3 and outputs the dependency score NN3(x3 k ).

[0354] In yet another example, the dependency function g h (·) can be represented as TIFF2025094018000025.tif6128, where g’ h (x h k ; θ’ his an affine function, a network function, etc. with a set of parameters θ’h, and represents the baseline probability of presentation for MHC allele h, the bias parameter θ in the set of parameters for the allelic interaction variable of the MHC allele h 0 is associated with.

[0355] In another embodiment, the bias parameter θ h 0 may be shared according to the gene family of MHC allele h. That is, the bias parameter θ for MHC allele h h 0 may be equivalent to θ 遺伝子(h) 0 and the gene (h) is the gene family of MHC allele h. For example, the class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family of "HLA-A", and the bias parameter θ for each of these MHC alleles h 0 may be shared. As another example, assign the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 to the gene family of "HLA-DRB", and the bias parameter θ for each of these MHC alleles h 0 can be shared.

[0356] As an example, returning to equation (2), the likelihood that peptide p h will be presented by MHC allele h = 3 among m = 4 different specified MHC alleles using the affine dependence function g k (·) is TIFF2025094018000026.tif6128 can generate, where x3k is the allelic interaction variable specified for MHC allele h = 3, and θ3 is the set of parameters determined for MHC allele h = 3 through loss function minimization.

[0357] As another example, among the m = 4 different specified MHC alleles using separate network transformation functions gh(·), the likelihood that peptide p k will be presented by MHC allele h = 3 is which can be generated by TIFF2025094018000027.tif6128, where x3 k is the allele interaction variable specified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.

[0358] Figure 7 illustrates the generation of the presentation likelihood of peptide p k associated with MHC allele h = 3 using an exemplary network model NN3(·). As shown in Figure 7, the network model NN3(·) receives the allele interaction variable x3 k for MHC allele h = 3 and generates an output NN3(x3 k ). The output is mapped by a function f(·) to generate an estimated presentation likelihood u k .

[0359] VIII.B.2. Per Allele with Allele Non-Interaction Variables In one implementation, the training module 316 incorporates allele non-interaction variables to model the estimated presentation likelihood uk of peptide p k by TIFF2025094018000028.tif7131, where w k represents the encoded allele non-interaction variable for peptide p k , and g w (·) is a function of the allele non-interaction variable w w based on the set of parameters θ k determined for the allele non-interaction variable. Specifically, the set of parameters θ h for each MHC allele h and the set of parameters θ w for the allele non-interaction variable have values of θ h and θw can be determined by minimizing the loss function with respect to, where i is each example in a subset S of the training data 170 generated from cells expressing a single MHC allele.

[0360] dependency function g w (w k ; θ w ) outputs a dependency score for the non-allelic interaction variables indicating whether the peptide p k is presented by one or more MHC alleles based on the influence of the non-allelic interaction variables. For example, the dependency score for the non-allelic interaction variables can be high if the C-terminal flanking sequence known to have a positive influence on the presentation of the peptide p k is bound to the peptide p k and can have a low value if the C-terminal flanking sequence known to have a negative influence on the presentation of the peptide p k is bound to the peptide p k .

[0361] According to equation (8), the allele-specific likelihood that the peptide sequence p k will be presented by the MHC allele h can be generated by applying the function g h (·) for the MHC allele h to the encoded version of the peptide sequence p k to generate the corresponding dependency scores for the allelic interaction variables. The function g w (·) for the non-allelic interaction variables is also applied to the encoded version of the non-allelic interaction variables to generate the dependency scores for the non-allelic interaction variables. Both scores are combined and the combined score is transformed by a transformation function f(·) to generate the allele-specific likelihood that the peptide sequence p k will be presented by the MHC allele h.

[0362] Alternatively, the training module 316 replaces the non-allelic interaction variables w k with the allelic interaction variables x hk By adding, it may include the allelic non - interaction variable \(w_k\) in the prediction. Thus, the presented likelihood can be given by TIFF2025094018000029.tif7128.

[0363] VIII.B.3 Dependency function for allelic non - interaction variables Dependency function \(g(\cdot)\) for allelic interaction variables h Similar to the dependency function \(g(\cdot)\) for allelic interaction variables, the dependency function \(g(\cdot)\) for allelic non - interaction variables w can be an affine function or a network function related to the allelic non - interaction variable \(w\) k in a network.

[0364] Specifically, the dependency function \(g(\cdot)\) w is an affine function given by TIFF2025094018000030.tif6128 that linearly combines the allelic non - interaction variables in \(w\) with the corresponding parameters in the set of parameters \(\theta\) k w TIFF2025094018000030.tif6128.

[0365] The dependency function \(g(\cdot)\) w can also be a network function represented by a network model \(NN(\cdot)\) with related parameters in the set of parameters \(\theta\) w w given by TIFF2025094018000031.tif6128. The network function may include one or more network models that each take different allelic non - interaction variables as inputs. TIFF2025094018000031.tif6128.

[0366] In another example, the dependency function \(g(\cdot)\) for allelic non - interaction variables w can be given by TIFF2025094018000032.tif6128, where \(g'(w;\theta')\) w (w k ;\theta' w) is the allelic non-interaction parameter θ' w with a set of affine functions, network functions, etc., where m k is the mRNA quantification measurement value for peptide p k , h(·) is a function that converts the quantification measurement value, and θ w m is a parameter in the set of parameters for allelic non-interaction variables that is combined with the mRNA quantification measurement value to generate a dependency score for the mRNA quantification measurement value. In one specific embodiment referred to throughout the remainder of this specification, h(·) is a log function, but in reality, h(·) can be any one of a variety of different functions.

[0367] In yet another example, the dependency function g w (·) is given by TIFF2025094018000033.tif6128, where g' w (w k ; θ' w ) is an affine function, network function, etc. with a set of allelic non-interaction parameters θ' w , o k is the index vector described in Section VII.C.2 that represents proteins and isoforms in the human proteome for peptide p k , and θ w o is the set of parameters in the set of parameters for allelic non-interaction variables that is combined with the index vector. In one variation, when the dimensions of the set of o k and the parameter θ w o are significantly high, TIFF2025094018000034.tif5128( TIFF2025094018000035.tif4128 can add parameter regularization terms such as the L1 norm, L2 norm, combination, etc. (representing the L1 norm, L2 norm, combination, etc.) to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined through an appropriate method.

[0368] In yet another example, the dependency function g for the allelic non-interaction variable w (·) is given by TIFF2025094018000036.tif14131, where g’ w (w k ;θ’ w ) is an affine function, network function, etc. with a set of allelic non-interaction parameters θ’ w , and TIFF2025094018000037.tif5128 is an indicator function equal to 1 when the peptide p k is derived from the source gene l for the allelic non-interaction variable as described above, and θ w l is a parameter indicating the "antigenicity" of the source gene l. In one variation, when L is large enough and thus the number of parameters θ w l=1, 2,...,L is large enough, a parameter regularization term such as TIFF2025094018000038.tif5128 (where TIFF2025094018000039.tif4128 is the L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined by an appropriate method.

[0369] In yet another example, the dependency function g for the allelic non-interaction variable w (·) is given by TIFF2025094018000040.tif14158, where g’ w (w k ;θ’ w) is the allergen non-interaction parameter θ' w and is an affine function, network function, etc. with a set of TIFF2025094018000041.tif5128 is the peptide p as described above with respect to the allergen non-interaction variable k when it is derived from the source gene l, and the peptide p k is an indicator function equal to 1 when it is derived from the tissue type m, and θ w lm is a parameter indicating the antigenicity of the combination of the source gene l and the tissue type m. Specifically, the antigenicity of the gene l of the tissue type m can indicate the residual tendency of the cells of the tissue type m to present the peptide derived from the gene l after regulation regarding RNA expression and peptide sequence context.

[0370] In one variation, when L or M is sufficiently large and thus the number of parameters θ w lm=1, 2,...,LM is sufficiently large, a parameter regularization term such as TIFF2025094018000042.tif5128 (where TIFF2025094018000043.tif4128 is the L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the value of the parameter. The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the value of the parameter so that the coefficients for the same source gene do not vary greatly between tissue types. For example, a penalty term such as TIFF2025094018000044.tif19128 (where TIFF2025094018000045.tif5128 is the average antigenicity across the tissue types of the source gene l) can add a penalty to the standard deviation of the antigenicity across different tissue types in the loss function.

[0371] In practice, the dependency function g w (·) regarding the allelic non-interaction variable can be generated by combining any additional terms of equations (10), (11), (12a), and (12b). For example, by adding together the term h(·) indicating the mRNA quantification value of equation (10) and the term indicating the antigenicity of the source gene of equation (12) with any other affine function or network function, the dependency function regarding the allelic non-interaction variable can be generated.

[0372] As an example, returning to equation (8), the likelihood that peptide p h (·), g w will be presented by MHC allele h = 3 among m = 4 different specified MHC alleles using the affine transformation functions g k is TIFF2025094018000046.tif6128, where w k is the allelic non-interaction variable specified for peptide p k , and θ w is a set of parameters determined for the allelic non-interaction variable.

[0373] As another example, the likelihood that peptide p h (·), g w will be presented by MHC allele h = 3 among m = 4 different specified MHC alleles using the network transformation functions g k is TIFF2025094018000047.tif6128, where w k is the allelic interaction variable specified for peptide p k , and θ w is a set of parameters determined for the allelic non-interaction variable.

[0374] Figure 8 shows peptide p w associated with MHC allele h = 3 using the exemplary network models NN3(·) and NN kThe generation of the presentation likelihood is described. As shown in FIG. 8, the network model NN3(·) receives the allele interaction variable x3 for the MHC allele h = 3 k and generates the output NN3(x3 k ). The network model NN w (·) receives the allele non-interaction variable w k for the peptide p k and generates the output NN w (w k ). The outputs are combined and mapped by the function f(·) to generate the estimated presentation likelihood u k .

[0375] VIII.C. Multiple Allele Model The training module 316 may also construct a presentation model for predicting the presentation likelihood of a peptide in a multiple allele setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on the data example S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof.

Example

[0376] VIII.C.1. Example 1: Maximum Value of Allele-by-Allele Model In one implementation, the training module 316 determines the estimated presentation likelihood u k of the peptide p k associated with the set of multiple MHC alleles H as the presentation likelihood determined for each of the MHC alleles h in the set H based on cells expressing a single allele, as described above together with equations (2)-(11) Modeled as a function of TIFF2025094018000048.tif4128. Specifically, the presentation likelihood u k can be TIFF2025094018000049.tif4128 any function. In one implementation, as shown in equation (12), the function is a maximum value function, and the presentation likelihood u kcan be determined as the maximum value of the presentation likelihood for each MHC allele h in set H. TIFF2025094018000050.tif6128

[0377] VIII.C.2. Example 2.1: Sum function model In one implementation, training module 316 models the estimated presentation likelihood u k of peptide p k as modeled by TIFF2025094018000051.tif14128, where element a h k is 1 for a plurality of MHC alleles H related to peptide sequence p k and x h k represents the encoded allele interaction variable for peptide p k and the corresponding MHC allele. The value of the set of parameters θ h for each MHC allele h can be determined by minimizing the loss function with respect to θ h , where i is each example in subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing a plurality of MHC alleles. The dependency function g h can be in any form of the dependency function g h introduced above in Section VIII.B.1.

[0378] According to equation (13), the presentation likelihood that peptide sequence p k will be presented by one or more MHC alleles h can be generated by applying the dependency function g h (·) to the encoded version of peptide sequence p k for each of the MHC alleles H to generate the corresponding scores for the allele interaction variables. The scores for each MHC allele h are combined and transformed by the transformation function f(·) to generate the presentation likelihood that peptide sequence p k will be presented by the set of MHC alleles H.

[0379] The presentation model of equation (13) is different from the per-allele model of equation (2) in that the number of related alleles for each peptide p k can be greater than 1. In other words, more than one element in a h k can have a value of 1 for multiple MHC alleles H related to the peptide sequence p k .

[0380] As an example, the likelihood that peptide p h will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the affinity transformation function g k (·) can be generated by TIFF2025094018000052.tif6128, where x2 , x3 k are allele interaction variables specified for MHC alleles h = 2, h = 3, and θ2, θ3 are sets of parameters determined for MHC alleles h = 2, h = 3. k

[0381] As another example, the likelihood that peptide p h will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the network transformation functions g w (·), g k (·) can be generated by TIFF2025094018000053.tif6128, where NN2(·), NN3(·) are network models specified for MHC alleles h = 2, h = 3, and θ2, θ3 are sets of parameters determined for MHC alleles h = 2, h = 3.

[0382] Figure 9 shows peptide p k related to MHC alleles h = 2, h = 3 using exemplary network models NN2(·) and NN3(·).​Describe the generation of the presentation likelihood. As shown in FIG. 9, the network model NN2(·) takes the allele interaction variable x2 for MHC allele h = 2 k and generates an output NN2(x2 k ). The network model NN3(·) takes the allele interaction variable x3 for MHC allele h = 3 k and generates an output NN3(x3 k ). The outputs are combined and mapped by the function f(·) to generate the estimated presentation likelihood u k .

[0383] VIII.C.3. Example 2.2: Sum Function Model with Allele Non-Interaction Variables In one implementation, the training module 316 incorporates allele non-interaction variables and models the estimated presentation likelihood u k of the peptide p k by TIFF2025094018000054.tif14142, where w k represents the encoded allele non-interaction variable for the peptide p k . Specifically, the values of the set of parameters θ h for each MHC allele h and the set of parameters θ w for the allele non-interaction variable can be determined by minimizing the loss function with respect to θ h and θ w , where i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The dependency function g w can be in any form of the dependency function g w introduced above in Section VIII.B.3.

[0384] Thus, according to equation (14), the presentation likelihood that the peptide sequence p k will be presented by one or more MHC alleles H is the function g h (·) for each of the MHC alleles H for the peptide sequence p kIt can be generated by applying it to the encoded version and generating the corresponding dependency scores for the allelic interaction variables of each MHC allele h. The function g w (·) is also applied to the encoded version of the allelic non-interaction variables to generate the dependency scores for the allelic non-interaction variables. The scores are combined, and the combined scores are transformed by the transformation function f(·) to generate the presentation likelihood that the peptide sequence p k will be presented by the MHC allele H.

[0385] In the presentation model of equation (14), the number of related alleles for each peptide p k can be greater than 1. In other words, a h k more than one element in can have a value of 1 for multiple MHC alleles H related to the peptide sequence p k .

[0386] As an example, for the affine transformation functions g h (·), g w (·), the likelihood that the peptide p k will be presented by the MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles is TIFF2025094018000055.tif6128 can generate, where w k is the allelic non-interaction variable specified for the peptide p k , and θ w is a set of parameters determined for the allelic non-interaction variables.

[0387] As another example, for the network transformation functions g h (·), g w (·), the likelihood that the peptide p k will be presented by the MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles is It can be generated by TIFF2025094018000056, where w k is the allelic interaction variable specified for peptide p k , and θ w is a set of parameters determined for the allelic non-interaction variable.

[0388] Figure 10 shows the generation of the presentation likelihood for peptide p w associated with MHC alleles h = 2, h = 3 using exemplary network models NN2(·), NN3(·), and NN k . As shown in Figure 10, the network model NN2(·) receives the allelic interaction variable x2 k for MHC allele h = 2 and generates the output NN2(x2 k ). The network model NN3(·) receives the allelic interaction variable x3 k for MHC allele h = 3 and generates the output NN3(x3 k ). The network model NN w (·) receives the allelic non-interaction variable w k for peptide p k and generates the output NN w (w k ). The outputs are combined and mapped by the function f(·) to generate the estimated presentation likelihood u k .

[0389] Alternatively, the training module 316 may include the allelic non-interaction variable w k in the prediction by adding the allelic non-interaction variable w h k to the allelic interaction variable x k in equation (15). Thus, the presentation likelihood can be given by TIFF2025094018000057.tif14134.

[0390] VIII.C.4. Example 3.1: Model Using Implicit Allele-Specific Likelihoods In another embodiment, the training module 316 models the estimated presentation likelihood u k of the peptide p k as, modeled by TIFF2025094018000058.tif7142, where the element a h k is 1 for a plurality of MHC alleles k associated with the peptide sequence p TIFF2025094018000059.tif4128, u’ k h is the implicit allele-specific presentation likelihood for the MHC allele h, the vector v is such that the element v h is a vector corresponding to a h k ·u’, k h s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the value of the input into a predetermined range. As described in more detail below, s(·) may be a sum function or a quadratic function, but it is recognized that in other embodiments, s(·) may be any function such as a maximum function. The set of values of the parameter θ for the implicit allele-specific likelihood can be determined by minimizing the loss function with respect to θ, and i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing a plurality of MHC alleles.

[0391] The presentation likelihood in the presentation model of equation (17) is modeled as a function of the implicit allele-specific presentation likelihood u’ k each corresponding to the likelihood by which the peptide p k h would be presented by an individual MHC allele h. The implicit allele-specific likelihood is different from the allele-specific presentation likelihood of Section VIII.B in that the parameters for the implicit allele-specific likelihood can be learned from a multiple allele setting where the direct relationship between the presented peptide and the corresponding MHC allele is unknown in addition to the single allele setting. Thus, in the multiple allele setting, the presentation model is the peptide p knot only be able to estimate whether it is presented as a whole by a set of MHC alleles H, but also the individual likelihood indicating which MHC allele h is most likely to have presented peptide p k can also be provided. The advantage of this is that the presentation model can generate implicit likelihoods without training data for cells expressing a single MHC allele. TIFF2025094018000060.tif4128

[0392] In one particular implementation referred to throughout the remainder of this specification, r(·) is a function having a range [0,1]. For example, r(·) can be a clip function: r(z)=min(max(z,0),1) where the minimum value between z and 1 is selected as the presentation likelihood u k In another implementation, r(·) is r(z)=tanh(z) which is the hyperbolic tangent function given, and the value of the domain z is 0 or greater.

[0393] VIII.C.5. Example 3.2: Sum Model of Functions In one particular implementation, s(·) is a sum function, and the presentation likelihood is given by summing the implicit presentation likelihoods for each allele. TIFF2025094018000061.tif16128

[0394] In one implementation, the implicit presentation likelihood for each allele of MHC allele h is generated by TIFF2025094018000062.tif7128 such that the presentation likelihood is estimated by TIFF2025094018000063.tif14129.

[0395] According to equation (19), the presentation likelihood that peptide sequence p k will be presented by one or more MHC alleles H is the function g hApply (·) to the encoded version of each peptide sequence p for each MHC allele H to generate the corresponding dependency score for the allele interaction variable, which can be generated by k Each dependency score is first transformed by a function f(·) to generate an implicit per-allele presentation likelihood u'. k h The per-allele likelihood u' k h is combined, and a clipping function is applied to the combined likelihood to clip the value into the range [0,1], and the presentation likelihood that the peptide sequence p k would be presented by the set of MHC alleles H can be generated. The dependency function g h can be in any form of the dependency function g h introduced above in Section VIII.B.1.

[0396] As an example, the likelihood that peptide p h would be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the affine transformation function g(·) is k which can be generated by TIFF2025094018000064.tif7128, where x2 k , x3 k are the allele interaction variables specified for MHC alleles h = 2, h = 3, and θ2, θ3 are the sets of parameters determined for MHC alleles h = 2, h = 3.

[0397] As another example, the likelihood that peptide p k h would be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the network transformation functions g(·), g(·) is w which can be generated by k It can be generated by TIFF2025094018000065.tif7128, where NN2(·) and NN3(·) are network models specified for MHC alleles h = 2 and h = 3, and θ2 and θ3 are sets of parameters determined for MHC alleles h = 2 and h = 3.

[0398] Figure 11 illustrates the generation of the presentation likelihood of peptide p related to MHC alleles h = 2 and h = 3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) receives the allele interaction variable x2 for MHC allele h = 2 k and generates an output NN2(x2 k ), and the network model NN3(·) receives the allele interaction variable x3 for MHC allele h = 3 k and generates an output NN3(x3 k ). Each output is mapped by the function f(·) and combined to generate an estimated presentation likelihood u k .

[0399] In another embodiment, when the prediction is made for the log of the mass spectrometry ion current, r(·) is the log function and f(·) is the exponential function.

[0400] VIII.C.6. Example 3.3: Sum Model of Functions with Allele Non-Interaction Variables In one embodiment, the implicit per-allele presentation likelihood for MHC allele h is generated by TIFF2025094018000066.tif7128 such that the presentation likelihood is generated by TIFF2025094018000067.tif14131, incorporating the influence of allele non-interaction variables on peptide presentation.

[0401] According to equation (21), the presentation likelihood that peptide sequence p k will be presented by one or more MHC alleles H is the function g hApply (·) to the encoded version of each peptide sequence p for each MHC allele H to generate the corresponding dependency score for the allele interaction variable of each MHC allele h. The function g k for the allele non-interaction variable is also applied to the encoded version of the allele non-interaction variable to generate a dependency score for the allele non-interaction variable. The score of the allele non-interaction variable is combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by the function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined, and a clipping function is applied to the combined output to clip the values into the range [0,1] to generate the presentation likelihood that peptide sequence p w will be presented by MHC allele H. The dependency function g k can be in any form of the dependency function g w introduced above in Section VIII.B.3. w

[0402] As an example, the likelihood that peptide p h will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the affine transformation functions g w (·), g k is generated by TIFF2025094018000068.tif7128, where w k is the allele non-interaction variable specified for peptide p k , and θ w is the set of parameters determined for the allele non-interaction variable.

[0403] As another example, the likelihood that peptide p h will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the network transformation functions g w (·), g k is ​Can be generated by TIFF2025094018000069.tif7139, where w k is the allelic interaction variable specified for peptide p k and θ w is a set of parameters determined for the allelic non - interaction variables.

[0404] Figure 12 illustrates the generation of the presentation likelihoods for peptides p w associated with MHC alleles h = 2, h = 3 using exemplary network models NN2(·), NN3(·), and NN k . As shown in Figure 12, the network model NN2(·) receives the allelic interaction variable x2 k for MHC allele h = 2 and generates the output NN2(x2 k ). The network model NN w (·) receives the allelic non - interaction variable w k for peptide p k and generates the output NN w (w k ). The outputs are combined and mapped by the function f(·). The network model NN3(·) receives the allelic interaction variable x3 k for MHC allele h = 3 and generates the output NN3(x3 k ), which is also combined with the output NN w (w w ) of the same network model NN k and mapped by the function f(·). Both outputs are combined to generate the estimated presentation likelihood u k .

[0405] In another embodiment, the implicit per - allele presentation likelihood for MHC allele h is generated by TIFF2025094018000070.tif7128 such that the presentation likelihood is generated by TIFF2025094018000071.tif14128.

[0406] VIII.C.7. Example 4: Quadratic Model In one implementation, s(·) is a quadratic function, and the estimated presentation likelihood u k of the peptide p k is given by TIFF2025094018000072.tif14137, where the element u’ k h is the implicit allele-specific presentation likelihood for the MHC allele h. The values of the set of parameters θ for the implicit allele-specific likelihood can be determined by minimizing the loss function with respect to θ, and i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit allele-specific presentation likelihood can be in any of the forms shown in equations (18), (20), and (22) above.

[0407] In one aspect, the model of equation (23) can imply that there is a possibility that the peptide sequence p k will be presented simultaneously by two MHC alleles, and the presentation by the two HLA alleles is statistically independent.

[0408] According to equation (23), the presentation likelihood that the peptide sequence p k will be presented by one or more MHC alleles H can be generated by combining the implicit allele-specific presentation likelihoods and subtracting from the sum the likelihood that each pair of MHC alleles will present the peptide p k simultaneously so as to generate the presentation likelihood that the peptide sequence p k will be presented by the MHC allele H.

[0409] As an example, the likelihood that the peptide p h will be presented by the HLA alleles h = 2, h = 3 among m = 4 different specified HLA alleles using the affine transformation function g k (·) can be generated by TIFF2025094018000073.tif6128, where x2k , x3 k is an allele interaction variable identified for HLA alleles h = 2 and h = 3, and θ2 and θ3 are sets of parameters determined for HLA alleles h = 2 and h = 3.

[0410] As another example, the network conversion function g h (·), g w (·) is used to calculate the likelihood that peptide p k will be presented by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles, which can be generated by TIFF2025094018000074.tif6137. In the formula, NN2(·) and NN3(·) are network models identified for HLA alleles h = 2 and h = 3, and θ2 and θ3 are sets of parameters determined for HLA alleles h = 2 and h = 3.

[0411] IX. Example 5: Prediction Module The prediction module 320 receives array data and selects candidate neoantigens in the array data using a presentation model. Specifically, the array data may be a DNA sequence, an RNA sequence, and / or a protein sequence extracted from a patient's tumor tissue cells. The prediction module 320 processes the array data into a plurality of peptide sequences p k having 8 - 15 amino acids for MHC-I or 6 - 30 amino acids for MHC-II. For example, the prediction module 320 can process TIFF2025094018000075.tif4128 into three peptide sequences TIFF2025094018000076.tif9128 having 9 amino acids. In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by comparing the array data extracted from the patient's normal tissue cells with the array data extracted from the patient's tumor tissue cells to identify portions having one or more mutations.

[0412] The prediction module 320 applies one or more of the presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation model to candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences having an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences having the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the candidate neoantigens selected for a given patient can be injected into the patient to induce an immune response.

[0413] X. Example 6: Patient Selection Module The patient selection module 324 selects a subset of patients for vaccine treatment based on whether the patients meet the inclusion criteria. In one implementation, the inclusion criteria are determined based on the presentation likelihood of the patient's neoantigen candidates generated by the presentation model. By adjusting the inclusion criteria, the patient selection module 324 can adjust the number of patients receiving vaccine administration based on the presentation likelihood of the patient's neoantigen candidates. Specifically, with strict inclusion criteria, the number of patients treated with the vaccine is smaller, but the proportion of vaccine-treated patients receiving an effective treatment (e.g., one or more tumor-specific neoantigens (TSNAs) are delivered) can be higher. In contrast, with loose inclusion criteria, the number of patients treated with the vaccine is larger, but the proportion of vaccine-treated patients receiving an effective treatment can be lower. The patient selection module 324 changes the inclusion criteria based on the desired balance between the target proportion of patients receiving vaccine administration and the result of the vaccine treatment, the proportion of patients receiving an effective treatment.

[0414] In one embodiment, a corresponding therapeutic subset of v types of neoantigen candidates that can potentially be included in an individualized vaccine for a patient having a vaccine volume v is associated with the patient. In one embodiment, the therapeutic subset for a given patient is the neoantigen candidate having the highest presentation likelihood as determined by a presentation model. For example, if the vaccine can include v = 20 types of epitopes, the vaccine can include the therapeutic subset for each patient having the highest presentation likelihood as determined by the presentation model. However, it is recognized that in other embodiments, the therapeutic subset for a given patient can also be determined based on other methods. For example, the therapeutic subset for a given patient can be randomly selected from the set of neoantigen candidates for that patient, or can be determined based in part on a combination of factors including a prior art model that models the binding affinity or stability of peptide sequences, or the presentation likelihood obtained from the presentation model and affinity or stability information regarding these peptide sequences.

[0415] In one embodiment, the patient selection module 324 determines that a patient meets the inclusion criteria if the tumor mutation burden of the patient is equal to or higher than the minimum mutation burden. The tumor mutation burden (TMB) of a given patient indicates the total number of non-synonymous mutations in the tumor exome. In one embodiment, the patient selection module 324 selects a patient suitable for vaccine treatment if the absolute number of the patient's TMB is equal to or higher than a predetermined threshold. In another implementation, the patient selection module 324 selects a patient suitable for vaccine treatment if the patient's TMB is within the threshold percentile among the TMBs determined for the set of patients.

[0416] In another embodiment, the patient selection module 324 determines that a patient meets the inclusion criteria if the utility score of the patient based on the therapeutic subset of the patient is equal to or higher than the minimum utility score. In one embodiment, the utility score is a measure of the estimated number of presented antigens from the therapeutic subset.

[0417] The estimated number of presented antigens can be predicted by modeling the presentation of neoantigens as a random variable of one or more probability distributions. In one implementation, the utility score of patient i is the expected number of presented neoantigen candidates from a treatment subset, or a specific function thereof. As an example, the presentation of each neoantigen can be modeled as a Bernoulli random variable where the probability of presentation (success) is given by the presentation likelihood of the neoantigen candidate. Specifically, for v types of neoantigen candidates p i1 , u i2 , …, u iv having v types of neoantigen candidates p i1 , p i2 , …, p iv in a treatment subset S i of the treatment subset S, the presentation of the neoantigen candidate p ij is given by the random variable A ij , where TIFF2025094018000077.tif6128The expected number of presented neoantigens is given by the sum of the presentation likelihoods of each neoantigen candidate. In other words, the utility score of patient i is expressed as: TIFF2025094018000078.tif16128The patient selection module 324 selects a subset of patients having a utility score equal to or higher than the minimum utility for vaccine treatment.

[0418] In another implementation, the utility score of patient i is the probability that at least a threshold number of neoantigens k are presented. In one example, the number of presented antigens in a treatment subset S i of neoantigen candidates is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the respective presentation likelihood of the epitope. Specifically, the number of presented antigens of patient i can be given by the random variable N i as: TIFF2025094018000079.tif14128where PBD(·) represents the Poisson binomial distribution. The probability that at least a threshold number of neoantigens k are presented is the number of presented antigens N iis given by an operation with a probability equal to or greater than k. In other words, the utility score of patient i is expressed as follows: The patient selection module 324 selects a subset of patients having a utility score equal to or higher than the minimum utility for the vaccine treatment.

[0419] In another embodiment, the number of neoantigens within a treatment subset for a patient need not be limited to the vaccine volume v, and the patient selection module 324 can select patients using a utility score determined based on any set of candidate neoantigens of that patient. For example, the utility score can be determined based on all mutations or candidate neoantigens identified for that patient. The utility score can be generated, for example, using the methods described with equations (24)-(27), where v is the variable v(i) that depends on patient i, and indicates the total number of mutations or candidate neoantigens identified for that patient.

[0420] In another implementation, the utility score of patient i is the number of neoantigens within a treatment subset S of neoantigen candidates having a binding affinity or predicted binding affinity lower than a fixed threshold (e.g., 500 nM) for one or more HLA alleles of that patient. In one example, the fixed threshold ranges from 1000 nM to 10 nM. Optionally, the utility score may also count only neoantigens detected as expressed by RNA-seq. i In another implementation, the utility score of patient i is the number of neoantigens within a treatment subset S of neoantigen candidates whose binding affinity for one or more HLA alleles of that patient is below the threshold percentile of the binding affinity of a random peptide for that HLA allele.

[0421] i iIt is the number of neoantigens within. In one example, the threshold percentile ranges from the 10th percentile to the 0.1st percentile. Optionally, the utility score may count only the neoantigens detected as being expressed by RNA-seq.

[0422] It should be recognized that the examples of utility scores described with respect to formulas (25) and (27) are merely illustrative, and the patient selection module 324 can also generate utility scores using other statistics or probability distributions.

[0423] XI. Example 7: Neoantigen Load for Immune Checkpoint Inhibitor Therapy and Other Immunotherapies The patient selection module 324 can also select patients who will undergo immune checkpoint inhibitor therapy (e.g., PD-1, CTLA4) or any other immunotherapy for which neoantigen load may be associated with efficacy, using the utility scores defined in section X above. Other immunotherapies include immunostimulants, immune-stimulatory molecule agonists (e.g., CD40), oncolytic viruses (e.g., T-VEC), neoantigen- or other cancer antigen-containing therapeutic vaccines, neoantigen- or other cancer antigen-targeted adoptive cell therapies, tumor microenvironment regulators (e.g., TGFβ), or any combination of these with immune checkpoint inhibitors.

[0424] For example, in some embodiments, the immune activator is an agent that blocks the signaling of inhibitory receptors of immune cells or their ligands. In some embodiments, the inhibitory receptor or ligand is selected from CTLA-4, PD-1, PD-L1, LAG-3, Tim3, TIGIT, neuritin, BTLA, KIR, and combinations thereof. In some aspects, the agent is selected from anti-PD-1 antibodies (e.g., pembrolizumab or nivolumab), anti-PD-L1 antibodies (e.g., atezolizumab), anti-CTLA-4 antibodies (e.g., ipilimumab), and combinations thereof. In some aspects, the agent is pembrolizumab. In some aspects, the agent is nivolumab. In some aspects, the agent is atezolizumab.

[0425] In some embodiments, the therapeutic agent is an agent that inhibits the interaction between PD-1 and PD-L1. In some aspects, additional agents that inhibit the interaction between PD-1 and PD-L1 are selected from antibodies, peptidomimetics, and small molecules. In some aspects, additional agents that inhibit the interaction between PD-1 and PD-L1 are selected from pembrolizumab, nivolumab, atezolizumab, avelumab, durvalumab, BMS-936559, sulfamethoxine 1, and sulfamethizole 2. In some embodiments, the additional agent that inhibits the interaction between PD-1 and PD-L1 is any therapeutic agent well known in the art having such activity, such as described in Weinmann et al., Chem Med Chem, 2016, 14:1576 (DOI: 10.1002 / cmdc.201500566), the entire contents of which are incorporated herein by reference.

[0426] In some embodiments, the immune activator is an agonist of a co-stimulatory receptor of immune cells. In some aspects, the co-stimulatory receptor is selected from OX40, ICOS, CD27, CD28, 4-1BB, and CD40. In some embodiments, the agonist is an antibody.

[0427] In some embodiments, the immunostimulant is a cytokine. In some aspects, the cytokine is selected from IL-2, IL-5, IL-7, IL-12, IL-15, IL-21, and combinations thereof.

[0428] In some embodiments, the immunostimulant is an oncolytic virus. In some aspects, the oncolytic virus is selected from herpes simplex virus, vesicular stomatitis virus, adenovirus, Newcastle disease virus, vaccinia virus, and Maraba virus.

[0429] In some embodiments, the immunostimulant is a T cell having a chimeric antigen receptor (CAR-T cell). In some embodiments, the immunostimulant is a bispecific or multispecific T cell-engaging antibody. In some embodiments, the immunostimulant is an anti-TGF-β antibody. In some embodiments, the immunostimulant is a TGF-β trap.

[0430] In some embodiments, the therapeutic agent is a vaccine against a tumor antigen. If the antigen is present in the tumor to be treated by the methods provided herein, the vaccine can be targeted to any suitable antigen. In some aspects, the tumor antigen is a tumor antigen that is overexpressed compared to its expression level in normal tissue. In some aspects, the tumor antigen is selected from cancer testis antigens, differentiation antigens, NY-ESO-1, MAGE-A1, MART, and combinations thereof. In some embodiments, the therapeutic agent is a vaccine against one or more neoantigens. The neoantigens in the vaccine can be identified by the methods provided herein.

[0431] Specifically, the patient selection module 324 determines a neoantigen load that indicates the total expected number of presented neoantigens for each patient. A checkpoint inhibitor therapy can be administered to a patient having a neoantigen load that meets the inclusion criteria. For example, a therapy can be administered to a patient having a neoantigen load that exceeds a predetermined threshold. In one embodiment, the neoantigen load is the utility score shown in Section XI, in which case v is the total number of mutations or candidate neoantigens identified for the patient, rather than a subset of candidate antigens for that patient.

[0432] If the neoantigen load in a particular tumor is high relative to the median, it may indicate that a subject having that tumor is likely to respond to treatment with a checkpoint inhibitor such as anti-CTLA4, anti-PD1, and / or anti-PDL1. For example, neoantigens are generally presented on the surface of tumor cells and are more likely to be recognized by T cells having higher activity against tumors after checkpoint inhibitor therapy, so the neoantigen load can be a better indicator of the effect of checkpoint inhibitors compared to the mutation load.

[0433] In another embodiment, the patient selection module 324 can use a utility score generated from a combination of one or more of the following characteristics: predicted HLA class I neoantigen load, predicted HLA class II neoantigen load, and tumor mutation load. The predicted HLA class I neoantigen load for a patient is the neoantigen load for the set of the patient's class I HLA alleles and represents the total expected number of neoantigens presented on the patient's class I HLA alleles. The predicted HLA class II neoantigen load for a patient is the neoantigen load for the set of the patient's class II HLA alleles and represents the total expected number of neoantigens presented on the patient's class II HLA alleles. For example, the utility score can be calculated as f(class I neoantigen load, class II neoantigen load, tumor mutation load; b) (where f(·) is a function parameterized by a set of machine-learned parameters b). The set of machine-learned parameters b can depend on the tumor type (e.g., b can be different for melanoma and non-small cell lung cancer).

[0434] In another embodiment, the patient selection module 324 can use a utility score incorporating information on immunogenic tumor antigens other than neoantigens. Examples of immunogenic tumor antigens other than neoantigens include cancer-germline antigens (CGA, such as MAGEA3), differentiation antigens (such as tyrosinase), and antigens overexpressed in tumors (such as CEA). The expression levels of these antigens can be determined using at least tumor RNA sequencing data, and the expected number of HLA class I or class II epitopes from these genes presented by the HLA alleles of the patient's tumor can be determined by applying a presentation model to each peptide from the set of tumor antigens using the RNA sequencing data for each gene. These presentation likelihoods can be incorporated into a utility score calculated as f(class I neoantigen load, class II neoantigen load, tumor mutation load, class I non-neoantigen tumor antigen load, class II non-neoantigen tumor antigen load; b) (where f(·) is a function parameterized by a set of machine-learned parameters b). The set of machine-learned parameters b can depend on the tumor type (e.g., b can be different for melanoma and non-small cell lung cancer).

[0435] When the utility score is higher, this indicates a tumor that presents more HLA epitopes recognized as foreign or non-self by the patient's immune system. Patients with tumors that present more non-self HLA epitopes are more likely to respond to checkpoint inhibitors or other immunotherapies because such tumors are more likely to be recognized by T cells with higher activity against the post-immunotherapy tumor.

[0436] The usefulness score described in section X above can also be adapted to select patients for treatment by adoptive cell therapy (e.g., expanded TIL, CAR-T, or engineered TCR) by using f(class I neoantigen load, class II neoantigen load, tumor mutation burden, class I non-neoantigen tumor antigen load, class II non-neoantigen tumor antigen load), provided that class I and class II neoantigens and non-neoantigens are considered only as those present or predicted to be present during adoptive immunotherapy. For example, when engineering a TCR therapy against a single neoantigen epitope, f can be reduced to the presentation likelihood of that single epitope.

[0437] XII. Example 8: Experimental results showing exemplary patient selection performance The validity of patient selection described in section X is verified by performing patient selection in a set of simulated patients associated with each of a test set of simulated neoantigen candidates, where it is known that a subset of the simulated neoantigens presented in the mass spectrometry data is presented. Specifically, each simulated neoantigen candidate in the test set is associated with a label indicating whether that neoantigen is presented in the mass spectrometry data set of multiple allele JY cell lines HLA-A*02:01 and HLA-B*07:02 from the Bassani-Sternberg data set (data set "D1") (data can be viewed at www.ebi.ac.uk / pride / archive / projects / PXD0000394). As described in detail below with FIG. 13A, a large number of neoantigen candidates for the simulated patients are sampled from the human proteome based on the known frequency distribution of mutation burdens in non-small cell lung cancer (NSCLC) patients.

[0438] For each allele presentation model for the same HLA allele, a training set is used which is a subset of the mass spectrometry data of the single alleles HLA-A*02:01 and HLA-B*07:02 from the IEDB dataset (dataset "D2") (the data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Specifically, for each allele presentation model, the N-terminal and C-terminal flanking sequences are used as allele non-interacting variables, and the network-dependent functions g h (·) and g w (·) together with the expit function f(·) are incorporated into the per-allele model shown in Equation (8). The presentation model for allele HLA-A*02:01 generates the presentation likelihood that a specific peptide is presented on allele HLA-A*02:01, with the peptide sequence as the allele interacting variable and the N-terminal and C-terminal flanking sequences as the allele non-interacting variables. The presentation model for allele HLA-B*07:02 generates the presentation likelihood that a specific peptide is presented on allele HLA-B*07:02, with the peptide sequence as the allele interacting variable and the N-terminal and C-terminal flanking sequences as the allele non-interacting variables.

[0439] As disclosed with reference to FIGS. 13A - 13G in the following examples, different models such as the presentation model trained for peptide binding prediction and prior art models are applied to a test set of neoantigen candidates for each simulated patient to identify different treatment subsets for the patient based on the prediction. Patients who meet the inclusion criteria for vaccine treatment are selected and associated with an individualized vaccine containing epitopes in the patient's treatment subset. The size of the treatment subset varies depending on the different vaccine volumes. No overlap is introduced between the training set used to train the presentation model and the test set of simulated neoantigen candidates.

[0440] In the following example, the proportion of selected patients having at least a certain number of presented neoantigens among the epitopes contained in the vaccine is analyzed. This statistic indicates the effectiveness of the simulated vaccine in delivering potential neoantigens that induce an immune response in patients. Specifically, the simulated neoantigens within a certain test set are presented if the neoantigen is presented in the mass spectrometry dataset D2. A high proportion of patients having the presented neoantigens indicates the likelihood of success of treatment with the neoantigen vaccine by inducing an immune response.

[0441] XII.A. Example 8A: Frequency Distribution of Tumor Mutation Burden in NSCLC Cancer Patients Figure 13A shows the specimen frequency distribution of mutation burden in NSCLC patients. The mutation burden and mutations in different tumor types including NSCLC can be found, for example, in the cancer genome atlas (TCGA) (https: / / cancergenome.nih.gov). The X-axis represents the number of non-synonymous mutations per patient, and the Y-axis represents the proportion of specimen patients having a certain number of non-synonymous mutations. The specimen frequency distribution in Figure 13A shows a range of 3 to 1786 mutations, and 30% of the patients have fewer than 100 mutations. Although not shown in Figure 13A, the mutation burden is higher in smokers compared to non-smokers, and studies have shown that the mutation burden can be a strong indicator of neoantigen load in patients.

[0442] As introduced at the beginning of section XI above, a test set of neoantigen candidates is associated with each of the simulated number of patients. The test set for each patient is the mutation burden m from the frequency distribution shown in Figure 13A for each patient. iIt is generated by sampling. For each mutation, a 21-mer peptide sequence from the human proteome is randomly selected to represent the mutated sequence to be simulated. A test set of neoantigen candidate sequences is generated for patient i by identifying the peptide sequences of each (8, 9, 10, 11)-mer over the mutations within the 21-mer. To each neoantigen candidate, a label is associated indicating whether the neoantigen candidate sequence is present in the mass spectrometry D1 dataset. For example, a label "1" can be associated with the neoantigen candidate sequences present in dataset D1, and a label "0" with the sequences not present in dataset D1. As described in more detail below, FIGS. 13B-13G show the experimental results of patient selection based on the presented neoantigens of patients in the test set.

[0443] XII.B. Example 8B: Ratio of selected patients with neoantigen presentation based on incorporation criteria of tumor mutation burden FIG. 13B shows the number of presented neoantigens in the simulated vaccine for patients selected based on the incorporation criteria of whether the patient meets the minimum tumor mutation burden. Identify the ratio of selected patients having at least a specific number of presented neoantigens in the corresponding test.

[0444] In FIG. 13B, the x-axis shows the proportion of patients excluded from vaccine treatment based on tumor mutation burden, indicated by the label "minimum number of mutations". For example, the data point at "minimum number of mutations" 200 indicates that the patient selection module 324 selected only a subset of simulated patients with a tumor mutation burden of at least 200 mutations. As another example, the data point at "minimum number of mutations" 300 indicates that the patient selection module 324 selected a lower proportion of simulated patients with at least 300 mutations. The y-axis shows the proportion of selected patients associated with at least a certain number of presented neoantigens within a test set without vaccine volume v. Specifically, the upper plot shows the proportion of selected patients presenting at least one type of neoantigen, the middle plot shows the proportion of selected patients presenting at least two types of antigens, and the lower plot shows the proportion of selected patients presenting at least three antigens.

[0445] As shown in FIG. 13B, the proportion of selected patients with presented neoantigens increases significantly as the tumor mutation burden increases. This indicates that tumor mutation burden as an inclusion criterion can be effective in selecting patients in whom neoantigen vaccines are likely to induce an effective immune response.

[0446] XII.C. Example 8C: Comparison of Neoantigen Presentation in Vaccines Identified by the Presentation Model and Prior Art Models FIG. 13C compares the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing treatment subsets identified based on the presentation model and selected patients associated with vaccines containing treatment subsets identified by prior art models. The left plot assumes a limited vaccine volume of v = 10, and the right plot assumes a limited vaccine volume of v = 20. Patients are selected based on a utility score indicating the expected number of presented neoantigens.

[0447] In FIG. 13C, the solid line indicates patients associated with a vaccine that includes a treatment subset identified based on the presentation models for alleles HLA-A*02:01 and HLA-B*07:02. The treatment subset for each patient is identified by applying each of the presentation models to the sequences within the test set and identifying v types of neoantigen candidates that have the highest presentation likelihood. The dotted line indicates patients associated with a vaccine that includes a treatment subset identified based on the prior art model NETMHCpan for the single allele HLA-A*02:01. Details of the implementation for NETMHCpan are provided at http: / / www.cbs.dtu.dk / services / NetMHCpan. The treatment subset for each patient is identified by applying the NETMHCpan model to the sequences within the test set and identifying v types of neoantigen candidates that have the highest estimated binding affinity. The x-axis of both graphs indicates the ratio of patients excluded from vaccine treatment based on an expected utility score that indicates the expected number of presented neoantigens in the treatment subset identified based on the presentation model. The expected utility score is determined as described in relation to Equation (25) in Section X. The y-axis indicates the ratio of selected patients who present at least a specific number of neoantigens (1, 2, or 3 types of neoantigens) included in the vaccine.

[0448] As shown in FIG. 13C, patients associated with a vaccine that includes a treatment subset based on the presentation model are administered a vaccine that includes presented neoantigens at a significantly higher rate than patients associated with a vaccine that includes a treatment subset based on the prior art model. For example, as shown in the graph on the right, 80% of the selected patients associated with a vaccine based on the presentation model are administered at least one presented neoantigen in the vaccine, compared to only 40% of the selected patients associated with a vaccine based on the prior art model. These results indicate that the presentation model described herein is effective in selecting neoantigen candidates for a vaccine that is likely to induce an immune response for treating tumors.

[0449] XII.D. Example 8D: Influence of HLA Coverage on Neoantigen Presentation of Vaccines Identified by the Presentation Model Figure 13D compares the number of neoantigens presented in simulated vaccines between selected patients associated with vaccines containing treatment subsets identified based on the single-allele presentation model for HLA-A*02:01 and selected patients associated with vaccines containing treatment subsets identified based on both allele presentation models for HLA-A*02:01 and HLA-B*07:02. The vaccine volume is set to v = 20 epitopes. For each experiment, patients are selected based on the expected utility score determined based on different treatment subsets.

[0450] In Figure 13D, the solid line indicates patients associated with vaccines containing treatment subsets based on both presentation models for the HLA alleles HLA-A*02:01 and HLA-B*07:02. The treatment subset for each patient is identified by applying each of the presentation models to the sequences within the test set and identifying v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with vaccines containing treatment subsets based on the single presentation model for the HLA allele HLA-A*02:01. The treatment subset for each patient is identified by applying the presentation model for only a single HLA allele to the sequences within the test set and identifying v neoantigen candidates with the highest presentation likelihood. In the solid line plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score for the treatment subset identified by both presentation models. In the dotted line plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score for the treatment subset identified by the single presentation model. The y-axis indicates the proportion of selected patients presenting at least a certain number of neoantigens (1, 2, or 3 neoantigens).

[0451] As shown in FIG. 13D, patients associated with a vaccine containing a treatment subset identified from the presentation models for both HLA alleles present neoantigens at a significantly higher rate than patients associated with a vaccine containing a treatment subset identified by a single presentation model. These results demonstrate the importance of establishing a presentation model with high HLA coverage.

[0452] XII.E. Example 8E: Comparison of Neoantigen Presentation in Patients Selected by Tumor Mutation Burden and Expected Number of Presented Neoantigens FIG. 13E compares the number of neoantigens presented in a simulated vaccine between patients selected based on tumor mutation burden and patients selected by an expected utility score. The expected utility score is determined based on the treatment subset identified by a presentation model having a size of v = 20 epitopes.

[0453] In FIG. 13E, the solid line indicates patients selected based on the expected utility score associated with a vaccine containing the treatment subset identified by the presentation model. The treatment subset for each patient is identified by applying each of the presentation models to the sequences within the test set and identifying v = 20 neoantigen candidates having the highest presentation likelihood. The treatment utility score is determined based on the presentation likelihood of the treatment subset identified based on Equation (25) in Section X. The dotted line indicates patients selected based on the tumor mutation burden associated with a vaccine containing the treatment subset identified by the presentation model. The x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score of the solid line plot and the proportion of patients excluded based on the tumor mutation burden of the dotted line plot. The y-axis indicates the proportion of selected patients to whom a vaccine containing at least a specific number of presented neoantigens (1, 2, or 3 neoantigens) is administered.

[0454] As shown in FIG. 13E, patients selected based on the expected utility score are administered vaccines containing presented neoantigens at a higher rate than patients selected based on tumor mutation burden. However, patients selected based on tumor mutation burden are administered vaccines containing presented neoantigens at a higher rate than patients who are not selected. Therefore, tumor mutation burden is an effective patient selection criterion in effective neoantigen vaccine therapy, but the expected utility score is more effective.

[0455] XIII. Exemplary Computer FIG. 14 illustrates an exemplary computer 1400 for implementing the entities shown in FIGS. 1 and 3. Computer 1400 includes at least one processor 1402 coupled to a chipset 1404. Chipset 1404 includes a memory controller hub 1420 and an input / output (I / O) controller hub 1422. Memory 1406 and graphics adapter 1412 are coupled to memory controller hub 1420, and display 1418 is coupled to graphics adapter 1412. Storage device 1408, input device 1414, and network adapter 1416 are coupled to I / O controller hub 1422. Other embodiments of computer 1400 have different architectures.

[0456] The memory device 1408 is a non - transitory computer - readable storage medium such as a hard drive, a compact disc read - only memory (CD - ROM), a DVD, or a solid - state memory device. The memory 1406 holds instructions and data used by the processor 1402. The input interface 1414 is a touch - screen interface, a mouse, a trackball, or other type of pointing device, a keyboard, or some combination thereof, and is used to input data into the computer 1400. In some embodiments, the computer 1400 may be configured to receive input (e.g., commands) from the input interface 1414 via gestures from a user. The graphics adapter 1412 displays images and other information on the display 1418. The network adapter 1416 couples the computer 1400 to one or more computer networks.

[0457] The computer 1400 is adapted to perform computer program modules for providing the functionality described herein. As used herein, the term "module" refers to computer program logic used to provide a particular functionality. Thus, a module can be executed in hardware, firmware, and / or software. In one embodiment, the program modules are stored on the memory device 1408, loaded into the memory 1406, and executed by the processor 1402.

[0458] The type of computer 1400 used by the entity of FIG. 1 can vary according to the embodiments and processing power required by the entity. For example, the presentation specific system 160 can be launched on a single computer 1400 or on multiple computers 1400 that communicate with each other through a network, such as in a server farm. The computer 1400 may lack some of the above components, such as the graphics adapter 1412 and the display 1418.

[0459] References TIFF2025094018000081.tif223158TIFF2025094018000082.tif238159TIFF2025094018000083.tif238159TIFF2025094018000084.tif233160TIFF2025094018000085.tif52153

[0460] Array information SEQUENCE LISTING <110> GRITSTONE BIO, INC. <120> NEOANTIGEN IDENTIFICATION, MANUFACTURE, AND USE <150> US 62 / 517,786 <151> 2017-06-09 <160> 22 <170> PatentIn version 3.5 <210> 1 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 1 Tyr Val Tyr Val Ala Asp Val Ala Ala Lys 1 5 10 <210> 2 <211> 17 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 2 Tyr Glu Met Phe Asn Asp Lys Ser Gln Arg Ala Pro Asp Asp Lys Met 1 5 10 15 Phe <210> 3 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 3 Tyr Glu Met Phe Asn Asp Lys Ser Phe 1 5 <210> 4 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (3)..(3) <223> Pyrrolysine <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <400> 4 His Arg Xaa Glu Ile Phe Ser His Asp Phe Xaa 1 5 10 <210> 5 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Ile or Leu <220> <221> MOD_RES <222> (5)..(5) <223> Ile or Leu <220> <221> MOD_RES <222> (7)..(7) <223> Pyrrolysine <400> 5 Phe Xaa Ile Glu Xaa Phe Xaa Glu Ser Ser 1 5 10 <210> 6 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Pyrrolysine <400> 6 Asn Glu Ile Xaa Arg Glu Ile Arg Glu Ile 1 5 10 <210> 7 <211> 27 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (1)..(1) <223> Ile or Leu <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <220> <221> MOD_RES <222> (15)..(15) <223> Selenocysteine <220> <221> MOD_RES <222> (21)..(21) <223> Ile or Leu <220> <221> MOD_RES <222> (27)..(27) <223> Ile or Leu <400> 7 Xaa Phe Lys Ser Ile Phe Glu Met Met Ser Xaa Asp Ser Ser Xaa Ile 1 5 10 15 Phe Leu Lys Ser Xaa Phe Ile Glu Ile Phe Xaa 20 25 <210> 8 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (11)..(11) <223> Pyrrolysine <400> 8 Lys Asn Phe Leu Glu Asn Phe Ile Glu Ser Xaa Phe Ile 1 5 10 <210> 9 <211> 15 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Pyrrolysine <220> <221> MOD_RES <222> (14)..(14) <223> Ile or Leu <400> 9 Phe Xaa Glu Ile Phe Asn Asp Lys Ser Leu Asp Lys Phe Xaa Ile 1 5 10 15 <210> 10 <211> 16 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (16)..(16) <223> Ile or Leu <400> 10 Gln Cys Glu Ile Xaa Trp Ala Arg Glu Phe Leu Lys Glu Ile Gly Xaa 1 5 10 15 <210> 11 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Selenocysteine <400> 11 Phe Ile Glu Xaa His Phe Trp Ile 1 5 <210> 12 <211> 12 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (7)..(7) <223> Ile or Leu <220> <221> MOD_RES <222> (10)..(10) <223> Selenocysteine <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <400> 12 Phe Glu Trp Arg His Arg Xaa Thr Arg Xaa Xaa Arg 1 5 10 <210> 13 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Ile or Leu <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (8)..(8) <223> Ile or Leu <400> 13 Gln Ile Glu Xaa Xaa Glu Ile Xaa Glu 1 5 <210> 14 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <400> 14 Gln Cys Glu Ile Xaa Trp Ala Arg Glu 1 5 <210> 15 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Ile or Leu <220> <221> MOD_RES <222> (9)..(9) <223> Pyrrolysine <220> <221> MOD_RES <222> (11)..(11) <223> Ile or Leu <400> 15 Phe Xaa Glu Leu Phe Ile Ser Asx Xaa Ser Xaa Phe Ile Glu 1 5 10 <210> 16 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (9)..(9) <223> Ile or Leu <400> 16 Ile Glu Phe Arg Xaa Glu Ile Phe Xaa Glu Phe 1 5 10 <210> 17 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (9)..(9) <223> Ile or Leu <400> 17 Ile Glu Phe Arg Xaa Glu Ile Phe Xaa 1 5 <210> 18 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Pyrrolysine <220> <221> MOD_RES <222> (8)..(8) <223> Ile or Leu <400> 18 Glu Phe Arg Xaa Glu Ile Phe Xaa Glu 1 5 <210> 19 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (3)..(3) <223> Pyrrolysine <220> <221> MOD_RES <222> (7)..(7) <223> Ile or Leu <400> 19 Phe Arg Xaa Glu Ile Phe Xaa Glu Phe 1 5 <210> 20 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 20 Tyr Glu Met Phe Asn Asp Lys Ser Phe Gln Arg Ala Pro Asp Asp Lys 1 5 10 15 Met Phe <210> 21 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (6)..(6) <223> Selenocysteine <220> <221> MOD_RES <222> (7)..(8) <223> Pyrrolysine <400> 21 Phe Glu Gly Arg Lys Xaa Xaa Xaa Ile 1 5 <210> 22 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Ile or Leu <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (7)..(7) <223> Ile or Leu <220> <221> MOD_RES <222> (8)..(8) <223> Pyrrolysine <220> <221> MOD_RES <222> (10)..(10) <223> Ile or Leu <220> <221> MOD_RES <222> (14)..(14) <223> Pyrrolysine <400> 22 Pro Xaa Phe Ile Xaa Glu Xaa Xaa Ile Xaa Gly Glu Ile Xaa 1 5 10

Claims

1. 1. A method for identifying a subset of patients suitable for treatment, comprising: obtaining, for each patient, at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from the patient's tumor cells and normal cells, wherein the tumor nucleotide sequencing data is used to obtain a peptide sequence for each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells to the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen for the patient includes at least one alteration that makes it different from a corresponding wild-type parent peptide sequence identified from the patient's normal cells; generating, for each patient, a set of numerical presentation likelihoods for the set of neoantigens for the patient by inputting the peptide sequences of each of the set of neoantigens into a machine-learned presentation model, each presentation likelihood representing the likelihood that a corresponding neoantigen will be presented by one or more MHC alleles on the surface of tumor cells of the patient, the set of presentation likelihoods having been determined based at least on mass spectrometry data; identifying for each patient one or more neoantigens from said set of neoantigens for said patient; determining for each patient a utility score indicative of the estimated number of neoantigens presented on the surface of tumor cells of said patient as determined by the corresponding presentation likelihood for said one or more neoantigens for said patient; selecting a subset of patients suitable for treatment, wherein each patient in said subset of patients is associated with a utility score that meets predetermined inclusion criteria; The method comprising:

2. 2. The method of claim 1, wherein identifying the one or more neoantigens for the patient comprises selecting a subset of neoantigens in the set of neoantigens for the patient.

3. 3. The method of claim 2, wherein said subset of neoantigens are neoantigens having the highest presentation likelihood among said set of presentation likelihoods for said patient.

4. 2. The method of claim 1, further comprising treating each patient in the selected subset of patients with a corresponding neoantigen vaccine that includes at least one of the one or more neoantigens identified for the patient.

5. 2. The method of claim 1, further comprising identifying, for each patient in the selected subset of patients, one or more T cells or T cell receptors that are antigen-specific for at least one of the one or more neoantigens identified for the patient.

6. 2. The method of claim 1, wherein identifying one or more neoantigens for the patient comprises selecting the entire set of identified neoantigens for the patient.

7. 7. The method of claim 6, further comprising administering checkpoint inhibitor therapy to each patient in said selected subset of patients.

8. 2. The method of claim 1, wherein selecting a subset of patients suitable for treatment comprises selecting a subset of patients having a tumor mutational burden (TMB) higher than a minimum threshold, the TMB of a patient indicating the number of neoantigens in a set of neoantigens associated with that patient.

9. Selecting the subset of patients suitable for treatment is key. Selecting a subset of patients having a utility score above a minimum threshold The method of claim 1 , comprising:

10. 2. The method of claim 1, wherein the utility score is the sum of the presentation likelihoods for each neoantigen within the identified subset of neoantigens for the patient.

11. The method of claim 1 , wherein the utility score is the probability that the number of presented neoantigens among the one or more identified neoantigens for the patient is above a minimum threshold.

12. The machine-learned presentation model is a label obtained by mass spectrometry to measure the presence of a peptide bound to at least one MHC allele identified as being present in at least one of the plurality of samples; a training peptide sequence including information on a set of amino acids constituting the training peptide sequence and positions of the amino acids within the training peptide sequence; at least one MHC allele associated with said training peptide sequence; a plurality of parameters determined based at least on the training data set, the plurality of parameters including: A function that represents the relationship between the peptide sequence and the likelihood of presentation based on the plurality of parameters. The method of claim 1 , comprising:

13. The training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of the isolated peptides; and (b) data relating to measurements of peptide-MHC binding stability for at least one of the isolated peptides; The method of claim 12 , further comprising at least one of:

14. The set of numerical likelihoods is (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; The method of claim 1 , further characterized by characteristics including at least one of:

15. 2. The method of claim 1, wherein the set of presentation likelihoods is further identified by at least the expression level of the one or more MHC alleles in the subject as measured by RNA-seq or mass spectrometry.

16. The set of presentation likelihoods is (a) the predicted affinity between neoantigens within said set of neoantigens and said one or more MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex. The method of claim 1 , further characterized by characteristics including at least one of:

17. inputting the peptide sequence into the machine-learned representation model; applying the machine learning presentation model to the peptide sequence of each neoantigen to generate, for each of the one or more MHC alleles, a dependency score indicating whether the MHC allele presents the neoantigen based on a particular amino acid at a particular position of the peptide sequence. The method of claim 1 , comprising:

18. inputting the peptide sequence into the machine-learned representation model; transforming the dependency scores to generate, for each MHC allele, a corresponding per-allele likelihood that indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining said per-allele likelihoods to generate a presentation likelihood for said neoantigen.

20. The method of claim 17, comprising:

19. 20. The method of claim 18, wherein transforming the dependency score models presentation of the neoantigen as mutually exclusive across the one or more class MHC alleles.

20. inputting the peptide sequence into the machine-learned representation model; Transforming the combination of dependency scores to generate a representation likelihood. and transforming the combination of dependency scores models presentation of the neoantigen as interfering between the one or more MHC alleles.

Citation Information

Patent Citations

  • Compositions and Methods of Individualized Neoplastic Vaccines

    JP2016518355A