Identification of neoantigens using hotspots

An optimized sequencing and machine-learned model for predicting neoantigen presentation enhances the positive predictive value, addressing inefficiencies in current methods by accurately identifying therapeutically relevant neoantigens for personalized cancer vaccines and T cell therapies.

JP7712970B2Active Publication Date: 2025-07-24GRITSTONE BIO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023018098
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-10-10
Filing Date
2023-02-09
Publication Date
2025-07-24
Estimated Expiration
2038-10-10

AI Technical Summary

Technical Problem

Current methods for identifying neoantigens and neoantigen-recognizing T cells in cancer treatment are time-consuming, laborious, or not sufficiently accurate, leading to low positive predictive value (PPV) and inefficiencies in neoantigen-based vaccine and T cell therapy, with many predicted peptides not being presented on tumor surfaces and existing approaches missing somatic mutations or promoting autoimmunity.

Method used

An optimized approach using next-generation sequencing and a machine-learned presentation model that co-models peptide-allele mapping and motif analysis to predict neoantigen presentation likelihood, considering k-mer blocks and MHC alleles, enhancing the positive predictive value (PPV) for identifying therapeutically relevant neoantigens.

Benefits of technology

The method significantly improves the identification of neoantigens with higher confidence, enabling more efficient and cost-effective personalized cancer vaccines and T cell therapies by reducing the number of peptides that need to be screened, thus accelerating the progression towards effective antigen-targeted immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712970000144
    Figure 0007712970000144
  • Figure 0007712970000145
    Figure 0007712970000145
  • Figure 0007712970000146
    Figure 0007712970000146
Patent Text Reader

Abstract

Methods are provided for identifying neoantigens that are likely to be presented on the surface of tumor cells in a subject. [Solution] Peptide sequences of tumor neoantigens are obtained by sequencing a subject's tumor cells. Each peptide sequence of these neoantigens is associated with one or more k-mer blocks from a plurality of k-mer blocks in the subject's nucleotide sequencing data. The peptide sequences and associated k-mer blocks are input into a machine-learned presentation model to generate presentation likelihoods for the tumor neoantigens, each representing the likelihood that the neoantigen will be presented on the surface of the subject's tumor cells by an MHC allele. A subset of neoantigens is selected based on the presentation likelihoods, returning the set of selected neoantigens.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cancer vaccines and T cell therapies based on tumor-specific neoantigens are extremely promising as next-generation personalized cancer immunotherapies 1~3 。Cancers with a high genetic mutation load, such as non-small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies because they are relatively likely to generate neoantigens 4,5 。Initial evidence indicates that vaccination with neoantigens induces T cell responses 6 and that T cell therapies targeting neoantigens can cause tumor regression in selected patients 7 。Both MHC class I and MHC class II influence T cell responses 70~71 。

[0002] However, the identification of neoantigens and neoantigen-recognizing T cells is a central challenge in assessing tumor responses 77,110 examining tumor evolution 111 and designing next-generation personalized therapies 112 。Current neoantigen identification methods are either time-consuming and laborious 84,96 or not sufficiently accurate 87,91-93 。Neoantigen-recognizing T cells are a major component of TILs 84,96,113,114 and circulate in the peripheral blood of cancer patients 107 as shown in recent years, but current methods for identifying neoantigen-reactive T cells have some combination of the following limitations: (1) they rely on clinically difficult-to-obtain specimens such as TILs 97,98 or leukapheresis 107 ; (2) they require screening of impractically large peptide libraries; or (3) they rely on MHC alleles that are only practically available for a small number of MHC alleles

[0003] Furthermore, initial methods incorporating mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides have been proposed8 However, in these proposed methods, many steps other than gene expression and MHC binding (e.g., TAP transport, proteasome cleavage, MHC binding, transport of the peptide-MHC complex to the cell surface, and / or recognition of MHC-I by the TCR; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsin), competition with CLIP peptides for HLA binding catalyzed by HLA-DM, transport of the peptide-MHC complex to the cell surface, and / or recognition of MHC-II by the TCR) are included 9 The entire epitope generation process cannot be modeled. Therefore, existing methods tend to have the problem that the positive predictive value (PPV) is low (Figure 1A).

[0004] Indeed, analysis of peptides presented by tumor cells, performed by multiple groups, has shown that less than 5% of the peptides predicted to be presented using gene expression and MHC binding affinity are found on MHCs on the tumor surface 10,11 (Figure 1B). This low correlation between binding prediction and MHC presentation is further indicated by the lack of improvement in the prediction accuracy of neoantigens restricted to binding for checkpoint inhibitor response with respect to the number of mutations alone 12 .

[0005] Such a low positive predictive value (PPV) of existing methods for predicting cues presents problems in neoantigen-based vaccine design and also in neoantigen-based T cell therapy. If a vaccine is designed using a low PPV prediction, it is likely that therapeutic neoantigens will be administered to a minority of patients, and even fewer patients will receive multiple neoantigens (even assuming that all of the presented peptides are immunogenic). Similarly, if therapeutic T cells are designed based on a low PPV prediction, it is likely that T cells reactive to tumor neoantigens will be administered to a minority of patients, and the time and physical resource costs of identifying predicted neoantigens using downstream assays after prediction can become prohibitively high. Thus, current methods of neoantigen vaccination and T cell therapy are likely to be ineffective in a significant number of subjects with tumors (Figure 1C).

[0006] Furthermore, previous approaches have generated candidate neoantigens using only cis-acting mutations, and have largely not considered additional sources of neoORFs, including mutations in splicing factors that occur in multiple tumor types and lead to aberrant splicing in many genes 13 , and mutations that create or remove protease cleavage sites.

[0007] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard approaches to tumor analysis may erroneously promote sequence artifacts or germline polymorphisms as neoantigens, leading to inefficient use of vaccine capacity or risk of autoimmunity, respectively. SUMMARY OF THE INVENTION

[0008] Disclosed herein is an optimized approach for identifying and selecting neoantigens for personalized cancer vaccines, for T cell therapy, or for both. First, efforts are made towards an optimized tumor exome and transcriptome analysis approach for identifying neoantigen candidates using next-generation sequencing (NGS) of neoantigens (NGS). These methods are based on the standard approach of tumor analysis by NGS such that the most sensitive and specific neoantigen candidates are developed across all classes of genomic changes. Second, a novel approach to high positive predictive value (PPV) neoantigen selection is provided to overcome the problem of specificity and to make neoantigens developed for vaccine addition and / or as targets for T cell therapy more likely to induce anti-tumor immunity. These approaches include, depending on the embodiment, a trained statistical regression or non-linear deep learning model that co-models peptide - allele mapping, and motif per allele for peptides of multiple lengths that share statistical power across peptides of different lengths. These deep learning models also utilize parameters that describe the presence or absence of presentation hotspots within k-mer blocks associated with the peptide sequence in determining the presentation likelihood of the peptide. In particular, the non-linear deep learning model can be designed and trained to treat different MHC alleles within the same cell as independent, thus solving the problems associated with linear models where linear models interfere with each other. Finally, further concerns in the design and manufacture of neoantigen-based personalized vaccines and in the manufacture of personalized neoantigen-specific T cells for T cell therapy are addressed.

[0009] The models disclosed herein outperform by up to an order of magnitude the performance of state-of-the-art prediction tools trained on binding affinity and initial prediction tools based on MS peptide data. By predicting peptide presentation with higher confidence, the models enable a more time- and cost-efficient identification of neoantigen-specific or tumor antigen-specific T cells for personalized therapy using limited amounts of a patient's peripheral blood, screening fewer peptides per patient, and not necessarily relying on MHC multimers. However, in another embodiment, by using the models disclosed herein, the number of peptides bound to MHC multimers that need to be screened to identify neoantigen- or tumor antigen-specific T cells is reduced, enabling a more time- and cost-efficient identification of tumor antigen-specific T cells using MHC multimers.

[0010] The predictive performance of the models disclosed herein in the task of identifying TIL neoepitope datasets and predicted neoantigen-reactive T cells indicates that it is now possible to obtain predictions of therapeutically useful neoepitopes by modeling HLA processing and presentation. In summary, this work accelerates the progression towards a patient's cure by enabling practical in-silico antigen identification for antigen-targeted immunotherapy. [Invention 1001] A method for identifying one or more neoantigens derived from one or more tumor cells of a subject that are likely to be presented on the surface of the tumor cells, comprising A step of obtaining at least one of exosome, transcriptome, or nucleotide sequencing data of the whole genome from the tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing peptide sequences of respective sets of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, and each peptide sequence of the neoantigen includes at least one change such that the peptide sequence is different from the corresponding wild-type peptide sequence identified from the normal cells of the subject, the step of obtaining; A step of encoding each peptide sequence of the neoantigen into a corresponding numerical vector, wherein each numerical vector includes information regarding a plurality of amino acids constituting the peptide sequence and a set of positions of the amino acids within the peptide sequence, the step of encoding; A step of associating each peptide sequence of the neoantigen with one or more k-mer blocks among a plurality of k-mer blocks of the nucleotide sequencing data of the subject; A step of inputting the numerical vector and the one or more associated k-mer blocks into a presentation model that has been machine-learned using a computer processor to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood within the set represents the likelihood that the corresponding neoantigen is presented on the surface of the tumor cells of the subject by one or more MHC alleles, and the machine-learned presentation model is A plurality of parameters identified based at least on a training data set, wherein the training data set is For each sample of a plurality of samples, a label obtained by mass spectrometry measuring the presence of a peptide bound to at least one MHC allele within a set of MHC alleles identified as present in the sample, For each of the samples, a training peptide sequence encoded as a numerical vector including information regarding a plurality of amino acids constituting the peptide and a set of positions of the amino acids within the peptide, and For each of the samples, for each of the training peptide sequences of the sample, the association between the training peptide sequence and one or more k-mer blocks of the nucleotide sequencing data of the training peptide sequence comprising a subset of the plurality of parameters representing the presence or absence of a presentation hot spot in the one or more k-mer blocks, the plurality of parameters, a function representing the relationship between the numerical vector received as input and the one or more k-mer blocks, and the presentation likelihood generated as output based on the numerical vector, the one or more k-mer blocks, and the parameters comprising the step of inputting; selecting a subset of the set of neoantigens to generate a selected set of neoantigens based on the set of presentation likelihoods; and returning the selected set of neoantigens comprising, the method. [Invention 1002] The step of inputting the numerical vector into the machine-learned presentation model is applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate a dependency score indicating whether the MHC allele presents the neoantigen based on a specific amino acid at a specific position of the peptide sequence for each of the one or more MHC alleles comprising, the method of Invention 1001. [Invention 1003] The step of inputting the numerical vector into the machine-learned presentation model is for each MHC allele, converting the dependency score to generate a corresponding allele likelihood indicating the likelihood that the corresponding MHC allele will present the corresponding neoantigen; and combining the allele likelihoods to generate a presentation likelihood of the neoantigen further comprising, the method of Invention 1002. [The present invention 1004] The method of the present invention 1003, wherein converting the dependency score models the presentation of the neoantigen as mutually exclusive across the one or more MHC alleles. [The present invention 1005] The step of inputting the numerical vector into the machine-learned presentation model further includes converting a combination of the dependency scores to generate the presentation likelihood, and converting the combination of the dependency scores models the presentation of the neoantigen as interfering between the one or more MHC alleles, the method of the present invention 1002. [The present invention 1006] The set of presentation likelihoods is further specified by at least one or more allele non-interaction characteristics, applying the machine-learned presentation model to the allele non-interaction characteristics to generate a dependency score for the allele non-interaction characteristics indicating whether the peptide sequence of the corresponding neoantigen is presented based on the allele non-interaction characteristics The method according to any one of the present inventions 1002 to 1005, further comprising. [The present invention 1007] combining the dependency scores for each MHC allele of the one or more MHC alleles with the dependency scores for the allele non-interaction characteristics; converting the combined dependency scores for each MHC allele to generate an allele-specific likelihood for each MHC allele indicating the likelihood that the corresponding MHC allele presents the corresponding neoantigen; combining the allele-specific likelihoods to generate the presentation likelihood; The method of the present invention 1006, further comprising. [The present invention 1008] combining the dependency scores for each of the MHC alleles with the dependency scores for the allele non-interaction characteristics; converting the combined dependency scores to generate the presentation likelihood; The method of the present invention 1006, further comprising. [Invention 1009] Any method of Inventions 1006 - 1008, wherein the at least one or more allergen non - interacting characteristics comprise an association between the peptide sequence of the neoantigen and one or more k - mer blocks among the plurality of k - mer blocks of the nucleotide sequencing data of the neoantigen. [Invention 1010] Any method of Inventions 1001 - 1009, wherein the one or more MHC alleles comprise two or more different MHC alleles. [Invention 1011] Any method of Inventions 1001 - 1010, wherein the peptide sequence comprises a peptide sequence having a length other than nine amino acids. [Invention 1012] Any method of Inventions 1001 - 1011, wherein encoding the peptide sequence comprises encoding the peptide sequence using a one - hot encoding scheme. [Invention 1013] The plurality of samples are (a) one or more cell lines engineered to express a single MHC allele, (b) one or more cell lines engineered to express multiple MHC alleles, (c) one or more human cell lines obtained from or derived from multiple patients, (d) fresh or frozen tumor samples obtained from multiple patients, and (e) fresh or frozen tissue samples obtained from multiple patients Any method of Inventions 1001 - 1012, comprising at least one of the above. [Invention 1014] The training dataset is (a) data related to measured values of peptide - MHC binding affinity for at least one of the peptides, and (b) data related to measured values of peptide - MHC binding stability for at least one of the peptides Any of the methods of the present invention 1001-1013, further comprising at least one of them. [The present invention 1015] Any of the methods of the present invention 1001-1014, wherein the set of presentation likelihoods is further specified by at least the expression level of the one or more MHC alleles in the subject, measured by RNA-seq or mass spectrometry. [The present invention 1016] The set of presentation likelihoods is (a) the predicted affinity between the neoantigens in the set of neoantigens and the one or more MHC alleles, and (b) the predicted stability of the neoantigen-encoding peptide-MHC complex Any of the methods of the present invention 1001-1015, further specified by a property comprising at least one of them. [The present invention 1017] The set of numerical likelihoods is (a) the C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within its source protein sequence, and (b) the N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within its source protein sequence Any of the methods of the present invention 1001-1016, further specified by a property comprising at least one of them. [The present invention 1018] Any of the methods of the present invention 1001-1017, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of being presented on the surface of the tumor cells compared to non-selected neoantigens, based on the machine-learned presentation model. [The present invention 1019] Any of the methods of the present invention 1001-1018, wherein selecting the set of selected neoantigens comprises selecting neoantigens that have an increased likelihood of inducing a tumor-specific immune response in the subject compared to non-selected neoantigens, based on the machine-learned presentation model. [The present invention 1020] Selecting the set of the selected neoantigens includes selecting neoantigens that, based on the presentation model, have an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) as compared to non-selected neoantigens, and optionally, wherein the APC is a dendritic cell (DC), according to any method of the present invention from 1001 to 1019. [Inventive concept 1021] Selecting the set of the selected neoantigens includes selecting neoantigens that, based on the machine-learned presentation model, have a decreased likelihood of being inhibited by central tolerance or peripheral tolerance as compared to non-selected neoantigens, according to any method of the present invention from 1001 to 1020. [Inventive concept 1022] Selecting the set of the selected neoantigens includes selecting neoantigens that, based on the machine-learned presentation model, have a decreased likelihood of inducing an autoimmune response against normal tissue in the subject as compared to non-selected neoantigens, according to any method of the present invention from 1001 to 1021. [Inventive concept 1023] The one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer, according to any method of the present invention from 1001 to 1022. [Inventive concept 1024] Further comprising generating an output for constructing an individualized cancer vaccine from the set of the selected neoantigens, according to any method of the present invention from 1001 to 1023. [Inventive concept 1025] The output for the individualized cancer vaccine includes at least one peptide sequence or at least one nucleotide sequence encoding the set of the selected neoantigens, according to the method of the present invention 1024. [Inventive concept 1026] The method according to any one of the present inventions 1001 to 1025, wherein the machine-learned presentation model is a neural network model. [The present invention 1027] The method according to the present invention 1026, wherein the neural network model includes a plurality of network models for MHC alleles, each network model is assigned to a corresponding MHC allele among the plurality of MHC alleles, and includes a series of nodes arranged in one or more layers. [The present invention 1028] The method according to the present invention 1027, wherein the neural network model is trained by updating parameters of the neural network model, and the parameters of at least two network models are updated together for at least one training iteration. [The present invention 1029] The method according to any one of the present inventions 1026 to 1028, wherein the machine-learned presentation model is a deep learning model including one or more layers of nodes. [The present invention 1030] The method according to any one of the present inventions 1001 to 1029, wherein the one or more MHC alleles are class I MHC alleles. [The present invention 1031] A computer processor, and When executed by the computer processor, the computer processor is caused to obtain at least one of exosome, transcriptome, or whole-genome nucleotide sequencing data from tumor cells and normal cells of a subject, wherein the nucleotide sequencing data is used to obtain data representing peptide sequences of each set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, and each peptide sequence of the neoantigens includes at least one change that is different from the corresponding wild-type peptide sequence identified from the normal cells of the subject; the obtaining; Encoding each peptide sequence of the neoantigen into a corresponding numerical vector, wherein each numerical vector includes information regarding a plurality of amino acids constituting the peptide sequence and a set of positions of the amino acids within the peptide sequence; Associating each peptide sequence of the neoantigen with one or more k-mer blocks among a plurality of k-mer blocks of the nucleotide sequencing data of the subject; Inputting the numerical vector and the one or more associated k-mer blocks into a presentation model that has been machine-learned to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood within the set represents the likelihood that a corresponding neoantigen is presented on the surface of the tumor cells of the subject by one or more MHC alleles, and the machine-learned presentation model is A plurality of parameters identified based at least on a training dataset, wherein the training dataset is For each sample of a plurality of samples, a label obtained by mass spectrometry measuring the presence of a peptide bound to at least one MHC allele within a set of MHC alleles identified as being present in the sample, For each of the samples, a training peptide sequence encoded as a numerical vector including information regarding a plurality of amino acids constituting the peptide and a set of positions of the amino acids within the peptide, and For each of the samples, for each of the training peptide sequences of the sample, an association between the training peptide sequence and one or more k-mer blocks among the k-mer blocks of the nucleotide sequencing data of the training peptide sequence including, A subset of the plurality of parameters represents the presence or absence of presentation hotspots in the one or more k-mer blocks, the plurality of parameters, A function representing the relationship between the numerical vector and the one or more k-mer blocks received as input and the presentation likelihood generated as output based on the numerical vector, the one or more k-mer blocks, and the parameters including said inputting; selecting, based on said set of presentation likelihoods, a subset of said set of neoantigens to generate a selected set of neoantigens; and returning said selected set of neoantigens a memory storing computer program instructions causing a computer system including.

Brief Description of the Drawings

[0011] These features, aspects, and facets of the invention, as well as other features, aspects, and facets, will be better understood with reference to the following description and the accompanying drawings.

[0012]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 1F

Figure 1G

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 15D

Figure 15E

Figure 15F

Figure 16

Figure 17A

Figure 17B

Figure 17C

Figure 18A-1

Figure 18A-2

Figure 18B-1

Figure 18B-2

Figure 19

Figure 20A

Figure 20B

Figure 20C

Figure 21

Figure 22

Figure 23-1

Figure 23-2

Figure 24

Figure 25-1

Figure 25-2

Figure 26-1

Figure 26-2

Figure 27-1

Figure 27-2

Figure 28

Figure 29

DETAILED DESCRIPTION OF THE INVENTION

[0013] I. Definitions In general, the terms used in the claims and the specification are to be construed as having the ordinary meaning understood by those skilled in the art. Specific terms are defined below for further clarity. In case of a conflict between the ordinary meaning and the given definition, the given definition shall be used.

[0014] As used herein, the term "antigen" refers to a substance that induces an immune response.

[0015] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes the antigen different from the corresponding wild-type parental antigen, for example, by a mutation in a tumor cell or a post-translational modification specific to a tumor cell. A neoantigen may comprise a polypeptide sequence or a nucleotide sequence. The mutation can include a frameshift or non-frameshift insertion or deletion (indel), a missense or nonsense substitution, a splice site change, a genomic rearrangement or gene fusion, or any genomic or expression change that results in a novel ORF. The mutation can also include splice variants. Post-translational modifications specific to tumor cells can include abnormal phosphorylation. Post-translational modifications specific to tumor cells can also include splice antigens generated by the proteasome. See Liepe et al., A large fraction of HLA class I ligands are proteasome-generated spliced peptides; Science. 2016 Oct 21;354(6310):354-358.

[0016] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in a subject's tumor cells or tissue but not in the subject's corresponding normal cells or tissue.

[0017] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct based on one or more neoantigens, for example, multiple neoantigens.

[0018] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.

[0019] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.

[0020] As used herein, the term "coding mutation" refers to a mutation that occurs in the coding region.

[0021] As used herein, the term "ORF" means open reading frame.

[0022] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that results from other abnormalities such as mutations or splicing.

[0023] As used herein, the term "missense mutation" refers to a mutation that causes a substitution of one amino acid for another.

[0024] As used herein, the term "nonsense mutation" refers to a mutation that causes a substitution of an amino acid for a stop codon.

[0025] As used herein, the term "frameshift mutation" refers to a mutation that causes a change in the protein frame.

[0026] As used herein, the term "indel" refers to an insertion or deletion of one or more nucleic acids.

[0027] As used herein, the term "identity" (%) in the context of the sequences of two or more nucleic acids or polypeptides refers to the percentage of specific nucleotides or amino acid residues that are the same when comparing and aligning the sequences using one of the following sequence comparison algorithms (e.g., BLASTP and BLASTN, or other algorithms available to those skilled in the art) or by visual inspection, for the greatest match. Depending on the application, the "identity" (%) can exist over the region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.

[0028] In array comparison, generally, one array functions as a reference array against which a test array is compared. When using an array comparison algorithm, the test array and the reference array are input into a computer, and if necessary, partial array coordinates are specified and the parameters of the array algorithm program are specified. Then, the array comparison algorithm calculates the percentage of array identity of the test array relative to the reference array based on the specified program parameters. Alternatively, the similarity or dissimilarity of arrays can also be established by the presence or absence combination of specific nucleotides at selected array positions (e.g., array motifs), or in the translated array, amino acids.

[0029] The optimal alignment of the arrays for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson & Lipman, Proc. Nat’l. Acad. Sci. USA 85:2444 (1988), by computer processing execution of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (generally, see Ausubel et al. below).

[0030] As an example of one algorithm suitable for determining the percentage of array identity and the percentage of array similarity, there is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information.

[0031] As used herein, the term "non-stop or read-through" refers to a mutation that causes the removal of a natural stop codon.

[0032] As used herein, the term "epitope" refers to the specific portion of an antigen to which an antibody or T cell receptor typically binds.

[0033] As used herein, the term "immunogenicity" refers to the ability to induce an immune response, for example, via T cells, B cells, or both.

[0034] As used herein, the terms "HLA binding affinity", "MHC binding affinity" refer to the affinity of binding of a specific antigen to a specific MHC allele.

[0035] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.

[0036] As used herein, the term "mutation" refers to a difference between a subject nucleic acid and a reference human genome used as a control.

[0037] As used herein, the term "mutation call" refers to an algorithmic determination of the presence of a mutation, typically from sequencing.

[0038] As used herein, the term "polymorphism" refers to a germline mutation, i.e., a mutation found in all DNA-bearing cells of an individual.

[0039] As used herein, the term "somatic mutation" refers to a mutation that occurs in non-germline cells of an individual.

[0040] As used herein, the term "allele" refers to one version of a gene, or one version of a gene sequence, or one version of a protein.

[0041] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.

[0042] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the cellular degradation of mRNA due to premature stop codons.

[0043] As used herein, the term "truncal mutation" refers to a mutation that occurs early in tumor development and is present in most of the tumor cells.

[0044] As used herein, the term "subclonal mutation" refers to a mutation that occurs late in tumor development and is present in only some of the tumor cells.

[0045] As used herein, the term "exome" refers to a subset of the genome that encodes proteins. The exome can be the collective exons of the genome.

[0046] As used herein, the term "logistic regression" refers to a regression model for binary data from statistics in which the logit of the probability that the dependent variable equals 1 is modeled as a linear function of the dependent variable.

[0047] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise non-linear transformations typically trained by stochastic gradient descent and backpropagation.

[0048] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, group of cells, or organism.

[0049] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. The peptidome may also refer to the properties of a cell or collection of cells (e.g., the tumor peptidome refers to the union of the peptidomes of all cells containing the tumor).

[0050] As used herein, the term "ELISPOT" means enzyme-linked immunosorbent spot assay, a common method for observing immune responses in humans and animals.

[0051] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.

[0052] As used herein, the term "MHC multimer" refers to a peptide-MHC complex composed of multiple peptide-MHC monomer units.

[0053] As used herein, the term "MHC tetramer" refers to a peptide-MHC complex composed of four peptide-MHC monomer units.

[0054] As used herein, the term "tolerance" or "immunological tolerance" refers to a state of immune non-responsiveness to one or more antigens, such as self-antigens.

[0055] As used herein, the term "central tolerance" is the tolerance imparted in the thymus by either deleting autoreactive T cell clones or promoting the differentiation of autoreactive T cell clones into immunosuppressive regulatory T cells (Tregs).

[0056] As used herein, the term "peripheral tolerance" is the tolerance imparted in the peripheral system by downregulating or anergizing autoreactive T cells that have survived central tolerance, or by promoting the differentiation of these T cells into Tregs.

[0057] The term "sample" can include a single cell, or a plurality of cells, or a fragment of a cell, or an aliquot of a body fluid, collected from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, wash sample, scraping, surgical incision, or intervention, or other means known in the art.

[0058] The term "subject" includes any cell, tissue, or organism, human or non-human, male or female, in vivo, ex vivo, or in vitro. The term "subject" includes mammals including humans.

[0059] The term "mammal" includes both human and non-human, including but not limited to humans, non-human primates, dogs, cats, mice, cows, horses, and pigs.

[0060] The term "clinical factor" refers to a measure of the state of a subject, such as the activity or severity of a disease. "Clinical factor" includes all markers of the health state of the subject, including non-sample markers, and / or other characteristics of the subject, such as, without limitation, age and gender. A clinical factor can be a score, value, or set of values that can be obtained from the assessment of a subject or a sample (or population of samples) from the subject under a given condition. A clinical factor can also be predicted by other parameters such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.

[0061] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.

[0062] Note that, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise.

[0063] Terms not directly defined in this specification should be understood to have the meanings customarily associated with them, as would be understood within the scope of the technical field of the present invention. Certain terms are considered in this specification for the purpose of providing further guidance to the practitioner in describing the compositions, devices, methods, etc. of aspects of the present invention, as well as the methods of their manufacture or use. It will be recognized that there may be multiple ways of referring to the same thing. Thus, alternative words and synonyms may be used for any one or more of the terms considered in this specification. Whether or not a term is detailed or considered in this specification should not be weighted. Some synonyms or alternative ways, materials, etc. are provided. The listing of one or several synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the aspects of the invention in this specification.

[0064] All references, issued patents, and patent applications cited in the text of this specification are hereby incorporated by reference in their entirety for all purposes.

[0065] II. Methods for Identifying Neoantigens This specification discloses a method for identifying neoantigens derived from a subject's tumor cells that are likely to be presented on the surface of the tumor cells. The method includes obtaining exosome, transcriptome, and / or whole genome nucleotide sequencing data from the subject's tumor cells and normal cells. Using this nucleotide sequencing data, the peptide sequence of each neoantigen within a set of neoantigens is obtained. The set of neoantigens is identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells. Specifically, the peptide sequence of each neoantigen within the set of neoantigens includes at least one change such that the peptide sequence is different from the corresponding wild-type parental peptide sequence identified from the subject's normal cells. The method further includes encoding the peptide sequence of each neoantigen within the set of neoantigens into a corresponding numerical vector. Each numerical vector includes information describing the amino acids that make up the peptide sequence and the positions of the amino acids within the peptide sequence. The method includes associating each peptide sequence of the neoantigens with one or more k-mer blocks of the plurality of k-mer blocks of the subject's nucleotide sequencing data. The method further includes generating a presentation likelihood for each neoantigen within the set of neoantigens by inputting the numerical vector and the associated k-mer blocks into a machine-learned presentation model. Each presentation likelihood represents the likelihood that the corresponding neoantigen is presented by an MHC allele on the surface of the subject's tumor cells. The machine-learned presentation model includes a plurality of parameters and functions. The plurality of parameters are identified based on a training dataset.The training dataset includes a label obtained by mass spectrometry that measures the presence of a peptide bound to at least one MHC allele within a set of MHC alleles identified as present in each of a plurality of samples, a training peptide sequence encoded as a numerical vector including information describing the amino acids constituting the peptide and / or the positions of the amino acids within the peptide, and for each of the training peptide sequences of the samples, an association between the training peptide sequence and one or more k-mer blocks among a plurality of k-mer blocks of the nucleotide sequencing data of the training peptide sequence. The function represents the relationship between a numerical vector and an associated k-mer block received as input by a machine-learned presentation model, and a presentation likelihood generated as output by the machine-learned presentation model based on the numerical vector, the associated k-mer block, and a plurality of parameters. The method further includes selecting a subset of the set of neoantigens based on the presentation likelihood to generate a selected set of neoantigens, and returning the selected set of neoantigens.

[0066] In some embodiments, inputting a numerical vector into the machine-learned presentation model comprises applying the machine-learned presentation model to the peptide sequences of neoantigens to generate a dependency score for each of the MHC alleles. The dependency score for a given MHC allele indicates whether that MHC allele presents a neoantigen based on a particular amino acid at a particular position in the peptide sequence. In further embodiments, inputting a numerical vector into the machine-learned presentation model further comprises converting the dependency scores and combining the per-allele likelihoods to generate a presentation likelihood for the neoantigen, wherein for each MHC allele, the per-allele likelihood indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen. In some embodiments, converting the dependency scores models the presentation of neoantigens as being mutually exclusive across MHC alleles. In alternative embodiments, inputting a numerical vector into the machine-learned presentation model further comprises converting a combination of dependency scores to generate a presentation likelihood. In such embodiments, converting the combination of dependency scores models the presentation of neoantigens as interference between MHC alleles.

[0067] In some embodiments, the set of presentation likelihoods is further specified by one or more allele non-interaction characteristics. In such embodiments, the method further comprises applying the machine-learned presentation model to the allele non-interaction characteristics to generate a dependency score for the allele non-interaction characteristics. The dependency score indicates whether the peptide sequence of the corresponding neoantigen is presented based on the allele non-interaction characteristics. In some embodiments, the one or more allele non-interaction characteristics include values indicating the presence or absence of presentation hotspots in respective k-mer blocks of the peptide sequence of each neoantigen.

[0068] In some embodiments, the method further includes combining the dependency score for each MHC allele with the dependency score for allele non-interaction characteristics, converting the combined dependency score for each MHC allele to generate an allele likelihood for each MHC allele, and combining the allele likelihoods to generate a presentation likelihood. The allele likelihood for a given MHC allele indicates the likelihood that the MHC allele presents the corresponding neoantigen. In an alternative embodiment, the method further includes combining the dependency score for the MHC allele with the dependency score for allele non-interaction characteristics and converting the combined dependency score to generate a presentation likelihood.

[0069] In some embodiments, the MHC allele includes two or more different MHC alleles.

[0070] In some embodiments, the peptide sequence includes peptide sequences having a length other than nine amino acids.

[0071] In some embodiments, encoding the peptide sequence includes encoding the peptide sequence using a one-hot encoding scheme.

[0072] In certain embodiments, the plurality of samples includes at least one of a cell line engineered to express a single MHC allele, a cell line engineered to express multiple MHC alleles, a human cell line obtained from or derived from multiple patients, and a fresh or frozen tissue sample obtained from multiple patients.

[0073] In some embodiments, the training dataset further includes at least one of data related to a measurement of peptide-MHC binding affinity for at least one of the peptides and data related to a measurement of peptide-MHC binding stability for at least one of the peptides.

[0074] In some embodiments, the set of presentation likelihoods is further specified by the expression levels of MHC alleles within the subject, as measured by RNA-seq or mass spectrometry.

[0075] In some embodiments, the set of presentation likelihoods is further specified by a property including at least one of the predicted affinity between a neoantigen and an MHC allele within the set of neoantigens, and the predicted stability of the neoantigen-encoding peptide-MHC complex.

[0076] In some embodiments, the set of numerical likelihoods is further specified by a property including at least one of the C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within its source protein sequence, and the N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within its source protein sequence.

[0077] In some embodiments, selecting the set of selected neoantigens includes selecting neoantigens that have an increased likelihood of being presented on the surface of tumor cells as compared to non-selected neoantigens, based on a machine-learned presentation model.

[0078] In some embodiments, selecting the set of selected neoantigens includes selecting neoantigens that have an increased likelihood of being able to induce a tumor-specific immune response in the subject as compared to non-selected neoantigens, based on a machine-learned presentation model.

[0079] In some embodiments, selecting the set of selected neoantigens includes selecting neoantigens that have an increased likelihood of being able to be presented to naive T cells by professional antigen-presenting cells (APCs) as compared to non-selected neoantigens, based on a presentation model. In such embodiments, the APC is optionally a dendritic cell (DC).

[0080] In some embodiments, selecting a set of selected neoantigens comprises selecting neoantigens that, based on a machine-learned presentation model, have a reduced likelihood of being inhibited by central or peripheral tolerance as compared to non-selected neoantigens.

[0081] In some embodiments, selecting a set of selected neoantigens comprises selecting neoantigens that, based on a machine-learned presentation model, have a reduced likelihood of inducing an autoimmune response against normal tissue in a subject as compared to non-selected neoantigens.

[0082] In some embodiments, the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0083] In some embodiments, the method further comprises generating an output for constructing an individualized cancer vaccine from the set of selected neoantigens. In such embodiments, the output for the individualized cancer vaccine can comprise at least one peptide sequence or at least one nucleotide sequence encoding the set of selected neoantigens.

[0084] In some embodiments, the machine-learned presentation model is a neural network model. In such embodiments, the neural network model can include a plurality of network models for MHC alleles, where each network model includes a series of nodes assigned to corresponding MHC alleles among the plurality of MHC alleles and arranged in one or more layers. In such embodiments, the neural network model can be trained by updating the parameters of the neural network model, and the parameters of at least two network models are updated together for at least one training iteration. In some embodiments, the machine-learned presentation model can be a deep learning model including one or more layers of nodes.

[0085] In some embodiments, the MHC allele is a class I MHC allele.

[0086] This specification also discloses a computer system including a computer processor and a memory storing computer program instructions. When the computer program instructions are executed by the computer processor, the instructions cause the computer processor to perform any of the methods described above.

[0087] III. Identification of Tumor-Specific Mutations in Neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells). In particular, these mutations can be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject having cancer, but not in normal tissue from the subject.

[0088] Gene mutations in tumors can be considered useful for tumor immunological targeting when they exclusively result in changes in the amino acid sequence of proteins in the tumor. Useful mutations include the following: (1) Nonsynonymous mutations that result in different amino acids in the protein; (2) Read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a novel tumor-specific sequence at the C-terminus; (3) Splice site mutations that result in the inclusion of introns in the mature mRNA and thus a unique tumor-specific protein sequence; (4) Chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) Frameshift mutations or deletions that result in a new open reading frame with a novel tumor-specific protein sequence. Mutations can also include one or more of non-frameshift insertions or deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in a de novo ORF.

[0089] For example, peptides or mutant polypeptides having mutations resulting from splice site, frameshift, read-through, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.

[0090] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.

[0091] A variety of methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. For example, several techniques have been described, including various DNA "chip" technologies such as dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, TaqMan systems, and Affymetrix SNP chips. These methods typically utilize amplification of the target gene region by PCR. Still other methods are based on the generation of small signal molecules by invasive cleavage and subsequent mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.

[0092] PCR-based detection means can simultaneously include multiplex amplification of a number of markers. For example, it is well known in the art to select PCR primers such that PCR products are generated that do not overlap in size and can be analyzed simultaneously. Alternatively, it is possible to amplify different markers with primers that are differentially labeled and thus can be differentially detected. Of course, hybridization-based detection means enable differential detection of multiple PCR products in a sample. Other techniques that enable multiplex analysis of multiple markers are known in the art.

[0093] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA. For example, single nucleotide polymorphisms can be detected by using specialized exonuclease-resistant nucleotides, such as those disclosed in, for example, Mundy, C.R. (U.S. Patent No. 4,656,127). According to this method, a primer complementary to the allelic sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a particular animal or human. If the polymorphic site on the target molecule contains a nucleotide that is complementary to a particular exonuclease-resistant nucleotide derivative present, the derivative is incorporated onto the end of the hybridized primer. Because of such incorporation, the primer becomes resistant to exonuclease, thereby allowing its detection. Since the identity of the exonuclease-resistant derivative in the sample is known, the finding that the primer has become resistant to exonuclease reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage of not requiring the determination of large amounts of exogenous sequence data.

[0094] To determine the identity of the nucleotide at the polymorphic site, solution-based methods can be used (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087). As in the method of Mundy of U.S. Patent No. 4,656,127, a primer that is complementary to the allelic sequence immediately 3' of the polymorphic site is used. This method determines the identity of the nucleotide at that site using a labeled dideoxynucleotide derivative that will be incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site. An alternative method, known as Genetic Bit Analysis or GBA, is described by Goelet, P. et al. (PCT Application No. 92 / 15712). The method of Goelet, P. et al. uses a mixture of a labeled terminator and a primer that is complementary to the 3' sequence of the polymorphic site. The method of Goelet, P. et al. uses a mixture of a labeled terminator and a primer that is complementary to the 3' sequence of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO91 / 02087), the method of Goelet, P. et al. can be a heterogeneous assay in which the primer or the target molecule is immobilized on a solid phase.

[0095] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J.S. et al., Nucl. Acids Res. 17:7779-7784 (1989); Sokolov, B.P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M.N. et al., Proc. Natl. Acad. Sci. (U.S.A.) 88:1143-1147 (1991); Prezant, T.R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, since the signal is proportional to the number of incorporated deoxynucleotides, polymorphisms occurring in the same nucleotide run can result in signals proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).

[0096] Numerous initiatives directly obtain sequence information in parallel from millions of individual molecules of DNA or RNA. Sequencing techniques by real-time single-molecule synthesis rely on the detection of fluorescent nucleotides when incorporated into the nascent DNA strand that is complementary to the template being sequenced. In one method, oligonucleotides 30 to 50 bases in length are covalently attached at their 5'-ends to a glass coverslip. These immobilized strands serve two functions. First, they act as capture sites for target template strands when the template is constructed with a capture tail that is complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis of sequence readout. The capture primer functions as a fixed-site for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and cleavage of the dye. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. When a nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other synthetic sequencing techniques also exist.

[0097] Any suitable sequencing platform by synthesis can be used to identify mutations. As noted above, four major sequencing platforms by synthesis are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing platforms by synthesis have also been described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the nucleic acid molecules to be sequenced are attached to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' end and / or 5' end of the template. The nucleic acid can be attached to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. The capture sequence (also referred to as a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to the support that can act as a universal primer in a dual capacity.

[0098] As an alternative to the capture sequence, members of a coupling pair (e.g., antibody / antigen, receptor / ligand, or, for example, an avidin-biotin pair as described in U.S. Patent Application No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.

[0099] Following capture, the array can be analyzed by, for example, single molecule detection / sequencing as described in the Examples and U.S. Patent No. 7,283,337, including, for example, sequencing by template-dependent synthesis. In sequencing by synthesis, the molecules bound to the surface are exposed to a large number of labeled nucleotide triphosphates in the presence of polymerase. The sequence of the template is determined by the order of the labeled nucleotides incorporated into the 3' end of the growing chain. This can be done in real time and in a step-and-repeat mode. For real-time analysis, different optical labels can be incorporated for each nucleotide, and multiple lasers can be utilized for the excitation of the incorporated nucleotides.

[0100] Sequencing can also include other massively parallel processing sequencing, or next generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel processing sequencing techniques and platforms are Illumina HiSeq or MiSeq, Thermo PGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel processing sequencing technologies, and future generations of these technologies, can be used.

[0101] Any cell type or tissue can be utilized to obtain a nucleic acid sample for use in the methods described herein. For example, a DNA or RNA sample can be obtained from a tumor or a body fluid such as blood or saliva obtained by known techniques (e.g., venipuncture). Alternatively, a nucleic acid test can be performed on a dried sample (e.g., hair or skin). In addition, a sample can be obtained from a tumor for sequencing, and another sample can be obtained from normal tissue for sequencing if the normal tissue is of the same tissue type as the tumor. A sample can be obtained from a tumor for sequencing, and another sample can be obtained from normal tissue for sequencing if the normal sample is of a tissue type distinct from the tumor.

[0102] The tumor can include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0103] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. The peptides can be acid eluted from tumor cells or from HLA molecules immunoprecipitated from tumors and then identified using mass spectrometry.

[0104] IV. Neoantigens Neoantigens can include nucleotides or polynucleotides. For example, a neoantigen can be an RNA sequence encoding a polypeptide sequence. Neoantigens useful in a vaccine can thus include a nucleotide sequence or a polypeptide sequence.

[0105] Isolated peptides containing tumor-specific mutations identified by the methods disclosed herein, peptides containing known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein are disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the nucleotide sequences (e.g., DNA or RNA) encoding the polypeptide sequences to which the neoantigens are related are included.

[0106] One or more polypeptides encoded by a neoantigen nucleotide sequence can comprise at least one of the following: binding affinity to MHC with an IC50 value of less than 1000 nM, for MHC class I peptides, a length of 8 to 15 amino acids, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids, the presence of a sequence motif within or near the peptide that promotes proteasome cleavage, and the presence of a sequence motif that promotes TAP transport. For MHC class II polypeptides, a length of 6 to 30 amino acids, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids, the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.

[0107] One or more neoantigens can be present on the surface of a tumor.

[0108] One or more neoantigens can be immunogenic in a subject having a tumor, e.g., can elicit a T cell response or a B cell response in the subject.

[0109] One or more neoantigens that induce an autoimmune response in a subject can be excluded from consideration in the context of vaccine generation for a subject having a tumor.

[0110] The size of at least one neoantigenic peptide molecule can include, but is not limited to, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derived from these ranges. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.

[0111] Neoantigenic peptides and polypeptides can be 15 residues or less in length for MHC class I, usually consisting of between about 8 and about 11 residues, and can in particular be 9 or 10 residues; for MHC class II, they can be 6 - 30 residues.

[0112] If desired, longer peptides can be designed in several ways. In one example, if the likelihood of peptide presentation on an HLA allele is predicted or known, longer peptides can consist of (1) individual presented peptides having an extension of 2-5 amino acids towards the N-terminus and C-terminus of each corresponding gene product; (2) any of some or all of the chains of presented peptides, each having an extended sequence. In another example, if sequencing reveals the presence in the tumor of long (longer than 10 residues) neoepitope sequences (e.g., due to frameshifts, read-throughs, or inclusion of introns that result in novel peptide sequences), longer peptides can consist of (3) the entire stretch of novel tumor-specific amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides presented by the strongest HLA. In any of the examples, the use of longer peptides allows for endogenous processing by patient cells and can result in more effective antigen presentation and induction of T cell responses.

[0113] Neoantigenic peptides and polypeptides can be presented on HLA proteins. In some embodiments, neoantigenic peptides and polypeptides are presented on HLA proteins with a stronger affinity than wild-type peptides. In some embodiments, a neoantigenic peptide or polypeptide can have an IC50 of at least less than 5000 nM, at least less than 1000 nM, at least less than 500 nM, at least less than 250 nM, at least less than 200 nM, at least less than 150 nM, at least less than 100 nM, at least less than 50 nM, or less than that.

[0114] In some embodiments, neoantigenic peptides and polypeptides, when administered to a subject, do not induce an autoimmune response and / or do not cause immune tolerance.

[0115] Also provided is a composition comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides are derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which the neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some aspects, the tumor-specific mutations are driver mutations for a particular cancer type.

[0116] Neoantigenic peptides and polypeptides having desirable activities or properties can be modified to provide certain desirable attributes, such as improved pharmacological characteristics, while increasing or at least substantially retaining all of the biological activities of the unmodified peptides that bind to desirable MHC molecules and activate appropriate T cells. By way of example, neoantigenic peptides and polypeptides can be further subjected to various modifications, such as conservative or non-conservative substitutions, which can provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions are meant to replace an amino acid residue with another that is biologically and / or chemically similar, e.g., replacing one hydrophobic residue with another hydrophobic residue or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effect of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (N.Y., Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2d Ed. (1984).

[0117] Modification of peptides and polypeptides with various amino acid mimics or unnatural amino acids can be particularly useful for increasing the stability of peptides and polypeptides in vivo. Stability can be assayed in a number of ways. By way of example, peptidases, as well as various biological media such as human plasma and serum, have been used to test stability. See, for example, Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). The half-life of a peptide can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally is as follows. Pooled human serum (type AB, non-heat inactivated) is defatted by centrifugation prior to use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The turbid reaction samples are cooled (4 °C) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse phase HPLC using stability-specific chromatographic conditions.

[0118] Peptides and polypeptides can be modified to provide desirable attributes other than an improved serum half-life. By way of example, the ability of a peptide to induce CTL activity can be enhanced by ligation to a sequence containing at least one epitope capable of inducing a T helper cell response. Immunogenic peptide / T helper conjugates can be linked by a spacer molecule. The spacer typically consists of relatively small neutral molecules such as amino acids or amino acid mimics that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that any optional spacer present need not be composed of the same residues and can thus be a hetero-oligomer or a homo-oligomer. When present, the spacer will usually be at least 1 or 2 residues, more usually 3 - 6 residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.

[0119] Neoantigenic peptides can be linked to T helper peptides either directly or via a spacer at either the amino or carboxy terminus of the peptide. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include 830 - 843 of tetanus toxin, 307 - 319 of influenza, around 382 - 398 and 378 - 389 of malaria sporozoite.

[0120] A protein or peptide can be made by any technique known to those of skill in the art, including expression of a protein, polypeptide, or peptide through standard molecular biology techniques, isolation of a protein or peptide from a natural source, or chemical synthesis of a protein or peptide. Nucleotide as well as protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located at the website of the National Institutes of Health. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.

[0121] In a further aspect, the neoantigen comprises a nucleic acid (e.g., polynucleotide) encoding the neoantigenic peptide or a part thereof. The polynucleotide can be, for example, single-stranded and / or double-stranded polynucleotides such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), polynucleotides having a phosphorothioate backbone, etc., either in their native form or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further aspect provides an expression vector capable of expressing the polypeptide or a part thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector such as a plasmid in the proper orientation and correct reading frame for expression. If necessary, the DNA can be ligated to appropriate transcriptional and translational regulatory nucleotide sequences recognized by the desired host, but such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Protocols can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.

[0122] IV. Vaccine Composition Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can elicit a specific immune response, e.g., a tumor-specific immune response. The vaccine composition typically comprises a number of neoantigens selected, for example, using the methods described herein. The vaccine composition can also be referred to as a vaccine.

[0123] The vaccine can contain 1 to 30 types of peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine can contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine can contain 1 to 30 types of neoantigen sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13 or 14 different neoantigen sequences, or 12, 13 or 14 different neoantigen sequences.

[0124] In one embodiment, the different peptides and / or polypeptides, or nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides can bind to different MHC molecules such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, one vaccine composition comprises a coding sequence of a peptide and / or polypeptide that can bind to the most frequently present MHC class I molecule and / or MHC class II molecule. Thus, the vaccine composition can contain different fragments that can bind to at least 2 preferred, at least 3 preferred, or at least 4 preferred MHC class I molecules and / or MHC class II molecules.

[0125] The vaccine composition can generate a specific cytotoxic T cell response and / or a specific helper T cell response.

[0126] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are shown below in this specification. The composition can bind to a carrier such as a protein, or an antigen-presenting cell such as a dendritic cell (DC) that can present a peptide to T cells, for example.

[0127] An adjuvant is any substance whose mixing into the vaccine composition increases or otherwise modifies the immune response to a neoantigen. A carrier can be a scaffold structure to which a neoantigen can bind, such as a polypeptide or a polysaccharide. Optionally, the adjuvant is conjugated covalently or non-covalently.

[0128] The ability of an adjuvant to increase the immune response to an antigen is typically demonstrated by a significant or substantial increase in an immune-mediated reaction, or a reduction in disease symptoms. For example, an increase in humoral immunity is typically demonstrated by a significant increase in the titer of antibodies generated against the antigen, and an increase in T cell activity is typically demonstrated in an increase in cell proliferation, or cytotoxicity, or cytokine secretion. An adjuvant can also change the immune response, for example, by changing a predominantly humoral or Th response to a predominantly cellular or Th response.

[0129] Suitable adjuvants include, but are not limited to, 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA206, Montanide ISA 50V, Montanide ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila’s QS21 stimulon derived from saponin (Aquila Biotech, Worcester, Mass., USA), bacterial extracts and synthetic bacterial cell wall mimetics, and other proprietary adjuvants such as Ribi’s Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are useful. Some immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison A C; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Some cytokines are directly linked to effects on the migration of dendritic cells to lymphoid tissues (e.g., TNF-α), acceleration of the maturation of dendritic cells into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Patent No. 5,849,589, which is specifically incorporated herein by reference in its entirety), and acting as immunological adjuvants (e.g., IL-12) (Gabrilovich D I, et al., J ImmunotherEmphasis Tumor Immunol. 1996(6):414-418).

[0130] CpG immunostimulatory oligonucleotides have also been reported to enhance the effect of adjuvants in vaccine settings. Other TLR-binding molecules such as RNA that bind to TLR 7, TLR 8, and / or TLR 9 may also be used.

[0131] Other examples of useful adjuvants include chemically modified CpGs (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies such as cyclophosphamide, sunitinib, bevacizumab, celecoxib, NCX-4016, sildenafil, tadalafil, vardenafil, sorafenib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175 that can act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony stimulating factors such as granulocyte macrophage colony stimulating factor (GM-CSF, sargramostim).

[0132] Vaccine compositions can contain more than one different adjuvant. Further, therapeutic compositions can contain any adjuvant substance including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.

[0133] The carrier (or excipient) can exist independently of the adjuvant. The function of the carrier can be, for example, to increase the activity or immunogenicity, to confer stability, to increase the biological activity, or to increase the serum half-life, particularly by increasing the molecular weight of the variant. Furthermore, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, keyhole limpet hemocyanin, serum proteins such as transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, immunoglobulins, or hormones such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is acceptable and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be dextran, such as sepharose.

[0134] Cytotoxic T lymphocytes (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. The MHC molecules themselves are located on the cell surface of antigen-presenting cells. Thus, activation of CTLs is possible when a trimeric complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when the peptide is used for activation of CTLs, but additionally when APCs having the respective MHC molecules are added, it can enhance the immune response. Thus, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.

[0135] Neoantigens can also be included in virus vector-based vaccine platforms, including but not limited to vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses of the second, third, or hybrid second / third generation, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61, Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18, Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690, Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880). Depending on the packaging capacity of the virus vector-based vaccine platforms described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The arrays may be adjacent non-mutated arrays, may be separated by linkers, or one or more arrays targeting intracellular compartments may precede (see, for example, Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8, Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41, Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express neoantigens, thereby eliciting a host immune (e.g., CTL) response to the peptides. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guerin). The BCG vector is described in Stover et al. (Nature 351:456-460 (1991)). A variety of other vaccine vectors useful for the therapeutic administration or immunization with neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.

[0136] IV.A. Further Considerations in Vaccine Design and Manufacture IV.A.1. Determination of a Set of Peptides that Cover All Tumor Subclones Truncal peptides, which mean those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides predicted to be presented with high probability and be immunogenic, or if the number of truncal peptides predicted to be presented with high probability and be immunogenic is so small that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .

[0137] IV.A.2. Neoantigen Prioritization After applying all of the above neoantigen filters, more candidate neoantigens may still be available for vaccine inclusion than the vaccine technology can accommodate. Additionally, there may be uncertainties remaining about various aspects of neoantigen analysis, and there may be trade-offs between various properties of candidate vaccine neoantigens. Thus, instead of pre-determined filters at each stage of the selection process, an integrated multidimensional model can be considered that places candidate neoantigens in a space with at least the following axes and optimizes the selection using an integrated approach. 1. Risk of autoimmunity or tolerance (germline risk) (lower risk of autoimmunity is typically preferred) 2. Probability of sequencing artifact (lower probability of artifact is typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferred) 5. Gene expression (higher expression is typically preferred) 6. Coverage of HLA genes (A larger number of HLA molecules involved in the presentation of the set of neoantigens may reduce the probability that the tumor evades immune attack via downregulation or mutation of HLA molecules.) 7. Coverage of HLA classes (Covering both HLA-I and HLA-II may increase the probability of treatment response and decrease the probability of tumor immune escape.)

[0138] V. Treatment and manufacturing methods There is also provided a method of vaccinating a subject against cancer, inducing a tumor-specific immune response in the subject, and treating and / or alleviating the symptoms of cancer in the subject by administering to the subject one or more neoantigens, such as a plurality of neoantigens identified using the methods disclosed herein.

[0139] In some embodiments, the subject is diagnosed with cancer or at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desirable. The tumor can be any solid tumor, such as a tumor of the breast, ovary, prostate, lung, kidney, stomach, colon, testis, head and neck, pancreas, brain, melanoma, and tumors of other tissues and organs, as well as hematological tumors, such as lymphoma and leukemia, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.

[0140] The neoantigen can be administered in an amount sufficient to induce a CTL response.

[0141] The neoantigen can be administered alone or in combination with other therapeutic substances. The therapeutic substance can be, for example, a chemotherapeutic agent, radiation, or immunotherapy. Any suitable therapeutic treatment for a particular cancer can be administered.

[0142] In addition, an anti-immunosuppressive / immunostimulatory substance such as a checkpoint inhibitor can be further administered to the subject. For example, an anti-CTLA antibody or anti-PD-1 or anti-PD-L1 can be further administered to the subject. Blocking of CTLA-4 or PD-L1 by the antibody can enhance the immune response against cancer cells in the patient. In particular, CTLA-4 blockade has been shown to be effective when a vaccination protocol is employed.

[0143] The optimal amount of each neoantigen to be included in the vaccine composition, and the optimal dosing regimen, can be determined. For example, the neoantigen or its variant can be prepared for intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal (i.p.) injection, intramuscular (i.m.) injection. The methods of injection include s.c., i.d., i.p., i.m., and i.v. The methods of DNA or RNA injection include i.d., i.m., s.c., i.p., and i.v. Other methods of administration of the vaccine composition are known to those skilled in the art.

[0144] The vaccine can be engineered such that the selection, number, and / or amount of neoantigens present in the composition are specific to the tissue, cancer, and / or patient. By way of example, the stringent selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. The selection can depend on the specific type of cancer, the disease state, earlier treatment regimens, the patient's immune status, and, of course, the patient's HLA haplotype. Further, the vaccine can contain individualized components according to the particular needs of a given patient. Examples include changing the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjustments for secondary treatment after the first round or scheme of treatment.

[0145] For a composition to be used as a vaccine for cancer, neoantigens having similar normal self - peptides that are highly expressed in normal tissues can be avoided or present in small amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a certain neoantigen in large amounts, each pharmaceutical composition for the treatment of this cancer can be present in large amounts and / or can include this specific neoantigen or more than one neoantigen specific to the pathway of this neoantigen.

[0146] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In a therapeutic application, the composition is administered to the patient in an amount sufficient to elicit an effective CTL response against tumor antigens and to cure or at least partially arrest symptoms and / or complications. The amount considered appropriate to achieve this is defined as the "therapeutically effective dose". The effective amount for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general state of health, as well as the judgment of the prescribing physician. It should be borne in mind that the composition can generally be used in severe disease states, i.e., life - threatening or potentially life - threatening situations, particularly when the cancer has metastasized. In such cases, taking into account the minimization of foreign substances and the relatively non - toxic nature of the neoantigens, it is possible to administer a substantially excessive amount of these compositions, and the treating physician may feel it desirable.

[0147] For therapeutic use, administration can begin at the time of tumor detection or surgical removal. This is followed by boost doses until at least the symptoms are substantially reduced and then for a certain period of time.

[0148] Pharmaceutical compositions (e.g., vaccine compositions) for therapeutic treatment are intended for parenteral, topical, nasal, oral, or local administration. The pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against the tumor. Parenteral administration compositions containing a solution of neoantigens are disclosed herein, and the vaccine compositions are dissolved or suspended in an acceptable carrier, for example, an aqueous carrier. Various aqueous carriers, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, etc. can be used. These compositions can be sterilized by conventional well-known sterilization techniques or can be sterile filtered. The resulting aqueous solution can be packaged as is for use or can be lyophilized, and the lyophilized preparation is combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances necessary to approximate physiological conditions, such as pH adjusters and buffers, tonicity agents, wetting agents, etc., for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, etc.

[0149] Neoantigens can also be administered via liposomes that target them to specific tissues such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigens to be delivered are incorporated as part of the liposome, either alone or together with a molecule that binds to a dominant receptor among, for example, lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Thus, liposomes filled with the desired neoantigens can be directed to the site of lymphoid cells, where the liposomes then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including phospholipids having neutral and negative charges, and sterols such as cholesterol. The choice of lipids is generally guided by considerations such as liposome size, acid lability, and stability of the liposome in the bloodstream. For example, various methods are available for preparing liposomes, as described in Szoka et al., Ann.Rev.Biophys.Bioeng.9;467 (1980), U.S. Patent Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.

[0150] For targeting to immune cells, the ligand to be incorporated into the liposome can include, for example, an antibody or fragment thereof specific for the cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at a dose that varies, inter alia, according to the mode of administration, the peptide being delivered, and the stage of the disease being treated.

[0151] For therapeutic or immunization purposes, the peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient. Numerous methods are conveniently used to deliver nucleic acids to a patient. By way of example, nucleic acids can be delivered directly as "naked DNA". This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), as well as in U.S. Patents Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery as described, for example, in U.S. Patent No. 5,204,253. Particles consisting of just DNA can be administered. Alternatively, DNA can be attached to particles such as gold particles. Approaches for delivering nucleic acid sequences can include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.

[0152] Nucleic acids can also be delivered complexed with cationic compounds such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372WOAWO 96 / 18372;9324640WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833;9106309WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).

[0153] Neoantigens can also be included in virus vector-based vaccine platforms, including but not limited to vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses of the second, third, or hybrid second / third generation, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61, Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18, Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690, Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880). Depending on the packaging capacity of the virus vector-based vaccine platforms described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The arrays may be adjacent non-mutated arrays, may be separated by linkers, or one or more arrays targeting intracellular compartments may precede (see, for example, Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8, Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41, Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express neoantigens, thereby eliciting a host immune (e.g., CTL) response to the peptides. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is BCG (Bacillus Calmette-Guerin). BCG vectors are described in Stover et al. (Nature 351:456-460 (1991)). A variety of other vaccine vectors useful for the therapeutic administration or immunization of neoantigens, such as typhoid vectors, will be apparent to those skilled in the art from the description herein.

[0154] The means of administering nucleic acids uses a mini-gene construct encoding one or more epitopes. To create a DNA sequence (mini-gene) encoding a selected CTL epitope for expression in human cells, the amino acid sequence of the epitope is reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are placed directly adjacent to create a continuous polypeptide sequence. Additional elements can be incorporated during mini-gene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the mini-gene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences proximal to the CTL epitope. The mini-gene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the mini-gene. Overlapping oligonucleotides (30-100 bases in length) are synthesized, phosphorylated, purified, and annealed using well-known techniques under appropriate conditions. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic mini-gene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.

[0155] Purified plasmid DNA can be prepared for injection using various formulations. The simplest of these is the reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively called protective, interactive, non-condensing (PINC) can also complex with purified plasmid DNA to affect variables such as stability, intramuscular dispersion, or transport to specific organs or cell types.

[0156] Also disclosed herein is a method for manufacturing a tumor vaccine, which includes performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising a plurality of neoantigens or a subset of a plurality of neoantigens.

[0157] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector (e.g., a vector containing at least one sequence encoding one or more neoantigens) disclosed herein can include culturing a host cell under conditions suitable for expressing the neoantigen or vector, wherein the host cell contains at least one polynucleotide encoding the neoantigen or vector, and can include the step of purifying the neoantigen or vector. Standard purification methods include chromatographic techniques, electrophoretic techniques, immunological techniques, precipitation techniques, dialysis techniques, filtration techniques, concentration techniques, and chromatofocusing techniques.

[0158] The host cell can include Chinese hamster ovary (CHO) cells, NS0 cells, yeast, or HEK293 cells. The host cell can be transformed with one or more polynucleotides containing at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further includes a promoter sequence operably linked to at least one nucleic acid sequence encoding a neoantigen or vector. In certain embodiments, the isolated polynucleotide can be cDNA.

[0159] VI. Identification of Neoantigens VI.A. Identification of Neoantigen Candidates Research methods for NGS analysis of tumor and normal exomes and transcriptomes are described and applied in the space of neoantigen identification. 6,14,15. The following examples consider certain optimizations for greater sensitivity and specificity in the identification of neoantigens in a clinical setting. These optimizations can be grouped into two areas, those related to laboratory processes and those related to NGS data analysis.

[0160] VI.A.1. Optimization of Laboratory Processes The process improvements presented herein were developed for the evaluation of reliable cancer driver genes in a targeted cancer panel 16 and extended to the whole exome and whole transcriptome settings required for neoantigen identification to address the challenges in high-precision neoantigen discovery from low tumor content and small volume clinical specimens. Specifically, these improvements include: 1. Targeting deep (greater than 500x) native mean coverage across the tumor exome to detect mutations present at low variant allele frequencies due to either low tumor content or subclonal status. 2. As an example, less than 5% of bases covered at less than 100x so that the likelihood of missing potential neoantigens is minimized, a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for regions that are not sufficiently covered 3. Targeting uniform coverage across the normal exome such that less than 5% of bases are covered at less than 20x so that the likelihood that potential neoantigens remain unclassified with respect to somatic / germline status (and thus not usable as TSNA) is minimized. 4. To minimize the total amount of sequencing required, the sequence capture probes are designed for only the coding regions of genes since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing18 。 b. Exclusion of genes predicted to produce few or no candidate neoantigens due to factors such as insufficient expression, suboptimal digestion by the proteasome, or aberrant sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples is extracted using probe-based enrichment with the same or similar probes used to capture exosomes in DNA. 19 using.

[0161] VI.A.2. Optimization of NGS Data Analysis Improvements in analysis methods address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically consider customizations relevant for the identification of neoantigens in a clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment (since it contains multiple MHC region assemblies that better reflect population polymorphisms as opposed to previous genome releases). 2. Overcoming the limitations of single variant caller 20 by merging results from various programs 5 . a. Single nucleotide variants and indels are detected from tumor DNA, tumor RNA, and normal DNA using a series of tools including: Strelka 21 and Mutec t22 programs based on the comparison of tumor and normal DNA; and programs that incorporate tumor DNA, tumor RNA, and normal DNA, such as 23 UNCeqR, which are particularly advantageous in low purity samples. b. Indels are determined using programs that perform local reassembly, such as Strelka and ABRA 24 . c. Structural rearrangements are detected using Pindel 25 or Breakseq26 It is determined using dedicated tools such as 3. To detect and block sample swaps, variant calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. As an example, extensive filtering of artificial calls is performed as follows: a. Removal of mutations found in normal DNA with lenient detection parameters in the case of potentially low coverage, and with permissive proximity criteria in the case of indels. b. Removal of mutations due to low mapping quality or low base quality 27 . c. Removal of mutations resulting from recurrent sequencing artifacts even if not observed in the corresponding normal. 27 . Examples include mutations detected primarily on one strand. d. Removal of mutations detected in an unrelated set of controls 27 . 5. seq2HLA 28 , ATHLATES 29 , or one of Optitype is used, and also exome and RNA sequencing data are combined 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing such as long-read DNA sequencing 30 , or the adaptation of methods for ligating RNA fragments to maintain continuity 31 . 6. Robust detection of neo-ORFs arising from tumor-specific splice variants is performed by assembling transcripts from RNA-seq data using CLASS 32 , Bayesembler 33 , StringTie 34 , or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate their entire transcripts from each experiment). Cufflinks 35is commonly used for this purpose, but it frequently generates a prohibitively large number of splice variants, many of which are much shorter than the full-length gene and may not be able to recover a simple positive control. The coding sequence and potential nonsense mutation-dependent decay mechanisms were determined using tools such as SpliceR 36 and MAMBA 37 . Gene expression was determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels were determined using tools developed for these purposes, such as ASE 38 or HTSeq 39 . Potential filtering steps include: a. Removal of candidate nascent ORFs that are considered to be poorly expressed. b. Removal of candidate nascent ORFs predicted to cause nonsense mutation-dependent decay (NMD). 7. Candidate neoantigens (e.g., nascent ORFs) that are only observed in RNA and cannot be directly verified as tumor-specific are classified as likely to be tumor-specific according to additional parameters by considering, for example: a. The presence of support for cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmation of trans-acting mutations in tumor DNA only in splicing factors. For example, in three independently published experiments with the R625 mutant SF3B1, the genes that exhibited the most differential splicing were consistent 40 even though one experiment examined uveal melanoma patients 41 , the second experiment examined uveal melanoma cell lines 42 , and the third experiment examined breast cancer patients. c. For novel splicing isoforms, the presence of confirmation of "novel" splice-junction reads in RNASeq data. d. For novel rearrangements, the presence of validation of exon - proximal reads in tumor DNA that is not present in normal DNA. e. GTEx 43 Absence from gene expression summaries such as (i.e., making the likelihood of germline origin lower). 8. To directly avoid alignment - and annotation - based errors and artifacts, complementation of reference - genome - alignment - based analysis (e.g., for somatic mutations occurring near germline mutations or repeat - context indels) by comparing tumor and normal reads of the assembled DNA (or k - mers derived from such reads).

[0162] In samples with polyadenylated RNA, the presence of viral and microbial RNA in RNA - seq data is evaluated using RNA CoMPASS44 or a similar method towards the identification of additional factors that can predict patient response.

[0163] VI.B. Isolation and Detection of HLA Peptides Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) methods after lysis and solubilization of tissue samples. 55~58 The clarified lysate was used for HLA - specific IP.

[0164] Immunoprecipitation was performed using antibodies coupled to beads that are specific for HLA molecules. For pan - class I HLA immunoprecipitation, pan - class I CR antibodies were used, and for class II HLA - DR, HLA - DR antibodies were used. Antibodies were covalently attached to NHS - Sepharose beads during an overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60Immunoprecipitation can also be performed using antibodies that are not covalently bound to beads. Generally, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to retain the antibody on the column. Some antibodies that can be used to selectively enrich MHC / peptide complexes are shown below. TIFF0007712970000001.tif42149

[0165] The clarified tissue lysate is added to the antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is stored for additional experiments including additional IP. Using standard techniques, the IP beads are washed to remove non-specific binding and the HLA / peptide complex is eluted from the beads. The protein components are removed from the peptides using a molecular weight spin column or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and in some cases stored at -20 °C prior to MS analysis.

[0166] The dried peptides are reconstituted in an HPLC buffer suitable for reverse-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution in a Fusion Lumos mass spectrometer (Thermo). The MS1 spectrum of the peptide mass / charge (m / z) is collected at high resolution in an Orbitrap detector and then the MS2 low-resolution scan is collected in an ion trap detector after HCD fragmentation of the selected ions. Additionally, the MS2 spectrum can be obtained using either the CID or ETD fragmentation method, or any combination of three techniques to obtain greater amino acid coverage of the peptide. The MS2 spectrum can also be measured at high-resolution mass accuracy in an Orbitrap detector.

[0167] The MS2 spectrum from each analysis is searched against a protein database using Comet 61、62 and peptide identification is performed using Percolator63~65 Score using it. Further sequencing can be performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or sequencing methods including spectral matching and de novo sequencing 75 can be used.

[0168] VI.B.1. Study of the MS detection limit for comprehensive HLA peptide sequencing Using the peptide YVYVADVAAK, what is the detection limit was determined using various amounts of peptide loaded on the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol. (Table 1) The results are shown in Figure 1F. These results show that the lowest limit of detection (LoD) is in the attomolar range (10 -18 ), the dynamic range extends over 5 digits, and the signal-to-noise appears to be sufficient for sequencing in the low femtomolar range (10 -15 ).

[0169] TIFF0007712970000002.tif56128

[0170] VII. Presentation model VII.A. Overview of the system Figure 2A is an overview of an environment 100 for determining the likelihood of peptide presentation in a patient according to one embodiment. The environment 100 provides a context for introducing a presentation specific system 160 that itself includes a presentation information storage device 165.

[0171] The presentation prediction system 160 is one or more computer models embodied in a computer computing system as discussed below with respect to FIG. 29, and receives a peptide sequence related to a set of MHC alleles, and determines the likelihood that the peptide sequence will be presented by one or more of the related set of MHC alleles. The presentation prediction system 160 can be applied to both class I and class II MHC alleles. This is useful in a variety of contexts. One example of a specific use of the presentation prediction system 160 is to receive a nucleotide sequence of a candidate neoantigen related to a set of MHC alleles from tumor cells of a patient 110, and determine that the candidate neoantigen will be presented by one or more of the related MHC alleles of the tumor and / or is likely to induce an immunogenic response in the immune system of patient 110. Those candidate neoantigens having a high likelihood when determined by system 160 can be selected for inclusion in vaccine 118, such that such an anti-tumor immune response can be elicited from the immune system of patient 110 that provides the tumor cells. Additionally, it is possible to generate T cells having a TCR reactive to a candidate neoantigen having a high presentation likelihood for use in T cell therapy, thereby also eliciting an anti-tumor immune response from the immune system of patient 110.

[0172] The presentation likelihood is determined by the presentation system 160 through one or more presentation models. Specifically, the presentation model generates a likelihood of whether a given peptide sequence will be presented for a set of related MHC alleles, and the likelihood is generated based on presentation information stored in the storage device 165. For example, the presentation model can generate a likelihood of whether the peptide sequence "YVYVADVAAK" will be presented for the set of alleles HLA-A*02:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:03, HLA-C*01:04 on the cell surface of a sample. The presentation information 165 contains information about whether these peptides bind to various types of MHC alleles such that the peptides are presented by the MHC alleles, which is determined in the model according to the positions of the amino acids in the peptide sequence. The presentation model can predict, based on the presentation information 165, whether an unrecognized peptide sequence will bind and be presented with a related set of MHC alleles. As described above, the presentation model can be applied to both class I and class II MHC alleles.

[0173] VII.B. Presentation Information FIG. 2 illustrates a method of obtaining presentation information according to one embodiment. The presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. The allele interaction information includes information that affects the presentation of peptide sequences and depends on the type of MHC allele. The allele non-interaction information includes information that affects the presentation of peptide sequences and is independent of the type of MHC allele.

[0174] VII.B.1. Allele Interaction Information The allergenic interaction information includes a specified peptide sequence, which is known to be presented mainly by one or more specified MHC molecules derived from humans, mice, etc. Notably, this may or may not include data obtained from tumor samples. The presented peptide sequence may be identified from cells expressing a single MHC allele. In this example, the presented peptide sequence is generally collected from a single allele cell line that has been engineered to express a pre-determined MHC allele and then exposed to a synthetic protein. The peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows an exemplary peptide presented on the pre-determined MHC allele HLA-DRB1*12:01 TIFF0007712970000003.tif4128 is isolated and identified by mass spectrometry, showing this example. In this context, since the peptide is identified through cells engineered to express a single pre-determined MHC protein, the direct relationship between the presented peptide and the MHC protein to which it binds is deterministically known.

[0175] The presented peptide sequence may also be collected from cells expressing multiple MHC alleles. Typically in humans, six different types of MHC I molecules and up to twelve different types of MHC II molecules are expressed in cells. Such presented peptide sequences may be identified from multiple allele cell lines that have been engineered to express multiple pre-determined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal tissue samples or tumor tissue samples. In this particular example, the MHC molecules can be immunoprecipitated from normal or tumor tissue. The peptides presented on multiple MHC alleles can likewise be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows six exemplary peptides TIFF0007712970000004.tif11128 is presented on the identified class I MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, HLA-B*08:01, and class II MHC alleles HLA-DRB1*10:01, HLA-DRB1:11:01, isolated, and identified by mass spectrometry, as shown in this example. In contrast to single allele cell lines, since the bound peptides are isolated from the MHC molecules before they are identified, the direct relationship between the presented peptides and the MHC proteins to which they are bound may be unknown.

[0176] Allele interaction information can also include mass spectrometry ion current, which depends on both the concentration of the peptide-MHC molecule complex and the ionization efficiency of the peptide. Ionization efficiency varies for each peptide in a sequence-dependent manner. Generally, the ion efficiency varies for each peptide over approximately two orders of magnitude, while the concentration of the peptide-MHC complex varies over a much larger range.

[0177] Allele interaction information can also include a measured or predicted value of the binding affinity between a given MHC allele and a given peptide. One or more affinity models can generate such predicted values (72,73,74). For example, returning to the example shown in FIG. 1D, the presentation information 165 can include a predicted binding affinity value of 1000 nM between the peptide YEMFNDKSF and the class I allele HLA-A * 01:01. Peptides with an IC50 > 1000 nm are presented by the MHC only rarely, and lower IC50 values increase the probability of presentation. The presentation information 165 can include a predicted binding affinity value between the peptide KNFLENFIESOFI and the class II allele HLA-DRB1:11:01.

[0178] Allergen interaction information can also include a measured or predicted value of the stability of the MHC complex. One or more stability models can generate such predicted values. More stable peptide-MHC complexes (i.e., complexes with a longer half-life) are more likely to be presented at high copy numbers on tumor cells and on antigen-presenting cells that encounter the vaccine antigen. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include a predicted stability value for a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a predicted stability value for the half-life of the class II molecule HLA-DRB1:11:01.

[0179] Allergen interaction information can also include the measured or predicted rate of the peptide-MHC complex formation reaction. Complexes that form at a faster rate are more likely to be presented on the cell surface at high concentrations.

[0180] Allergen interaction information can also include the sequence and length of the peptide. MHC class I molecules typically prefer to present peptides having a length of 8-15 peptides. 60-80% of the presented peptides have a length of 9. MHC class II molecules generally tend to present peptides having a length of 6-30 peptides.

[0181] Allergen interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoding peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoding peptide. The presence of kinase motifs affects the probability of post-translational modifications that can enhance or interfere with MHC binding.

[0182] Allergen interaction information can also include the expression or activity levels of proteins involved in the post-translational modification process (when measured or predicted by RNA seq, mass spectrometry, or other methods), such as kinases.

[0183] Allele interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing a particular MHC allele, as evaluated by mass spectrometry proteomics or other means.

[0184] Allele interaction information can also include the expression level of a particular MHC allele in the individual in question (e.g., as measured by RNA-seq or mass spectrometry). Peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.

[0185] Allele interaction information can also include the overall probability of presentation by a particular MHC allele, independent of the neoantigen-encoding peptide sequence, in other individuals expressing that particular MHC allele.

[0186] Allele interaction information can also include the overall probability of presentation by MHC alleles of molecules of the same family (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals, independent of the peptide sequence. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and thus the presentation of peptides by HLA-C is a priori less probable than presentation by HLA-A or HLA-B II. As another example, since HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, it is presumed that the presentation of peptides by HLA-DP is less probable than presentation by HLA-DR or HLA-DQ.

[0187] Allele interaction information can also include the protein sequence of a particular MHC allele.

[0188] Any MHC allele non-interaction information listed in the sections below can also be modeled as MHC allele interaction information.

[0189] VII.B.2. Allele-noninteracting information Allele-noninteracting information can include the C-terminal sequence adjacent to the neoantigen-encoding peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence can affect the proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters the MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and thus the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include the C-terminal flanking sequence FOEIFNDKSLDKFJI of the presented peptide FJIEJFOESS identified from the source protein of the peptide.

[0190] Allele-noninteracting information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide the mass spectrometry training data. As described later with respect to FIG. 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, the mRNA quantification measurements are identified from the software tool RSEM. Details of the RSEM software tool implementation can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).

[0191] Allele-noninteracting information can also include the N-terminal sequence adjacent to the peptide within its source protein sequence.

[0192] The allelic non-interaction information can also include the source gene of the peptide sequence. The source gene can be defined as the Ensembl protein family of the peptide sequence. In other examples, the source gene can be defined as the source DNA or source RNA of the peptide sequence. The source gene can be represented, for example, as a string of nucleotides encoding a protein, or alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode a particular protein. In another example, the allelic non-interaction information can also include the source transcript or isoform of the peptide sequence or a set of potential source transcripts or isoforms extracted from a database such as Ensembl or RefSeq.

[0193] The allelic non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.

[0194] The allelic non-interaction information can also include the presence of protease cleavage motifs in the peptide, optionally weighted according to the expression of the corresponding protease in tumor cells (when measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and thus are likely to be less stable intracellularly and therefore less likely to be presented.

[0195] The allelic non-interaction information can also include the metabolic turnover rate of the source protein when measured in an appropriate cell type. A faster metabolic turnover rate (i.e., a lower half-life) increases the probability of presentation, but this property has low predictive power when measured in dissimilar cell types.

[0196] Allele non-interaction information can also include the length of the source protein, optionally considering the specific splice variant ("isoform") that is most highly expressed in tumor cells when measured by RNA-seq or proteomic mass spectrometry, or when predicted from the annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.

[0197] Allele non-interaction information can also include the level of expression of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry). Different proteasomes have different preferences for cleavage sites. A greater weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.

[0198] Allele non-interaction information can also include the expression of the source gene of the peptide (e.g., when measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in the tumor sample. Peptides derived from more highly expressed genes are more likely to be presented. Peptides derived from genes with undetectable levels of expression can be excluded from consideration.

[0199] Allele non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to a nonsense-mediated decay mechanism, such as predicted by a model from Rivas et al, Science 2015.

[0200] Allergen non-interaction information can also include the typical tissue-specific expression of the source gene of the peptide during various stages of the cell cycle. Genes that are expressed at generally low levels (when measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are presented than genes that are stably expressed at very low levels.

[0201] Allergen non-interaction information can also include a comprehensive catalog of the properties of the source protein, such as those provided in UniProt or PDB http: / / www.rcsb.org / pdb / home / home.do. These properties can include, among others, the secondary and tertiary structure of the protein, intracellular localization 11, gene ontology (GO) terms. Specifically, this information can contain annotations that act at the protein level, such as the 5'UTR length, and annotations that act at the level of specific residues, such as a helix motif at residues 300-310. These properties can also include turn motifs, sheet motifs, and disordered residues.

[0202] Allergen non-interaction information can also include properties that describe the nature of the domain of the source protein that contains the peptide, such as secondary or tertiary structure (e.g., α-helix vs. β-sheet); alternative splicing.

[0203] The allergen non-interaction information can also include an association between the peptide sequence of the neoantigen and one or more k-mer blocks of the source gene of the neoantigen (as present in the nucleotide sequencing data of interest). During training of the presentation model, these associations between the peptide sequence of the neoantigen and the k-mer blocks of the nucleotide sequencing data of the neoantigen are input into the model, and the model uses a part of them to learn model parameters representing the presence or absence of presentation hotspots in the k-mer blocks associated with the training peptide sequence. Then, during use of the trained model, an association between the test peptide sequence and one or more k-mer blocks of the source gene of the test peptide sequence is input into the model, and the presentation model can make a more accurate prediction of the presentation likelihood of the test peptide sequence based on the parameters learned by the model during training.

[0204] Generally, the parameter of the model representing the presence or absence of a presentation hot spot in a k-mer block represents the residual tendency for the k-mer block to produce a presented peptide after controlling for all other variables (e.g., peptide sequence, RNA expression, amino acids commonly found in HLA-binding peptides). The parameter representing the presence or absence of a presentation hot spot in a k-mer block can be a binary coefficient (e.g., 0 or 1), or an analog coefficient along a specific scale (e.g., inclusively 0 to 1). In either case, the larger the coefficient (e.g., closer to 1 or 1), the higher the likelihood of producing a presented peptide that controls for other factors, and the smaller the coefficient (e.g., closer to 0 or 0), the lower the likelihood that the k-mer block will produce a presented peptide. For example, a k-mer block with a low hot spot coefficient may be a k-mer block from a gene that shows high RNA expression for amino acids commonly found in HLA-binding peptides, in which case the source gene produces many other presented peptides, but few presented peptides are seen within the k-mer block. Since other sources of peptide presence are already accounted for by other parameters (e.g., k-mer blocks commonly found in HLA-binding peptides or RNA expression at a larger unit), these hot spot parameters provide distinct new information that does not "double count" the information captured by other parameters.

[0205] Allele non-interaction information can also include the probability of presentation of peptides from the source protein of the peptide in question in other individuals (after adjusting for the expression levels of the source protein in those individuals and the effects of the various HLA types of those individuals).

[0206] Allele non-interaction information can also include the probability that a peptide will not be detected by mass spectrometry or will be overrepresented due to technical bias.

[0207] Expression of various gene modules / pathways (which need not contain the source protein of the peptide) as measured by gene expression assays such as RNASeq, microarrays, targeted panels such as Nanostring, or single / multiple genes representative of gene modules measured by assays such as RT-PCR, providing information about the state of tumor cells, stroma, or tumor infiltrating lymphocytes (TIL).

[0208] Allele non-interaction information can also include the copy number of the source gene of the peptide in tumor cells. For example, peptides derived from genes that are subject to homozygous deletion in tumor cells can be assigned a presentation probability of zero.

[0209] Allele non-interaction information can also include the probability that a peptide binds to TAP, or the measured or predicted binding affinity of the peptide for TAP. Peptides that are more likely to bind to TAP, or that bind to TAP with a higher affinity, are more likely to be presented by MHC-I.

[0210] Allele non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, immunohistochemistry). In MHC-I, higher TAP expression levels increase the probability of presentation of all peptides.

[0211] Allele non-interaction information can also include, but is not limited to, the presence or absence of tumor mutations: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, NTRK3. ii. in a gene encoding a protein involved in the antigen presentation machinery (for example, any of the genes encoding B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or a component of the proteasome or immunoproteasome). Peptides whose presentation depends on components of the antigen presentation machinery that are under the influence of loss-of-function mutations in the tumor have a reduced probability of presentation.

[0212] The presence or absence of functional germline polymorphisms, including but not limited to: i. in a gene encoding a protein involved in the antigen presentation machinery (for example, any of the genes encoding B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or a component of the proteasome or immunoproteasome).

[0213] Allele non-interaction information can also include tumor type (for example, NSCLC, melanoma).

[0214] The allelic non-interaction information can also include known functionality of HLA alleles, such as that reflected by HLA allele suffixes. For example, the suffix N in the allele name HLA-A*24:09N indicates a null allele that does not express and thus has a low likelihood of presenting epitopes; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.

[0215] The allelic non-interaction information can also include clinical tumor subtypes (e.g., squamous cell lung cancer vs. non-squamous cell).

[0216] The allelic non-interaction information can also include smoking history.

[0217] The allelic non-interaction information can also include history of sunburn, sunlight exposure, or exposure to other mutagens.

[0218] The allelic non-interaction information can also include local expression of the source gene of the peptide in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be presented.

[0219] The allelic non-interaction information can also include the frequency of mutations in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.

[0220] In the case of an example of a mutated tumor-specific peptide, a list of the characteristics used to predict the probability of presentation may also include the annotation of the mutation (e.g., missense, read-through, frameshift, fusion, etc.), or whether the mutation is predicted to result in a nonsense-mediated decay (NMD) mechanism. For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in a decrease in mRNA translation, which decreases the probability of presentation.

[0221] VII.C. Presentation-Specific System FIG. 3 is a high-level block diagram illustrating the computer logic components of a presentation-specific system 160 according to one embodiment. In this exemplary embodiment, the presentation-specific system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. The presentation-specific system 160 is also composed of a training data storage device 170 and a presentation model storage device 175. Some embodiments of the model management system 160 have modules different from those described herein. Similarly, the functions can be distributed among the modules in a manner different from that described herein.

[0222] VII.C.1. Data Management Module The data management module 312 generates a set of training data 170 from the presentation information 165. Each set of training data contains a number of data examples, and each data example i contains at least a peptide sequence p that is presented or not presented i and one or more associated MHC alleles a i combined with the peptide sequence p i and a dependent variable y representing the information that the presentation-specific system 160 is interested in predicting a new value of the independent variable i and contains a set of independent variables z i .

[0223] In one particular implementation referred to throughout the remainder of this specification, the dependent variable y i is a binary label indicating whether the peptide p i is presented by one or more associated MHC alleles a i . However, in other implementations, it is recognized that the dependent variable y i can represent any other type of information that the presentation prediction system 160 is interested in predicting depending on the independent variable z i . For example, in another implementation, the dependent variable y i may also be a numerical value indicating a mass spectrometry ion current specified for a data example.

[0224] The peptide sequence p for data example i i is a sequence of k i amino acids, where k i can vary within a certain range among data examples i. For example, the range can be 8 - 15 for MHC class I or 6 - 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p i in the training data set can have the same length, e.g., 9. The number of amino acids in the peptide sequence can vary depending on the type of MHC allele (e.g., MHC alleles in humans, etc.). The MHC allele a for data example i i indicates which MHC allele was present in combination with the corresponding peptide sequence p i .

[0225] The data management module 312 can also include additional allele interaction variables such as predicted values of the binding affinity b i and stability s i along with the peptide sequence p i and the bound MHC allele a i contained in the training data 170. For example, the training data 170 contains the predicted binding affinity values b i for each of the bound MHC molecules shown in a i with the peptide p imay contain. As another example, the training data 170 may contain a i stability prediction value s for each of the MHC alleles shown in i .

[0226] The data management module 312 may also include the peptide sequence p i along with allelic non-interaction variables w such as the C-terminal flanking sequence and mRNA quantification values i .

[0227] The data management module 312 also identifies peptide sequences not presented by the MHC allele and generates training data 170. Generally, this involves identifying the "longer" sequence of the source protein containing the peptide sequence to be presented, prior to presentation. If the presentation information contains an engineered cell line, the data management module 312 identifies a series of peptide sequences in the synthetic protein to which the cell was exposed but which were not presented on the MHC alleles of the cell. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein that is the origin of the presented peptide sequence and identifies a series of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.

[0228] The data management module 312 also artificially generates peptides having random sequences of amino acids and identifies the generated sequences as peptides not presented on the MHC allele. This can be achieved by randomly generating peptide sequences and enables the data management module 312 to easily generate large amounts of synthetic data for peptides not presented on the MHC allele. In practice, since a small percentage of peptide sequences are presented by the MHC allele, synthetically generated peptide sequences are very likely not to be presented by the MHC allele, even if they are contained in proteins processed by the cell.

[0229] FIG. 4 illustrates an exemplary set of training data 170A according to one embodiment. Specifically, the first three data examples in the training data 170A are a single allele cell line containing allele HLA-C*01:03, as well as three peptide sequences showing peptide presentation information from TIFF0007712970000005.tif4128. The fourth data example in the training data 170A shows a multiple allele cell line containing alleles HLA-B*07:02, HLA-C*01:03, HLA-A*01:01, and peptide information from the peptide sequence QIEJOEIJE. The first data example shows that the peptide sequence QCEIOWARE was not presented by allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, the negatively labeled peptide sequences may be randomly generated by the data management module 312 or identified from the source protein of the presented peptides. The training data 170A also includes binding affinity prediction values of 1000 nM and stability prediction values of a half-life of 1 hour for the peptide sequence-allele pairs. The training data 170A also includes the C-terminal flanking sequence of the peptide FJELFISBOSJFIE, and 10 2 also includes allele non-interaction variables such as mRNA quantification measurements of TPM. The fourth data example shows that the peptide sequence QIEJOEIJE was presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. The training data 170A also includes binding affinity prediction values and stability prediction values for each of the alleles, as well as the C-terminal flanking sequence of the peptide and mRNA quantification measurements for the peptide.

[0230] VII.C.2. Encoding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more presentation models. In one implementation, the encoding module 314 encodes an array (e.g., a peptide array or a C-terminal flanking array) in one-hot for a predefined 20-letter amino acid alphabet. Specifically, a peptide sequence p i having k i amino acids is represented as a row vector of 20·k i elements, and the single element in p i 20·(j-1)+1 , p i 20·(j-1)+2 ,..., p i 20·j corresponding to the amino acid at the j-th position of the peptide sequence has a value of 1. All other remaining elements have a value of 0. As an example, for a given alphabet {A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y}, the peptide sequence EAF of 3 amino acids for data example i can be represented as a row vector of 60 elements represented by TIFF0007712970000006.tif18147. The C-terminal flanking sequence c i , as well as the protein sequence d h for the MHC allele, and other array data in the presentation information can be encoded in the same manner as described above.

[0231] If the training data 170 contains amino acid sequences of different lengths, the encoding module 314 can further encode the peptides into vectors of equal length by adding PAD characters to extend the predefined alphabet. For example, this can be done by left-padding the peptide sequences with PAD characters until the length of the peptide sequence reaches the length of the peptide sequence with the maximum length in the training data 170. Thus, if the peptide sequence with the maximum length has k 最大 amino acids, the encoding module 314 can encode each sequence into a vector of (20 + 1)·k 最大Numerically represent as a row vector of elements. As an example, the extended alphabet {PAD, A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y} and k 最大 = 5 for the maximum amino acid length, for the same exemplary peptide sequence EAF of 3 amino acids, is a row vector of 105 elements TIFF0007712970000007.tif18147. The C-terminal flanking sequence c i or other sequence data can, similarly, be encoded as described above. Thus, for the peptide sequence p i or c i each independent variable or column in represents the presence of a particular amino acid at a particular position in the array.

[0232] The above-described method of encoding array data has been described with respect to arrays having amino acid sequences, but the method can, similarly, be extended to other types of array data, such as, for example, DNA or RNA sequence data.

[0233] The encoding module 314 also encodes one or more MHC alleles a i for data example i into a row vector of m elements, where each element h = 1, 2,..., m corresponds to a unique specified MHC allele. The element corresponding to the MHC allele specified for data example i has a value of 1. The remaining elements, the rest, have a value of 0. As an example, for alleles HLA - B*07:02 and HLA - DRB1*10:01 for data example i corresponding to a plurality of allele cell lines among the unique specified MHC allele types {HLA - A*01:01, HLA - C*01:08, HLA - B*07:02, HLA - DRB1*10:01} with m = 4, can be represented by a row vector a i = [0 0 1 1], where a3 i = 1 and a4 i = 1. Examples for four specified MHC allele types are described herein, but the number of MHC allele types can, in reality, be in the hundreds or thousands. As described above, each data example i typically has a peptide sequence pi includes up to six different MHC allele types in relation to

[0234] The encoding module 314 also encodes the label y for each data example i i as a binary variable having a value from the set {0,1}, where a value of 1 indicates that the peptide x i is presented by one of the associated MHC alleles a i and a value of 0 indicates that the peptide x i is not presented by any of the associated MHC alleles a i . When the dependent variable y i represents a mass spectrometry ion current, the encoding module 314 can additionally scale the values using various functions such as a log function having a range of [-∞,∞] for ion current values between [0,∞].

[0235] The encoding module 314 can represent the pair of the peptide p i and the allele interaction variable x for the associated MHC allele h h i as a row vector in which the numerical representations of the allele interaction variables are concatenated one after another. For example, the encoding module 314 can represent x h i as a row vector equivalent to [p i , [p i b h i , [p i s h i , or [p i b h i s h i , where b h i is the predicted binding affinity value for the peptide pi and the associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of the allele interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0236] In one example, the encoding module 314 represents binding affinity information by incorporating values measured or predicted for binding affinity into the allelic interaction variable x h i

[0237] In one example, the encoding module 314 represents binding stability information by incorporating values measured or predicted for binding stability into the allelic interaction variable x h i

[0238] In one example, the encoding module 314 represents binding on-rate information by incorporating values measured or predicted for binding on-rate into the allelic interaction variable x h i

[0239] In one example, for a peptide presented by a class I MHC molecule, the encoding module 314 represents the peptide length as a vector TIFF0007712970000008.tif4128 (where TIFF0007712970000009.tif3128 is an indicator function and L k means the length of peptide p k ). The vector T k can be included in the allelic interaction variable x h i In another example, for a peptide presented by a class II MHC molecule, the encoding module 314 represents the peptide length as a vector TIFF0007712970000010.tif18146 (where TIFF0007712970000011.tif3128 is an indicator function and L k means the length of peptide p k ). The vector T k can be included in the allelic interaction variable x h ​​​i can be included.

[0240] In one example, the encoding module 314 represents the RNA expression information of MHC alleles by incorporating the RNA-seq-based expression levels of MHC alleles into the allelic interaction variable xhi.

[0241] Similarly, the encoding module 314 can represent the allelic non-interaction variable w i as a row vector in which the numerical representations of the allelic non-interaction variables are chained one after another. For example, w i can be a row vector equivalent to [c i or [c i m i w i , and w i is a row vector representing the C-terminal flanking sequence of the peptide pi and any other allelic non-interaction variables in addition to the mRNA quantification measurement value m i related to the peptide. Alternatively, one or more combinations of allelic non-interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0242] In one example, the encoding module 314 represents the metabolic turnover rate of the source protein for the peptide sequence by incorporating the metabolic turnover rate or half-life into the allelic non-interaction variable w i .

[0243] In one example, the encoding module 314 represents the length of the source protein or isoform by incorporating the protein length into the allelic non-interaction variable w i .

[0244] In one example, the encoding module 314 represents the average expression of immunoproteasome-specific proteasome subunits including the β1 i , β2 i , β5 i subunits into the allelic non-interaction variable w iBy incorporating it, it represents the activation of immunoproteasome.

[0245] In one example, the coding module 314 represents the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM), or the source protein of the gene or transcript of the peptide, by incorporating the abundance of the source protein into the allelic non-interaction variable w i By incorporating it.

[0246] In one example, the coding module 314 represents the probability that the transcript of origin of a peptide will undergo nonsense-mediated decay (NMD), as estimated by a model such as in Rivas et.al. Science, 2015, by incorporating this probability into the allelic non-interaction variable w i By incorporating it.

[0247] In one example, the coding module 314 represents the activation status of a gene module or pathway evaluated via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM for each gene in the pathway, and then computationally calculating summary statistics, such as the mean, across the genes in the pathway. The mean can be incorporated into the allelic non-interaction variable w i Can be incorporated into.

[0248] In one example, the coding module 314 represents the copy number of the source gene by incorporating the copy number into the allelic non-interaction variable w i By incorporating it.

[0249] In one example, the coding module 314 represents the measured or predicted TAP binding affinity (e.g., in nanomolar units) by including it in the allelic non-interaction variable w i To represent the TAP binding affinity.

[0250] In one example, the encoding module 314 represents the TAP expression level by including the TAP expression level measured by RNA-seq (and quantified, e.g., in units of TPM by RSEM) as an allelic non-interaction variable w i in it.

[0251] In one example, the encoding module 314 represents the tumor mutation as a vector of indicator variables in the allelic non-interaction variable w i (i.e., if the peptide p k is derived from a sample having a KRAS G12D mutation, then d k = 1, otherwise 0).

[0252] In one example, the encoding module 314 represents the germline polymorphism in the antigen presentation gene as a vector of indicator variables (i.e., if the peptide p k is derived from a sample having a germline polymorphism specific to TAP, then d k = 1). These indicator variables can be included in the allelic non-interaction variable w i .

[0253] In one example, the encoding module 314 represents the tumor type as a length-1 one-hot encoding vector for the alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot encoding variables can be included in the allelic non-interaction variable w i .

[0254] In one example, the encoding module 314 represents the MHC allele suffix by processing the 4-digit HLA allele with various suffixes. For example, HLA-A*24:09N is considered a different allele than HLA-A*24:09 for the purposes of the model. Alternatively, since HLA alleles ending with the N suffix do not express, the probability of presentation by MHC alleles with the N suffix can be set to zero for all peptides.

[0255] In one example, the encoding module 314 represents the tumor subtype as a length-1 one-hot encoding vector for the alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot encoding variables can be included in the allelic non-interaction variable w i can be included.

[0256] In one example, the encoding module 314 can include the smoking history in the allelic non-interaction variable w i as a binary indicator variable (d k = 1 if the patient has a smoking history, 0 otherwise). Alternatively, the smoking history can be encoded as a length-1 one-hot encoding variable for the alphabet of smoking severity. For example, the smoking status can be rated on a scale of 1 - 5, where 1 indicates a non-smoker and 5 indicates a current heavy smoker. Since the smoking history is mainly associated with lung tumors, when training a model for multiple tumor types, this variable can also be defined as 1 if the patient has a smoking history and the tumor type is a lung tumor, and 0 otherwise.

[0257] In one example, the encoding module 314 can include the sunburn history in the allelic non-interaction variable w i as a binary indicator variable (d k = 1 if the patient has a history of severe sunburn, 0 otherwise). Since severe sunburn is mainly associated with melanoma, when training a model for multiple tumor types, this variable can also be defined as 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and 0 otherwise.

[0258] In one example, the encoding module 314 represents the distribution of the expression levels of a particular gene or transcript for each gene or transcript in the human genome as summary statistics (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, for peptide p in a sample having the tumor type melanoma, the measured expression level of the gene or transcript of origin of peptide p can be included in the allelic non-interaction variable w, and the mean and / or median gene or transcript expression of the gene or transcript of origin of peptide p in melanoma as measured by TCGA can also be included. k For peptide p k the measured expression level of the gene or transcript of origin of peptide p can be included in the allelic non-interaction variable w i and the mean and / or median gene or transcript expression of the gene or transcript of origin of peptide p in melanoma as measured by TCGA can also be included. k

[0259] In one example, the encoding module 314 represents the mutation type as a one-hot encoding variable of length 1 for the alphabet of mutation types (e.g., missense, frameshift, NMD-inducing, etc.). These one-hot encoding variables can be included in the allelic non-interaction variable w i .

[0260] In one example, the encoding module 314 represents the protein-level characteristics of a protein as values of the annotation of the source protein (e.g., 5’UTR length) in the allelic non-interaction variable w i . In another example, the encoding module 314 represents the residue-level annotation of the source protein for peptide p i by including an indicator variable that is equivalent to 1 if peptide p i overlaps with the helix motif and 0 otherwise, or is equivalent to 1 if peptide p i is completely contained within the helix motif in the allelic non-interaction variable wi. In another example, the encoding module 314 represents the characteristic representing the proportion of residues in peptide p i contained within the helix motif annotation as an allelic non-interaction variable wi can be included in.

[0261] In one example, the encoding module 314 represents the type of protein or isoform in the human proteome as an indicator vector o having the same length as the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is 1 if the peptide p k is derived from protein i and 0 otherwise.

[0262] In one example, the encoding module 314 represents the source gene G = gene(p i ) of the peptide p as a categorical variable having L possible categories (where L indicates the upper limit of the number of indexed source genes 1, 2,..., L). i

[0263] In one example, the encoding module 314 represents the tissue type, cell type, tumor type, or tumor histology type T = tissue(p i ) of the peptide p as a categorical variable having M possible categories (where M indicates the upper limit of the number of indexed types 1, 2,..., M). Examples of tissue types include, for example, lung tissue, heart tissue, intestinal tissue, nerve tissue, etc. Examples of cell types include dendritic cells, macrophages, CD4 T cells, etc. Examples include lung adenocarcinoma, squamous cell carcinoma of the lung, melanoma, non-Hodgkin lymphoma, etc. i

[0264] The encoding module 314 can also represent the entire set of variables z i for the peptide p i and the related MHC allele h as a row vector in which the numerical representations of the allele interaction variable x i and the allele non-interaction variable w i are chained one after another. For example, the encoding module 314 represents z h ibe represented as a row vector equivalent to [x h i w i or [w i x h i .

[0265] VIII. Training Module The training module 316 constructs one or more presentation models that generate the likelihood that a peptide sequence will be presented by an MHC allele related to the peptide sequence. Specifically, for a peptide sequence p k and an MHC allele a k related to the peptide sequence p k is given, and each presentation model generates an estimated value u k indicating the likelihood that the peptide sequence p k will be presented by one or more of the related MHC alleles a k .

[0266] VIII.A. Overview The training module 316 constructs one or more presentation models based on a training data set stored in the storage device 170, which is generated from the presentation information stored in 165. Generally, regardless of the specific type of the presentation model, all of the presentation models capture the dependency between the independent variable and the dependent variable in the training data 170 such that the loss function is minimized. Specifically, the loss function TIFF0007712970000012.tif4128 represents the contradiction between the value of the dependent variable y i∈S for one or more data examples S in the training data 170 and the estimated likelihood u i∈S for the data example S generated by the presentation model. In one particular implementation referred to throughout the remainder of this specification, the loss function TIFF0007712970000013.tif4128 is the negative log-likelihood function given by the following equation (1a). TIFF0007712970000014.tif10128However, in practice, other loss functions may be used. For example, when predicting mass spectrometry ion currents, the loss function is the mean squared loss given by Equation 1b below. TIFF0007712970000015.tif10128

[0267] The presentation model can be a parametric model where one or more parameters θ mathematically specify the dependence between the independent and dependent variables. Typically, the loss function TIFF0007712970000016.tif4128The various parameters of the parametric type presentation model that minimizes the loss function are determined through gradient-based numerical optimization algorithms such as, for example, batch gradient algorithms, stochastic gradient algorithms, etc. Alternatively, the presentation model can be a non-parametric model where the model structure is determined from the training data 170 and is not strictly based on a fixed set of parameters.

[0268] VIII.B. Per-allele model The training module 316 can construct a presentation model for predicting the presentation likelihood of peptides on a per-allele basis. In this example, the training module 316 can train the presentation model based on the data example S in the training data 170 generated from cells expressing a single MHC allele.

[0269] In one implementation, the training module 316 TIFF0007712970000017.tif7128models the estimated presentation likelihood u k of the peptide p k for a particular allele h, where the peptide sequence x h k represents the peptide p k and the encoded allele interaction variable for the corresponding MHC allele h, f(·) is an arbitrary function, and for convenience of description, is referred to as a transformation function throughout this specification. Further, g h(·) is an arbitrary function, which for convenience of description is referred to as a dependency function throughout this specification, and the parameter θ determined for the MHC allele h h Based on the set of, the allele interaction variable x h k generates a dependency score for. The value of the set of parameters θ for each MHC allele h h can be determined by minimizing a loss function with respect to θ h where i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele h.

[0270] The dependency function g h (x h k ; θ h ) outputs a dependency score for the MHC allele h, indicating whether the MHC allele h presents the corresponding neoantigen based on at least the allele interaction characteristic x h k and in particular based on the position of the amino acid in the peptide sequence of the peptide p k . For example, the dependency score for the MHC allele h may have a high value when the MHC allele h is likely to present the peptide p k and may have a low value when the likelihood of presentation is low. The conversion function f(·) converts the input, and more specifically, in this example, the dependency score generated by g h (x h k ; θ h ) into an appropriate value indicating the likelihood that the peptide p k will be presented by the MHC allele.

[0271] In one particular implementation referred to throughout the remainder of this specification, f(·) is a function having a range within [0,1] for an appropriate domain range. In one example, f(·) is the expit function given by TIFF0007712970000018.tif10128. As another example, f(·) can also be the hyperbolic tangent function given by TIFF0007712970000019.tif4128. Alternatively, if the prediction is made for a mass spectrometry ion current having a value outside the range [0,1], f(·) can be any function such as, for example, the identity function, an exponential function, a log function, etc.

[0272] Thus, the per - allele likelihood that a peptide sequence p k will be presented by an MHC allele h can be generated by applying a dependency function g h (·) to the encoded version of the peptide sequence p k to generate a corresponding dependency score. The dependency score may be transformed by a transformation function f(·) such that it generates the per - allele likelihood that the peptide sequence p k will be presented by the MHC allele h.

[0273] VIII.B.1 Dependency Function for Allele Interaction Variables In one particular implementation referred to throughout this specification, the dependency function g h (·) is an affine function given by h k linearly combining each allele interaction variable in x with the corresponding parameter in a set of parameters θ h determined for the associated MHC allele h. TIFF0007712970000020.tif5128.

[0274] In another particular implementation referred to throughout this specification, the dependency function g h (·) is a network function given by h a network model NN TIFF0007712970000021.tif5128 having a series of nodes arranged in one or more layers. The nodes have parameters θh It can be connected to other nodes through connections each having related parameters in the set. The value at one particular node can be expressed as the sum of the values of the nodes connected to the particular node, weighted by the related parameters mapped by the activation function related to the particular node. In contrast to an affine function, the network model is advantageous because the presentation model can incorporate non-linearity and process data having amino acid sequences of different lengths. Specifically, through non-linear modeling, the network model can capture the interactions between amino acids at different positions in the peptide sequence and how this interaction affects peptide presentation.

[0275] Generally, the network model NN h (·) can be structured as a feed-forward network such as an artificial neural network (ANN), a convolutional neural network (CNN), a deep neural network (DNN), etc., and / or a recurrent network such as a long short-term memory network (LSTM), a bidirectional recurrent network, a deep bidirectional recurrent network, etc.

[0276] In one example referred to throughout the remainder of this specification, each MHC allele at h = 1, 2,..., m is associated with a separate network model, and NN h (·) means the output from the network model associated with MHC allele h.

[0277] Figure 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h = 3. As shown in Figure 5, the network model NN3(·) for MHC allele h = 3 includes three types of input nodes at layer l = 1, four types of nodes at layer l = 2, two types of nodes at layer l = 3, and one type of output node at layer l = 4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2),..., θ3(10). The network model NN3(·) receives input values (individual data examples including encoded polypeptide sequence data and any other training data used) for three allelic interaction variables x3 k (1), x3 k (2), and x3 k (3), and outputs the value NN3(x3 k ). The network function may include one or more network models that each take a different allelic interaction variable as input.

[0278] In another example, the specified MHC alleles h = 1, 2,..., m are associated with a single network model NN H (·), where NN h (·) represents one or more outputs of a single network model associated with MHC allele h. In such an example, the set of parameters θ h may correspond to the set of parameters for a single network model, and thus the set of parameters θ h may be shared by all MHC alleles.

[0279] Figure 6A illustrates an exemplary network model NN H (·) shared by MHC alleles h = 1, 2,..., m. As shown in Figure 6A, the network model NN H (·) includes m output nodes, each corresponding to an MHC allele. The network model NN3(·) receives the allelic interaction variable x3 k for MHC allele h = 3 and outputs the value NN3(x3k Outputs m values, including

[0280] In yet another example, a single network model NN H (·) is a network model that takes as input the allelic interaction variable x of MHC allele h h k and the encoded protein sequence d h and outputs a dependency score. In such an example, the set of parameters θ h can again correspond to the set of parameters for a single network model, and thus the set of parameters θ h can be shared by all MHC alleles. Thus, in such an example, NNh(·) is the output of the single network model NN h k d h given the input [x H (·). Such a network model is advantageous because it can correctly predict the peptide presentation probability for MHC alleles that were unknown in the training data simply by identifying their protein sequences.

[0281] Figure 6B illustrates an exemplary network model NN H (·) shared by MHC alleles. As shown in Figure 6B, the network model NN H (·) receives as input the allelic interaction variable and protein sequence of MHC allele h = 3 and outputs the dependency score NN3(x3 k ) corresponding to MHC allele h = 3.

[0282] In yet another example, the dependency function g h (·) can be represented as TIFF0007712970000022.tif5128, where g’ h (x h k ;θ’ his an affine function, a network function, etc. with a set of parameters θ’h, and represents the baseline probability of presentation for MHC allele h, the bias parameter θ in the set of parameters for the allelic interaction variable of the MHC allele h 0 is associated with.

[0283] In another embodiment, the bias parameter θ h 0 may be shared according to the gene family of MHC allele h. That is, the bias parameter θ for MHC allele h h 0 may be equivalent to θ 遺伝子(h) 0 and gene (h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family of "HLA-A", and the bias parameter θ for each of these MHC alleles h 0 may be shared. As another example, assign class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 to the gene family of "HLA-DRB", and the bias parameter θ for each of these MHC alleles h 0 can be shared.

[0284] As an example, returning to equation (2), the likelihood that peptide p h will be presented by MHC allele h = 3 among m = 4 different specified MHC alleles using the affinity-dependent function g k (·) is TIFF0007712970000023.tif5128 can be generated, where x3 k is the allelic interaction variable specified for MHC allele h = 3, and θ3 is the set of parameters determined for MHC allele h = 3 through loss function minimization.

[0285] As another example, for a peptide p among m = 4 different specified MHC alleles using separate network transformation functions gh(·), the likelihood that it will be presented by MHC allele h = 3 k would be can be generated by TIFF0007712970000024.tif5128, where x3 k is the allele interaction variable specified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.

[0286] Figure 7 illustrates the generation of the presentation likelihood of peptide p associated with MHC allele h = 3 using an exemplary network model NN3(·). As shown in Figure 7, the network model NN3(·) receives the allele interaction variable x3 k for MHC allele h = 3 and generates an output NN3(x3 k ). The output is mapped by a function f(·) to generate an estimated presentation likelihood u k . k

[0287] VIII.B.2. Per Allele with Allele Non-Interaction Variables In one implementation, the training module 316 incorporates allele non-interaction variables and models the estimated presentation likelihood uk of peptide p by TIFF0007712970000025.tif7128, where w k represents the encoded allele non-interaction variable for peptide p k and g k (·) is a function for allele non-interaction variable w w based on the set of parameters θ w determined for the allele non-interaction variable. Specifically, the set of parameters θ k for each MHC allele h and the set of parameters θ h for the allele non-interaction variable have their values as θ w h ​​and θ w can be determined by minimizing a loss function with respect to, where i is each example in a subset S of the training data 170 generated from cells expressing a single MHC allele.

[0288] dependency function g w (w k ; θ w ) outputs a dependency score for the allele-noninteracting variables that indicates whether the peptide p k is presented by one or more MHC alleles, based on the influence of the allele-noninteracting variables. For example, the dependency score for the allele-noninteracting variables can have a high value if a C-terminal flanking sequence known to have a positive influence on the presentation of the peptide p k is bound to the peptide p k and can have a low value if a C-terminal flanking sequence known to have a negative influence on the presentation of the peptide p k is bound to the peptide p k .

[0289] According to equation (8), the per-allele likelihood that the peptide sequence p k will be presented by the MHC allele h can be generated by applying the function g h (·) for the MHC allele h to the encoded version of the peptide sequence p k to generate the corresponding dependency scores for the allele-interacting variables. The function g w (·) for the allele-noninteracting variables is also applied to the encoded version of the allele-noninteracting variables to generate the dependency scores for the allele-noninteracting variables. Both scores are combined and the combined score is converted by a conversion function f(·) to generate the per-allele likelihood that the peptide sequence p k will be presented by the MHC allele h.

[0290] Alternatively, the training module 316 replaces the allele-noninteracting variables w k with the allele-interacting variables x in equation (2)h k By adding, it may include the allelic non - interaction variable \(w_k\) in the prediction. Thus, the presented likelihood can be given by TIFF0007712970000026.tif7128.

[0291] VIII.B.3 Dependency Function for Allelic Non - interaction Variables Dependency function \(g\) for allelic interaction variables h Similar to the dependency function \(g\) w (·) for allelic non - interaction variables, the dependency function \(g\) k (·) can be an affine function or a network function related to the allelic non - interaction variable \(w\)

[0292] Specifically, the dependency function \(g\) w (·) linearly combines the allelic non - interaction variables in \(w\) with the corresponding parameters in the set of parameters \(\theta\) k and is an affine function given by w TIFF0007712970000027.tif5128. TIFF0007712970000027.tif5128.

[0293] The dependency function \(g\) w (·) can also be a network function represented by a network model \(NN\) w (·) having related parameters in the set of parameters \(\theta\) w and is given by TIFF0007712970000028.tif5128. The network function may include one or more network models that each take different allelic non - interaction variables as inputs.

[0294] In another example, the dependency function \(g\) w (·) can be given by TIFF0007712970000029.tif5128, where \(g'\) w (w k ;\(\theta'\)w ) is an affine function, network function, etc. with a set of allelic non-interaction parameters θ' w , where m k is the mRNA quantification measurement value for peptide p k , h(·) is a function that converts the quantification measurement value, and θ w m is a parameter in the set of parameters for allelic non-interaction variables that is combined with the mRNA quantification measurement value to generate a dependence score for the mRNA quantification measurement value. In one particular embodiment referred to throughout the remainder of this specification, h(·) is a log function, but in reality, h(·) can be any one of a variety of different functions.

[0295] In yet another example, the dependence function g w (·) for allelic non-interaction variables is given by TIFF0007712970000030.tif5128, where g' w (w k ;θ' w ) is an affine function, network function, etc. with a set of allelic non-interaction parameters θ' w , where o k is the index vector described in Section VII.C.2 that represents proteins and isoforms in the human proteome for peptide p k , and θ w o is the set of parameters in the set of parameters for allelic non-interaction variables that is combined with the index vector. In one variation, when the dimensions of the set of o k and parameter θ w o are significantly high, TIFF0007712970000031.tif4128( TIFF0007712970000032.tif4128 can add parameter regularization terms such as the L1 norm, L2 norm, combination, etc. (representing the L1 norm, L2 norm, combination, etc.) to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined through an appropriate method.

[0296] In yet another example, the dependency function g for the allelic non-interaction variable w (·) is given by the following formula. That is, TIFF0007712970000033.tif13128 However, g’ w (w k ;θ’ w ) is an affine function, network function, etc. with a set of allelic non-interaction parameters θ’ w , and TIFF0007712970000034.tif4128 is an indicator function equal to 1 when the peptide p k is derived from the source gene l for the allelic non-interaction variable as described above, and θ w l is a parameter indicating the "antigenicity" of the source gene l. In one variation, when L is large enough, and thus the number of parameters θ w l=1, 2,...,L is large enough, TIFF0007712970000035.tif5128 such parameter regularization terms (however, TIFF0007712970000036.tif4128 is the L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined by an appropriate method.

[0297] In yet another example, the dependency function g for the allelic non-interaction variable w (·) is given by the following formula. That is, TIFF0007712970000037.tif13146 However, g’ w (w k ;θ’w ) is an affine function, network function, etc. with a set of allelic non-interaction parameters θ' w and is, for example, TIFF0007712970000038.tif4128 is an indicator function equal to 1 when the peptide p k is derived from the source gene l and the peptide p k is derived from the tissue type m, and θ w lm is a parameter indicating the antigenicity of the combination of the source gene l and the tissue type m. Specifically, the antigenicity of gene l of tissue type m can indicate the remaining tendency of cells of tissue type m to present peptides derived from gene l after regulation regarding RNA expression and peptide sequence context.

[0298] In one variation, when L or M is large enough and thus the number of parameters θ w lm=1, 2,...,LM is large enough, a parameter regularization term such as TIFF0007712970000039.tif5128 (where TIFF0007712970000040.tif4128 is the L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the value of the parameter. The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the value of the parameter so that the coefficients for the same source gene do not vary greatly between tissue types. For example, a penalty term such as TIFF0007712970000041.tif18128 (where TIFF0007712970000042.tif5128 is the average antigenicity across tissue types of the source gene l) can add a penalty to the standard deviation of antigenicity across different tissue types in the loss function.

[0299] In yet another example, the dependence function g on the allele non-interaction variable w (·) is given by the following equation. That is, TIFF0007712970000043.tif29128where g’ w (w k ;θ’ w ) is an affine function and the network function w with a set of allele non-interaction parameters θ’ TIFF0007712970000044.tif4128is an indicator function equal to 1 when the peptide p k is derived from the source gene l described above with respect to the allele non-interaction variable, and θ w l is a parameter indicating the "antigenicity" of the source gene l, TIFF0007712970000045.tif4128is an indicator function equal to 1 when the peptide p k is from the proteome position m, and θ w m is a parameter indicating the degree to which the proteome position m is a presented "hot spot". In one embodiment, the proteome position may include a block of n adjacent peptides from the same protein, and n is a hyperparameter of the model determined by a suitable method such as grid search cross-validation.

[0300] In practice, the dependence function g on the allele non-interaction variable can be generated by combining any additional terms of equations (10), (11), (12a), (12b), and (12c). For example, the dependence function on the allele non-interaction variable can be generated by adding together the term h(·) showing the mRNA quantification measurement value of equation (10) and the term showing the antigenicity of the source gene of equation (12) with any other affine function or network function. w (·) can be generated. For example, by adding together the term h(·) showing the mRNA quantification measurement value of equation (10) and the term showing the antigenicity of the source gene of equation (12) with any other affine function or network function, the dependence function on the allele non-interaction variable can be generated.

[0301] As an example, returning to equation (8), the affine transformation functions g h (·), gw The likelihood that peptide p will be presented by MHC allele h = 3 among m = 4 different specified MHC alleles using (·) k will be, can be generated by TIFF0007712970000046.tif5128, where w k is the allele non - interaction variable specified for peptide p k and θ w is a set of parameters determined for the allele non - interaction variable.

[0302] As another example, the likelihood that peptide p will be presented by MHC allele h = 3 among m = 4 different specified MHC alleles using the network transformation functions g h (·), g w (·) will be, k can be generated by TIFF0007712970000047.tif5128, where w is the allele interaction variable specified for peptide p k and θ k is a set of parameters determined for the allele non - interaction variable. w w

[0303] Figure 8 illustrates the generation of the presentation likelihood of peptide p associated with MHC allele h = 3 using the exemplary network models NN3(·) and NN w (·). As shown in Figure 8, the network model NN3(·) receives the allele interaction variable x3 k for MHC allele h = 3 and generates the output NN3(x3 k ). The network model NN k (·) receives the allele non - interaction variable w w for peptide p k and generates the output NN k (w w ). The outputs are combined and mapped by the function f(·) to generate the estimated presentation likelihood u k . k

[0304] VIII.C. Multiple Allele Model The training module 316 may also construct a presentation model for predicting the likelihood of peptide presentation in a multiple allele setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on data examples S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof.

Example

[0305] VIII.C.1. Example 1: Maximum Value of Allele-by-Allele Model In one implementation, the training module 316 determines the estimated presentation likelihood u k of a peptide p k for each MHC allele h in the set H determined based on cells expressing a single allele, as described above together with equations (2)-(11). Model it as a function of TIFF0007712970000048.tif4128. Specifically, the presentation likelihood u k is can be any function of TIFF0007712970000049.tif4128. In one implementation, as shown in equation (12), the function is a maximum value function, and the presentation likelihood u k can be determined as the maximum value of the presentation likelihood for each MHC allele h in the set H. TIFF0007712970000050.tif5128

[0306] VIII.C.2. Example 2.1: Sum Function Model In one implementation, the training module 316 models the estimated presentation likelihood u k of a peptide p k as modeled by TIFF0007712970000051.tif13128, where the element a h k is the peptide sequence pk is 1 for a plurality of MHC alleles H associated with, and x h k is the peptide p k and the encoded allele interaction variable for the corresponding MHC allele. The parameter θ h for each set of values of the MHC allele h can be determined by minimizing the loss function with respect to θ h where i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing a plurality of MHC alleles. The dependency function g h is the dependency function g introduced above in Section VIII.B.1 h and can be in any form thereof.

[0307] According to equation (13), the presentation likelihood that the peptide sequence p k will be presented by one or more MHC alleles h can be generated by applying the dependency function g h (·) to the encoded version of the peptide sequence p k for each of the MHC alleles H to generate the corresponding scores for the allele interaction variables. The scores for each MHC allele h are combined and transformed by the transformation function f(·) to generate the presentation likelihood that the peptide sequence p k will be presented by the set of MHC alleles H.

[0308] The presentation model of equation (13) differs from the per-allele model of equation (2) in that the number of associated alleles for each peptide p k can be greater than 1. In other words, more than one element in a h k can have a value of 1 for the plurality of MHC alleles H associated with the peptide sequence p k .

[0309] As an example, the affine transformation function g hThe likelihood that peptide p will be presented by MHC alleles h = 2 and h = 3 among m = 4 different specified MHC alleles using (·) k can be generated by TIFF0007712970000052.tif5128, where x2 k , x3 k are allelic interaction variables specified for MHC alleles h = 2 and h = 3, and θ2, θ3 are sets of parameters determined for MHC alleles h = 2 and h = 3.

[0310] As another example, the likelihood that peptide p will be presented by MHC alleles h = 2 and h = 3 among m = 4 different specified MHC alleles using the network transformation functions g h (·), g w (·) k can be generated by TIFF0007712970000053.tif5128, where NN2(·), NN3(·) are network models specified for MHC alleles h = 2 and h = 3, and θ2, θ3 are sets of parameters determined for MHC alleles h = 2 and h = 3.

[0311] Figure 9 illustrates the generation of the presentation likelihood of peptide p associated with MHC alleles h = 2 and h = 3 using exemplary network models NN2(·) and NN3(·). As shown in Figure 9, the network model NN2(·) receives the allelic interaction variable x2 k for MHC allele h = 2 and generates the output NN2(x2 k ), and the network model NN3(·) receives the allelic interaction variable x3 k for MHC allele h = 3 and generates the output NN3(x3 k ). The outputs are combined and mapped by the function f(·) to generate the estimated presentation likelihood u k k k .

[0312] VIII.C.3. Example 2.2: Sum Function Model with Allele Non-Interaction Variables In one implementation, the training module 316 incorporates allele non-interaction variables to model the estimated presentation likelihood u k of peptide p k by TIFF0007712970000054.tif13130, where w k represents the encoded allele non-interaction variables for peptide p k . Specifically, the values of the set of parameters θ h for each MHC allele h and the set of parameters θ w for the allele non-interaction variables are determined by minimizing the loss function with respect to θ h and θ w , where i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The dependency function g w can be in any form of the dependency function g w introduced above in Section VIII.B.3.

[0313] Thus, according to equation (14), the presentation likelihood that peptide sequence p k will be presented by one or more MHC alleles H can be generated by applying the function g h (·) to the encoded version of peptide sequence p k for each of the MHC alleles H to generate the corresponding dependency scores for the allele interaction variables of each MHC allele h. The function g w (·) for the allele non-interaction variables is also applied to the encoded version of the allele non-interaction variables to generate the dependency scores for the allele non-interaction variables. The scores are combined and the combined scores are transformed by the transformation function f(·) to generate the presentation likelihood that peptide sequence p k will be presented by MHC allele H.

[0314] In the presentation model of equation (14), for each peptide p k the number of related alleles can be greater than 1. In other words, more than one element in a h k can have a value of 1 for multiple MHC alleles H related to the peptide sequence p k

[0315] As an example, the likelihood that peptide p will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the affinity conversion functions g h (·), g w (·) can be generated by TIFF0007712970000055.tif5128, where w k is an allele non-interaction variable specified for peptide p, and θ is a set of parameters determined for the allele non-interaction variables. k is the allele non-interaction variable specified for peptide p, and θ k is a set of parameters determined for the allele non-interaction variables. w is a set of parameters determined for the allele non-interaction variables.

[0316] As another example, the likelihood that peptide p will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the network conversion functions g h (·), g w (·) can be generated by TIFF0007712970000056.tif5128, where w k is an allele interaction variable specified for peptide p, and θ is a set of parameters determined for the allele non-interaction variables. k is the allele interaction variable specified for peptide p, and θ k is a set of parameters determined for the allele non-interaction variables. w is a set of parameters determined for the allele non-interaction variables.

[0317] Figure 10 shows peptide p related to MHC alleles h = 2, h = 3 using exemplary network models NN2(·), NN3(·), and NN w (·), k ​The generation of the presentation likelihood is described. As shown in FIG. 10, the network model NN2(·) receives the allele interaction variable x2 for MHC allele h = 2 k and generates the output NN2(x2 k ). The network model NN3(·) receives the allele interaction variable x3 for MHC allele h = 3 k and generates the output NN3(x3 k ). The network model NN w (·) receives the allele non - interaction variable w k for peptide p k and generates the output NN w (w k ). The outputs are combined and mapped by the function f(·) to generate the estimated presentation likelihood u k .

[0318] Alternatively, the training module 316 may include the allele non - interaction variable w k in the prediction by adding the allele non - interaction variable w h k to the allele interaction variable x k . Thus, the presentation likelihood can be given by TIFF0007712970000057.tif13128.

[0319] VIII.C.4. Example 3.1: Model Using Implicit Allele - Specific Likelihoods In another embodiment, the training module 316 models the estimated presentation likelihood u k of peptide p k by TIFF0007712970000058.tif7128, where the element a h k is 1 for multiple MHC alleles h ∈ H related to the peptide sequence p k , u’ k h is the implicit allele - specific presentation likelihood for MHC allele h, and the vector v has elements v h such that a hk ·u' k h is the corresponding vector, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the value of the input into a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, but it is recognized that in other embodiments, s(·) may be any function such as a maximum value function. The set of values of the parameter θ for the implicit allele-by-allele likelihood can be determined by minimizing the loss function with respect to θ, and i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.

[0320] The presentation likelihood in the presentation model of equation (17) is such that each corresponds to the likelihood that the peptide p k will be presented by the individual MHC allele h, and is modeled as a function of the implicit allele-by-allele presentation likelihood u' k h . The implicit allele-by-allele likelihood is different from the allele-by-allele presentation likelihood in Section VIII.B in that the parameters for the implicit allele-by-allele likelihood can be learned from a multiple allele setting where the direct relationship between the presented peptide and the corresponding MHC allele is unknown, in addition to the single allele setting. Thus, in a multiple allele setting, the presentation model can not only estimate whether the peptide p k is presented as a whole by the set of MHC alleles H, but also the individual likelihood u' k indicating which MHC allele h is most likely to have presented the peptide p k h∈H can also be provided. The advantage of this is that the presentation model can generate implicit likelihoods without training data for cells expressing a single MHC allele.

[0321] In one particular implementation referred to throughout the remainder of this specification, r(·) is a function having a range [0,1]. For example, r(·) is a clip function: r(z)=min(max(z,0),1) may also be such that the minimum value between z and 1 is the presentation likelihood u k is selected. In another embodiment, r(·) is r(z)=tanh(z) the hyperbolic tangent function given as, and the value of the domain z is 0 or more.

[0322] VIII.C.5. Example 3.2: Sum Model of Functions In one particular embodiment, s(·) is a sum function, and the presentation likelihood is given by summing the per-allele presentation likelihoods implicitly. TIFF0007712970000059.tif15128

[0323] In one embodiment, the per-allele presentation likelihood for MHC allele h is generated by TIFF0007712970000060.tif7128 such that the presentation likelihood is estimated by TIFF0007712970000061.tif13128.

[0324] According to equation (19), the presentation likelihood that one or more MHC alleles H will present peptide sequence p k can be generated by applying the function g h (·) to the encoded version of peptide sequence p k for each of the MHC alleles H to generate the corresponding dependency scores for the allele interaction variables. Each dependency score is first converted by the function f(·) so as to generate the per-allele likelihood u’ k h . The per-allele likelihoods u’ k h are combined, and a clipping function is applied to the combined likelihood to clip the value within the range [0,1] to generate the presentation likelihood that peptide sequence p k will be presented by the set of MHC alleles H. The dependency function g his the dependency function g introduced above in Section VIII.B.1 h and can be in any of the forms

[0325] As an example, the likelihood that peptide p h will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the affine transformation function g k can be generated by TIFF0007712970000062.tif7128, where x2 k , x3 k are allele interaction variables specified for MHC alleles h = 2, h = 3, and θ2, θ3 are sets of parameters determined for MHC alleles h = 2, h = 3

[0326] As another example, the likelihood that peptide p h will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles using the network transformation functions g w (·), g k can be generated by TIFF0007712970000063.tif7128, where NN2(·), NN3(·) are network models specified for MHC alleles h = 2, h = 3, and θ2, θ3 are sets of parameters determined for MHC alleles h = 2, h = 3

[0327] Figure 11 illustrates the generation of the presentation likelihood of peptide p k associated with MHC alleles h = 2, h = 3 using exemplary network models NN2(·) and NN3(·). As shown in Figure 9, the network model NN2(·) receives the allele interaction variable x2 k for MHC allele h = 2 and generates the output NN2(x2 k ), and the network model NN3(·) receives the allele interaction variable x3 k for MHC allele h = 3 and generates the output NN3(x3 kGenerate (). Each output is mapped by the function f(·) and combined to generate the estimated presentation likelihood u k Generate.

[0328] In another embodiment, when the prediction is made about the log of the mass spectrometry ion current, r(·) is the log function and f(·) is the exponential function.

[0329] VIII.C.6. Example 3.3: Sum model of functions with allele non-interaction variables In one embodiment, the implicit per-allele presentation likelihood for MHC allele h is Generated by TIFF0007712970000064.tif7128 so that the presentation likelihood is Generated by TIFF0007712970000065.tif13129 to incorporate the influence of allele non-interaction variables on peptide presentation.

[0330] According to equation (21), the presentation likelihood that peptide sequence p k Will be presented by one or more MHC alleles H can be generated by applying the function g h (·) to the encoded version of peptide sequence p k For each of the MHC alleles H to generate the corresponding dependency scores for the allele interaction variables of each MHC allele h. The function g w (·) for the allele non-interaction variables is also applied to the encoded version of the allele non-interaction variables to generate the dependency scores for the allele non-interaction variables. The scores of the allele non-interaction variables are combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is converted by the function f(·) to generate the implicit per-allele presentation likelihood. The implicit likelihoods are combined and a clipping function is applied to the combined output to clip the values into the range [0,1] to generate the presentation likelihood that peptide sequence p k Will be presented by MHC allele H. The dependency function g wis the dependency function g introduced above in Section VIII.B.3 w and can be in any of the forms

[0331] As an example, for the affine transformation functions g h (·), g w (·), the likelihood that peptide p k will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles can be generated by TIFF0007712970000066.tif7128, where w k is the allele non - interaction variable specified for peptide p k and θw is the set of parameters determined for the allele non - interaction variable

[0332] As another example, for the network transformation functions g h (·), g w (·), the likelihood that peptide p k will be presented by MHC alleles h = 2, h = 3 among m = 4 different specified MHC alleles can be generated by TIFF0007712970000067.tif7128, where w k is the allele interaction variable specified for peptide p k and θ w is the set of parameters determined for the allele non - interaction variable

[0333] Figure 12 illustrates the generation of the presentation likelihood of peptide p w associated with MHC alleles h = 2, h = 3 using exemplary network models NN2(·), NN3(·), and NN k (·). As shown in Figure 12, the network model NN2(·) receives the allele interaction variable x2 k for MHC allele h = 2 and generates the output NN2(x2 k ). The network model NN w (·) takes peptide p kThe allele non - interaction variable w for k is received, and the output NN w (w k ) is generated. The output is combined and mapped by the function f(·). The network model NN3(·) receives the allele interaction variable x3 k for MHC allele h = 3 and generates the output NN3(x3 k ), which is also combined with the output NN w (·) of the same network model NN w (w k ) and mapped by the function f(·). The two outputs are combined to generate the estimated presentation likelihood u k .

[0334] In another embodiment, the implicit per - allele presentation likelihood for MHC allele h is generated by TIFF0007712970000068.tif7128 such that the presentation likelihood is generated by TIFF0007712970000069.tif13128.

[0335] VIII.C.7. Example 4: Quadratic Model In one embodiment, s(·) is a quadratic function, and the estimated presentation likelihood u k of the peptide p k is given by TIFF0007712970000070.tif13128, where the element u’ k h is the implicit per - allele presentation likelihood for MHC allele h. The set of values of the parameter θ for the implicit per - allele likelihood can be determined by minimizing the loss function with respect to θ, and i is each example in the subset S of the training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit per - allele presentation likelihood can be in any of the forms shown in the above equations (18), (20), and (22).

[0336] In one aspect, the model of equation (23) is for the peptide sequence p k There may be a possibility that it will be presented simultaneously by two MHC alleles, which may imply that the presentation by the two HLA alleles is statistically independent.

[0337] According to equation (23), the presentation likelihood that the peptide sequence p k will be presented by one or more MHC alleles H can be generated by combining the implicit allele-by-allele presentation likelihoods and subtracting from the sum the likelihood that each pair of MHC alleles will present the peptide p k simultaneously. k

[0338] As an example, for m = 4 different specified HLA alleles using the affinity transformation function g h (·), the likelihood that the peptide p k will be presented by the HLA alleles h = 2, h = 3 can be generated by TIFF0007712970000071.tif5128, where x2 k , x3 k are the allele interaction variables specified for the HLA alleles h = 2, h = 3, and θ2, θ3 are the sets of parameters determined for the HLA alleles h = 2, h = 3.

[0339] As another example, for m = 4 different specified HLA alleles using the network transformation functions g h (·), g w (·), the likelihood that the peptide p k will be presented by the HLA alleles h = 2, h = 3 can be generated by TIFF0007712970000072.tif5128, where NN2(·), NN3(·) are the network models specified for the HLA alleles h = 2, h = 3, and θ2, θ3 are the sets of parameters determined for the HLA alleles h = 2, h = 3. ​

[0340] IX. Example 5: Prediction Module The prediction module 320 receives array data and selects candidate neoantigens in the array data using a presentation model. Specifically, the array data may be a DNA sequence, an RNA sequence, and / or a protein sequence extracted from a patient's tumor tissue cells. The prediction module 320 processes the array data into a plurality of peptide sequences p having 8 to 15 amino acids for MHC-I or 6 to 30 amino acids for MHC-II k For example, the prediction module 320 can process a predetermined sequence IEFROEIFJEF into three peptide sequences having 9 amino acids TIFF0007712970000073.tif4128. In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by identifying portions having one or more mutations by comparing the array data extracted from the patient's normal tissue cells with the array data extracted from the patient's tumor tissue cells

[0341] The prediction module 320 applies one or more of the presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation model to the candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences having an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences having the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the candidate neoantigens selected for a given patient can be injected into the patient to induce an immune response

[0342] X. Example 6: Patient Selection Module The patient selection module 324 selects a subset of patients for vaccine therapy and / or T cell therapy based on whether the patients meet the selection criteria. In one embodiment, the selection criteria are determined based on the likelihood of presentation of the patient's neoantigen candidates generated by the presentation model. By adjusting the selection criteria, the patient selection module 324 can adjust the number of patients receiving vaccine administration and / or T cell therapy based on the likelihood of presentation of the patient's neoantigen candidates. Specifically, with strict selection criteria, the number of patients treated with vaccine and / or T cell therapy will be fewer, but the proportion of treated patients receiving effective treatment (e.g., one or more tumor-specific neoantigens (TSNAs) and / or one or more neoantigen-responsive T cells) by vaccine and / or T cell therapy can be higher. In contrast, with loose selection criteria, the number of patients treated with vaccine and / or T cell therapy will be more, but the proportion of treated patients receiving effective treatment by vaccine and / or T cell therapy can be lower. The patient selection module 324 changes the selection criteria based on the desired balance between the target proportion of patients receiving treatment and the proportion of patients receiving effective treatment.

[0343] In some embodiments, the selection criteria for selecting patients for vaccine therapy are the same as the selection criteria for selecting patients for T cell therapy. However, in alternative embodiments, the selection criteria for selecting patients for vaccine therapy can be different from the selection criteria for selecting patients for T cell therapy. In Sections X.A and X.B below, the selection criteria for selecting patients for vaccine therapy and the selection criteria for selecting patients for T cell therapy are considered respectively.

[0344] X.A. Selection of Patients for Vaccine Therapy In one embodiment, a corresponding therapeutic subset of v neoantigen candidates that can potentially be included in an individualized vaccine for a patient having a vaccine volume v is associated with the patient. In one embodiment, the therapeutic subset for a patient is the neoantigen candidate having the highest presentation likelihood determined by a presentation model. For example, if the vaccine can contain v = 20 epitopes, the vaccine can include the therapeutic subset for each patient having the highest presentation likelihood determined by the presentation model. However, it is recognized that in other embodiments, the therapeutic subset for a patient can also be determined based on other methods. For example, the therapeutic subset for a patient can be randomly selected from the set of neoantigen candidates for that patient, or can be determined based in part on a combination of specific factors including a prior art model that models the binding affinity or stability of peptide sequences, or the presentation likelihood obtained from the presentation model and affinity or stability information regarding these peptide sequences.

[0345] In one embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the tumor mutation burden of the patient is equal to or higher than the minimum mutation burden. The tumor mutation burden (TMB) of a patient indicates the total number of non-synonymous mutations in the tumor exome. In one embodiment, the patient selection module 324 selects a patient for vaccine treatment if the absolute number of the patient's TMB is equal to or higher than a predetermined threshold. In another implementation, the patient selection module 324 selects a patient for vaccine treatment if the patient's TMB is within the threshold percentile among the TMBs determined for the set of patients.

[0346] In another embodiment, the patient selection module 324 determines that a patient meets the selection criteria if the utility value score of the patient based on the therapeutic subset of the patient is equal to or higher than the minimum utility value score. In one embodiment, the utility value score is a measure of the estimated number of presented antigens from the therapeutic subset.

[0347] The estimated number of presented antigens can be predicted by modeling the presentation of neoantigens as a random variable of one or more probability distributions. In one implementation, the utility score for patient i is the expected number of presented neoantigen candidates from a subset of treatments, or a specific function thereof. As an example, the presentation of each neoantigen can be modeled as a Bernoulli random variable where the probability of presentation (success) is given by the presentation likelihood of the neoantigen candidate. Specifically, for a subset S of v neoantigen candidates p with the highest presentation likelihoods u, u, …, u respectively, the presentation of neoantigen candidate p is given by a random variable A, where i1 u i2 u iv and so on for v neoantigen candidates p, p, …, p having the highest presentation likelihoods u, u, …, u respectively, the presentation of neoantigen candidate p is given by a random variable A, where i1 p i2 p iv and so on for v neoantigen candidates p, p, …, p having the highest presentation likelihoods u, u, …, u respectively, the presentation of neoantigen candidate p is given by a random variable A, where i for a subset S of v neoantigen candidates p, p, …, p having the highest presentation likelihoods u, u, …, u respectively, the presentation of neoantigen candidate p is given by a random variable A, where ij A ij and where TIFF0007712970000074.tif5128. The expected number of presented neoantigens is given by the sum of the presentation likelihoods of each neoantigen candidate. In other words, the utility score for patient i is TIFF0007712970000075.tif15128. The patient selection module 324 selects a subset of patients having a utility score equal to or higher than the minimum utility value for the vaccine treatment.

[0348] In another implementation, the utility score for patient i is the probability that at least a threshold number of neoantigens k are presented. In one example, the number of presented antigens within a subset S of neoantigen candidates is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the respective presentation likelihood of each epitope. Specifically, the number of presented antigens for patient i can be given by a random variable N, where i and where the number of presented antigens within a subset S of neoantigen candidates is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the respective presentation likelihood of each epitope. Specifically, the number of presented antigens for patient i can be given by a random variable N, where i N TIFF0007712970000076.tif13128 and where PBD(·) denotes the Poisson binomial distribution. The probability that at least a threshold number of neoantigens k are presented is the number of presented antigens N iis given by an operation with a probability equal to or greater than k. In other words, the utility value score of patient i is is represented as TIFF0007712970000077.tif13128. The patient selection module 324 selects a subset of patients having a utility value score equal to or higher than the minimum utility value for the vaccine treatment.

[0349] In another embodiment, the utility value score of patient i is the number of neoantigens in a treatment subset S of neoantigen candidates having a binding affinity or predicted binding affinity lower than a fixed threshold (e.g., 500 nM) for one or more HLA alleles of the patient. In one example, the fixed threshold ranges from 1000 nM to 10 nM. Optionally, the utility value score may also count only the neoantigens detected as being expressed by RNA-seq. i In another embodiment, the utility value score of patient i is the number of neoantigens in a treatment subset S of neoantigen candidates whose binding affinity for one or more HLA alleles of that patient is below the threshold percentile of the binding affinity of a random peptide for that HLA allele. In one example, the threshold percentile ranges from the 10th percentile to the 0.1st percentile. Optionally, the utility value score may count only the neoantigens detected as being expressed by RNA-seq.

[0350] In another embodiment, the utility value score of patient i is the number of neoantigens in a treatment subset S of neoantigen candidates whose binding affinity for one or more HLA alleles of that patient is below the threshold percentile of the binding affinity of a random peptide for that HLA allele. i In one example, the threshold percentile ranges from the 10th percentile to the 0.1st percentile. Optionally, the utility value score may count only the neoantigens detected as being expressed by RNA-seq.

[0351] It should be recognized that the examples of utility value scores described with respect to equations (25) and (27) are merely illustrative, and the patient selection module 324 can also generate utility value scores using other statistical or probability distributions.

[0352] X.B. Selection of Patients for T Cell Therapy In another embodiment, instead of or in addition to receiving a vaccine treatment, a patient can receive a T cell therapy. Similar to the vaccine treatment, in embodiments where the patient receives a T cell therapy, the patient can be associated with a corresponding treatment subset of v neoantigen candidates as described above. This treatment subset of v neoantigen candidates can be used for in vitro identification of patient-derived T cells that have reactivity to one or more of the v neoantigen candidates. These identified T cells can then be expanded and injected into the patient in an individualized T cell therapy.

[0353] Patients who receive T cell therapy at two different time points can be selected. The first time point is after the treatment subset of v neoantigen candidates for the patient has been predicted using a model, but before in vitro screening of T cells specific to the predicted treatment subset of v neoantigen candidates is performed. The second time point is after in vitro screening of T cells specific to the predicted treatment subset of v neoantigen candidates has been performed.

[0354] First, after predicting a therapeutic subset of v neoantigen candidates for a patient and before performing in vitro identification of patient-derived T cells that are specific for the predicted therapeutic subset of v neoantigen candidates, a patient to receive T cell therapy can be selected. Specifically, since in vitro screening of patient-derived neoantigen-specific T cells can be costly, it is considered desirable to select a patient to screen for neoantigen-specific T cells only if the patient is likely to have neoantigen-specific T cells. To select a patient prior to the in vitro T cell screening step, the same criteria used to select a patient to receive vaccine therapy can be used. Specifically, in some embodiments, the patient selection module 324 can select a patient to receive T cell therapy if the tumor mutation burden of the patient is equal to or higher than the minimum mutation burden as described above. In another embodiment, the patient selection module 324 can select a patient to receive T cell therapy if the utility score of the patient based on the therapeutic subset of v neoantigen candidates for that patient is equal to or higher than the minimum utility score as described above.

[0355] Second, in addition to, or alternatively to, selecting a patient to receive T cell therapy before performing in vitro identification of patient-derived T cells that are specific for the predicted therapeutic subset of v neoantigen candidates, a patient to receive T cell therapy can also be selected after performing in vitro identification of T cells that are specific for the predicted therapeutic subset of v neoantigen candidates. Specifically, a patient can be selected to receive T cell therapy if at least a threshold amount of neoantigen-specific TCRs for that patient are identified in the in vitro screening of the patient's T cells for neoantigen recognition. For example, a patient can be selected to receive T cell therapy only if at least two neoantigen-specific TCRs are identified for that patient, or only if neoantigen-specific TCRs are identified for two different neoantigens.

[0356] In another embodiment, a patient can be selected to receive T cell therapy only if a threshold amount of neoantigens of a therapeutic subset of v neoantigen candidates for that patient are recognized by the patient's TCRs. For example, a patient can be selected to receive T cell therapy only if at least one neoantigen of a therapeutic subset of v neoantigen candidates for that patient is recognized by the patient's TCRs. In a further embodiment, a patient can be selected to receive T cell therapy only if at least a threshold amount of the patient's TCRs are identified as being neoantigen-specific for neoantigen peptides of a particular HLA restriction class. For example, a patient can be selected to receive T cell therapy only if at least one of the patient's TCRs is identified as being a neoantigen-specific HLA class I-restricted neoantigen peptide.

[0357] In yet a further embodiment, a patient can be selected to receive T cell therapy only if at least a threshold amount of neoantigen peptides of a particular HLA restriction class are recognized by the patient's TCRs. For example, a patient can be selected to receive T cell therapy only if at least one HLA class I-restricted neoantigen peptide is recognized by the patient's TCRs. In another example, a patient can be selected to receive T cell therapy only if at least two HLA class II-restricted neoantigen peptides are recognized by the patient's TCRs. Any combination of the above criteria can also be used to select a patient to receive T cell therapy after in vitro identification of T cells that are specific for a therapeutic subset of v neoantigen candidates predicted for that patient.

[0358] XI. Example 7: Experimental Results Demonstrating Exemplary Patient Selection Performance The validity of patient selection described in Section X is verified by performing patient selection on a set of simulated patients associated with each test set of simulated neoantigen candidates, where it has been found that a subset of simulated neoantigens presented in the mass spectrometry data is presented. Specifically, each simulated neoantigen candidate in the test set is associated with a label indicating whether the neoantigen is presented in the mass spectrometry data set of multiple allele JY cell lines HLA-A*02:01 and HLA-B*07:02 from the Bassani-Sternberg data set (data set "D1") (the data can be found at www.ebi.ac.uk / pride / archive / projects / PXD0000394). As described in detail below with Figure 13A, a large number of neoantigen candidates for the simulated patients are sampled from the human proteome based on the known frequency distribution of mutation loads in non-small cell lung cancer (NSCLC) patients.

[0359] The allele-specific presentation model for the same HLA allele is trained using a training set that is a subset of the mass spectrometry data of single allele HLA-A*02:01 and HLA-B*07:02 from the IEDB data set (data set "D2") (the data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Specifically, the presentation model for each allele is a network-dependent function g h (·) and g wThe per-allele model shown in formula (8) incorporating (·) and the expit function f(·) was used. The presentation model for allele HLA-A*02:01 generates the presentation likelihood that a specific peptide is presented on allele HLA-A*02:01, assuming that the peptide sequence is given as the allele interaction variable and the N-terminal and C-terminal flanking sequences are given as the allele non-interaction variables. The presentation model for allele HLA-B*07:02 generates the presentation likelihood that a specific peptide is presented on allele HLA-B*07:02, assuming that the peptide sequence is given as the allele interaction variable and the N-terminal and C-terminal flanking sequences are given as the allele non-interaction variables.

[0360] As disclosed with reference to FIGS. 13A-13E in the following example, different models such as presentation models trained for peptide binding prediction and prior art models are applied to a test set of neoantigen candidates for each simulated patient to identify different treatment subsets for the patient based on the prediction. Patients meeting the selection criteria for vaccine treatment are selected and associated with an individualized vaccine containing an epitope in the patient's treatment subset. The size of the treatment subset varies depending on the different vaccine volumes. No overlap is introduced between the training set used to train the presentation model and the test set of simulated neoantigen candidates.

[0361] In the following example, the ratio of selected patients having at least a certain number of presented neoantigens among the epitopes included in the vaccine is analyzed. This statistic indicates the effectiveness of the simulated vaccine in delivering potential neoantigens that induce an immune response in the patient. Specifically, the simulated neoantigens within a certain test set are presented if the neoantigen is presented in the mass spectrometry dataset D2. A high ratio of patients with presented neoantigens indicates the likelihood of success of treatment with a neoantigen vaccine by inducing an immune response.

[0362] XI.A. Example 7A: Frequency Distribution of Mutation Load in NSCLC Cancer Patients Figure 13A shows the specimen frequency distribution of mutation load in NSCLC patients. The mutation load and mutations in different tumor types including NSCLC can be seen, for example, in the cancer genome atlas (TCGA) (https: / / cancergenome.nih.gov). The X-axis represents the number of non-synonymous mutations per patient, and the Y-axis represents the ratio of specimen patients having a specific number of non-synonymous mutations. The specimen frequency distribution in Figure 13A shows a range of 3 to 1786 mutations, and 30% of the patients have fewer than 100 mutations. Although not shown in Figure 13A, the mutation load is higher in smokers compared to non-smokers, and studies have shown that the mutation load can be a strong indicator of neoantigen load in patients.

[0363] As introduced at the beginning of Section XI above, a test set of neoantigen candidates is associated with each of the simulated numbers of patients. The test set for each patient is generated by sampling a mutation load m from the frequency distribution shown in Figure 13A for each patient. i For each mutation, a 21-mer peptide sequence from the human proteome is randomly selected to represent the mutated sequence being simulated. A test set of neoantigen candidate sequences is generated for patient i by identifying the peptide sequences of each (8, 9, 10, 11)-mer over the mutations within the 21-mer. A label indicating whether the neoantigen candidate sequence is present in the mass spectrometry D1 dataset is associated with each neoantigen candidate. For example, a label "1" can be associated with neoantigen candidate sequences present in dataset D1, and a label "0" can be associated with sequences not present in dataset D1. As described in more detail below, Figures 13B - 13E show the experimental results of patient selection based on the presented neoantigens of patients in the test set.

[0364] XI.B. Example 7B: Ratio of selected patients having neoantigen presentation based on mutation load selection criteria Figure 13B shows the number of presented neoantigens in the simulated vaccine for patients selected based on the selection criteria of whether the patient meets the minimum mutation load. Identify the proportion of selected patients having at least a specific number of presented neoantigens in the corresponding test.

[0365] In Figure 13B, the x-axis shows the proportion of patients excluded from vaccine treatment based on tumor mutation load, labeled "minimum number of mutations". For example, the data point at "minimum number of mutations" 200 indicates that the patient selection module 324 selected only a subset of the simulated patients having a mutation load of at least 200 mutations. As another example, the data point at "minimum number of mutations" 300 indicates that the patient selection module 324 selected a lower proportion of the simulated patients having at least 300 mutations. The y-axis shows the proportion of selected patients associated with at least a specific number of presented neoantigens within the test set without vaccine volume v. Specifically, the upper plot shows the proportion of selected patients presenting at least 1 neoantigen, the middle plot shows the proportion of selected patients presenting at least 2 antigens, and the lower plot shows the proportion of selected patients presenting at least 3 antigens.

[0366] As shown in Figure 13B, the proportion of patients with presented neoantigens increases significantly as the mutation load increases. This indicates that the mutation load as a selection criterion can be effective in selecting patients in whom the neoantigen vaccine is likely to induce an effective immune response.

[0367] XI.C. Example 7C: Comparison of Neoantigen Presentation in Vaccines Identified by the Presentation Model and the Prior Art Model FIG. 13C compares the number of presented neoantigens in a simulated vaccine between selected patients associated with a vaccine containing a treatment subset identified based on a presentation model and selected patients associated with a vaccine containing a treatment subset identified by a prior art model. The left plot assumes a limited vaccine volume of v = 10, and the right plot assumes a limited vaccine volume of v = 20. Patients are selected based on a utility score indicating the expected number of presented neoantigens.

[0368] In FIG. 13C, the solid line indicates patients associated with a vaccine containing a treatment subset identified based on the presentation model for alleles HLA-A*02:01 and HLA-B*07:02. The treatment subset for each patient is identified by applying each of the presentation models to the sequences within the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with a vaccine containing a treatment subset identified based on the prior art model NETMHCpan for the single allele HLA-A*02:01. Details of the implementation for NETMHCpan are provided at http: / / www.cbs.dtu.dk / services / NetMHCpan. The treatment subset for each patient is identified by applying the NETMHCpan model to the sequences within the test set and identifying the v neoantigen candidates with the highest estimated binding affinity. The x-axis of both graphs indicates the ratio of patients excluded from vaccine treatment based on an expected utility score indicating the expected number of presented neoantigens within the treatment subset identified based on the presentation model. The expected utility score is determined as described in relation to Equation (25) in Section X. The y-axis indicates the ratio of selected patients presenting at least a specific number of neoantigens (1, 2, or 3 neoantigens) included in the vaccine.

[0369] As shown in Figure 13C, patients associated with a vaccine containing a treatment subset based on a presentation model are administered a vaccine containing presented neoantigens at a significantly higher rate than patients associated with a vaccine containing a treatment subset based on a prior art model. For example, as shown in the graph on the right, 80% of the selected patients associated with a vaccine based on the presentation model, compared to only 40% of the selected patients associated with a vaccine based on a prior art model, are administered at least one presented neoantigen in the vaccine. These results indicate that the presentation model described herein is effective in selecting neoantigen candidates for vaccines that are likely to induce an immune response for treating tumors.

[0370] XI.D. Example 7D: Effect of HLA Coverage on Neoantigen Presentation of Vaccines Identified by the Presentation Model Figure 13D compares the number of presented neoantigens in a simulated vaccine between selected patients associated with a vaccine containing a treatment subset identified based on a single-allele-per-presentation model for HLA-A*02:01 and selected patients associated with a vaccine containing a treatment subset identified based on both allele-per-presentation models for HLA-A*02:01 and HLA-B*07:02. The vaccine volume is set to v = 20 epitopes. For each experiment, patients are selected based on an expected utility value score determined based on different treatment subsets.

[0371] In Figure 13D, the solid line indicates patients associated with a vaccine that includes a treatment subset based on both presentation models for the HLA alleles HLA-A*02:01 and HLA-B*07:02. The treatment subset for each patient is identified by applying each of the presentation models to the sequences within the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with a vaccine that includes a treatment subset based on a single presentation model for the HLA allele HLA-A*02:01. The treatment subset for each patient is identified by applying the presentation model for only a single HLA allele to the sequences within the test set and identifying the v neoantigen candidates with the highest presentation likelihood. In the solid line plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility value score for the treatment subset identified by both presentation models. In the dotted line plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility value score for the treatment subset identified by the single presentation model. The y-axis indicates the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens).

[0372] As shown in Figure 13D, patients associated with a vaccine that includes a treatment subset identified by presentation models for both HLA alleles present neoantigens at a significantly higher rate than patients associated with a vaccine that includes a treatment subset identified by a single presentation model. These results demonstrate the importance of establishing presentation models with high HLA coverage.

[0373] XI.E. Example 7E: Comparison of Neoantigen Presentation in Patients Selected by Mutation Burden and Expected Number of Presented Neoantigens Figure 13E compares the number of presented neoantigens in a simulated vaccine between patients selected based on mutation burden and patients selected by expected utility value score. The expected utility value score is determined based on a treatment subset identified by a presentation model having an epitope size of v = 20.

[0374] In Figure 13E, the solid line shows patients selected based on the expected utility value scores associated with vaccines that include treatment subsets identified by the presentation model. The treatment subset for each patient is identified by applying each of the presentation models to the sequences within the test set and identifying the v = 20 neoantigen candidates with the highest presentation likelihood. The treatment utility value score is determined based on the presentation likelihood of the treatment subset identified based on Equation (25) in Section X. The dotted line shows patients selected based on the mutation load associated with vaccines that include treatment subsets identified by the presentation model. The x-axis shows the proportion of patients excluded from vaccine treatment based on the expected utility value scores of the solid line plot and the proportion of patients excluded based on the mutation load of the dotted line plot. The y-axis shows the proportion of selected patients to whom a vaccine containing at least a specified number of presented neoantigens (1, 2, or 3 neoantigens) is administered.

[0375] As shown in Figure 13E, patients selected based on the expected utility value scores are administered vaccines containing presented neoantigens at a higher rate than patients selected based on the mutation load. However, patients selected based on the mutation load are administered vaccines containing presented neoantigens at a higher rate than patients who are not selected. Thus, the mutation load is an effective patient selection criterion in effective neoantigen vaccine omission, but the expected utility value score is more effective.

[0376] XII. Example 8: Evaluation of a Mass Spectrometry Training Model Against Held-Out Mass Spectrometry Data Since HLA peptide presentation by tumor cells is a major requirement in anti-tumor immunity 91,96,97 , a large-scale (N = 74 patients) integrated dataset of human tumor and normal tissue samples containing paired class I HLA peptide sequences, HLA types, and transcriptome RNA-seq (methods) is used with these data and published data 92,98,99 to train a novel deep learning model for predicting antigen presentation in human cancer 100It was generated for the purpose of. Samples were selected based on the ease of tissue acquisition from among several tumor types targeted for the development of immunotherapy. By mass spectrometry, an average of 3,704 peptides per sample were identified at a peptide-level FDR of less than 0.1 (range of 344 - 11,301 species). These peptides were 8 - 15 amino acids in length and followed the characteristic class I HLA length distribution with a modal length of 9 (56% of the peptides). Consistent with previous reports, the majority of the peptides (median 79%) were predicted by MHCflurry to bind to at least one patient's HLA allele at a standard affinity threshold of 500 nM. 90 However, there was significant variation between samples (e.g., in one sample, the predicted affinity for 33% of the peptides was greater than 500 nM). The commonly used 101 threshold of 50 nM for "strong binders" captured only a median of 42% of the presented peptides. With transcriptome sequencing, an average of 131M unique reads were obtained per sample, and 68% of the genes were expressed at a level of at least 1 TPM (transcript per million) in at least one sample, which emphasizes the value of the large and diverse samples that were set up so that the maximum number of gene expressions would be observed. Peptide presentation by HLA was strongly correlated with mRNA expression. Differences in the proportion of peptide presentation that were significant and reproducible and greater than could be explained by differences in RNA expression or sequence alone were observed between genes. The observed HLA types were mostly consistent with the predictions for samples from patient groups with mainly European ancestry.

[0377] These, and published HLA peptide data 92,98,99A neural network (NN) model was trained using [the method] to predict HLA antigen presentation. From the tumor mass spectrometry data, a novel network architecture (method) was developed that can jointly learn allele-peptide mapping and allele-specific presentation motifs to learn allele-specific models where each peptide may be presented by any one of six HLA alleles. For each patient, the data points labeled as positive were the peptides detected by mass spectrometry, and the data points labeled as negative were peptides from the reference proteome (SwissProt) that were not detected by mass spectrometry in that sample. The data was split into training, validation, and test sets [using the method]. The training set consisted of 142,844 HLA-presented peptides (FDR < ~0.02) obtained from 101 samples (69 newly described in this study and 32 previously published). The validation set (used for early stopping) consisted of 18,004 presented peptides from the same 101 samples. The following two mass spectrometry data sets were used for testing. Namely, (1) a tumor sample test set consisting of 571 presented peptides obtained from five additional tumor samples (2 lung, 2 colon, 1 ovary) removed from the training data, and (2) a single allele cell line test set consisting of 2,128 presented peptides from genomic position windows (blocks) adjacent (but different) to the positions of single allele peptides included in the training data (see "Methods" for further details regarding the training / test split).

[0378] The training data identified prediction models for 53 HLA alleles. Previous studies 92,104Unlike these, these models captured the dependence of HLA presentation on each array position in peptides of multiple lengths. This model properly learned the critical dependence on gene RNA expression and gene-specific presentation tendencies, and when independently combining mRNA abundance and the learned tendency of presentation per gene, it produced a difference in presentation ratio of up to approximately 60-fold between the gene with the lowest expression and the least presentable and the gene with the highest expression and the most presentable. This model predicts the measured stability of HLA / peptide complexes in the IEDB even after controlling the predicted binding affinity 88 was further observed (p < 1e-10 for 10 alleles; p < 0.05 for 8 out of 10 alleles tested). Collectively, these properties form the basis for an improved prediction of immunogenic HLA class I peptides.

[0379] The performance of this NN model as a predictor of HLA presentation for the exclusion mass spectrometry test set was evaluated. Specifically, Figure 14 compares the positive predictive value (PPV) at a 40% recall rate of different versions of the MS model and a recently published approach (MixMHCPred) that models eluted peptides from mass spectrometry when each model was tested with five different exclusion test samples. Figure 14 also shows the average PPV at a 40% recall rate of each model for the five test samples.

[0380] Each model tested in Figure 14 is (from left to right), the "complete MS model" (the complete NN model described in the "Method"), the "MS model, no flanking array" (the same as the complete NN model except for excluding the flanking array characteristics), the "MS model, no flanking array and no gene-specific parameters" (the same as the complete NN model except for excluding the flanking array and gene-specific parameter characteristics), the "peptide-only MS model, all lengths trained together" (the same as the complete NN model except that the characteristics used are only the peptide sequence and HLA type), the "peptide-only MS model, each length trained separately" (in this model, the model structure is the same as the peptide-only MS model except that separate models were trained for 9-mers and 10-mers), the "linear peptide-only MS model (with ensembling)" (the same as the peptide-only MS model with each peptide length trained separately except that instead of using a neural network to model the peptide sequence, an ensemble of linear models trained using the same optimization procedure as used in the complete model and described in the "Method" was used); "MixMHCPred 1.1" is MixMHCPred with default settings; "binding affinity" is the same MHCflurry 1.2.0.

[0381] The "complete MS model", the "MS model, no flanking array", the "MS model, no flanking array and no gene-specific parameters", the "peptide-only MS model, all lengths trained together", the "peptide-only MS model, each length trained separately", and the "linear peptide-only MS model" are all neural network models trained with the mass spectrometry data described above. However, each model is trained and tested using different characteristics of the samples. The "MixMHCPred 1.1" model and the "binding affinity" model are initial approaches for modeling HLA-presented peptides 104。Since MixMHCPred does not currently model peptides of lengths other than 9 and 10, only 9-mers and 10-mers were used for comparison. The last five models ("Peptide-only MS model, trained together all lengths" to "Binding affinity") have the same input of only peptide sequence and HLA type. Specifically, none of the last five models uses RNA abundance for prediction.

[0382] The best-performing peptide-only model ("Peptide-only MS model, trained together all lengths") gave an average PPV of 0.41 at a recall of 40%, while the worst-performing peptide-only model trained on mass spectrometry data ("Linear peptide-only MS model") had an average PPV of only 28% (only slightly higher than the 18% average PPV of MixMHCpred1.1), highlighting the value of improved NN modeling of peptide sequences. MixMHCpred1.1 is trained on different data from the linear peptide-only MS model but has many of the same modeling characteristics (e.g., it is a linear model where each peptide length is trained separately).

[0383] Overall, the NN models achieved significantly improved predictions of HLA peptide presentation, with PPV up to 9-fold higher than standard binding affinity + gene expression in the tumor test set. The major advantage of the PPV of the NN models using MS was maintained across different recall thresholds and was statistically significant (p < 10 -6 ) for all tumor samples. The positive median rate of standard binding affinity + gene expression for HLA peptide presentation reached a low value of 6%, consistent with previous estimates. 87,93 However, it should be noted that this ~6% PPV represents an improvement of over 100-fold compared to the baseline prevalence rate, as only a small percentage of peptides are detected as being presented (e.g., about 1 out of 2500 in the tumor MS test dataset).

[0384] By comparing the reduced model trained with mass spectrometry data using only HLA type and peptide sequence as input with the complete MS model, it was determined that approximately 30% of the increase in PPV for the prediction of binding affinity was due to the modeling of exogenous characteristics (RNA abundance, flanking sequences, gene - specific parameters) of peptides that can be captured by mass spectrometry but not by binding affinity assays. The remaining approximately 70% of the increase was due to the improved modeling of the peptide sequence. This modeling showed higher performance than the initial approach in the modeling of HLA - presented peptides in human tumors, indicating that the improvement in performance was due not only to the nature of the training dataset (HLA - presented peptides) but also to the overall model architecture. With this new model architecture, it became possible to learn allele - specific models through an end - to - end training process that does not require prior assignment of peptides to the putative presenting alleles using a binding affinity prediction or hard - clustering approach. Importantly, this model architecture does not impose limitations that reduce the accuracy of allele - specific sub - models as an essential condition for convolution such as linear convolution, nor does it consider each peptide length separately. The complete model outperforms the performance of several simplified models and previously published approaches subject to these limitations. 104 Since it showed higher performance than 104~106 the initial approach, the improvement in performance was due not only to the nature of the training dataset (HLA - presented peptides) but also to the overall model architecture. With this new model architecture, it became possible to learn allele - specific models through an end - to - end training process that does not require prior assignment of peptides to the putative presenting alleles using a binding affinity prediction or hard - clustering approach. 104 This new model architecture does not impose limitations that reduce the accuracy of allele - specific sub - models as an essential condition for convolution such as linear convolution, nor does it consider each peptide length separately. The complete model outperforms the performance of several simplified models and previously published approaches subject to these limitations.

[0385] XIII. Example 9: Experimental Results including Presentation Hotspot Modeling To specifically evaluate the advantages of using presentation hot spot parameters in modeling HLA presentation, we compared the performance of a neural network presentation model incorporating presentation hot spot parameters with that of a neural network presentation model without incorporating presentation hot spot parameters. The basic neural network architecture was the same for both models and was the same as the presentation model described above in Section VII. Briefly, the model included peptide and flanking amino acid sequence parameters, RNA sequencing transcript data (TPM), protein family data, a sample identifier for each sample, and HLA-A, B, C types. An ensemble of five networks was used for each model. For the model including presentation hot spot parameters, Equation 12c described above in Section VIII.B.3 was used with a proteome block size of 10 per gene and a peptide length of 8 - 12.

[0386] The two models were compared by conducting experiments using the mass spectrometry dataset described above in Section XII. Specifically, five samples were excluded (held - out) from model training and evaluation for the purpose of fairly evaluating competing models. Ninety percent of the remaining samples were randomly divided for model training and 10% for training validation.

[0387] Figure 15A compares the average positive predictive value (PPV) across the recall rates of the presentation model using presentation hot spot parameters and the presentation model not using presentation hot spot parameters when each model was tested with five excluded test samples. The model incorporating presentation hot spot parameters outperformed the model without incorporating presentation hot spot parameters individually for each sample, with an average precision of 0.82 with presentation hot spot parameters and 0.77 without presentation hot spot parameters.

[0388] Figures 15B to F compare the precision-recall curves of a presentation model using presentation hot spot parameters and a presentation model not using presentation hot spot parameters when each model was tested with each of five exclusion test samples.

[0389] XIV. Example 10: Evaluation of Presentation Hot Spot Parameters for Identifying T Cell Epitopes The advantage of using presentation hot spot parameters in modeling HLA presentation to identify human tumor CD8 T cell epitopes (i.e., targets of immunotherapy) was further directly tested. Defining an appropriate test data set for this evaluation is difficult because the test data set must contain peptides that are recognized by T cells and presented on the tumor cell surface by HLA. Furthermore, a formal performance evaluation requires not only positively labeled (i.e., recognized by T cells) peptides but also a sufficient number of negatively labeled (i.e., tested but not recognized) peptides. Mass spectrometry data sets correspond to tumor presentation but not to T cell recognition, and conversely, priming after vaccination or T cell assays correspond to T cell recognition but not to tumor presentation.

[0390] To obtain an appropriate data set, the inventors collected published CD8 T cell epitopes from the following five recent studies that meet the required criteria. That is, Study A 96 examined TILs in nine patients with gastrointestinal tumors and reported recognition by 12 T cells out of 1,053 somatic SNV mutations tested by IFN-γ ELISPOT using the tandem mini gene (TMG) method in autologous DCs. Study B 84 also used TMG and reported recognition by 6 T cells out of 574 SNVs by CD8+PD-1+ circulating lymphocytes obtained from five melanoma patients. Study C 97 evaluated TILs obtained from three melanoma patients using pulsed peptide stimulation and found responses to 5 out of 381 tested SNV mutations. Study D 108Using a combination of TMG assays, TILs obtained from one breast cancer patient pulsed with minimal epitope peptides were evaluated, reporting recognition of 2 out of 62 SNVs. Study E 160 evaluated TILs in 17 patients obtained from the National Cancer Institute with 52 types of TSNA. The combined dataset included 4,843 assayed SNVs from 33 patients, including 75 types of TSNA showing existing T cell responses. Importantly, since the dataset consisted mostly of neoantigen recognition by tumor-infiltrating lymphocytes, an effective prediction for this dataset indicates that this model has the ability to identify not only neoantigens that can prime T cells as described in the above section, but also neoantigens presented by tumors to T cells.

[0391] To stimulate the selection of antigens for personalized immunotherapy, somatic mutations were ranked in order of presentation probability using two methods. Namely, (1) an MS model including hotspot characteristics (described by Equation 12c with block size n = 10), and (2) a conventional MS model without hotspot characteristics. Since the ability of antigen-specific immunotherapy is restricted in the number of specificities targeted (e.g., current personalized vaccines encode approximately 10 - 20 mutations 6、81~82 ), the prediction methods were compared by counting the number of existing T cell responses in peptides ranked 5th, 10th, 20th, or 30th for each patient. The results are shown in Table 16.

[0392] Specifically, FIG. 16 compares the percentage of peptides recognized by T cells across somatic mutations in peptides ranked 5th, 10th, 20th, and 30th by a presentation model using presentation hot spot parameters and a presentation model not using presentation hot spot parameters, for a test set consisting of test samples taken from patients with at least one existing T cell response. As shown in FIG. 16, the model including hot spot characteristics showed performance equivalent to that of the model without hot spot characteristics, and both models predicted 45 and 31 T cell responses, respectively, in the peptides ranked 20th and 10th. However, the hot spot model showed improvement when predicting the 30th and 5th ranked peptides, with the hot spot model including 6 and 4 more T cell responses, respectively.

[0393] XIII.A. Data The inventors obtained mutation calls, HLA types, and T cell recognition data from the supplementary information of Gros et al. 84 , Tran et al. 140 , Stronen et al. 141 and Zacharakis et al., and Kosaloglu-Yalcin et al. 160 For the analysis of the mutation level (FIG. 16), Gros et al., Tran et al., Zacharakis et al.

[0394] , and Kosaloglu-Yalcin et al. 108 160Data points shown as positive in [study name] were defined as mutations recognized by patient T cells in both the TMG assay and the minimal epitope peptide pulse assay. Data points shown as negative were defined as all other mutations tested in the TMG assay. In Stronen et al., mutations shown as positive were defined as mutations spanning at least one recognized peptide, and negative data points were defined as all mutations that were tested but not recognized in the tetramer assay. The mutated 25-mer TMG assay tests T cell recognition of all peptides spanning the mutation, so for the data of Gros, Tran, and Zacharakis, mutations were ranked by summing the probabilities of presentation or taking the minimum binding affinity across all peptides spanning the mutation. For the data of Stronen, mutations were ranked by summing the probabilities of presentation or taking the minimum binding affinity across all peptides spanning the mutations tested in the tetramer assay. A complete list of mutations and characteristics is shown in Supplementary Table 1.

[0395] At the epitope level of analysis, data points shown as positive were defined as all minimal epitopes recognized by patient T cells in the peptide pulse assay or the tetramer assay, and negative data points were defined as all minimal epitopes not recognized by T cells in the peptide pulse assay or the tetramer assay, and all peptides spanning mutations from the tested TMGs not recognized by patient T cells. For Gros et al., Tran et al., and Zacharakis et al., minimal epitope peptides spanning mutations recognized in the TMG analysis but not tested by the peptide pulse assay were excluded from the analysis because the T cell recognition status of these peptides could not be experimentally examined.

[0396] XV. Example 11: Identification of Neoantigen-Responsive T Cells in Cancer Patients In this example, we demonstrate that the improved predictions enable the identification of neoantigens from normal patient samples. To do this, archived FFPE tumor biopsies and 5 - 30 ml of peripheral blood were analyzed in nine patients with metastatic NSCLC undergoing anti - PD(L)1 therapy (Supplementary Table 2: Patient demographics and treatment information for N = 9 patients examined in Figure 17A - C. The main fields include tumor stage and subtype, anti - PD1 therapy performed, and an overview of the NGS results). By tumor whole - exome sequencing, tumor transcriptome sequencing, and matched normal exome sequencing, an average of 198 somatic mutations (SNVs and short indels) were obtained per patient, of which an average of 118 were expressed (\"Methods\", Supplementary Table 2). Twenty neoepitopes per patient were prioritized to test against existing anti - tumor T - cell responses by applying the complete MS model. To focus the analysis on likely CD8 responses, the prioritized peptides were synthesized as 8 - 11 - mer minimal epitopes (\"Methods\"), and then peripheral blood mononuclear cells (PBMCs) were cultured with the synthesized peptides in short - term in vitro stimulation (IVS) cultures to expand neoantigen - reactive T cells (Supplementary Table 3). After two weeks, the presence of antigen - specific T cells was evaluated using IFN - γ ELISpot against the prioritized neoepitopes. Separate experiments were further performed in seven patients with sufficient PBMCs available to perform a complete or partial de - convolution of the recognized specific antigens. These results are shown in Figures 17A - C and Figures 18A - 21.

[0397] Figure 17A shows the detection of T - cell responses to patient - specific neoantigen peptide pools in nine patients. For each patient, the predicted neoantigens were combined into two pools of 10 peptides each according to model ranking and any sequence homology (homologous peptides were split into different pools). Then, for each patient, the PBMCs expanded in vitro for that patient were stimulated with the two patient - specific neoantigen peptide pools in an IFN - γ ELISpot. The data in Figure 17A are seeded cells 10 minus the background (corresponding DMSO negative control) 5Shown as the spot-forming units (SFUs) per well. The background measurements (DMSO negative control) are shown in Figure 21. For patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, and CU05, the responses of a single well (patients 1-038-001, CU02, CU03, and 1-050-001) or replicates (all other patients) including the mean and standard deviation to the cognate peptide pools #1 and #2 are shown. In patients CU02 and CU03, due to cell numbers, testing was only possible against the specific peptide pool #1. Samples with a doubling rate value more than twice the background were considered positive and are indicated with an asterisk (responsive donors include patients 1-038-001, CU04, 1-024-001, 1-024-002, and CU02). Non-responsive donors include patients 1-050-001, 1-001-002, CU05, and CU03. Figure 17C shows a photograph of an ELISpot well containing in vitro-expanded PBMC from patient CU04 stimulated with DMSO negative control, PHA positive control, CU04-specific neoantigen peptide pool #1, CU04-specific peptide 1, CU04-specific peptide 6, and CU04-specific peptide 8 in an IFN-γ ELISpot.

[0398] Figures 18A - B show the results of control experiments using patient neoantigens in HLA-matched healthy donors. The results of these experiments indicate that the in vitro culture conditions did not enable de novo priming in vitro but only expanded existing in vivo-primed memory T cells.

[0399] Figure 19 shows the detection of T cell responses to the PHA positive control for each donor and each in vitro expansion shown in Figure 17A. For each donor and each in vitro expansion in Figure 17A, the in vitro-expanded patient PBMC were stimulated with PHA for maximum T cell activation. The data in Figure 19 are the seeded cells 10 minus the background (corresponding DMSO negative control) 5Shown as the spot forming units (SFUs) per well. The responses of a single well or biological replicate are shown for patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, CU05, and CU03. In patient CU02, the test with PHA was not performed. Since the positive response to peptide pool #1 (Figure 17A) indicated viable and functional T cells, cells from patient CU02 were included in the analysis. As shown in Figure 17A, donors responsive to the peptide pool included patients 1-038-001, CU04, 1-024-001, and 1-024-002. Also as shown in Figure 17A, donors not responsive to the peptide pool included patients 1-050-001, 1-001-002, CU05, and CU03.

[0400] Figure 20A shows the detection of T cell responses to each individual patient-specific neoantigen peptide in pool #2 in patient CU04. Figure 20A also shows the detection of T cell responses to the PHA positive control in patient CU04. (This positive control data is also shown in Figure 19.) In patient CU04, the patient's in vitro-expanded PBMC were stimulated in an IFN-γ ELISpot with individual patient-specific neoantigen peptides from pool #2 for patient CU04. The patient's in vitro-expanded PBMC were also stimulated in an IFN-γ ELISpot with PHA as a positive control. The data are the seeded cells 10 minus the background (corresponding DMSO negative control) 5 Shown as the spot forming units (SFUs) per well.

[0401] Figure 20B shows the detection of T cell responses to individual patient-specific neoantigen peptides at each of the three visits of patient CU04 and at each of the two visits of patient 1-024-002 (each visit is performed at a different time point). In both patients, the patient's in vitro-expanded PBMC were stimulated in an IFN-γ ELISpot with individual patient-specific neoantigen peptides. For each patient, the data for each visit are the seeded cells 10 minus the background (corresponding DMSO control)5 is shown as the cumulative (sum) spot-forming units (SFUs) per visit. The data for patient CU04 are shown as the cumulative SFUs for three visits after subtracting the background. For patient CU04, the SFUs after subtracting the background are shown for the first visit (T0) and subsequent visits at 2 months (T0 + 2 months) and 14 months (T0 + 14 months) after the first visit (T0). The data for patient 1-024-002 are shown as the cumulative SFUs for two visits after subtracting the background. For patient 1-024-002, the SFUs after subtracting the background are shown for the first visit (T0) and subsequent visits at 1 month (T0 + 1 month) after the first visit (T0). Samples with a doubling rate value more than twice higher than the background were considered positive and are indicated by an asterisk.

[0402] Figure 20C shows the detection of T cell responses against individual patient-specific neoantigen peptides and patient-specific neoantigen peptide pools in each of two visits for patient CU04 and in each of two visits for patient 1-024-002 (each visit is performed at a different time point). In both patients, the in vitro-expanded PBMC of that patient were stimulated with patient-specific individual neoantigen peptides and patient-specific neoantigen peptide pools in IFN-γ ELISpot. Specifically, for patient CU04, the in vitro-expanded PBMC of patient CU04 were stimulated with CU04-specific individual neoantigen peptides 6 and 8 and the CU04-specific neoantigen peptide pool in IFN-γ ELISpot, and for patient 1-024-002, the in vitro-expanded PBMC of patient 1-024-002 were stimulated with 1-024-002-specific individual neoantigen peptides 16 and the 1-024-002-specific neoantigen peptide pool in IFN-γ ELISpot. The data in Figure 20C are for each technical replicate with mean and range, seeded cells after subtracting the background (corresponding DMSO control) 10 5 ​Shown as the spot-forming units (SFUs) per well. The data for patient CU04 are shown as the SFUs for two visits after subtracting the background. For patient CU04, the SFUs after subtracting the background are shown for the first visit (T0, technical triplicate) and the subsequent visit two months after the first visit (T0 + 2 months, technical triplicate). The data for patient 1-024-002 are shown as the SFUs for two visits after subtracting the background. For patient 1-024-002, the SFUs after subtracting the background are shown for the first visit (T0, technical triplicate) and the subsequent visit one month after the first visit (T0 + 1 month, technical duplicate excluding samples stimulated with the patient 1-024-002-specific neoantigen peptide pool).

[0403] Figure 21 shows the detection of T cell responses to two patient-specific neoantigen peptide pools and a DMSO negative control for the patient of Figure 17A. For each patient, PBMCs expanded in vitro for that patient were stimulated with two patient-specific neoantigen peptide pools by IFN-γ ELISpot. For each donor and each in vitro expansion, the patient PBMCs expanded in vitro were also stimulated with DMSO as a negative control in IFN-γ ELISpot. The data in Figure 21 are the seeded cells 10 including the background (corresponding DMSO negative control) for the patient-specific neoantigen peptide pools and the corresponding DMSO controls 5Shown as the spot-forming units (SFUs) per individual. For patients 1-038-001, 1-050-001, 1-001-002, CU04, 1-024-001, 1-024-002, and CU05, the responses of a single well (patients 1-038-001, CU02, CU03, and 1-050-001) or the mean including the standard deviation of biological duplicates to the cognate peptide pools #1 and #2 are shown. For patients CU02 and CU03, due to cell numbers, testing was only possible against the specific peptide pool #1. Samples with a doubling rate value more than twice higher than the background were considered positive and are indicated by an asterisk (responsive donors include patients 1-038-001, CU04, 1-024-001, 1-024-002, and CU02). Non-responsive donors include patients 1-050-001, 1-001-002, CU05, and CU03.

[0404] As briefly described above with respect to FIGS. 18A - B, to confirm that the in vitro culture conditions did not enable de novo priming in vitro but only expanded existing in vivo-primed memory T cells, a series of control experiments were performed using neoantigens with HLA-matched healthy donors. The results of these experiments are shown in FIGS. 18A - B and Supplementary Table 5. These experimental results confirmed that no de novo priming occurred and no detectable neoantigen-specific T cell responses occurred in healthy donors using the IVS culture method.

[0405] In contrast, existing neoantigen-responsive T cells were identified in the majority (5 / 9, 56%) of patients tested with patient-specific peptide pools using IFN-γ ELISpot (Figures 17A and 19 - 21). Of the seven patients in whom cell numbers enabled complete or partial testing of individual neoantigen cognate peptides, four patients responded to at least one of the neoantigen peptides tested, and all of these patients showed a response to the corresponding pool (Figure 17B). The remaining three patients (patients 1-001-002, 1-050-001, and CU05) tested with individual neoantigens showed no detectable response to a single peptide (data not shown), and it was confirmed that there was no response to the neoantigen pool in these patients (Figure 17A). Of the four responding patients, samples from one visit were obtained in two patients who showed a response (patients 1-024-001 and 1-038-001), and samples from multiple visits were obtained in the remaining two patients who showed a response (CU04 and 1-024-002). For the two patients with samples from multiple visits, the cumulative (sum) spot-forming units (SFU) from three visits (patient CU04) and two visits (patient 1-024-002) are shown in Figure 17B, and the breakdown by visit is shown in Figure 20B. Additional PBMC samples from the same visit were also obtained in patients 1-024-002 and CU04, and responses to patient-specific neoantigens were confirmed by repeated IVS culture and ELISpot (Figure 20C).

[0406] Overall, among the patients in whom at least one T cell - recognized neo - epitope was identified as shown by the response to the pool of 10 peptides in Figure 17A, the number of recognized neo - epitopes was at least 2 per patient on average (counting recognized pools that could not be back - folded as 1 recognized peptide, with a minimum of 10 epitopes identified in 5 patients). In addition to testing for IFN - γ responses by ELISpot, culture supernatants were tested for granzyme B by ELISA, and for TNF - α, IL - 2, and IL - 5 by MSD cytokine multiplex assay. Among the 5 patients who showed positive ELISpots, cells from 4 patients secreted more than 3 test substances including granzyme B (Supplementary Table 4), indicating the multifunctionality of neo - antigen - specific T cells. Importantly, the combination of the prediction and IVS methods does not rely on a limited set of available MHC multimers, so the responses were widely tested across restricted HLA alleles. Furthermore, this approach directly identifies the minimal epitope, unlike tandem mini - gene screening which requires another back - folding step to identify the recognized mutations and the minimal epitope. Overall, the specific yield of neo - antigens was equivalent to the best conventional method of testing TILs against all mutations using apheresis samples 96 while only screening 20 synthetic peptides using only 5 - 30 mL of normal whole blood.

[0407] XV.A. Peptide Custom - made recombinant lyophilized peptides were purchased from JPT Peptide Technologies (Berlin, Germany) or Genscript (Piscataway, NJ, USA), reconstituted at 10 - 50 mM in sterile DMSO (VWR International, Pittsburgh, PA, USA), aliquoted, and stored at - 80 °C.

[0408] XV.B. Human Peripheral Blood Mononuclear Cells (PBMC) Lyophilized HLA-typed PBMC from healthy donors (confirmed to be seronegative for HIV, HCV, and HBV) were purchased from Precision for Medicine (Gladstone, NJ, USA) or Cellular Technology, Ltd. (Cleveland, OH, USA) and stored in liquid nitrogen until use. Fresh blood samples were purchased from Research Blood Components (Boston, MA, USA), leukopaks were purchased from AllCells (Boston, MA, USA), PBMC were isolated by Ficoll-Paque density gradient (GE Healthcare Bio, Marlborough, MA, USA), and then cryopreserved. Patient PBMC were processed according to a protocol approved by the local clinical standard operating procedure (SOP) and IRB at the local clinical processing center. The approved IRBs were Quorum Review IRB, Comitato Etico Interaziendale A.O.U. San Luigi Gonzaga di Orbassano, and Comite Etico de la Investigacion del Grupo Hospitalario Quiron en Barcelona.

[0409] Briefly, PBMC were isolated by density gradient centrifugation, washed, counted, and resuspended at 5x10 6They were cryopreserved at cells / ml. The cryopreserved cells were shipped by cryoport, transferred after arrival, and stored in LN2. The patient demographics are shown in Supplementary Table 2. The cryopreserved cells were thawed, washed twice in OpTmizer T-cell Expansion Basal Medium (Gibco, Gaithersburg, MD, USA) supplemented with Benzonase (EMD Millipore, Billerica, MA, USA), and once without Benzonase. Cell count and viability were evaluated using the modules on the Guava ViaCount reagent and Guava easyCyte HT cytometer (EMD Millipore). The cells were then resuspended in a medium suitable for the assay at a concentration suitable for the subsequent assay (see the next section).

[0410] XV.C. In Vitro Stimulation (IVS) Culture Existing T cells obtained from healthy donor or patient samples were expanded in the presence of cognate peptides and IL-2 using an approach similar to that applied by Ott et al 81 Briefly, the thawed PBMC were allowed to rest overnight and stimulated for 14 days in ImmunoCult™-XF T-cell Expansion Medium (STEMCELL Technologies) supplemented with 10 IU / ml of rhIL-2 (R&D Systems Inc., Minneapolis, MN) in a 24-well tissue culture plate in the presence of a peptide pool (10 μM per peptide, 10 peptides per pool). The cells were seeded at 2x10 6 cells / well and cultured by replacing 2 / 3 of the medium every 2 - 3 days. One patient sample showed a deviation from the protocol and was considered a potential false negative. Patient CU03 did not yield a sufficient number of cells after thawing, so the cells were seeded at 2x10 5 cells per peptide pool (10-fold less than described in the protocol).

[0411] XV.D. IFNγ Enzyme-Linked Immunospot (ELISpot) Assay Detection of IFNγ-producing T cells was performed by ELISpot assay 142 Briefly, PBMC (after ex vivo or in vitro expansion) were harvested, washed in serum-free RPMI (VWR International), and cultured in OpTmizer T-cell Expansion Basal Medium (ex vivo) or ImmunoCult™-XF T-cell Expansion Medium (expanded cultures) in the presence of control or cognate peptides in ELISpot Multiscreen plates (EMD Millipore) coated with anti-human IFNγ capture antibody (Mabtech, Cincinatti, OH, USA). After incubation for 18 hours in a humidified incubator at 5% CO2, 37 °C, the cells were removed from the plates and IFNγ bound to the membrane was detected using anti-human IFNγ detection antibody (Mabtech), Vectastain Avidin peroxidase conjugate (Vector Labs, Burlingame, CA, USA), and AEC Substrate (BD Biosciences, San Jose, CA, USA). The ELISpot plates were dried, stored in the dark, and sent to Zellnet Consulting, Inc., Fort Lee, NJ, USA) for standardized evaluation 143. Data are presented as spot forming units (SFU) per number of cells seeded in the plates.

[0412] XV.E. Granzyme B ELISA and MSD multiplex assay Detection of IL-2, IL-5, and TNF-α secreted into the ELISpot supernatant was performed using the MSD U-PLEX Biomarker assay (Catalog number K15067L-2), a multiplex assay. The assay was performed according to the manufacturer's instructions. The analyte concentration (pg / ml) was calculated for each cytokine using serial dilutions of known standards. To graph the data, values below the minimum range of the standard curve were set equal to 0. Detection of granzyme B in the ELISpot supernatant was performed using GranzymeB DuoSet® ELISA (R&D Systems, Minneapolis, MN) according to the manufacturer's instructions. Briefly, the ELISpot supernatant was diluted 1:4 in sample diluent and run in parallel with serial dilutions of granzyme B standards to calculate the concentration (pg / ml). To graph the data, values below the minimum range of the standard curve were set equal to 0.

[0413] XV.F.IVS assay negative control experiment - neoantigens derived from tumor cell lines tested in healthy donors Figure 18A shows the negative control experiment of the IVS assay for neoantigens derived from tumor cell lines tested in healthy donors. PBMCs from healthy donors were stimulated during IVS culture with a peptide pool containing a positive control peptide (previously exposed to an infectious disease), a neoantigen with HLA matching that of the tumor cell line (not exposed), and a peptide derived from a pathogen for which the donor was seronegative. The expanded cells were then stimulated with DMSO (negative control, black circles), PHA and a general infectious disease peptide (positive control, red circles), neoantigen (not exposed, light blue circles), or HIV and HCV peptides (from pathogens for which the donor was seronegative. Dark blue, A and B), and then analyzed by IFNγ ELISpot (10 5 cells / well). The data are shown as spot-forming units (SFUs) per 10 5 seeded cells. Biological replicates including the mean and SEM are shown. No response was observed to the neoantigen or peptides derived from pathogens to which the donor was not exposed (seronegative).

[0414] XV.G.IVS assay negative control experiment - Neoantigens derived from patients tested in healthy donors Figure 18A shows the negative control experiment of the IVS assay for neoantigens derived from patients tested for responsiveness in healthy donors. Evaluation of T cell responses in healthy donors against HLA-matched neoantigen peptide pools. Left panel: PBMCs from healthy donors were stimulated with controls (DMSO, CEF, and PHA) or HLA-matched patient-derived neoantigen peptides in ex vivo IFNγ ELISpot. Data are shown as spot-forming units (SFU) per 2x10 5 cells seeded per well of the plate for triplicate wells. Right panel: PBMCs from healthy donors after IVS culture grown in the presence of neoantigen pool or CEF pool were stimulated with controls (DMSO, CEF, and PHA) or HLA-matched patient-derived neoantigen peptide pools in IFNγ ELISpot. Data are shown as SFU per 1x10 5 cells seeded per well of the plate for triplicate wells. No response to neoantigens was observed in healthy donors.

[0415] XV.H. Supplementary Table 3: Peptides tested for T cell recognition in NSCLC patients Details of the neoantigen peptides tested in N = 9 patients examined in Figures 17A - C (Identification of neoantigen-responsive T cells from NSCLC patients). The main fields include the source mutation, peptide sequence, and the observed pool and individual peptide sequences. The column "most_probable_restriction" indicates which allele the model predicted was most likely to present each peptide. Also included is the rank of these peptides among all mutant peptides of each patient calculated by the binding affinity prediction ("method").

[0416] Four peptides were ranked highly by the complete MS model and recognized by CD8 T cells with low predicted binding affinity or ranked low by binding affinity prediction.

[0417] For three of these peptides, this is due to differences in HLA coverage between the model and MHCflurry 1.2.0. Peptide YEHEDVKEA is predicted to be presented by HLA-B*49:01 which is not covered by MHCflurry 1.2.0. Similarly, peptides SSAAAPFPL and FVSTSDIKSM are predicted to be presented by HLA-C*03:04 which is also not covered by MHCflurry 1.2.0. The online NetMHCpan 4.0 (BA) prediction tool, a pan-allele covering predictive binding affinity tool that in principle covers all alleles, ranks SSAAAPFPL as a strong binder to HLA-C*03:04 (23.2 nM, ranked 2nd in patient 1-024-002), predicts weak binding of FVSTSDIKSM to HLA-C*03:04 (943.4 nM, ranked 39th in patient 1-024-002) and weak binding of YEHEDVKEA to HLA-B*49:01 (3387.8 nM), and also predicts stronger binding of YEHEDVKEA to HLA-B*41:01 (208.9 nM, ranked 11th in patient 1-038-001) which is also present in this patient but not covered by the model. Thus, among these three peptides, FVSTSDIKSM would have been missed by binding affinity prediction, SSAAAPFPL would have been captured, and the HLA restriction of YEHEDVKEA is unclear.

[0418] The remaining five peptides with peptide-specific T cell responses that were reverse-folded were from patients where the most likely presenting alleles determined by the model were also covered by MHCflurry 1.2.0. Four out of five (4 / 5) of these peptides had predicted binding affinities stronger than the standard 500 nM threshold and were ranked in the top 20, but were ranked somewhat lower than by the model (peptide TIFF0007712970000078.tif4128 was ranked 2nd, 14th, 7th, and 9th by MHCflurry, while it was ranked 0th, 4th, 5th, and 7th by the model, respectively. The peptide GTKKDVDVLK was recognized by CD8 T cells and ranked 1st by the model, but its rank by MHCflurry was 70th, and the predicted binding affinity was 2169 nM.

[0419] Overall, 6 out of 8 (6 / 8) of the individually recognized peptides that were highly ranked by the complete MS model were also highly ranked when using binding affinity prediction, and the predicted binding affinity was less than 500 nM. In contrast, 2 out of 8 (2 / 8) of the individually recognized peptides would likely have been missed if binding affinity prediction had been used instead of the complete MS model.

[0420] XV.I. Supplementary Table 4: MSD Cytokine Multiplex and ELISA Assays for ELISpot Supernatants Obtained from NSCLC Neoantigen Peptides The analytes detected in the supernatants obtained from positive ELISpot (IFNγ) wells are shown for Granzyme B (ELISA), TNFα, IL-2, and IL-5 (MSD). Values are shown as the average pg / ml from technical replicates. Positive values are shown in italics. Granzyme B ELISA: Values 1.5-fold or more above the DMSO background were considered positive. U-Plex MSD assay: Values 1.5-fold or more above the DMSO background were considered positive.

[0421] XV.J. Supplementary Table 5: Neoantigens and Infectious Disease Epitopes in the IVS Control Experiment The details of the tumor cell line neoantigens and viral peptides tested in the IVS control experiment shown in Figures 18A - B include the source cell line or virus, the peptide sequence, and the predicted presenting HLA allele.

[0422] XV.K. Data The MS peptide dataset (Figure 16) used to train and test the prediction model is available from the MassIVE archive (massive.ucsd.edu) with accession number MSV000082648. Neoantigen peptides tested by ELISpot (Figures 17A - C and 18A - B) are included with the manuscript (Supplementary Tables 3 and 5).

[0423] XVI. Methods of Examples 8 - 11 XVI.A. Mass Spectrometry XVI.A.1. Samples Archived frozen tissue samples for mass spectrometry analysis were obtained from vendors including BioServe (Beltsville, MD), ProteoGenex (Culver City, CA), iSpecimen (Lexington, MA), and Indivumed (Hamburg, Germany). A subset of the samples was also pre - collected from patients at Hopital Marie Lannelongue (Le Plessis - Robinson, France) based on a research protocol approved by the Comite de Protection des Personnes, Ile - de - France VII.

[0424] XVI.A.2. HLA Immunoprecipitation Isolation of HLA peptide molecules was performed using immunoprecipitation (IP) method after lysis and solubilization of tissue samples 87,124-126 Fresh frozen tissue was pulverized (CryoPrep; Covaris, Woburn, MA), lysis buffer (1% CHAPS, 20 mM Tris - HCl, 150 mM NaCl, protease and phosphatase inhibitors, pH = 8) was added to solubilize the tissue, and the resulting solution was centrifuged at 4°C for 2 hours to pellet debris. The clarified lysate was used for HLA - specific IP. Immunoprecipitation was performed using the antibody W6 / 32 as described previously 127The lysate was added to antibody beads and rotated overnight at 4°C for immunoprecipitation. After immunoprecipitation, the beads were removed from the lysate. The IP beads were washed to remove non-specific binding, and the HLA / peptide complex was eluted from the beads with 2N acetic acid. The protein components were removed from the peptides using a molecular weight spin column. The resulting peptides were dried by SpeedVac evaporation and stored at -20°C until MS analysis was performed.

[0425] XVI.A.3 Peptide Sequencing The dried peptides were resuspended in HPLC buffer A, loaded onto a C-18 microcapillary HPLC column and gradient eluted into a mass spectrometer. The peptides were eluted into a Fusion Lumos mass spectrometer (Thermo) using a 180-minute gradient of 0 - 40% B (solvent A: 0.1% formic acid, solvent B: 0.1% formic acid in 80% acetonitrile). The MS1 spectrum of the peptide mass / charge (m / z) was collected in an Orbitrap detector with a resolution of 120,000, followed by 20 low-resolution MS2 scans in an Orbitrap or ion trap detector after HCD fragmentation of the selected ions. Selection of MS2 ions was performed using data-dependent acquisition mode and 30 seconds of dynamic exclusion after MS2 selection of the ions. The automatic gain control (AGC) for MS1 scans was set to 4x10 5 and for MS2 scans it was set to 1x10 4 For HLA peptide sequencing, charge states of +1, +2, and +3 can be selected for MS2 fragmentation.

[0426] The MS2 spectra from each analysis were searched against a protein database using Comet 128,129 and peptide identifications were scored using Percolator 130~132

[0427] XVI.B. Machine Learning XVI.B.1. Encoding of Data ​For each sample, the training data points were all 8- to 11-mer (inclusive) peptides from the reference proteome that mapped to exactly one gene expressed in the sample. The overall training data set was generated by concatenating the training data sets from each training sample. The range of 8 to 11 was chosen to capture approximately 95% of all HLA class I presented peptides, although adding lengths of 12 to 15 could also be achieved using the same method at the cost of a moderate increase in computational requirements. Peptides and flanking sequences were vectorized using a one-hot encoding scheme. Peptides of multiple lengths (8 to 11) were represented as fixed-length vectors by padding the amino acid alphabet with padding characters and padding all peptides to a maximum length of 11. The RNA abundance of the source protein of the training peptides was represented as the logarithm of the isoform-level TPM (transcripts per million) estimates obtained from RSEM 133 and used as the logarithm of the isoform-level TPM (transcripts per million) estimates obtained from RSEM. For each peptide, the per-peptide TPM was calculated as the sum of the per-isoform TPM estimates for each isoform containing the peptide. Peptides from genes expressed at 0 TPM were excluded from the training data, and at test time, peptides from non-expressed genes were assigned a presentation probability of 0. Finally, each peptide was assigned an Ensembl protein family ID, and each unique Ensembl protein family ID corresponded to a gene-specific presentation propensity slice (see the next section).

[0428] XVI.B.2. Specification of the Model Architecture The full presentation model has the following functional form. That is, TIFF0007712970000079.tif5128where k is the index of the HLA allele within the data set from 1 to m, TIFF0007712970000080.tif5128is the label variable, which has a value of 1 if allele k is present in the sample from which peptide i is derived and 0 otherwise. For a particular peptide i, Of the 5128 in TIFF0007712970000081, all except for the maximum 6 (corresponding to the HLA type of the sample from which peptide i is derived) are 0. The sum of the probabilities is, for example, 4128 in TIFF0007712970000082, Clip at 3128 in TIFF0007712970000083.

[0429] Model the per-allele probability of presentation as follows. That is, Pr(presentation of peptide i by allele a) = sigmoid{NN a (peptide i ) + NN フランキング (flanking i ) + NN RNA (log(TPM i )) + α 試料(i) + β タンパク質(i)} In the formula, the variables have the following meanings. sigmoid is the sigmoid (also known as expit) function, peptide i is the one-hot encoded and middle-padded amino acid sequence of peptide i, NN a is a neural network with a linear final layer activation that models the contribution of the peptide sequence to the presentation probability, flanking i is the one-hot encoded flanking sequence of peptide i in the source protein, NN フランキング is a neural network with a linear final layer activation that models the contribution of the flanking sequence to the presentation probability, TPM i is the expression of the source mRNA of peptide i within TPM units, sample(i) is the sample (i.e., patient) from which peptide i is derived, α 試料(i) is the section for each sample, protein(i) is the source protein of peptide i, and β タンパク質(i) is the section for each protein (also known as the per-gene trend of presentation).

[0430] In the model described in the Results section, the neural network for each component has the following architecture. That is, · NN a each of which is a single output node of a single-hidden layer multi-layer perceptron (MLP) with an input dimension of 231 (11 residues × 21 possible characters per residue (including the pad character)), a width of 256, a rectified linear unit (ReLU) activation in the hidden layer, a linear activation in the output layer, and one output node for ea...

Claims

1. A method for identifying one or more neoantigens derived from one or more tumor cells of a subject that are likely to be presented on the surface of the tumor cells, comprising: obtaining data representing the peptide sequence of each of the sets of neoantigens; encoding each of the peptide sequences of the neoantigens into a corresponding numerical vector, each numerical vector including information regarding a plurality of amino acids constituting the peptide sequence and a set of positions of the amino acids within the peptide sequence, the encoding step; inputting the numerical vectors into a presentation model that has been machine-learned using a computer processor to generate a set of presentation likelihoods for the set of neoantigens, wherein each presentation likelihood within the set represents the likelihood that the corresponding neoantigen is presented on the surface of the tumor cells of the subject by one or more MHC alleles, and the machine-learned presentation model is a label obtained by mass spectrometry that measures the presence of a peptide bound to at least one MHC allele within a set of MHC alleles identified as present in each of a plurality of samples, for each of the samples, a training peptide sequence encoded as a numerical vector including information regarding a plurality of amino acids constituting the peptide and a set of positions of the amino acids within the peptide, identified based at least on a training dataset including the plurality of parameters including one or more hotspot parameters representing the presence or absence of presentation hotspots, the inputting step; selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a selected set of neoantigens; and returning the selected set of neoantigens The method comprising.

2. The step of inputting the numerical vectors into the machine-learned presentation model is applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate a dependency score for each of the one or more MHC alleles indicating whether the MHC allele presents the neoantigen based on a specific amino acid at a specific position of the peptide sequence The method according to claim 1, comprising.

3. The step of inputting the numerical vectors into the machine-learned presentation model is For each MHC allele, converting the dependency score to generate a per-allele likelihood corresponding to the likelihood that the corresponding MHC allele will present the corresponding neoantigen; further comprising combining the per-allele likelihoods to generate a presentation likelihood of the neoantigen; converting the dependency score models the presentation of the neoantigen as mutually exclusive across the one or more MHC alleles or as interfering between the one or more MHC alleles; The method according to claim 2.

4. The set of presentation likelihoods is further specified by at least one or more allele non-interaction characteristics; applying the machine-learned presentation model to the allele non-interaction characteristics to generate a dependency score for the allele non-interaction characteristics indicating whether the peptide sequence of the corresponding neoantigen is presented based on the allele non-interaction characteristics; The method according to claim 2 or 3, further comprising.

5. (A) combining the dependency score for each MHC allele of the one or more MHC alleles with the dependency score for the allele non-interaction characteristics; converting the combined dependency score for each MHC allele to generate a per-allele likelihood for each MHC allele indicating the likelihood that the corresponding MHC allele will present the corresponding neoantigen; combining the per-allele likelihoods to generate the presentation likelihood; further comprising; or (B) combining the dependency score for each of the MHC alleles with the dependency score for the allele non-interaction characteristics; converting the combined dependency score to generate the presentation likelihood; further comprising; The method according to claim 4.

6. (A) the one or more MHC alleles include two or more different MHC alleles; (B) the peptide sequence includes peptide sequences having lengths other than 9 amino acids; (C) encoding the peptide sequence includes encoding the peptide sequence using a one-hot encoding scheme; (D) the plurality of samples includes (a) one or more cell lines engineered to express a single MHC allele; (b) one or more cell lines engineered to express multiple MHC alleles; (c) one or more human cell lines obtained from or derived from multiple patients, (d) fresh or frozen tumor samples obtained from multiple patients, and (e) fresh or frozen tissue samples obtained from multiple patients including at least one of; and / or (E) the training data set is (a) data related to the measured peptide-MHC binding affinity for at least one of the peptides, and (b) data related to the measured peptide-MHC binding stability for at least one of the peptides further including at least one of, The method according to any one of claims 1 to 5.

7. (A) the set of presentation likelihoods is further specified by the expression level of at least one of the one or more MHC alleles in the subject, measured by RNA-seq or mass spectrometry; (B) the set of presentation likelihoods is (a) the predicted affinity between neoantigens within the set of neoantigens and the one or more MHC alleles, and (b) the predicted stability of the neoantigen-encoding peptide-MHC complex further specified by at least one of the properties including; and / or (C) the set of presentation likelihoods is (a) the C-terminal sequence adjacent to the neoantigen-encoding peptide sequence, and (b) the N-terminal sequence adjacent to the neoantigen-encoding peptide sequence further specified by at least one of the properties including, The method according to any one of claims 1 to 6.

8. Selecting the set of selected neoantigens is (A) based on the machine-learned presentation model, neoantigens with an increased likelihood of being presented on the surface of the tumor cells compared to non-selected neoantigens; (B) based on the machine-learned presentation model, neoantigens with an increased likelihood of inducing a tumor-specific immune response in the subject compared to non-selected neoantigens; (C) based on the presentation model, neoantigens with an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) compared to non-selected neoantigens; (D) based on the machine-learned presentation model, neoantigens with a decreased likelihood of being inhibited by central or peripheral tolerance compared to non-selected neoantigens; and / or (E) neoantigens with a reduced likelihood of inducing an autoimmune response against normal tissue in the subject compared to non - selected neoantigens, based on the machine - learned presentation model The method according to any one of claims 1 to 7, comprising selecting such neoantigens **Claim 9** (A) the one or more tumor cells are selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B - cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T - cell lymphocytic leukemia, non - small cell lung cancer, and small cell lung cancer; and / or (B) the method further comprises generating an output for constructing an individualized cancer vaccine from the set of selected neoantigens, the output for the individualized cancer vaccine comprising at least one peptide sequence or at least one nucleotide sequence encoding the set of selected neoantigens The method according to any one of claims 1 to 8 **Claim 10** The method according to any one of claims 1 to 9, wherein the machine - learned presentation model is a neural network model **Claim 11** The method according to claim 10, wherein the neural network model comprises a plurality of network models for MHC alleles, each network model being assigned to a corresponding MHC allele among the plurality of MHC alleles and comprising a series of nodes arranged in one or more layers **Claim 12** The method according to claim 10 or 11, wherein the machine - learned presentation model is a deep learning model comprising one or more layers of nodes **Claim 13** The method according to any one of claims 1 to 12, wherein the one or more MHC alleles are class I MHC alleles **Claim 14** The method according to any one of claims 1 to 13, wherein the plurality of parameters of the machine - learned presentation model comprise one or more hotspot parameters **Claim 15** The method according to any one of claims 1 to 14, wherein the training dataset further comprises an association between training peptide sequences and one or more hotspot characteristics

Citation Information

Patent Citations

  • Mathematical processes for determination of peptidase cleavage

    US20160117441A1

  • Compositions And Methods For Viral Cancer Neoepitopes

    US20170028044A1

  • Neoantigen Identification, Manufacture, and Use

    US20170212984A1