Identification, production, and use of neoantigens
By employing next-generation sequencing and machine-learned models to identify tumor-specific peptides, the method enhances the accuracy of neoantigen selection for personalized cancer vaccines, addressing the low PPV issue and improving therapeutic efficacy.
Patent Information
- Application Number
- JP2019567557
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-06-09
- Filing Date
- 2018-06-08
- Publication Date
- 2026-01-07
- Estimated Expiration
- 2038-06-08
AI Technical Summary
Current methods for identifying neoantigens for personalized cancer vaccines suffer from low positive predictive value (PPV), failing to accurately predict which tumor-specific peptides will be presented on the cell surface, leading to ineffective vaccination strategies.
An optimized approach using next-generation sequencing and machine-learned presentation models to identify neoantigens, considering various genomic alterations and MHC binding, coupled with deep learning to enhance prediction accuracy.
Improves the positive predictive value of neoantigen selection, ensuring that vaccines target effective tumor-specific peptides, thereby increasing the likelihood of eliciting anti-tumor immunity.
Smart Images

Figure 0007795291000085 
Figure 0007795291000086 
Figure 0007795291000087
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 517,786, filed June 9, 2017, which is incorporated herein by reference in its entirety. [Background technology]
[0002] background Therapeutic vaccines based on tumor-specific neoantigens hold great promise as the next generation of personalized cancer immunotherapy. 1~3 Cancers with high mutational burden, such as non-small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies due to their relatively high likelihood of generating neoantigens. 4,5 Early evidence suggests that neoantigen-based vaccination induces T cell responses. 6 , Cell therapy targeting neoantigens can induce tumor regression in selected patients 7 Both MHC class I and MHC class II influence T cell responses. 70~71 .
[0003] One question regarding neoantigen vaccine design is which of the many coding mutations present in the target tumor can give rise to the "best" therapeutic neoantigen (e.g., an antigen capable of eliciting antitumor immunity and causing tumor regression).
[0004] Early methods have been proposed that incorporate mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides. 8However, these proposed methods involve many steps other than gene expression and MHC binding (e.g., TAP transport, proteasomal cleavage, MHC binding, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-I; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsins), competition with CLIP peptides for HLA binding catalyzed by HLA-DM, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-II). 9 The entire epitope generation process cannot be modeled, and therefore existing methods tend to suffer from low positive predictive value (PPV) (Figure 1A).
[0005] Indeed, analyses of peptides presented by tumor cells conducted by multiple groups have shown that less than 5% of peptides predicted to be presented using gene expression and MHC binding affinity are found on tumor surface MHC. 10,11 (Figure 1B). This low correlation between binding prediction and MHC presentation is further supported by the lack of improved prediction accuracy for binding-restricted neoantigens in response to checkpoint inhibitors relative to the number of mutations alone. 12 .
[0006] Such low positive predictive values (PPV) of existing methods for predicting presentation present a problem in the design of neoantigen-based vaccines. If vaccines are designed using low PPV predictions, most patients are unlikely to receive therapeutic neoantigens, and even fewer will receive multiple neoantigens (even assuming all presented peptides are immunogenic). Thus, neoantigen vaccination using current methods is unlikely to be successful in a significant number of tumor-bearing subjects (Figure 1C).
[0007] Furthermore, previous approaches have only used cis-acting mutations to generate candidate neoantigens, including splicing factor mutations that occur in multiple tumor types and lead to aberrant splicing of many genes. 13 , and additional sources of nascent ORFs, including mutations that create or remove protease cleavage sites, were not considered in most cases.
[0008] Finally, standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard tumor analysis approaches may erroneously promote sequence artifacts or germline polymorphisms as neoantigens, which can lead to inefficient utilization of vaccine doses or the risk of autoimmunity, respectively. Summary of the Invention
[0009] overview Optimized approaches for identifying and selecting neoantigens for personalized cancer vaccines are disclosed herein. First, we address optimized tumor exome and transcriptome analysis approaches to identify neoantigen candidates using next-generation sequencing (NGS). These methods build on standard approaches for tumor analysis by NGS so that the most sensitive and specific neoantigen candidates are developed across all classes of genomic alterations. Second, novel approaches for high-PPV neoantigen selection are provided to overcome specificity issues and ensure that neoantigens developed for vaccine administration are more likely to elicit anti-tumor immunity. Depending on the embodiment, these approaches include trained statistical regression or nonlinear deep learning models that jointly model peptide-allele mapping and motifs for each allele of peptides of multiple lengths that share statistical power across peptides of different lengths. In particular, nonlinear deep learning models can be designed and trained to treat different MHC alleles within the same cell as independent, thereby overcoming the problem associated with linear models where linear models interfere with each other. Finally, a further concern regarding the design and production of personalized neoantigen-based vaccines is addressed.
[0010] Also disclosed herein is a method for identifying a subset of patients suitable for treatment. Tumor nucleotide sequencing data for at least one of the exome, transcriptome, or whole genome is obtained from tumor cells and normal cells for each patient. Using the tumor nucleotide sequencing data, peptide sequences for each of a set of neoantigens identified by comparing nucleotide sequencing data from tumor cells with nucleotide sequencing data from normal cells are obtained. The peptide sequence for each neoantigen in the patient contains at least one change that differentiates it from the corresponding wild-type parent peptide sequence identified from the patient's normal cells. Each peptide sequence in the set of neoantigens is input into a machine-learned presentation model to generate a set of numerical presentation likelihoods for the set of neoantigens for each patient. Each presentation likelihood represents the likelihood that the corresponding neoantigen will be presented by one or more MHC alleles on the surface of the patient's tumor cells. The set of presentation likelihoods is determined based at least on mass spectrometry data. One or more neoantigens are identified from the set of neoantigens in the patient. A utility score is determined for each patient, indicating the estimated number of neoantigens displayed on the surface of the patient's tumor cells, as determined by the corresponding likelihood of display of one or more neoantigens for the patient. A subset of patients is selected for treatment. Each patient within this patient subset is associated with a utility score that meets predetermined inclusion criteria. The selected subset of patients can be treated with a treatment such as a neoantigen vaccine or checkpoint inhibitor therapy. [The present invention 1001] 1. A method for identifying a subset of patients suitable for treatment, comprising: obtaining, for each patient, at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from the patient's tumor cells and normal cells, wherein the tumor nucleotide sequencing data is used to obtain a peptide sequence for each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen for the patient includes at least one alteration that makes it different from a corresponding wild-type parent peptide sequence identified from the patient's normal cells; generating, for each patient, a set of numerical presentation likelihoods for the set of neoantigens for the patient by inputting the peptide sequence of each of the set of neoantigens into a machine-learned presentation model, wherein each presentation likelihood represents the likelihood that the corresponding neoantigen will be presented by one or more MHC alleles on the surface of tumor cells in the patient, and wherein the set of presentation likelihoods has been determined based at least on mass spectrometry data; identifying for each patient one or more neoantigens from said set of neoantigens for said patient; determining for each patient a utility score indicative of the estimated number of neoantigens presented on the surface of tumor cells of said patient, as determined by the corresponding presentation likelihoods for said one or more neoantigens for said patient; selecting a subset of patients suitable for treatment, wherein each patient in said subset of patients is associated with a utility score that meets predetermined inclusion criteria; The method comprising: [The present invention 1002] 1002. The method of claim 1001, wherein identifying said one or more neoantigens for said patient comprises selecting a subset of neoantigens in said set of neoantigens for said patient. [The present invention 1003] 1003. The method of claim 1002, wherein said subset of neoantigens are neoantigens having the highest presentation likelihood among said set of presentation likelihoods for said patient. [The present invention 1004] The method of claim 1001, further comprising treating each patient within said selected subset of patients with a corresponding neoantigen vaccine comprising at least one of said one or more neoantigens identified for said patient. [The present invention 1005] The method of claim 1001, further comprising identifying, for each patient within said selected subset of patients, one or more T cells or T cell receptors that are antigen-specific for at least one of said one or more neoantigens identified for said patient. [The present invention 1006] 1002. The method of claim 1001, wherein identifying one or more neoantigens for said patient comprises selecting the entire set of identified neoantigens for said patient. [The present invention 1007] The method of claim 1006, further comprising administering checkpoint inhibitor therapy to each patient within said selected subset of patients. [The present invention 1008] The method of the present invention 1001, wherein selecting a subset of patients suitable for treatment comprises selecting a subset of patients having a tumor mutational burden (TMB) higher than a minimum threshold, wherein a patient's TMB indicates the number of neoantigens in a set of neoantigens associated with that patient. [The present invention 1009] Selecting a subset of patients suitable for treatment is crucial. Selecting a subset of patients with a utility score above a minimum threshold The method of the present invention 1001, comprising: [The present invention 1010] 1002. The method of claim 1001, wherein said utility score is the sum of the likelihoods of presentation for each neoantigen within said identified subset of neoantigens for said patient. [The present invention 1011] 1002. The method of claim 1001, wherein said usefulness score is the probability that the number of presented neoantigens among said identified one or more neoantigens for said patient is above a minimum threshold. [The present invention 1012] The machine-learned presentation model is a label obtained by mass spectrometry to measure the presence of a peptide bound to at least one MHC allele identified as being present in at least one of the plurality of samples; a training peptide sequence containing information about a set of amino acids constituting the training peptide sequence and the positions of the amino acids within the training peptide sequence; at least one MHC allele associated with said training peptide sequence; a plurality of parameters determined based at least on the training data set, the plurality of parameters including: A function that represents the relationship between the peptide sequence and the likelihood of presentation based on the plurality of parameters. The method of the present invention 1001, comprising: [The present invention 1013] the training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of the isolated peptides; and (b) data relating to measurements of peptide-MHC binding stability for at least one of the isolated peptides; The method of the present invention 1012 further comprising at least one of: [The present invention 1014] The set of numerical likelihoods is: (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; The method of the present invention 1001 further characterized by the characteristics including at least one of: [The present invention 1015] 1001. The method of claim 1001, wherein said set of presentation likelihoods is further identified by at least the expression level of said one or more MHC alleles in said subject as measured by RNA-seq or mass spectrometry. [The present invention 1016] The set of presentation likelihoods is: (a) predicted affinity between neoantigens within said set of neoantigens and said one or more MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex; The method of the present invention 1001 further characterized by the characteristics including at least one of: [The present invention 1017] inputting the peptide sequence into the machine-learned presentation model; applying the machine-learned presentation model to the peptide sequence of each neoantigen to generate, for each of the one or more MHC alleles, a dependency score indicating whether the MHC allele presents the neoantigen based on a particular amino acid at a particular position in the peptide sequence. The method of the present invention 1001, comprising: [The present invention 1018] inputting the peptide sequence into the machine-learned presentation model; transforming the dependency scores to generate, for each MHC allele, a corresponding per-allele likelihood that indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining the per-allele likelihoods to generate a presentation likelihood for the neoantigen. The method of the present invention 1017, comprising: [The present invention 1019] 1019. The method of claim 1018, wherein transforming said dependency score models presentation of said neoantigen as mutually exclusive across said one or more class MHC alleles. [The present invention 1020] inputting the peptide sequence into the machine-learned presentation model; Transforming the combination of dependency scores to generate a presentation likelihood. 1017. The method of claim 1017, comprising: transforming said combination of dependency scores to model presentation of said neoantigen as interfering between said one or more MHC alleles. [Brief explanation of the drawings]
[0011] These and other features, aspects, and aspects of the present invention will become better understood with regard to the following description and accompanying drawings.
[0012] [Figure 1A] Current clinical approaches to neoantigen identification are presented. [Figure 1B] It shows that less than 5% of the predicted binding peptides are displayed on tumor cells. [Figure 1C] Illustrates the impact of specificity issues on neoantigen prediction. [Figure 1D] This shows that binding prediction is not sufficient to identify neoantigens. [Figure 1E] Probability of MHC-I presentation as a function of peptide length. [Figure 1F] An exemplary peptide spectrum generated from a Promega dynamic range standard is shown. [Figure 1G] We show how adding features increases the positive predictive value of the model. [Figure 2A] 1 is a schematic of an environment for identifying the likelihood of peptide presentation in a patient, according to one embodiment. [Figure 2B] A method for obtaining presentation information according to one embodiment is described. Figure 2B discloses SEQ ID NO:20. [Figure 2C] A method for obtaining presentation information according to one embodiment is illustrated in Figure 2C, which discloses SEQ ID NOs: 3 to 8, respectively, in order of appearance. [Figure 3] FIG. 1 is a high-level block diagram illustrating computer logic components of a presentation identification system, according to one embodiment. [Figure 4] Illustrating an exemplary set of training data, according to one embodiment, Figure 4 discloses the peptide sequences as SEQ ID NOS: 10-13 and the C-terminal flanking sequences as SEQ ID NOS: 15, 21-22, and 22, respectively, in order of appearance. [Figure 5] 1 illustrates an exemplary network model related to MHC alleles. [Figure 6A] An exemplary network model NNH(·) shared by MHC alleles according to one embodiment [Figure 6B] 1. An exemplary network model NNH(·) shared by MHC alleles according to another embodiment [Figure 7] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 8] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 9] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 10]1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 11] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 12] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 13A] 1 shows the sample frequency distribution of tumor mutation burden in NSCLC patients. [Figure 13B] 1 shows the number of neoantigens presented in a simulated vaccine for patients selected based on the inclusion criterion of whether the patient met a minimum tumor mutational burden, according to one embodiment. [Figure 13C] 10A-10C show a comparison of the number of neoantigens presented in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on a presented model and selected patients associated with vaccines containing therapeutic subsets identified by a prior art model, according to one embodiment. [Figure 13D] 10A-B compares the number of neoantigens presented in simulated vaccines between selected patients associated with a vaccine containing therapeutic subsets identified based on a single allele-per-presentation model for HLA-A*02:01 and selected patients associated with a vaccine containing therapeutic subsets identified based on both an allele-per-presentation model for HLA-A*02:01 and HLA-B*07:02. According to one embodiment, the vaccine volume is set at v=20 epitopes. [Figure 13E] FIG. 10 compares the number of neoantigens presented in simulated vaccines between patients selected based on tumor mutation burden and patients selected by expected utility score, according to one embodiment. [Figure 14] An exemplary computer for implementing the entities shown in FIGS. 1 and 3 will now be described. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description I. Definition In general, terms used in the claims and the specification shall be interpreted as having their ordinary meaning as understood by one of ordinary skill in the art. Certain terms are defined below to provide further clarity. If there is a conflict between the ordinary meaning and a given definition, the given definition shall control.
[0014] As used herein, the term "antigen" refers to a substance that induces an immune response.
[0015] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes it different from its corresponding wild-type parent antigen, for example, due to a tumor cell mutation or tumor cell-specific post-translational modification. Neoantigens may include polypeptide or nucleotide sequences. Mutations can include frameshift or non-frameshift insertion / deletions (indels), missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression change that results in a new ORF. Mutations can also include splice variants. Tumor cell-specific post-translational modifications can include aberrant phosphorylation. Tumor cell-specific post-translational modifications can also include splice antigens generated by the proteasome. See Liepe et al., "A large fraction of HLA class I ligands are proteasome-generated spliced peptides"; Science. 2016 Oct 21;354(6310):354-358.
[0016] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in tumor cells or tissues of a subject, but is not present in the corresponding normal cells or tissues of the subject.
[0017] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct that is based on one or more neoantigens, eg, multiple neoantigens.
[0018] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.
[0019] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.
[0020] As used herein, the term "coding mutation" refers to a mutation that occurs in a coding region.
[0021] As used herein, the term "ORF" means open reading frame.
[0022] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that arises due to mutation or other abnormalities such as splicing.
[0023] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.
[0024] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon.
[0025] As used herein, the term "frameshift mutation" is a mutation that causes an alteration in the frame of a protein.
[0026] As used herein, the term "indel" refers to the insertion or deletion of one or more nucleic acids.
[0027] As used herein, the term "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences in which a certain percentage of nucleotides or amino acid residues are the same when compared and aligned for maximum correspondence, as measured using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN, or other algorithms available to those of skill in the art), or by visual inspection. Depending on the application, the "percent identity" can exist over a region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.
[0028] In sequence comparison, generally, one sequence serves as a reference sequence to which test sequences are compared.When using a sequence comparison algorithm, test sequences and reference sequences are input into a computer, subsequence coordinates are designated if necessary, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the sequence identity (%) of the test sequence to the reference sequence based on the designated program parameters.Alternatively, sequence similarity or difference can also be established by the combination of the presence or absence of a specific nucleotide at a selected sequence position (e.g., sequence motif) or an amino acid in a translated sequence.
[0029] Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (see generally Ausubel et al., infra).
[0030] One example of an algorithm that is suitable for determining percent sequence identity and percent sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.
[0031] As used herein, the term "non-stop or read-through" refers to a mutation that results in the removal of the natural stop codon.
[0032] As used herein, the term "epitope" refers to a specific portion of an antigen that is typically bound by an antibody or T-cell receptor.
[0033] As used herein, the term "immunogenic" refers to the ability to elicit an immune response, for example, via T cells, B cells, or both.
[0034] As used herein, the terms "HLA binding affinity" and "MHC binding affinity" refer to the affinity of binding between a specific antigen and a specific MHC allele.
[0035] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.
[0036] As used herein, the term "mutation" is a difference between the nucleic acid of a subject and a reference human genome used as a control.
[0037] As used herein, the term "variant calling" is the algorithmic determination, typically from sequencing, of the presence of a mutation.
[0038] As used herein, the term "polymorphism" refers to a germline mutation, ie, a mutation found in all DNA-bearing cells of an individual.
[0039] As used herein, the term "somatic mutation" is a mutation that occurs in a non-germline cell of an individual.
[0040] As used herein, the term "allele" refers to one version of a gene or one version of a gene sequence or one version of a protein.
[0041] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.
[0042] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the degradation of mRNA by the cell due to a premature stop codon.
[0043] As used herein, the term "truncal mutation" is a mutation that occurs early in the development of a tumor and is present in the majority of the cells of the tumor.
[0044] As used herein, the term "subclonal mutation" is a mutation that occurs late in the development of a tumor and is present in only a portion of the cells of the tumor.
[0045] As used herein, the term "exome" refers to the subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.
[0046] As used herein, the term "logistic regression" is a regression model for binary data from statistics in which the logit of the probability that the dependent variable is equal to 1 is modeled as a linear function of the dependent variable.
[0047] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinear transformations typically trained by stochastic gradient descent and backpropagation.
[0048] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.
[0049] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome can also refer to the properties of a cell or a collection of cells (e.g., a tumor peptidome refers to the union of the peptidomes of all cells that comprise a tumor).
[0050] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.
[0051] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.
[0052] As used herein, the term "tolerance or immune tolerance" refers to a state of immune unresponsiveness to one or more antigens, eg, self-antigens.
[0053] As used herein, the term "central tolerance" is tolerance conferred in the thymus by either deleting autoreactive T cell clones or promoting their differentiation into immunosuppressive regulatory T cells (Tregs).
[0054] As used herein, the term "peripheral tolerance" refers to tolerance conferred in the peripheral system by downregulating or anergizing autoreactive T cells that survive central tolerance or by promoting the differentiation of these T cells into Tregs.
[0055] The term "sample" can include a single cell, or multiple cells, or fragments of cells, or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.
[0056] The term "subject" includes cells, tissues, or organisms, human or non-human, whether male or female, in vivo, ex vivo, or in vitro. The term subject includes mammals, including humans.
[0057] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.
[0058] The term "clinical factor" refers to a measurement of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values that can be obtained from assessing a subject or a sample (or a population of samples) from a subject under a given condition. A clinical factor can also be predicted by other parameters, such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.
[0059] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.
[0060] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0061] Terms not directly defined herein should be understood to have the meanings generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide further guidance to the practitioner in describing the compositions, devices, methods, etc. of embodiments of the present invention, as well as how to make or use them. It will be recognized that multiple ways of saying the same thing may be used. Accordingly, alternative terms and synonyms may be used for any one or more of the terms discussed herein. No weight should be placed on whether a term is detailed or discussed herein. Several synonyms or alternative methods, materials, etc. are provided. The recitation of one or more synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the inventive embodiments herein.
[0062] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.
[0063] II. Methods for identifying neoantigens Disclosed herein are methods for identifying neoantigens from a subject's tumor that are likely to be presented on the tumor cell surface and / or likely to be immunogenic. For example, one such method includes: acquiring at least one of tumor exome, transcriptome, or whole genome nucleotide sequencing data from tumor cells of the subject, wherein the tumor nucleotide sequencing data is used to obtain data representing each peptide sequence of a set of neoantigens, where each neoantigen peptide sequence contains at least one alteration that makes the peptide sequence different from a corresponding wild-type parent peptide sequence; inputting the peptide sequence of each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each neoantigen will be presented by one or more MHC alleles on the tumor cell surface of the subject's tumor cells or by cells present within the tumor, where the set of numerical likelihoods has been identified based at least on received mass spectrometry data; and selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens.
[0064] The proposed model can include a statistical regression or machine learning (e.g., deep learning) model trained with a set of reference data (also referred to as a training dataset) including a corresponding set of labels, the set of reference data being obtained from each of a plurality of separate subjects, some of whom may have tumors, and the set of reference data including at least one of data representing exome nucleotide sequences from tumor tissue, data representing exome nucleotide sequences from normal tissue, data representing transcriptome nucleotide sequences from tumor tissue, data representing proteome sequences from tumor tissue, and data representing MHC peptidome sequences from tumor tissue, and data representing MHC peptidome sequences from normal tissue. The reference data can further include mass spectrometry data, sequencing data, RNA sequencing data, and proteomics data of synthetic proteins, normal and tumor human cell lines, and single-allelic cell lines engineered to express a predetermined MHC allele that are subsequently exposed to fresh and frozen primary samples, as well as T cell assays (e.g., ELISPOT). In certain embodiments, the set of reference data includes each form of reference data.
[0065] The proposed model can include a set of features derived at least in part from a set of reference data, the set of features including at least one of an allele-dependent feature and an allele-independent feature. In certain embodiments, each feature is included.
[0066] Characteristics of dendritic cell presentation to naive T cells can include at least one of the above characteristics: the dose and type of antigen in the vaccine (e.g., peptide, mRNA, virus, etc.): (1) the route by which dendritic cells (DCs) take up the antigen type (e.g., endocytosis, micropinocytosis); and / or (2) the efficiency with which the antigen is taken up by DCs; the dose and type of adjuvant in the vaccine; the length of the vaccine antigen sequence; the number and site of vaccine administration; and the patient's baseline immune function (e.g., as measured by a history of recent infections, blood counts, etc.). For RNA vaccines, the following characteristics may be considered: (1) the turnover rate of mRNA-protein products within dendritic cells; (2) the rate of translation of mRNA after uptake by dendritic cells, as measured by in vitro or in vivo experiments; and / or (3) the number or rounds of translation of mRNA after uptake by dendritic cells, as measured by in vivo or in vitro experiments. The presence of protease cleavage motifs in the peptide, optionally giving additional weight to proteases typically expressed in dendritic cells (e.g., as measured by RNA-seq or mass spectrometry). The level of proteasome and immunoproteasome expression in typical activated dendritic cells (which can be measured by RNA-seq, mass spectrometry, immunohistochemistry, or other standard techniques). The expression level of a particular MHC allele in the individual of interest (e.g., as measured by RNA-seq or mass spectrometry), optionally specifically measured in activated dendritic cells or other immune cells. The probability of peptide presentation by a particular MHC allele in other individuals that express the particular MHC allele, optionally specifically measured in activated dendritic cells or other immune cells. The probability of peptide presentation by MHC alleles of the same molecular family (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals, optionally specifically measured in activated dendritic cells or other immune cells.
[0067] The immune tolerance escape trait can include at least one of the following: direct measurement of the self-peptidome by protein mass spectrometry performed on one or several cell types; estimation of the self-peptidome by taking the union of all k-mer (e.g., 5-25) substrings of the self-protein; estimation of the self-peptidome using a model of presentation similar to the presentation model described above applied to all non-mutated self-proteins, optionally accounting for germline mutations.
[0068] Ranking can be performed using a plurality of neoantigens provided by at least one model based at least in part on numerical likelihood. After ranking, selection can be performed to select a subset of the ranked neoantigens according to a selection criterion. After selection, the subset of ranked peptides can be provided as an output.
[0069] The set of selected neoantigens may be 20 in number.
[0070] The presentation model can represent the dependency between the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in a peptide sequence and the likelihood of presentation of such a peptide sequence containing a particular amino acid at a particular position on the surface of a tumor cell by a particular one of the paired MHC alleles.
[0071] The methods disclosed herein may also include applying one or more presentation models to the peptide sequences of the corresponding neoantigens to generate, for each of one or more MHC alleles, a dependency score indicating whether the MHC allele presents the corresponding neoantigen based on at least the position of an amino acid in the peptide sequence of the corresponding neoantigen.
[0072] The methods disclosed herein may also include transforming the dependency scores to generate corresponding per-allele likelihoods for each MHC allele, which indicate the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining the per-allele likelihoods to generate a numerical likelihood.
[0073] The process of transforming the dependency scores can model the presentation of peptide sequences of corresponding neoantigens as mutually exclusive.
[0074] The methods disclosed herein may also further include transforming the combination of dependency scores to generate a numerical likelihood.
[0075] Transforming the combination of dependency scores can model the presentation of the corresponding neoantigen peptide sequences as interference between MHC alleles.
[0076] The set of numerical likelihoods can be further identified by at least the allele non-interaction feature, and the methods disclosed herein can also include applying an allele non-interaction model of the one or more presentation models to the allele non-interaction feature to generate a dependency score for the allele non-interaction feature that indicates whether the peptide sequence of the corresponding neoantigen is presented based on the allele non-interaction feature.
[0077] The methods disclosed herein may also include combining the dependency score for each MHC allele with the dependency score for the allele-non-interacting property in one or more MHC alleles; transforming the combined dependency score for each MHC allele to generate a corresponding per-allele likelihood for the MHC allele, which indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining the per-allele likelihoods to generate a numerical likelihood.
[0078] The methods disclosed herein may also include generating a numerical likelihood by transforming a combination of the dependency scores for each of the MHC alleles and the dependency scores for the allele-non-interacting trait.
[0079] The set of numerical parameters for the displayed model can be trained based on a training dataset that includes at least a set of training peptide sequences identified as present in a plurality of samples and one or more MHC alleles associated with each training peptide sequence, where the training peptide sequences are identified by mass spectrometry of isolated peptides eluted from MHC alleles from the plurality of samples.
[0080] The sample may also include a cell line engineered to express a single MHC class I or class II allele.
[0081] The sample may also include cell lines engineered to express multiple MHC class I or class II alleles.
[0082] The sample may also include human cell lines obtained or derived from multiple patients.
[0083] Samples may also include fresh or frozen tumor samples obtained from multiple patients.
[0084] Samples may also include fresh or frozen tissue samples obtained from multiple patients.
[0085] The sample may also contain peptides identified using a T cell assay.
[0086] The training dataset may further include data relating to the peptide abundance of the set of training peptides present in the sample; the peptide length of the set of training peptides in the sample.
[0087] The training dataset can be generated by comparing a set of training peptide sequences by alignment with a database containing a set of known protein sequences, where the set of training protein sequences is longer than and includes the training peptide sequences.
[0088] The training dataset may be generated by performing nucleotide sequencing on the cell line or by previously performing nucleotide sequencing to obtain at least one of exome, transcriptome, or whole genome sequencing data from the cell line, wherein the sequencing data includes at least one nucleotide sequence that includes the variation.
[0089] The training dataset may be generated based on obtaining at least one of normal nucleotide sequencing data of the exome, transcriptome, or whole genome from a normal tissue sample.
[0090] The training dataset may further include data relating to proteome sequences associated with the sample.
[0091] The training dataset may further include data relating to MHC peptidome sequences associated with the sample.
[0092] The training dataset may further include data relating to measurements of peptide-MHC binding affinity for at least one of the isolated peptides.
[0093] The training dataset may further include data relating to a measure of peptide-MHC binding stability for at least one of the isolated peptides.
[0094] The training dataset may further include data relating to the transcriptome associated with the sample.
[0095] The training dataset may further include data related to the genome associated with the sample.
[0096] Training peptide sequences can range in length from k-mers (k is between 8 and 15 for MHC class I, or between 6 and 30 for MHC class II).
[0097] The methods disclosed herein may also include encoding the peptide sequence using a one-hot encoding scheme.
[0098] The methods disclosed herein may also include encoding the training peptide sequences using a left-padded one-hot encoding scheme.
[0099] A method for treating a subject with a tumor, comprising performing the steps described in claim 1, and further comprising obtaining a tumor vaccine comprising a set of selected neoantigens, and administering the tumor vaccine to the subject.
[0100] Also disclosed herein is a method for producing a tumor vaccine, the method comprising: obtaining at least one of tumor nucleotide sequencing data of an exome, a transcriptome, or a whole genome from tumor cells of a subject; using the tumor nucleotide sequencing data to obtain data representing a peptide sequence for each of a set of neoantigens, wherein the peptide sequence for each neoantigen comprises at least one mutation that makes the peptide sequence different from a corresponding wild-type parent peptide sequence; inputting the peptide sequence for each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each of the neoantigens will be presented by one or more MHC alleles on the tumor cell surface of the tumor cells of the subject, the set of numerical likelihoods being identified based at least on received mass spectrometry data; selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens; and producing or having previously produced a tumor vaccine comprising the selected set of neoantigens.
[0101] Also provided herein is a tumor vaccine comprising a set of selected neoantigens selected by performing a method comprising: obtaining at least one of tumor nucleotide sequencing data of an exome, transcriptome, or whole genome from tumor cells of a subject, wherein the tumor nucleotide sequencing data is used to obtain data representing a peptide sequence for each of a set of neoantigens, wherein the peptide sequence for each neoantigen comprises at least one mutation that makes the peptide sequence different from a corresponding wild-type parent peptide sequence; inputting the peptide sequence for each neoantigen into one or more presentation models to generate a set of numerical likelihoods that each of the neoantigens will be presented by one or more MHC alleles on the tumor cell surface of the tumor cells of the subject, wherein the set of numerical likelihoods is determined based at least on received mass spectrometry data; selecting a subset of the set of neoantigens based on the set of numerical likelihoods to generate a set of selected neoantigens; and producing or having previously produced a tumor vaccine comprising the set of selected neoantigens.
[0102] The tumor vaccine may comprise one or more of a nucleotide sequence, a polypeptide sequence, RNA, DNA, a cell, a plasmid, or a vector.
[0103] A tumor vaccine may comprise one or more neoantigens displayed on the surface of tumor cells.
[0104] A tumor vaccine may comprise one or more neoantigens that are immunogenic in a subject.
[0105] A tumor vaccine may not include one or more neoantigens that induce an autoimmune response against normal tissue in a subject.
[0106] The tumor vaccine may include an adjuvant.
[0107] The tumor vaccine may include an excipient.
[0108] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of being presented on the tumor cell surface relative to neoantigens that are not selected based on the presentation model.
[0109] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of inducing a tumor-specific immune response in a subject against neoantigens that are not selected based on the presentation model.
[0110] The methods disclosed herein may also include selecting neoantigens that have an increased likelihood of being presented to naive T cells by professional antigen-presenting cells (APCs) relative to neoantigens not selected based on a presentation model, optionally where the APCs are dendritic cells (DCs).
[0111] The methods disclosed herein may also include selecting neoantigens that have a reduced likelihood of being inhibited by central or peripheral tolerance to neoantigens that are not selected based on the presentation model.
[0112] The methods disclosed herein may also include selecting neoantigens that have a reduced likelihood of inducing an autoimmune response against normal tissue in a subject relative to neoantigens that are not selected based on the presentation model.
[0113] Nucleotide sequencing data of the exome or transcriptome can be obtained by performing sequencing on tumor tissue.
[0114] Sequencing may be next generation sequencing (NGS) or any massively parallel sequencing approach.
[0115] The set of numerical likelihoods can be further specified by at least MHC allele interaction properties, including at least one of the following: predicted affinity of binding between the MHC allele and the neoantigen-encoded peptide; predicted stability of the neoantigen-encoded peptide-MHC complex; sequence and length of the neoantigen-encoded peptide; probability of presentation of a neoantigen-encoded peptide with a similar sequence in cells from other individuals expressing the particular MHC allele, as assessed by mass spectrometry proteomics or other means; expression level of the particular MHC allele in the subject of interest (e.g., as measured by RNA-seq or mass spectrometry); probability of presentation by a particular MHC allele in other distinct individuals expressing the particular MHC allele, independent of the overall neoantigen-encoded peptide sequence; probability of presentation by MHC alleles of the same molecular family (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other distinct subjects, independent of the overall neoantigen-encoded peptide sequence.
[0116] The set of numerical likelihoods is further specified by at least MHC allele non-interacting properties, including at least one of the following: C-terminal and N-terminal sequences flanking the neoantigen-encoded peptide within its source protein sequence; the presence of protease cleavage motifs within the neoantigen-encoded peptide, optionally weighted according to the expression of the corresponding protease in tumor cells (as measured by RNA-seq or mass spectrometry); the turnover rate of the source protein measured in the appropriate cell type; measured by RNA-seq or proteomic mass spectrometry or predicted from the annotation of germline or somatic splicing variants detected in DNA or RNA sequence data. the length of the source protein, possibly taking into account the particular splice variants ("isoforms") that are most highly expressed in tumor cells; the level of expression of the proteasome, immunoproteasome, thymoproteasome, or other proteases in the tumor cells (which can be measured by RNA-seq, proteomic mass spectrometry, or immunohistochemistry); the expression of the source gene of the neoantigen-encoding peptide (e.g., measured by RNA-seq or mass spectrometry); the typical tissue-specific expression of the source gene of the neoantigen-encoding peptide in different phases of the cell cycle; e.g., as determined by uniProt or PDB a comprehensive catalog of the properties of the source protein and / or its domains, such as can be found at http: / / www.rcsb.org / pdb / home / home.do; properties describing the nature of the domain of the source protein containing the peptide, e.g., secondary or tertiary structure (e.g., α-helix versus β-sheet); alternative splicing; the probability of presentation of a peptide derived from the source protein of the neoantigen-encoded peptide of interest in other distinct subjects; the probability that the peptide will be undetected or over-represented by mass spectrometry due to technical bias; the expression of various gene modules / pathways (not necessarily including the source protein of the peptide) measured by RNASeq, which informs about the status of tumor cells, stroma, or tumor-infiltrating lymphocytes (TILs); the copy number of the source gene of the neoantigen-encoded peptide in tumor cells;The probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP; the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, immunohistochemistry); the presence or absence of tumor mutations, including, but not limited to, driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, NTRK3, and genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA- A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome. Peptides whose presentation depends on components of the antigen presentation machinery that carry loss-of-function mutations in the tumor have a low probability of presentation; including, but not limited to, the presence or absence of functional germline polymorphisms in genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-D ... polymorphisms in HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome; tumor type (e.g., NSCLC, melanoma); clinical tumor subtype (e.g., squamous cell lung cancer vs. non-squamous); smoking history;Typical expression of the peptide's source gene in relevant tumor types or clinical subtypes, possibly stratified by driver mutations;
[0117] The at least one mutation may be a frameshift or non-frameshift indel, a missense or nonsense substitution, a splice site alteration, a genomic rearrangement or gene fusion, or any genomic or expression alteration that results in a de novo ORF.
[0118] The tumor cells can be selected from the group consisting of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0119] The methods disclosed herein may also include obtaining a tumor vaccine comprising the selected set of neoantigens or a subset thereof, and optionally further comprising administering the tumor vaccine to a subject.
[0120] At least one of the neoantigens in the set of selected neoantigens, when in polypeptide form, can comprise at least one of the following: a binding affinity to MHC with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I polypeptides, or a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II polypeptides; the presence of a sequence motif within or near the polypeptide in the parent protein sequence that promotes proteasomal cleavage; and the presence of a sequence motif that promotes TAP transport. In MHC class II, the presence of sequence motifs within or near the peptide that promote HLA binding catalyzed by cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA-DM.
[0121] Also disclosed herein is a method for generating a model for identifying one or more neoantigens likely to be presented on the tumor cell surface of tumor cells, the method comprising the steps of receiving mass spectrometry data including data relating to a plurality of isolated peptides eluted from major histocompatibility complexes (MHC) from a plurality of samples; obtaining a training dataset by at least identifying a set of training peptide sequences present in the samples and one or more MHC alleles associated with each training peptide sequence; and training a set of numerical parameters of a presentation model using the training dataset including the training peptide sequences, wherein the presentation model provides a plurality of numerical likelihoods that a peptide sequence derived from a tumor cell will be presented by one or more MHC alleles on the tumor cell surface.
[0122] The presentation model can represent the dependency between the presence of a particular amino acid at a particular position in a peptide sequence and the likelihood of presentation of a peptide sequence having a particular amino acid at a particular position by one of the MHC alleles on a tumor cell.
[0123] The sample may also include a cell line engineered to express a single MHC class I or class II allele.
[0124] The sample may also include cell lines engineered to express multiple MHC class I or class II alleles.
[0125] The sample may also include human cell lines obtained or derived from multiple patients.
[0126] Samples may also include fresh or frozen tumor samples obtained from multiple patients.
[0127] The sample may also contain peptides identified using a T cell assay.
[0128] The training dataset may further include data relating to the peptide abundance of the set of training peptides present in the sample; the peptide length of the set of training peptides in the sample.
[0129] The methods disclosed herein may also include obtaining, based on the training peptide sequences, a set of training protein sequences that are longer than and include the training peptide sequences by comparing the set of training peptide sequences by alignment with a database comprising a set of known protein sequences.
[0130] The methods disclosed herein may also include performing or having previously performed mass spectrometry on the cell line to obtain at least one of exome, transcriptome, or whole genome nucleotide sequencing data from the cell line, wherein the nucleotide sequencing data comprises at least one protein sequence that includes the mutation.
[0131] The methods disclosed herein may also include encoding the training peptide sequences using a one-hot encoding scheme.
[0132] The methods disclosed herein can also include obtaining at least one of normal nucleotide sequencing data of the exome, transcriptome, and whole genome from a normal tissue sample, and training a set of parameters of the proposed model using the normal nucleotide sequencing data.
[0133] The training dataset may further include data relating to proteome sequences associated with the sample.
[0134] The training dataset may further include data relating to MHC peptidome sequences associated with the sample.
[0135] The training dataset may further include data relating to measurements of peptide-MHC binding affinity for at least one of the isolated peptides.
[0136] The training dataset may further include data relating to a measure of peptide-MHC binding stability for at least one of the isolated peptides.
[0137] The training dataset may further include data relating to the transcriptome associated with the sample.
[0138] The training dataset may further include data related to the genome associated with the sample.
[0139] The methods disclosed herein may also include performing a logistic regression of the set of parameters.
[0140] Training peptide sequences can range in length from k-mers (k is between 8 and 15 for MHC class I, or between 6 and 30 for MHC class II).
[0141] The methods disclosed herein may also include encoding the training peptide sequences using a left-padded one-hot encoding scheme.
[0142] The methods disclosed herein may also include determining values for the set of parameters using a deep learning algorithm.
[0143] Disclosed herein is a method for identifying one or more neoantigens likely to be presented on the tumor cell surface of tumor cells, the method comprising the steps of: receiving mass spectrometry data including data relating to a plurality of isolated peptides eluted from major histocompatibility complex (MHC) antigens from a plurality of fresh or frozen samples; obtaining a training dataset by identifying at least a set of training peptide sequences present in the tumor samples and presented on one or more MHC alleles associated with each training peptide sequence; obtaining a set of training protein sequences based on the training peptide sequences; and training a set of numerical parameters of a presentation model using the training protein sequences and the training peptide sequences, wherein the presentation model provides a plurality of numerical likelihoods that a peptide sequence from the tumor cell will be presented by one or more MHC alleles on the tumor cell surface.
[0144] The presentation model can represent the dependency between the presence of a pair of a particular one of the MHC alleles with a particular amino acid at a particular position in a peptide sequence, and the likelihood of presentation of such a peptide sequence comprising said particular amino acid at said particular position on the surface of a tumor cell by said particular one of the MHC alleles of said pair.
[0145] The methods disclosed herein may also include selecting a subset of neoantigens, each of which has an increased likelihood of being presented on the cell surface of a tumor relative to one or more distinct tumor neoantigens.
[0146] The methods disclosed herein may also include selecting a subset of neoantigens, each of which has an increased likelihood of being able to induce a tumor-specific immune response in a subject against one or more distinct tumor neoantigens.
[0147] The methods disclosed herein may also include selecting a subset of neoantigens, each selected for having an increased likelihood relative to one or more distinct tumor neoantigens that they can be presented to naive T cells by professional antigen-presenting cells (APCs), optionally the APCs being dendritic cells (DCs).
[0148] The methods disclosed herein may also include selecting a subset of neoantigens, each of which has a reduced likelihood of being inhibited by central or peripheral tolerance to one or more distinct tumor neoantigens.
[0149] The methods disclosed herein may also include selecting a subset of neoantigens, each of which has a reduced likelihood of inducing an autoimmune response against normal tissue in a subject against one or more distinct tumor neoantigens.
[0150] The methods disclosed herein may also include selecting a subset of neoantigens, each of which has a reduced likelihood of being differentially post-translationally modified in tumor cells relative to APCs, and optionally, the APCs are dendritic cells (DCs).
[0151] The practice of the methods herein will employ, unless otherwise indicated, conventional methods of protein chemistry, biochemistry, recombinant DNA technology, and pharmacology, within the skill of the art. Such techniques are fully explained in the literature. See, e.g., T.E. Creighton, Proteins: Structures and Molecular Properties (W.H. Freeman and Company, 1993); A.L. Lehninger, Biochemistry (Worth Publishers, Inc., current addition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (2nd Edition, 1989); Methods In Enzymology (S. Colowick and N. Kaplan eds., Academic Press, Inc.); Remington's Pharmaceutical Sciences, 18th Edition (Easton, Pennsylvania: Mack Publishing Company, 1990); Carey and Sundberg Advanced Organic Chemistry, 3rd Ed. (Plenum Press), Vols. A and B (1992).
[0152] A set of presentation likelihoods can also be generated based on the source genes of the set of neoantigens.
[0153] A set of presentation likelihoods can also be generated based on the source genes and source tissue types of the set of neoantigens.
[0154] The methods disclosed herein may include identifying a subset of patients suitable for treatment with a neoantigen vaccine, comprising the steps of: obtaining, for each patient, at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from the patient's tumor cells, wherein the tumor nucleotide sequencing data is used to obtain peptide sequences for each of a set of neoantigens, wherein the peptide sequence of each neoantigen includes at least one alteration that makes it different from a corresponding wild-type parent peptide sequence; generating, for each patient, a set of numerical presentation likelihoods for the set of neoantigens for the patient by inputting the peptide sequences of each of the set of neoantigens into one or more presentation models, wherein the set of presentation likelihoods represent the likelihood that each of the set of neoantigens will be presented by one or more MHC alleles on the surface of tumor cells of the patient, the set of presentation likelihoods being determined based at least on the received mass spectrometry data; identifying for each patient a therapeutic subset of neoantigens from the patient's set of neoantigens, the therapeutic subset corresponding to a predetermined number of neoantigens with the highest presentation likelihood within the set of presentation likelihoods generated for that patient; Selecting a subset of patients suitable for treatment with a neoantigen vaccine, wherein the selected subset of patients meets inclusion criteria based on a set of neoantigens obtained for each patient in the selected subset or based on tumor nucleotide sequencing data.
[0155] The methods disclosed herein may include treating each patient in a selected subset of patients with a corresponding neoantigen vaccine, wherein the neoantigen vaccine for the patient comprises a therapeutic subset identified by a set of presentation likelihoods for the patient.
[0156] The methods disclosed herein may include selecting a subset of patients having a tumor mutational burden (TMB) above a minimum threshold, where a patient's TMB indicates the number of neoantigens in a set of neoantigens associated with that patient.
[0157] The methods disclosed herein may include identifying for each patient a utility score indicating a measure of the estimated number of neoantigens presented from a treatment subset of patients; and selecting a subset of patients having a utility score higher than a minimum threshold.
[0158] Presentation of neoantigens can be modeled as a Bernoulli random variable, and the utility score can represent the expected number of presented neoantigens in a therapeutic subset for a patient, and the utility score can be given by the sum of the likelihood of presentation for each neoantigen in the therapeutic subset of the patient.
[0159] Neoantigen presentation can also be modeled as a Poisson binomial random variable, and the utility score can be the probability that the number of presented neoantigens in the treatment subset for a patient is above a minimum threshold.
[0160] III. Identification of tumor-specific mutations in neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells). In particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject with cancer, but may not be present in normal tissues from the subject.
[0161] Genetic mutations in tumors can be considered useful for immunological targeting of tumors if they result in changes in the amino acid sequence of proteins exclusively in tumors. Useful mutations include: (1) non-synonymous mutations that result in different amino acids in proteins; (2) read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; (3) splice site mutations that result in the inclusion of an intron in mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) frameshift mutations or deletions that result in new open reading frames with new tumor-specific protein sequences. Mutations can also include one or more of non-frameshift insertions / deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in new ORFs.
[0162] For example, mutated peptides or mutated polypeptides resulting from splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.
[0163] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0164] Various methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. Several techniques have been described, including dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize amplification of target gene regions, typically by PCR. Still other methods rely on the generation of small signal molecules by invasive cleavage followed by mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.
[0165] PCR-based detection means can involve the multiplex amplification of multiple markers simultaneously.For example, it is well known in the art to select PCR primers so as to generate PCR products that do not overlap in size and can be analyzed simultaneously.Alternatively, it is possible to amplify different markers with primers that are differentially labeled and therefore can be differentially detected.Of course, hybridization-based detection means allows the differential detection of multiple PCR products in a sample.Other techniques that allow multiplex analysis of multiple markers are known in the art.
[0166] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA.For example, single nucleotide polymorphisms can be detected by using special exonuclease-resistant nucleotides, as disclosed in Mundy, CR (US Patent No. 4,656,127).According to this method, a primer complementary to the allele sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a specific animal or human.If the polymorphic site on the target molecule contains a nucleotide that is complementary to the specific exonuclease-resistant nucleotide derivative present, this derivative will be incorporated onto the end of the hybridized primer.This incorporation makes the primer resistant to exonucleases, thereby enabling its detection.Since the identity of the exonuclease-resistant derivative of the sample is known, the knowledge that the primer has become resistant to exonucleases reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of exogenous sequence data.
[0167] To determine the identity of the nucleotide at a polymorphic site, a solution-based method can be used (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087)). As in the method of Mundy, U.S. Pat. No. 4,656,127, a primer is used that is complementary to the allelic sequence immediately 3' to the polymorphic site. This method uses a labeled dideoxynucleotide derivative that becomes incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site to determine the identity of the nucleotide at that site. An alternative method, known as Genetic Bit Analysis or GBA, has been described by Goelet, P. et al. (PCT Application No. 92 / 15712). The Goelet, P. et al. method uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' to the polymorphic site. Goelet, P. et al. The method of Goelet, P. et al. uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which either the primer or the target molecule is immobilized on a solid phase.
[0168] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. et al., Nucl. Acids. Res. 17:7779-7784 (1989); Sokolov, B. P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M. et al., Proc. Natl. Acad. Sci. (USA) 88:1143-1147 (1991); Prezant, T. R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al. al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, signal is proportional to the number of incorporated deoxynucleotides, so that polymorphisms occurring in runs of the same nucleotide can result in a signal proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).
[0169] Numerous initiatives obtain sequence information directly from millions of individual molecules of DNA or RNA in parallel. Real-time single-molecule sequencing by synthesis techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent strands of DNA complementary to the template being sequenced. In one method, oligonucleotides 30–50 bases in length are covalently anchored at their 5′ ends to glass coverslips. These anchored strands serve two functions. First, they act as capture sites for the target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis for sequence reading. The capture primers serve as fixed-location sites for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and dye cleavage. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. As the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing-by-synthesis techniques also exist.
[0170] Any suitable sequencing-by-synthesis platform can be used to identify mutations. As mentioned above, four major sequencing-by-synthesis platforms are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing-by-synthesis platforms have also been described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the multiple nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. A capture sequence (also called a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to a support that can double as a universal primer.
[0171] As an alternative to capture sequences, a member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair, e.g., as described in U.S. Patent Application Publication No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.
[0172] Following capture, the sequence can be analyzed by single-molecule detection / sequencing, including, for example, template-dependent sequencing by synthesis, as described, for example, in the Examples and in U.S. Patent No. 7,283,337. In sequencing by synthesis, surface-bound molecules are exposed to a multitude of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time, in a step-and-repeat mode. For real-time analysis, a different optical label can be incorporated for each nucleotide, and multiple lasers can be utilized for stimulation of the incorporated nucleotides.
[0173] Sequencing can also include other massively parallel sequencing or next-generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, ThermoPGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel sequencing technologies, and future generations of these technologies, can be used.
[0174] Any cell type or tissue can be used to obtain nucleic acid samples for use in the methods described herein.For example, DNA or RNA samples can be obtained from tumor or body fluids, for example, blood obtained by known techniques (for example, venipuncture) or saliva.Alternatively, nucleic acid testing can be performed on dry samples (for example, hair or skin).In addition, a sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal tissue is of the same tissue type as tumor.A sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal sample is of a different tissue type from tumor.
[0175] The tumor may include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
[0176] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. Peptides can be acid-eluted from tumor cells or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.
[0177] IV. Neoantigens Neoantigens can comprise nucleotides or polynucleotides. For example, neoantigens can be RNA sequences that encode polypeptide sequences. Neoantigens useful in vaccines can therefore comprise nucleotide sequences or polypeptide sequences.
[0178] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the neoantigen comprises nucleotide sequences (e.g., DNA or RNA) that encode the associated polypeptide sequence.
[0179] The one or more polypeptides encoded by the neoantigen nucleotide sequences can comprise at least one of the following: a binding affinity to MHC with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides; the presence of a sequence motif within or near the peptide that promotes proteasomal cleavage; and the presence of a sequence motif within or near the peptide that promotes TAP transport; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II peptides; and the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.
[0180] One or more neoantigens can be present on the surface of a tumor.
[0181] The one or more neoantigens can be immunogenic in a tumor-bearing subject, for example, capable of eliciting a T cell or B cell response in the subject.
[0182] One or more neoantigens that induce an autoimmune response in a subject can be eliminated from consideration in the context of generating a vaccine for a tumor-bearing subject.
[0183] The size of the at least one neoantigenic peptide molecule is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35 , about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derivable therein. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.
[0184] Neoantigenic peptides and polypeptides can be 15 residues or less in length for MHC class I, usually between about 8 and about 11 residues, particularly 9 or 10 residues; for MHC class II, they can be 6 to 30 residues in length.
[0185] If desired, longer peptides can be designed in several ways. In one example, if the likelihood of peptide presentation on HLA alleles is predicted or known, the longer peptides can consist of either (1) individual presented peptides with extensions of 2-5 amino acids toward the N- and C-termini of each corresponding gene product; or (2) a concatenation of some or all of the presented peptides, each with its extended sequence. In another example, if sequencing reveals long (more than 10 residues) neo-epitope sequences present in the tumor (e.g., due to frameshifts, readthrough, or intron inclusion resulting in novel peptide sequences), the longer peptides would (3) consist of the entire novel tumor-specific stretch of amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides that are presented to the strongest HLA alleles. In either example, the use of longer peptides may allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.
[0186] Neoantigenic peptides and polypeptides can be presented on HLA proteins. In some embodiments, the neoantigenic peptides and polypeptides are presented on HLA proteins with greater affinity than wild-type peptides. In some embodiments, the neoantigenic peptide or polypeptide can have an IC50 of at least 5000 nM or less, at least 1000 nM or less, at least 500 nM or less, at least 250 nM or less, at least 200 nM or less, at least 150 nM or less, at least 100 nM or less, at least 50 nM or less, or even less.
[0187] In some embodiments, the neoantigenic peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.
[0188] Also provided are compositions comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides can be derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some embodiments, the tumor-specific mutations are driver mutations for a particular cancer type.
[0189] Neoantigenic peptides and polypeptides with desired activities or properties can be modified to confer certain desirable attributes, e.g., improved pharmacological characteristics, while enhancing or at least retaining substantially all of the biological activity of the unmodified peptide, which binds to desired MHC molecules and activates appropriate T cells. For example, neoantigenic peptides and polypeptides can be further subjected to various modifications, such as conservative or non-conservative substitutions, which may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions refer to the replacement of an amino acid residue with another that is biologically and / or chemically similar, e.g., one hydrophobic residue with another hydrophobic residue, or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effects of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2nd Ed. (1984).
[0190] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful for increasing peptide and polypeptide stability in vivo. Stability can be assayed in a number of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally follows: Pooled human serum (type AB, non-heat-inactivated) is defatted by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.
[0191] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is typically composed of relatively small, neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that the optional spacer need not be composed of the same residues and can therefore be a hetero- or homo-oligomer. If present, the spacer will usually be at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.
[0192] The neoantigenic peptide can be linked to a T helper peptide at either the amino or carboxy terminus of the peptide, either directly or via a spacer. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxoid (830-843), influenza (307-319), and malaria sporozoite (around 382-398 and 378-389).
[0193] Proteins or peptides can be produced by any technique known to those of skill in the art, including expressing proteins, polypeptides, or peptides through standard molecular biology techniques, isolating proteins or peptides from natural sources, or chemically synthesizing proteins or peptides. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located on the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.
[0194] In a further embodiment, the neoantigen comprises a nucleic acid (e.g., a polynucleotide) encoding a neoantigenic peptide or a portion thereof. The polynucleotide can be, for example, a single-stranded and / or double-stranded polynucleotide, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), or a polynucleotide having a phosphorothioate backbone, either in a natural or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further embodiment provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host; such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.
[0195] IV. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., a tumor-specific immune response. Vaccine compositions typically include multiple neoantigens selected, e.g., using the methods described herein. Vaccine compositions may also be referred to as vaccines.
[0196] The vaccine can contain 1 to 30 different peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine may contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, It may contain 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine contains 1-30 neoantigen sequences: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 It can contain 6, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.
[0197] In one embodiment, the different peptides and / or polypeptides, or the nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides are capable of binding to different MHC molecules, such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, a vaccine composition comprises coding sequences for peptides and / or polypeptides capable of binding to the most frequently occurring MHC class I molecules and / or MHC class II molecules. Thus, the vaccine composition can comprise different fragments capable of binding to at least two preferred, at least three preferred, or at least four preferred MHC class I molecules and / or MHC class II molecules.
[0198] The vaccine composition may generate a specific cytotoxic T cell response and / or a specific helper T cell response.
[0199] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are provided herein below. The composition can be combined with a carrier, such as a protein, or an antigen-presenting cell, such as a dendritic cell (DC), which can present peptides to T cells.
[0200] An adjuvant is any substance whose incorporation into a vaccine composition enhances or otherwise modifies the immune response to a neoantigen. The carrier can be a scaffold, such as a polypeptide or polysaccharide, to which the neoantigen can be bound. Optionally, the adjuvant is covalently or non-covalently conjugated.
[0201] The ability of adjuvant to increase the immune response to antigen is typically manifested by a significant or substantial increase in immune-mediated reaction or a reduction in disease symptoms.For example, the increase in humoral immunity is typically manifested by a significant increase in the titer of antibody produced against antigen, and the increase in T cell activity is typically manifested in an increase in cell proliferation, or cellular cytotoxicity, or cytokine secretion.Adjuvant can also change immune response, for example, by changing mainly humoral or Th response to mainly cellular or Th response.
[0202] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA206, Montanide ISA 50V, and Montanide. Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila's QS21 stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponins, mycobacterial extracts and synthetic bacterial cell wall mimics, and other proprietary adjuvants such as Ribi's Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are also useful. Several immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison AC; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Several cytokines have been directly linked to influencing dendritic cell migration to lymphoid tissues (e.g., TNF-α), accelerating dendritic cell maturation into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J ImmunotherEmphasis Tumor Immunol. 1996(6):414-418).
[0203] CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in vaccine settings. Other TLR-binding molecules, such as RNA that binds to TLR 7, TLR 8, and / or TLR 9, may also be used.
[0204] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies, such as cyclophosphamide, sunitinib, bevacizumab, Celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony-stimulating factors such as granulocyte-macrophage colony-stimulating factor (GM-CSF, sargramostim).
[0205] A vaccine composition can include more than one different adjuvant. Additionally, a therapeutic composition can include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.
[0206] The carrier (or excipient) can exist independently of the adjuvant. The function of the carrier can be, for example, to increase activity or immunogenicity, to provide stability, to increase biological activity, or to increase serum half-life, particularly to increase the molecular weight of the variant. Furthermore, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, a serum protein such as keyhole limpet hemocyanin, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is tolerated and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be a dextran, such as Sepharose.
[0207] Cytotoxic T cells (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. MHC molecules themselves are located on the cell surface of antigen-presenting cells. Therefore, CTL activation is possible when a trimeric complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when peptides are used to activate CTLs, but also when APCs bearing the respective MHC molecules are added, it can enhance the immune response. Therefore, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.
[0208] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenovirus (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880), etc. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0209] IV.A. Additional Considerations for Vaccine Design and Manufacturing IV.A.1. Determining a set of peptides covering all tumor subclones Truncal peptides, meaning those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides that are predicted to be highly likely to be presented and immunogenic, or if the number of truncal peptides that are predicted to be highly likely to be presented and immunogenic is small enough that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .
[0210] IV.A.2. Neoantigen Prioritization After applying all of the above neoantigen filters, it is possible that more candidate neoantigens remain available for vaccine inclusion than vaccine technology can accommodate. Additionally, uncertainty about various aspects of neoantigen analysis may remain, and trade-offs may exist between various attributes of candidate vaccine neoantigens. Therefore, instead of predetermined filters at each stage of the selection process, an integral multidimensional model can be considered, in which candidate neoantigens are placed in a space with at least the following axes, and selection is optimized using an integral approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmune risk is typically preferred) 2. Probability of sequencing artifacts (lower artifact probabilities are typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferable) 5. Gene Expression (higher expression is typically preferred) 6. HLA gene coverage (a greater number of HLA molecules involved in presenting a set of neoantigens may decrease the probability that tumors will evade immune attack through downregulation or mutation of HLA molecules) 7. HLA class coverage (covering both HLA-I and HLA-II may increase the probability of therapeutic response and decrease the probability of tumor immune evasion)
[0211] V. Methods of Treatment and Preparation Also provided are methods for inducing a tumor-specific immune response in a subject, vaccinating against the tumor, and treating and / or alleviating symptoms of cancer in a subject by administering to the subject one or more neoantigens, such as multiple neoantigens identified using the methods disclosed herein.
[0212] In some embodiments, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desired. The tumor can be any solid tumor, such as breast, ovarian, prostate, lung, kidney, stomach, colon, testicular, head and neck, pancreas, brain, melanoma, and other tissue organ tumors, as well as hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.
[0213] The neoantigen can be administered in an amount sufficient to induce a CTL response.
[0214] The neoantigen can be administered alone or in combination with other therapeutic agents, such as chemotherapeutic agents, radiation, or immunotherapy. Any suitable therapeutic treatment for the particular cancer can be administered.
[0215] In addition, the subject can be further administered with an anti-immunosuppressive / immunostimulatory substance, such as a checkpoint inhibitor.For example, the subject can be further administered with an anti-CTLA antibody, or anti-PD-1 or anti-PD-L1.Blocking CTLA-4 or PD-L1 with an antibody can enhance the immune response against cancerous cells in patients.In particular, blocking CTLA-4 has been shown to be effective when used in vaccination protocols.
[0216] The optimal amount of each neoantigen to be included in the vaccine composition and the optimal dosing regimen can be determined. For example, the neoantigen or its variants can be formulated for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Methods of injection include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administering vaccine compositions are known to those skilled in the art.
[0217] Vaccines can be edited so that the selection, number, and / or amount of neoantigens present in the composition are tissue-, cancer-, and / or patient-specific. For example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. Selection can depend on the specific type of cancer, the state of the disease, earlier treatment regimens, the patient's immune status, and, of course, the patient's HLA haplotype. Furthermore, vaccines can contain components that are personalized according to the individual needs of a particular patient. Examples include altering the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjusting for secondary treatments after a first round or scheme of treatment.
[0218] For compositions to be used as vaccines for cancer, neoantigens with similar normal self-peptides that are abundantly expressed in normal tissues can be avoided or present in low amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a high amount of a particular neoantigen, the respective pharmaceutical composition for treating that cancer can be present in high amount and / or can include more than one neoantigen specific to that particular neoantigen or pathway of that neoantigen.
[0219] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In therapeutic applications, the compositions are administered to patients in an amount sufficient to elicit an effective CTL response against the tumor antigen and cure or at least partially halt symptoms and / or complications. An amount adequate to achieve this is defined as a "therapeutically effective dose." An amount effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general health, and the judgment of the prescribing physician. It should be kept in mind that compositions can generally be used in severe disease states, i.e., life-threatening or potentially life-threatening situations, particularly when cancer has metastasized. In such instances, it is possible, and the treating physician may find it desirable, to administer substantial excesses of these compositions, taking into account the minimization of adventitious substances and the relatively non-toxic nature of the neoantigens.
[0220] For therapeutic use, administration can begin at the time of detection or surgical removal of a tumor, followed by boosting doses until at least symptoms are substantially abated, and for a period thereafter.
[0221] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against tumors. Disclosed herein are compositions for parenteral administration that contain a solution of a neoantigen, where the vaccine composition is dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. Various aqueous carriers can be used, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional, well-known sterilization techniques or sterile filtered. The resulting aqueous solutions can be packaged for use as is or lyophilized, with the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances required to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, and the like, for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.
[0222] Neoantigens can also be administered via liposomes, which target them to specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigen to be delivered is incorporated as part of the liposome, either alone or in combination with a molecule that binds to a receptor dominant among lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Liposomes filled with the desired neoantigen can thus be directed to the site of lymphoid cells, where they then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid lability, and stability of the liposomes in the bloodstream. Various methods are available for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.
[0223] For targeting to immune cells, the ligand to be incorporated into the liposome can include, for example, an antibody or fragment thereof specific for a cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at doses that vary according to, inter alia, the mode of administration, the peptide being delivered, and the stage of the disease being treated.
[0224] The peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient for therapeutic or immunization purposes. Numerous methods are conveniently used to deliver nucleic acids to a patient. For example, nucleic acids can be delivered directly as "naked DNA." This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), and U.S. Pat. Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery, as described, for example, in U.S. Pat. No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, DNA can be attached to particles, such as gold particles. Approaches for delivering nucleic acid sequences include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.
[0225] Nucleic acids can also be delivered by complexing them with cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO 96 / 18372; 9324640 WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833; 9106309 WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).
[0226] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter can also be included in a viral vector-based vaccine platform, such as the ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20(13):3401-10). Upon introduction into the host, infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.
[0227] A means of administering nucleic acids uses minigene constructs encoding one or more epitopes. To generate DNA sequences (minigenes) encoding selected CTL epitopes for expression in human cells, the amino acid sequences of the epitopes are reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are then directly adjacent to generate a continuous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the minigene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.
[0228] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) compounds can also be complexed with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.
[0229] Also disclosed herein is a method of producing a tumor vaccine, comprising performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising multiple neoantigens or a subset of multiple neoantigens.
[0230] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector (e.g., a vector comprising at least one sequence encoding one or more neoantigens) disclosed herein can include culturing host cells under conditions suitable for expression of the neoantigen or vector, wherein the host cells comprise at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatographic, electrophoretic, immunological, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.
[0231] The host cell can comprise a Chinese hamster ovary (CHO) cell, an NS0 cell, yeast, or an HEK293 cell. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to the at least one nucleic acid sequence encoding the neoantigen or vector. In certain embodiments, the isolated polynucleotide can be a cDNA.
[0232] VI. Identification of neoantigens VI.A. Identification of Candidate Neoantigens A research method for NGS analysis of tumor and normal exomes and transcriptomes is described and applied in the specific space of neoantigens. 6,14,15The examples below consider certain optimizations for greater sensitivity and specificity for identifying neoantigens in a clinical setting. These optimizations can be grouped into two areas: those related to laboratory processes and those related to NGS data analysis.
[0233] VI.A.1. Laboratory Process Optimization The process improvements presented herein build on the concepts developed for reliable assessment of cancer driver genes in targeted cancer panels. 16 This addresses the challenges in high-precision neoantigen discovery from clinical specimens with low tumor content and small volumes by expanding the method to the whole-exome and whole-transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) unique average coverage across the tumor exome to detect mutations present at low mutant allele frequency due to either low tumor content or subclonal status. 2. Fewer than 5% of bases are covered at less than 100x to minimize missed potential neoantigens, e.g. a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for areas that are not sufficiently covered 3. Targeting uniform coverage across the normal exome, with less than 5% of bases covered below 20x, to minimize the chance of potential neoantigens remaining unclassified for somatic / germline status (and therefore unusable as TSNAs). 4. To minimize the total amount of sequencing required, sequence capture probes are designed only for the coding regions of the gene, since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing18 . b. Elimination of genes predicted to produce few or no candidate neoantigens due to factors such as poor expression, suboptimal digestion by the proteasome, or atypical sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples can be subjected to probe-based enrichment with the same or similar probes used to capture the exome in DNA. 19 It is extracted using
[0234] VI.A.2. Optimizing NGS Data Analysis Analytical method improvements address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically allow for customization relevant for identifying neoantigens in the clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment, as it contains multiple MHC region assemblies that better reflect population polymorphism, as opposed to earlier genome releases. 2. Various programs 5 Overcoming the limitations of single mutation callers 20 by merging results from a. Single nucleotide mutations and indels are detected in tumor DNA, tumor RNA, and normal DNA with a range of tools including: Strelka 21 and Mutec t22 and programs based on comparison of tumor and normal DNA, such as; and 23 , UNCeqR, and other programs that incorporate tumor DNA, tumor RNA, and normal DNA. b. Indels are found in Strelka and ABRA 24 This is determined by a program that performs local reassembly, such as c. Structural rearrangements are 25 or Breakseq26 It is determined using specialized tools such as 3. To detect and prevent sample swapping, mutation calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. Extensive filtering of artificial calls is performed, for example, by: a. Removal of mutations found in normal DNA, potentially with relaxed detection parameters in the case of low coverage and permissive proximity criteria in the case of indels. b. Removal of mutations due to poor mapping quality or poor base quality 27 . c. Elimination of mutations resulting from re-emerging sequencing artifacts, even if not observed in the corresponding normal 27 Examples include mutations that are detected primarily on one strand. d. Removal of mutations detected in a set of unrelated controls 27 . 5.seq2HLA 28 , ATHLATES 29 or Optitype, and also combine exome and RNA sequencing data 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing, such as long-read DNA sequencing. 30 or adaptation of methods for linking RNA fragments to maintain continuity. 31 Includes: 6. Robust Detection of Nascent ORFs Arising from Tumor-specific Splice Variants in CLASS 32 , Bayesembler 33 , StringTie 34 This is done by assembling transcripts from RNA-seq data using Cufflinks, or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate the entire transcript from each experiment). 35Although commonly used for this purpose, it frequently produces an incredibly large number of splice variants, many of which are much shorter than the full-length gene, and may not be able to recover a simple positive control. The coding sequence and potential nonsense-mediated decay mechanisms reintroduced the mutant sequence, SpliceR. 36 and MAMBA 37 Gene expression is determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels are determined using ASE. 38 or HTSeq 39 Potential filtering steps include: a. Removal of candidate nascent ORFs that are thought to be poorly expressed. b. Removal of candidate nascent ORFs predicted to trigger nonsense-mediated decay (NMD). 7. Candidate neoantigens observed only in RNA (e.g., neo-ORFs) that cannot be directly validated as tumor-specific are classified as likely to be tumor-specific according to additional parameters, for example, by considering the following: a. Presence of supporting cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmed trans-acting mutations in splicing factors in tumor DNA only. As an example, the gene that exhibited the most differential splicing in three independently published experiments with R625 mutant SF3B1 was 1. One experiment examined patients with uveal melanoma. 40 The second experiment examined uveal melanoma cell lines. 41 , and a third study looked at breast cancer patients. 42 Nevertheless, there was agreement. c. For novel splicing isoforms, the presence of confirmatory "novel" splice-junction reads in the RNASeq data. d. For de novo rearrangements, the presence of confirmatory exon-proximal reads in tumor DNA that are not present in normal DNA. e.GTEx 43 and absence from the gene expression compendium (i.e., making germline origin less likely). 8. Complementing reference genome alignment-based analyses by comparing tumor and normal reads (or k-mers derived from such reads) of assembled DNA to directly avoid alignment- and annotation-based errors and artifacts (e.g., for somatic mutations occurring near germline mutations or repeat-context indels).
[0235] In samples with polyadenylated RNA, the presence of viral and microbial RNA in the RNA-seq data will be assessed using RNA CoMPASS44 or similar methods to identify additional factors that may predict patient response.
[0236] VI.B. HLA Peptide Isolation and Detection Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) techniques after lysis and solubilization of tissue samples. 55~58 The clarified lysates were used for HLA-specific IP.
[0237] Immunoprecipitation was performed using antibodies coupled to beads, where the antibodies are specific for HLA molecules. For pan-class I HLA immunoprecipitation, a pan-class I CR antibody is used, and for class II HLA-DR, an HLA-DR antibody is used. The antibodies are covalently attached to NHS-Sepharose beads during overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60Immunoprecipitation can also be performed using antibodies that are not covalently attached to beads. Typically, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to retain the antibodies on the column. Some antibodies that can be used to selectively enrich MHC / peptide complexes are listed below. TIFF0007795291000001.tif46160
[0238] The clarified tissue lysate is added to antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is saved for further experiments, including additional IPs. The IP beads are washed to remove nonspecific binding, and the HLA / peptide complexes are eluted from the beads using standard techniques. Protein components are removed from the peptides using molecular weight spin columns or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and, in some cases, stored at -20°C prior to MS analysis.
[0239] The dried peptides were reconstituted in an HPLC buffer suitable for reversed-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution on a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of peptide mass / charge (m / z) were collected at high resolution on an Orbitrap detector, followed by MS2 low-resolution scans on an ion trap detector after HCD fragmentation of selected ions. Additionally, MS2 spectra can be acquired using either CID or ETD fragmentation methods, or any combination of the three techniques to obtain greater amino acid coverage of the peptide. MS2 spectra can also be measured with high-resolution mass accuracy on an Orbitrap detector.
[0240] The MS2 spectra from each analysis were analyzed using Comet 61、62 and peptide identifications were analyzed using Percolator63~65 Further sequencing is performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or spectral matching and de novo sequencing are performed. 75 Sequencing methods including:
[0241] VI.B.1. Investigation of MS detection limits for comprehensive HLA peptide sequencing Peptide YVYVADVAAK (SEQ ID NO: 1) Using this method, the limit of detection was determined using various amounts of peptide loaded onto the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol (Table 1). The results are shown in Figure 1F. These results indicate that the lowest limit of detection (LoD) was in the attomole range (10 -18 ), a dynamic range spanning five orders of magnitude, and a signal-to-noise ratio in the low femtomole range (10 -15 ) appears to be sufficient for sequencing.
[0242] TIFF0007795291000002.tif61128
[0243] VII. Presented Model VII.A. System Overview 2A is an overview of an environment 100 for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for implementing a presentation identification system 160, which itself includes a presentation information store 165.
[0244] The presentation identification system 160 is a computer model, embodied in a computational system such as that discussed below with respect to FIG. 14, that receives a peptide sequence associated with a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the set of associated MHC alleles. The presentation identification system 160 can be applied to both class I and class II MHC alleles, making it useful in a variety of contexts. One specific example application of the presentation identification system 160 is to receive the nucleotide sequence of a candidate neoantigen associated with a set of MHC alleles from tumor cells of a patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the tumor's associated MHC alleles and / or induce an immunogenic response in the patient's 110 immune system. Those candidate neoantigens with a high likelihood, as determined by the system 160, can be selected for inclusion in a vaccine 118, such that an anti-tumor immune response can be elicited from the immune system of the patient 110 that provided the tumor cells.
[0245] The presentation identification system 160 determines the presentation likelihood through one or more presentation models. Specifically, the presentation models generate a likelihood that a given peptide sequence will be presented for a set of associated MHC alleles, the likelihood being generated based on the presentation information stored in the storage device 165. For example, a presentation model may be generated for the peptide sequence "YVYVADVAAK" (SEQ ID NO: 1)" may generate a likelihood of whether a peptide sequence will be presented for the set of alleles HLA-A*02:01, HLA-A*03:01, HLA-B*07:02, HLA-B*08:03, HLA-C*01:04 on the cell surface of the sample. Presentation information 165 contains information about whether these peptides bind to various types of MHC alleles such that the peptides are presented by the MHC alleles, which is determined in the model according to the position of the amino acid in the peptide sequence. Based on presentation information 165, the presentation model can predict whether an unrecognized peptide sequence will be presented in association with the associated set of MHC alleles. As noted above, the presentation model can be applied to both class I and class II MHC alleles.
[0246] VII.B. Presentation information 2 illustrates a method for obtaining presentation information according to one embodiment. Presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. Allele interaction information includes information that affects presentation of peptide sequences that is dependent on the type of MHC allele. Allele non-interaction information includes information that affects presentation of peptide sequences that is independent of the type of MHC allele.
[0247] VII.B.1. Allelic Interaction Information The allele interaction information primarily includes identified peptide sequences known to be presented by one or more identified MHC molecules from humans, mice, etc. Of note, this may or may not include data obtained from tumor samples. Presented peptide sequences may be identified from cells expressing a single MHC allele. In this example, presented peptide sequences are generally collected from a monoallelic cell line engineered to express a predetermined MHC allele and then exposed to a synthetic protein. Peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows an exemplary peptide presented on the predetermined MHC allele HLA-DRB1*12:01. An example of this is shown in TIFF0007795291000003.tif5128, where a peptide is isolated and identified by mass spectrometry. In this situation, the direct association between the presented peptide and the MHC protein to which it binds is definitively known, since the peptide is identified through cells engineered to express a single, predetermined MHC protein.
[0248] Presented peptide sequences may also be collected from cells expressing multiple MHC alleles. Typically, in humans, cells express six different types of MHC-I molecules and up to 12 different types of MHC-II molecules. Such presented peptide sequences may be identified from multi-allelic cell lines engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal or tumor tissue samples. In this particular example, MHC molecules can be immunoprecipitated from normal or tumor tissue. Peptides presented on multiple MHC alleles can similarly be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows six exemplary peptides. An example of this is shown in TIFF0007795291000004.tif17166, which is presented on the identified class I MHC alleles HLA-A*01:01, HLA-A*02:01, HLA-B*07:02, and HLA-B*08:01, and the class II MHC alleles HLA-DRB1*10:01 and HLA-DRB1:11:01, isolated, and characterized by mass spectrometry. In contrast to monoallelic cell lines, because the bound peptide is isolated from the MHC molecule prior to its identification, the direct association between the presented peptide and the MHC protein to which it is bound may be unknown.
[0249] Allele interaction information can also include mass spectrometry ion currents, which depend on both the concentration of peptide-MHC molecule complexes and the ionization efficiency of the peptides. Ionization efficiency varies from peptide to peptide in a sequence-dependent manner. Generally, ionization efficiency varies from peptide to peptide over approximately two orders of magnitude, while the concentration of peptide-MHC complexes varies over an even larger range.
[0250] Allele interaction information can also include measured or predicted binding affinities between a given MHC allele and a given peptide. One or more affinity models can generate such predictions (72, 73, 74). For example, returning to the example shown in Figure 1D, representation 165 may represent a sequence of peptides YEMFNDKSF (SEQ ID NO:3) and class I alleles of HLA-A * 01:01. Few peptides with IC50>1000 nM are presented by the MHC, and a lower IC50 value increases the probability of presentation. Presentation information 165 can include a predicted binding affinity value of 1000 nM between 01:01 and 1000 nM. The predicted binding affinity between TIFF0007795291000005.tif4128 and the class II allele HLA-DRB1:11:01 may be included.
[0251] The allele interaction information can also include measured or predicted stability values for MHC complexes. One or more stability models can generate such predictions. More stable peptide-MHC complexes (i.e., complexes with longer half-lives) are more likely to be presented in high copy number on tumor cells and on antigen-presenting cells that encounter vaccine antigens. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include a predicted stability value for a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a predicted stability value for the half-life of the class II molecule HLA-DRB1:11:01.
[0252] Allele interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at faster rates are more likely to be presented at high concentrations on the cell surface.
[0253] Allele interaction information can also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of presented peptides have a length of 9 peptides. MHC class II molecules generally tend to present peptides with a length of 6-30 peptides.
[0254] Allele interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoded peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoded peptide. The presence of kinase motifs influences the probability of post-translational modifications that may enhance or interfere with MHC binding.
[0255] Allelic interaction information can also include expression or activity levels of proteins involved in post-translational modification processes, such as kinases (as measured or predicted by RNA-seq, mass spectrometry, or other methods).
[0256] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.
[0257] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual in question (e.g., as measured by RNA-seq or mass spectrometry): peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.
[0258] Allelic interaction information can also include the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals that express that particular MHC allele.
[0259] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and therefore, presentation of peptides by HLA-C is a priori less likely than presentation by HLA-A or HLA-B II. As another example, because HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, presentation of peptides by HLA-DP is predicted to be less likely than presentation by HLA-DR or HLA-DQ.
[0260] The allele interaction information can also include the protein sequence of a particular MHC allele.
[0261] Any of the MHC allele non-interacting information listed in the section below can also be modeled as MHC allele interacting information.
[0262] VII.B.2. Allelic Non-Interaction Information Non-allele-interacting information can include the C-terminal sequence adjacent to the neoantigen-encoded peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence can affect proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters an MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and therefore, the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in Figure 2C, presentation information 165 is the presented peptide FJIEJFOESS identified from the peptide's source protein. (SEQ ID NO:5) C-terminal flanking sequence of TIFF0007795291000006.tif5128.
[0263] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide mass spectrometry training data. As described later with respect to Figure 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, mRNA quantification measurements are determined from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).
[0264] The allele non-interacting information can also include N-terminal sequences adjacent to the peptide within its source protein sequence.
[0265] The allelic non-interaction information can also include a source gene for the peptide sequence. The source gene can be defined as an Ensembl protein family for the peptide sequence. In another example, the source gene can be defined as a source DNA or source RNA for the peptide sequence. The source gene can be represented, for example, as a string of nucleotides that encodes a protein, or alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode specific proteins. In another example, the allelic non-interaction information can also include a source transcript or isoform or a set of potential source transcripts or isoforms for the peptide sequence extracted from a database such as Ensembl or RefSeq.
[0266] The allelic non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.
[0267] The allele non-interaction information can also include the presence of protease cleavage motifs in peptides, optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and therefore less stable in cells, and therefore less likely to be presented.
[0268] Allelic non-interaction information can also include the turnover rate of the source protein when measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation, but this characteristic has low predictive power when measured in dissimilar cell types.
[0269] The allelic non-interaction information can also include the length of the source protein, optionally taking into account the specific splice variants ("isoforms") that are most highly expressed in tumor cells, as measured by RNA-seq or proteome mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.
[0270] Allele-free interaction information can also include the expression level of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Different proteasomes have different cleavage site preferences. More weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.
[0271] Allele-free interaction information can also include the expression of the peptide's source gene (e.g., as measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in the tumor sample. Peptides from genes with higher expression are more likely to be presented. Peptides from genes with undetectable levels of expression can be eliminated from consideration.
[0272] Allelic non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to nonsense-mediated decay as predicted by a model of nonsense-mediated decay, e.g., the model from Rivas et al., Science 2015.
[0273] Allelic non-interaction information can also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle. Genes that are expressed at low levels overall (as measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are displayed than genes that are stably expressed at very low levels.
[0274] Allelic non-interaction information can also include a comprehensive catalog of source protein properties, such as those provided in uniProt or the PDB (http: / / www.rcsb.org / pdb / home / home.do). These properties can include, among others, protein secondary and tertiary structure, subcellular localization, and Gene Ontology (GO) terms. Specifically, this information can include annotations operating at the protein level, e.g., 5'UTR length, and annotations operating at the level of specific residues, e.g., a helix motif between residues 300 and 310. These properties can also include turn motifs, sheet motifs, and disordered residues.
[0275] Allelic non-interacting information can also include features that describe the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (eg, alpha helix versus beta sheet); alternative splicing.
[0276] The allelic non-interaction information can also include properties that describe the presence or absence of presentation hotspots at the peptide's position in its source protein.
[0277] Allelic non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression level of the source protein in those individuals and the influence of the various HLA types of those individuals).
[0278] Allelic non-interaction information can also include the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias.
[0279] Expression of various gene modules / pathways (not necessarily containing the source protein of the peptides) as measured by gene expression assays such as RNASeq, microarrays, targeted panels such as Nanostring, or single / multiple genes representing gene modules measured by assays such as RT-PCR, that inform on the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs).
[0280] Allele non-interaction information can also include the copy number of the peptide's source gene in the tumor cell. For example, a peptide derived from a gene that is subject to homozygous deletion in the tumor cell can be assigned a presentation probability of zero.
[0281] The allele-non-interaction information can also include the probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP or that bind with higher affinity to TAP are more likely to be presented by MHC-I.
[0282] Allelic non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Higher TAP expression levels at MHC-I increase the probability of presentation of all peptides.
[0283] Allelic non-interaction information can also include the presence or absence of tumor mutations, including but not limited to: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3. ii. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery affected by loss-of-function mutations in the tumor have a reduced probability of presentation.
[0284] Presence or absence of functional germline polymorphisms, including but not limited to: i. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome).
[0285] The allelic non-interaction information can also include tumor type (eg, NSCLC, melanoma).
[0286] Allele non-interaction information can also include the known functionality of the HLA allele, e.g., as reflected by the HLA allele suffix. For example, the N suffix in the allele name HLA-A*24:09N indicates a null allele that is not expressed and therefore unlikely to present an epitope; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.
[0287] Allelic non-interaction information can also include clinical tumor subtype (eg, squamous cell lung cancer vs. non-squamous).
[0288] The allele non-interaction information can also include smoking history.
[0289] Allele non-interaction information can also include a history of sunburn, sun exposure, or exposure to other mutagens.
[0290] The allelic non-interaction information can also include regional expression of the peptide's source gene in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be represented.
[0291] The allelic non-interaction information can also include the frequency of the mutation in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.
[0292] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation can also include the mutation's annotation (e.g., missense, readthrough, frameshift, fusion, etc.) or whether the mutation is predicted to result in nonsense-mediated decay (NMD). For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in reduced mRNA translation, which reduces the probability of presentation.
[0293] VII.C. Presentation Identification System 3 is a high-level block diagram illustrating the computer logic components of presentation identification system 160, according to one embodiment. In this exemplary embodiment, presentation identification system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. Presentation identification system 160 also comprises a training data store 170 and a presentation model store 175. Some embodiments of model management system 160 have different modules than those described herein. Likewise, functionality may be distributed among the modules in a manner different from that described herein.
[0294] VII.C.1. Data Management Module The data management module 312 generates sets of training data 170 from the representation information 165. Each training data set contains a number of data examples, each of which contains at least one of the represented or unrepresented peptide sequences p i and the peptide sequence p i one or more relevant MHC alleles combined with i and the dependent variable y, which represents information that the presentation identification system 160 is interested in predicting new values of the independent variables. i and the independent variable z i Contains a set of
[0295] In one particular implementation referred to throughout the remainder of this specification, the dependent variable y i is the peptide p i but one or more associated MHC alleles a i However, in other implementations, the dependent variable y i is the result of the presentation identification system 160 determining the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that one is interested in predicting. For example, in another implementation, the dependent variable y i σ may also be a numerical value indicating the mass analysis ion current determined for the example data.
[0296] Peptide sequence p for data example i i is k i is a sequence of k amino acids, i can vary within a range among data instances i. For example, the range can be 8 to 15 for MHC class I, or 6 to 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p in the training data set are i may have the same length, e.g., 9. The number of amino acids in a peptide sequence may vary depending on the type of MHC allele (e.g., MHC allele in humans). MHC allele a for data example i i is the peptide sequence p i indicates whether it existed in combination with
[0297] The data management module 312 also manages the peptide sequences p contained in the training data 170. i and bound MHC allele a i Together with the binding affinity b i and stability i For example, the training data 170 may include a predictor of the peptide p i and a i The predicted binding affinity b between each of the bound MHC molecules shown in iAs another example, the training data 170 may contain a i The predicted stability value s for each of the MHC alleles shown in i may contain
[0298] The data management module 312 also receives the peptide sequence p i along with non-allele interacting variables such as C-terminal flanking sequences and mRNA quantification measurements. i It may also include.
[0299] The data management module 312 also identifies peptide sequences that are not presented by MHC alleles to generate the training data 170. Generally, this involves identifying a "longer" sequence of the source protein that contains the peptide sequence to be presented prior to presentation. If the presentation information contains an engineered cell line, the data management module 312 identifies a set of peptide sequences in the synthetic protein to which the cell was exposed that were not presented on the MHC alleles of the cell. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein from which the presented peptide sequence originated and identifies a set of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.
[0300] The data management module 312 also artificially generates peptides with random sequences of amino acids and identifies the generated sequences as peptides that are not presented on MHC alleles. This can be achieved by randomly generating peptide sequences, allowing the data management module 312 to easily generate large amounts of synthetic data for peptides that are not presented on MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, synthetically generated peptide sequences are very likely not presented by MHC alleles, even if they are included in proteins processed by cells.
[0301] 4 illustrates an exemplary set of training data 170A, according to one embodiment. Specifically, the first three data examples in training data 170A are a monoallelic cell line containing the allele HLA-C*01:03, and three peptide sequences: The fourth example data in training data 170A shows peptide presentation information from TIFF0007795291000007.tif9128. The fourth example data in training data 170A shows a multi-allelic cell line containing the alleles HLA-B*07:02, HLA-C*01:03, and HLA-A*01:01, and the peptide sequence QIEJOEIJE (SEQ ID NO:13) The first data example shows peptide information from the peptide sequence QCEIOWARE (SEQ ID NO:14) was not presented by the allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, the negatively labeled peptide sequences may be randomly generated by the data management module 312 or may be identified from the source protein of the presented peptide. The training data 170A also includes a predicted binding affinity of 1000 nM and a predicted stability with a half-life of 1 hour for the peptide sequence-allele pair. The training data 170A also includes a predicted binding affinity of 1000 nM and a predicted stability with a half-life of 1 hour for the peptide FJELFISBOSJFIE. (SEQ ID NO: 15) and the C-terminal flanking sequence of 10 2 It also includes non-allele-interacting variables, such as the mRNA quantification measurements of TPM. The fourth data example is the peptide sequence QIEJOEIJE (SEQ ID NO:13) was presented by one of the alleles HLA-B*07:02, HLA-C*01:03, or HLA-A*01:01. Training data 170A also includes predicted binding affinity and stability values for each of the alleles, as well as the C-terminal flanking sequence of the peptide and mRNA quantification measurements for the peptide.
[0302] VII.C.2. Coding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more representation models. In one implementation, the encoding module 314 one-hot encodes sequences (e.g., peptide sequences or C-terminal flanking sequences) for a predetermined 20-letter amino acid alphabet. Specifically, k i Peptide sequence p having amino acids i is 20·k i p, which is represented as a row vector of elements, corresponding to the alphabet of the amino acid at the jth position of the peptide sequence. i 20·(j-1)+1 ,p i 20·(j-1)+2 ,...,p i 20·j A single element in has a value of 1. The remaining elements have a value of 0. As an example, for a given alphabet {A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y}, the three amino acid peptide sequence EAF of data example i is a 60-element row vector The C-terminal flanking sequence c can be represented by TIFF0007795291000008.tif12138. i , and the protein sequence for the MHC allele d h , and other sequence data in the presentation information can be similarly coded as above.
[0303] If the training data 170 contains sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equivalent length by adding PAD characters to extend the predetermined alphabet. For example, this may be done by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the peptide sequence with the longest length in the training data 170. Thus, if the peptide sequence with the longest length is k 最大 amino acids, the encoding module 314 encodes each sequence as (20+1) k 最大It is represented numerically as a row vector of elements. For example, consider the extended alphabet {PAD,A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y} and k 最大 For a maximum amino acid length of ≡5, the same exemplary peptide sequence EAF of 3 amino acids is represented as a 105-element row vector The C-terminal flanking sequence c i or other sequence data can be similarly encoded as above. Thus, the peptide sequence p i or c i Each argument or row in represents the occurrence of a particular amino acid at a particular position in the sequence.
[0304] Although the above method for encoding sequence data has been described with respect to sequences having amino acid sequences, the method can be similarly extended to other types of sequence data, such as, for example, DNA or RNA sequence data.
[0305] The encoding module 314 also encodes one or more MHC alleles a for data instance i. i is encoded into an m-element row vector, where each element h=1,2,...,m corresponds to a uniquely identified MHC allele. The element corresponding to the identified MHC allele for data instance i has a value of 1. The remaining elements have a value of 0. As an example, among the m=4 uniquely identified MHC allele types {HLA-A*01:01, HLA-C*01:08, HLA-B*07:02, HLA-DRB1*10:01}, the alleles HLA-B*07:02 and HLA-DRB1*10:01 for data instance i corresponding to a multi-allelic cell line are encoded into a four-element row vector a i = [0 0 1 1], and a3 i =1 and a4 i = 1. An example with four identified MHC allele types is described herein, but the number of MHC allele types can actually be hundreds or thousands. As noted above, each data instance i typically contains a peptide sequence pi It contains up to six different MHC allele types associated with
[0306] The encoding module 314 also generates a label y for each data instance i. i We code, as a binary variable with values from the set {0,1}, where a value of 1 indicates that the peptide x i However, the associated MHC allele a i a value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele a i The dependent variable y i If represents the mass analysis ion current, the encoding module 314 may additionally scale the value using various functions, such as a log function with a range of [-∞,∞] for ion current values between [0,∞].
[0307] The coding module 314 encodes the peptide p i and the allele interaction variable x for the associated MHC allele h h i pairs as row vectors in which the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], where b h i is the predicted binding affinity for peptide p and associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).
[0308] In one example, the encoding module 314 encodes the measured or predicted values for binding affinity as a function of the allele interaction variable x h i The binding affinity information is expressed by incorporating the
[0309] In one example, the encoding module 314 encodes the measured or predicted values for binding stability as allele interaction variables x h i By incorporating it into
[0310] In one example, the encoding module 314 encodes the measured or predicted values for the binding on-rate as a function of the allele interaction variable x h i The combined on-rate information is expressed by incorporating
[0311] In one example, for peptides presented by class I MHC molecules, the encoding module 314 encodes the peptide length in the vector TIFF0007795291000010.tif5132 (However, TIFF0007795291000011.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i In another example, for peptides presented by class II MHC molecules, the encoding module 314 may include the peptide length in the vector TIFF0007795291000012.tif19161 (However, TIFF0007795291000013.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x hi can be included in
[0312] In one example, the encoding module 314 represents the RNA expression information of MHC alleles by incorporating the RNA-seq-based expression levels of the MHC alleles into an allele interaction variable xhi.
[0313] Similarly, the encoding module 314 encodes the allele non-interacting variable w i can be represented as a row vector in which the numerical representations of the allele-non-interacting variables are concatenated one after the other. For example, w i is [c i ] or [c i m i w i ], and w i is the C-terminal flanking sequence of peptide pi and the mRNA quantification measurement m associated with the peptide. i Alternatively, one or more combinations of allele non-interacting variables may be stored individually (e.g., as individual vectors or matrices).
[0314] In one example, the encoding module 314 encodes the turnover rate or half-life as a function of the allele non-interacting variable w i represents the turnover rate of the source protein for the peptide sequence.
[0315] In one example, the encoding module 314 encodes the protein length as a function of the allele non-interacting variable w i represents the length of the source protein or isoform by incorporating
[0316] In one example, the encoding module 314 generates β1 i , β2 i , β5 i The mean expression of immunoproteasome-specific proteasome subunits, including the subunits, was calculated using the allele-noninteracting variable w iIncorporation into the IL-1 protein results in activation of the immunoproteasome.
[0317] In one example, the encoding module 314 encodes the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM) or a source protein of a gene or transcript of the peptide, by correlating the source protein abundance with an allele-non-interacting variable w i This is expressed by incorporating it into
[0318] In one example, the encoding module 314 calculates the probability that the transcript of the peptide's origin will undergo nonsense-mediated decay (NMD), for example, as estimated by the model in Rivas et al. Science, 2015, and calculates this probability as a function of the allele non-interaction variable w i This is expressed by incorporating it into
[0319] In one example, encoding module 314 represents the activation status of a gene module or pathway assessed via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM, for each of the genes in the pathway, and then computing a summary statistic, such as a mean, across the genes in the pathway. The mean is calculated using the allele-non-interaction variable w i can be incorporated into.
[0320] In one example, the encoding module 314 encodes the copy number of the source gene by dividing the copy number by the allele non-interacting variable w i This is expressed by incorporating it into
[0321] In one example, the encoding module 314 encodes the measured or predicted TAP binding affinity (e.g., in nanomolar units) relative to the allele-non-interacting variable w i The TAP binding affinity is expressed by including
[0322] In one example, the encoding module 314 encodes TAP expression levels measured by RNA-seq (and quantified, for example, by RSEM in units of TPM) as a function of the allele-non-interacting variable w i The expression level of TAP is represented by the inclusion of
[0323] In one example, the encoding module 314 encodes the tumor mutations as allele-non-interacting variables w i vector of indicator variables in (i.e., peptide p k is derived from a sample with a KRAS G12D mutation, d k = 1, otherwise 0).
[0324] In one example, the encoding module 314 encodes germline polymorphisms in antigen-presenting genes as a vector of indicator variables (i.e., peptide p k If is derived from a sample with a specific germline polymorphism in TAP, then d k = 1). These indicator variables are expressed as the allele non-interaction variables w i can be included in
[0325] In one example, the encoding module 314 represents tumor types as one-hot coded vectors of length 1 for an alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot coded variables are combined into an allelic non-interaction variable w i can be included in
[0326] In one example, the encoding module 314 represents MHC allele suffixes by processing four-digit HLA alleles with various suffixes. For example, HLA-A*24:09N is considered a different allele from HLA-A*24:09 for purposes of the model. Alternatively, because HLA alleles ending in an N suffix are not expressed, the probability of presentation by an MHC allele with an N suffix can be set to zero for all peptides.
[0327] In one example, the encoding module 314 represents tumor subtypes as one-hot coded vectors of length 1 for an alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot coded variables are combined into an allelic non-interaction variable w i can be included in
[0328] In one example, the encoding module 314 encodes smoking history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a smoking history) can be included in k = 1 otherwise 0). Alternatively, smoking history can be coded as a one-hot coded variable of length 1 for the alphabet of smoking severity. For example, smoking status can be assessed on a 1-5 scale, with 1 indicating non-smoker and 5 indicating current heavy smoker. Because smoking history is primarily associated with lung tumors, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a smoking history and the tumor type is lung tumor, and zero otherwise.
[0329] In one example, the encoding module 314 encodes sunburn history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a history of severe sunburn) can be included in k = 1 otherwise 0). Because severe sunburn is primarily associated with melanoma, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and zero otherwise.
[0330] In one example, the encoding module 314 represents the distribution of expression levels of a particular gene or transcript for each gene or transcript in the human genome as a summary statistic (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, the expression of peptide p in samples with the tumor type melanoma is expressed as a summary statistic (e.g., mean, median) of the distribution of expression levels. k Regarding peptide p k The measured gene or transcript expression levels of the genes or transcripts of origin are compared with the allele-non-interacting variable w i Not only can it be included in the peptide p in melanoma as measured by TCGA, k The mean and / or median gene or transcript expression of the genes or transcripts of a given source may also be included.
[0331] In one example, the encoding module 314 represents the variant types as one-hot coded variables of length 1 for an alphabet of variant types (e.g., missense, frameshift, NMD-induced, etc.). These one-hot coded variables are grouped together into allelic non-interaction variables w i can be included in
[0332] In one example, the encoding module 314 encodes the protein-level characteristics of the protein as values of the source protein annotation (e.g., 5′ UTR length) and the allele-non-interacting variable w i In another example, the encoding module 314 encodes the peptide p i The residue-level annotation of the source protein for peptide p i is equal to 1 if overlaps with the helical motif, otherwise it is equal to 0, or i The allele-non-interaction variable wi represents the allele-non-interaction variable wi by including an indicator variable equal to 1 if p is completely contained within the helix motif. i The property that represents the proportion of residues in ican be included in
[0333] In one example, the encoding module 314 encodes the types of proteins or isoforms in the human proteome into an index vector o having a length equivalent to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is the peptide p k is 1 if comes from protein i, and 0 otherwise.
[0334] In one example, the encoding module 314 encodes the peptide p i Source gene G = gene(p i ) as a categorical variable with L possible categories (where L denotes an upper bound on the number of subscripted source genes 1, 2, ..., L).
[0335] In one example, the encoding module 314 encodes the peptide p i T = tissue type, cell type, tumor type, or tumor histology type of T = tissue (p i ) as a categorical variable with M possible categories (where M denotes an upper limit on the number of subscripted types 1, 2, ..., M). Tissue types can include, for example, lung tissue, cardiac tissue, intestinal tissue, and neural tissue. Cell types can include, for example, dendritic cells, macrophages, and CD4 T cells. Cancers can include, for example, lung adenocarcinoma, lung squamous cell carcinoma, melanoma, and non-Hodgkin's lymphoma.
[0336] The encoding module 314 also encodes the peptide p i and the variable z for the associated MHC allele h i The entire set of alleles is expressed as the allele interaction variable x i and the allele non-interaction variable w i For example, the encoding module 314 may represent z h i [xh i w i ] or [w i x h i ] can be represented as a row vector equivalent to
[0337] VIII. Training Module The training module 316 constructs one or more presentation models that generate a likelihood of whether a peptide sequence will be presented by an MHC allele associated with the peptide sequence. k and peptide sequence p k MHC alleles associated with a k Given a set of peptide sequences, each proposed model k However, the associated MHC allele a k the estimate u, which indicates the likelihood that one or more of k Generate.
[0338] VIII.A. Overview The training module 316 constructs one or more representation models based on a training data set stored in storage 170, which is generated from the representation information stored in 165. Generally, regardless of the specific type of representation model, all representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function TIFF0007795291000014.tif5128 is a graph of the dependent variable y for one or more data examples S in the training data 170. i∈S and the estimated likelihood u for the data example S generated by the proposed model. i∈S In one particular implementation, which will be mentioned throughout the remainder of this document, the loss function TIFF0007795291000015.tif4128 is the negative log likelihood function given by equation (1a) as follows: However, in practice, another loss function may be used. For example, if the prediction is made for mass spectrometry ion current, the loss function is the mean square loss given by Equation 1b as follows: TIFF0007795291000017.tif11128
[0339] The proposed model may be a parametric model, where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function The various parameters of the proposed parametric model that minimizes TIFF0007795291000018.tif4128 are determined through a gradient-based numerical optimization algorithm, such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the proposed model may be a non-parametric model, in which the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.
[0340] VIII.B. Allele-by-Allele Model The training module 316 may build a presentation model to predict the presentation likelihood of a peptide on an allele-by-allele basis. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele.
[0341] In one implementation, the training module 316: TIFF0007795291000019.tif7128 estimates the likelihood of peptide pk being presented for a particular allele h. k where the peptide sequence x h k is the peptide p k and the coded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, which for convenience of description will be referred to as a transformation function throughout this specification. h(·) is an arbitrary function, which for convenience of description will be referred to as the dependence function throughout this specification, and the parameter θ determined for the MHC allele h h Based on the set of allele interaction variables x h k Generate a dependency score for the parameter θ for each MHC allele h. h The set of values is θ h where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele h.
[0342] Dependence function g h (x h k ;θ h ) output is the MHC allele h with at least the allele interaction characteristic x h k and in particular the peptide p k The dependency score for MHC allele h indicates whether the corresponding neoantigen is presented based on the amino acid position of the peptide sequence of p. For example, the dependency score for MHC allele h is determined by the relationship between the MHC allele h and the peptide p. k The transformation function f(·) transforms the input, more specifically, g in this example. h (x h k ;θ h ) is used to calculate the dependency score for peptide p k is converted to an appropriate value indicating the likelihood that it will be presented by the MHC allele.
[0343] In one particular implementation referred to throughout the remainder of this specification, f(·) is a function with range in [0,1] for the appropriate domain range. In one example, f(·) is The expit function is given by TIFF0007795291000020.tif10128. As another example, f(·) also has the following meaning: for values in the domain z greater than or equal to 0, It can also be the hyperbolic tangent function given by TIFF0007795291000021.tif5128. Alternatively, if the prediction is made for mass analysis ion currents with values outside the range [0,1], f(·) can be any function, for example, the identity function, the exponential function, the log function, etc.
[0344] Therefore, the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is given by the dependency function g h (·) is the peptide sequence p k to generate a corresponding dependency score. The dependency score can be generated by applying k may be transformed by a transformation function f(·) to generate the allele-specific likelihood that h will be presented by MHC allele h.
[0345] VIII.B.1 Dependence Functions for Allelic Interaction Variables In one particular implementation mentioned throughout this specification, the dependency function g h (·) is x h k Each allele interaction variable in is compared with the parameter θ determined for the relevant MHC allele h. h with the corresponding parameters in the set is an affine function given by TIFF0007795291000022.tif6128.
[0346] In another specific implementation mentioned throughout this specification, the dependency function g h (·) is a network model NN with a set of nodes arranged in one or more layers. h (·), The network function is given by TIFF0007795291000023.tif6128. The nodes are connected by parameters θh A node may be connected to other nodes through connections, each having an associated parameter in a set of . The value at one particular node may be represented as the sum of the values of the nodes connected to the particular node, weighted by the associated parameter mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and process data having amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how these interactions affect peptide presentation.
[0347] Generally speaking, the network model NN h (·) can be structured as feedforward networks such as artificial neural networks (ANNs), convolutional neural networks (CNNs), deep neural networks (DNNs), and / or recurrent networks such as long short-term memory networks (LSTMs), bidirectional recurrent networks, and deep bidirectional recurrent networks.
[0348] In one example, which will be mentioned throughout the remainder of this specification, each MHC allele in h=1, 2,..., m is associated with a separate network model, NN h (·) denotes the output from the network model related to MHC allele h.
[0349] FIG. 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h=3. As shown in FIG. 5, the network model NN3(·) for MHC allele h=3 includes three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2), ..., θ3(10). The network model NN3(·) includes three allele interaction variables x3 for MHC allele h=3. k (1), x3 k (2) and x3 k (3) receives input values (individual data examples, including encoded polypeptide sequence data and any other training data used) and generates the value NN3(x3 k ) The network function may include one or more network models, each taking a different allele interaction variable as input.
[0350] In another example, the identified MHC alleles h=1,2,...,m are used to model the MHC alleles in a single network NN H (·) and NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θ h may correspond to the set of parameters for a single network model, and thus the parameters θ h The set of can be shared by all MHC alleles.
[0351] Figure 6A shows an exemplary network model NN shared by MHC alleles h = 1, 2, ..., m. H (·). As shown in Figure 6A, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. The network model NN3(·) has an allele interaction variable x3 for MHC allele h=3. k , and the value NN3(x3k ) and outputs m values.
[0352] In yet another example, a single network model NN H (·) is the allele interaction variable x for MHC allele h h k and the encoded protein sequence d h In such an example, the parameter θ h may again correspond to the set of parameters for a single network model, and thus the parameters θ h The set of MHC alleles can be shared by all MHC alleles. Therefore, in such an example, NNh(·) is a single network model with inputs [x h k d h ], a single network model NN H Such a network model is advantageous because it can correctly predict peptide presentation probabilities for MHC alleles that were unknown in the training data simply by identifying their protein sequences.
[0353] Figure 6B shows an exemplary network model NN shared by MHC alleles. H (·). As shown in Figure 6B, the network model NN H (·) takes as input the allele interaction variables and protein sequence of MHC allele h=3 and calculates the dependency score NN3 (x3 k ) is output.
[0354] In yet another example, the dependency function g h (·)teeth, TIFF0007795291000024.tif6128, where g' h (x h k ;θ' h) is an affine function, network function, etc., with a set of parameters θ'h, which represents the baseline probability of presentation for MHC allele h, and the bias parameter θ in the set of parameters for the allele interaction variables of the MHC alleles. h 0 accompanied by.
[0355] In another implementation, the bias parameter θ h 0 may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ for the MHC allele h h 0 is θ 遺伝子(h) 0 where gene (h) is the gene family of MHC allele h. For example, class I MHC alleles HLA-A*02:01, HLA-A*02:02, and HLA-A*02:03 may be assigned to the gene family "HLA-A," and the bias parameter θ for each of these MHC alleles may be h 0 As another example, if the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 are assigned to the "HLA-DRB" gene family, and the bias parameters θ for each of these MHC alleles are h 0 can be shared.
[0356] As an example, returning to equation (2), the affine dependency function g h Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000025.tif6128, where x3k is the allele interaction variable identified for MHC allele h=3 and θ3 is the set of parameters determined for MHC allele h=3 through loss function minimization.
[0357] As another example, let us consider the MHC allele h = 3 for peptide p among m = 4 different identified MHC alleles using separate network transformation functions gh(·). k The likelihood that will be presented is TIFF0007795291000026.tif6128 can be generated by k is the allele interaction variable identified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.
[0358] Figure 7 shows the correlation coefficients of peptide p associated with MHC allele h=3 using the exemplary network model NN3(·). k As shown in Figure 7, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0359] VIII.B.2. Per allele with non-allele interacting variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF0007795291000027.tif7131, peptide p k We model the estimated presentation likelihood uk, where w k is the peptide p k means the coded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interacting variable w Based on the set of allele-non-interacting variables w k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values of θ h and θw where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele.
[0360] Dependence function g w (w k ;θ w ) output is a measure of the peptide p expression by one or more MHC alleles based on the influence of allele-non-interacting variables. k represents a dependency score for an allele-non-interacting variable, indicating whether peptide p k The C-terminal flanking sequences and peptide p, which are known to positively influence the presentation of k If peptide p is bound, it may have a high value k The C-terminal flanking sequences and peptide p k If bound, it may have a low value.
[0361] According to equation (8), the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is the function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. w (·) is also applied to the coded version of the allele-non-interacting variable to generate a dependency score for the allele-non-interacting variable. Both scores are combined, and the combined score is used to estimate the association of peptide sequence p with MHC allele h. k is transformed by a transformation function f(·) to produce the allele-specific likelihoods that will be presented.
[0362] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (2) k the allele interaction variable x hk One may include the allele non-interacting variable wk in the prediction by adding The image can be given by TIFF0007795291000028.tif7128.
[0363] VIII.B.3 Dependence Functions for Allelic Non-Interacting Variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g for allelic non-interacting variables w (·) is an affine function, or a separate network model for the allelic non-interaction variables w k It can be a network function related to
[0364] Specifically, the dependency function g w (·) is w k The allele non-interaction variables in w with the corresponding parameters in the set It is an affine function given by TIFF0007795291000029.tif6128.
[0365] Dependence function g w (·) also corresponds to the parameter θ w the network model NN with relevant parameters in the set w (·), TIFF0007795291000030.tif6128. The network function may include one or more network models, each taking different allele-non-interacting variables as input.
[0366] In another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF0007795291000031.tif6128, where g' w (w k ;θ' w) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of m k is the peptide p k is the mRNA quantitative measurement for , h(·) is a function that transforms the quantitative measurement, and θ w m is a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measurement to generate a dependency score for the mRNA quantification measurement. In one particular embodiment mentioned throughout the remainder of this specification, h(·) is a log function, although in practice h(·) can be any one of a variety of different functions.
[0367] In yet another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF0007795291000032.tif6128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of k is the peptide p k is the indicator vector described in Section VII.C.2, which represents proteins and isoforms in the human proteome, and θ w o is the set of parameters in the set of parameters for the allele non-interacting variables that are combined with the indicator vector. k and parameter θ w o If the dimension of the set is significantly higher, TIFF0007795291000033.tif5128( Parameter regularization terms such as λ (representing L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined through an appropriate method.
[0368] In yet another example, the dependence function g on the allelic non-interacting variables w (·)teeth, TIFF0007795291000035.tif14131, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF0007795291000036.tif5128 is peptide p k is an indicator function equal to 1 if θ is derived from the source gene l as described above for allele-non-interacting variables, and θ w l is a parameter that indicates the "antigenicity" of the source gene l. In one variation, L is sufficiently large, and therefore the number of parameters θ w l=1, 2,...,L If is large enough, Parameter regularization term such as TIFF0007795291000037.tif5128 (where, TIFF0007795291000038.tif4128 can be added to the loss function when determining the parameter value (such as L1 norm, L2 norm, or a combination). The optimal value of the hyperparameter λ can be determined by an appropriate method.
[0369] In yet another example, the dependence function g on the allelic non-interacting variables w (·)teeth, TIFF0007795291000039.tif14158, where g' w (w k ;θ' w) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF0007795291000040.tif5128 shows peptide p as described above for allele-non-interacting variables. k is derived from source gene l, and peptide p k is an indicator function that is equal to 1 if originates from tissue type m, and θ w lm is a parameter indicating the antigenicity of the combination of source gene l and tissue type m. Specifically, the antigenicity of gene l in tissue type m may indicate the residual tendency of cells of tissue type m to present peptides derived from gene l after adjustment for RNA expression and peptide sequence context.
[0370] In one variation, L or M is sufficiently large, so that the number of parameters θ w lm=1, 2,...,LM If is large enough, Parameter regularization term such as TIFF0007795291000041.tif5128 (where, TIFF0007795291000042.tif4128 can be added to the loss function when determining the parameter values (such as L1 norm, L2 norm, or combination). The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the parameter values so that the coefficients for the same source gene do not differ significantly between tissue types. For example, a penalty term such as: TIFF0007795291000043.tif19128 (in the formula, TIFF0007795291000044.tif5128 is the average antigenicity across tissue types for source gene l) can add a penalty to the standard deviation of antigenicity across different tissue types in the loss function.
[0371] In practice, the dependence function g on the allelic non-interacting variables can be calculated by combining any of the additional terms in equations (10), (11), (12a) and (12b). w For example, the term h(·) representing the mRNA quantification measurement in equation (10) and the term representing the antigenicity of the source gene in equation (12) can be added together along with any other affine or network functions to generate a dependency function for the allele-non-interacting variables.
[0372] As an example, returning to equation (8), the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000045.tif6128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0373] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000046.tif6128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0374] FIG. 8 shows exemplary network models NN3(·) and NN w Peptide p associated with MHC allele h=3 using (·) kAs shown in Figure 8, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0375] VIII.C. Multi-Allele Models The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof. [Example]
[0376] VIII.C.1. Example 1: Maximum Per Allele Model In one implementation, the training module 316 trains peptides p associated with a set of MHC alleles H. k The estimated presentation likelihood u k is the presentation likelihood determined for each of the MHC alleles h in set H determined based on cells expressing a single allele, as explained above in conjunction with equations (2)-(11). TIFF0007795291000047.tif4128. Specifically, the presentation likelihood u k teeth, In one implementation, the function is a maximum function, as shown in equation (12), where the proposed likelihood u kcan be determined as the maximum of the presentation likelihood for each MHC allele h in set H. TIFF0007795291000049.tif6128
[0377] VIII.C.2. Example 2.1: Sum Function Model In one implementation, the training module 316 trains peptides p k The estimated presentation likelihood u k of, TIFF0007795291000050.tif14128, where element a h k is the peptide sequence p k 1 for multiple MHC alleles H associated with x h k is the peptide p k and the coded allele interaction variables for the corresponding MHC alleles. The parameter θ for each MHC allele h h The set of values is θ h The dependence function g can be determined by minimizing a loss function with respect to i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:
[0378] According to equation (13), the peptide sequence p k The likelihood that a given allele will be presented by one or more MHC alleles h is given by the dependency function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate a corresponding score for the allele interaction variable. The scores for each MHC allele h are combined to generate a corresponding score for the peptide sequence p k is transformed by a transformation function f(·) to produce the presentation likelihood that MHC allele H will be presented by the set of MHC alleles H.
[0379] The model presented in equation (13) is that for each peptide p k It differs from the allele-by-allele model of equation (2) in that the number of relevant alleles for a can be greater than 1. In other words, h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0380] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000051.tif6128 can be generated by k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0381] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000052.tif6128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.
[0382] FIG. 9 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). kAs shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0383] VIII.C.3. Example 2.2: Sum Function Model with Allelic Non-Interacting Variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF0007795291000053.tif14142, peptide p k The estimated presentation likelihood u k where w k is the peptide p k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values of θ h and θ w The dependence function g can be determined by minimizing a loss function with respect to i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:
[0384] Therefore, according to equation (14), one or more MHC alleles H can bind to a peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p kto generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores are combined and the combined score is used to estimate the association of the peptide sequence p with the MHC allele H. k is transformed by a transformation function f(·) to produce the presentation likelihood that
[0385] In the model presented in equation (14), each peptide p k The number of relevant alleles for a can be greater than 1. h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with
[0386] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000054.tif6128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0387] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000055.tif6128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0388] FIG. 10 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 10, the network model NN2(·) generates a representation likelihood for the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. The network model NN3(·) generates the allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.
[0389] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (15) k the allele interaction variable x h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF0007795291000056.tif14134.
[0390] VIII.C.4. Example 3.1: Model with Implicit Allele-by-Allele Likelihood In another implementation, the training module 316 k The estimated presentation likelihood u k of, TIFF0007795291000057.tif7142, where element a h k is the peptide sequence p k Multiple MHC alleles associated with TIFF0007795291000058.tif4128 is 1, and u' k h is the implicit allele-specific presentation likelihood for MHC allele h, and vector v has elements v h But, a h k ·u' k h where s(·) is a vector corresponding to v, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values of the input within a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, although it will be recognized that in other embodiments, s(·) may be any function, such as a maximum function. A set of values for the parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.
[0391] The presentation likelihood in the presentation model of equation (17) is the likelihood that each peptide p is presented by an individual MHC allele h. k The implicit allele-specific presentation likelihood u' corresponds to the likelihood that k h The implicit per-allele likelihood differs from the per-allele presentation likelihood of Section VIII.B in that the parameters for the implicit per-allele likelihood can be learned from a multi-allelic setting, in addition to a single-allelic setting, where the direct association between the presented peptide and the corresponding MHC allele is unknown. Thus, in a multi-allelic setting, the presentation model is based on the likelihood of the peptide p kNot only can we estimate whether peptide p is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are present in peptide p k individual likelihoods indicating which TIFF0007795291000059.tif4128 can also be provided. The advantage of this is that the proposed model can generate implicit likelihoods without training data for cells expressing a single MHC allele.
[0392] In one particular implementation that will be mentioned throughout the remainder of this specification, r(·) is a function with range [0,1]. For example, r(·) is a clip function: r(z)=min(max(z,0),1) may be the minimum value between z and 1, which represents the likelihood u k In another implementation, r(·) is chosen as r(z)=tanh(z) where the domain z is greater than or equal to 0.
[0393] VIII.C.5. Example 3.2: Sum of Functions Model In one particular implementation, s(·) is a summation function, and the presentation likelihood is given by summing the implicit per-allele presentation likelihoods. TIFF0007795291000060.tif16128
[0394] In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: The likelihood of the proposed model is shown in Figure 1. Let it be estimated by TIFF0007795291000062.tif14129.
[0395] According to equation (19), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h(·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. Each dependency score can be generated by first applying the implicit per-allele presentation likelihood u' k h The allele likelihood u' is transformed by the function f(·) to generate k h are combined and a clipping function is applied to the combined likelihood to clip the values into the range [0,1] to produce a peptide sequence p k A presentation likelihood can be generated that g will be presented by a set of MHC alleles H. The dependency function g h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:
[0396] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000063.tif7128, where x2 k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.
[0397] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000064.tif7128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.
[0398] FIG. 11 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) are then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.
[0399] In another implementation, if the prediction is made in terms of the log of the mass analysis ion current, then r(·) is the log function and f(·) is the exponential function.
[0400] VIII.C.6. Example 3.3: Sum of Functions Model with Allelic Non-Interacting Variables In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: The likelihood of the image generated by TIFF0007795291000065.tif7128 is As generated by TIFF0007795291000066.tif14131, the effects of allelic non-interacting variables are incorporated into peptide presentation.
[0401] According to equation (21), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h(·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele non-interaction variables to generate dependency scores for the allele non-interaction variables. The scores of the allele non-interaction variables are combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by the function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined, and a clipping function is applied to the combined output to clip values into the range [0,1] to determine the likelihood of presentation of the peptide sequence p by the MHC allele H. k A likelihood of being presented can be generated. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:
[0402] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000067.tif7128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0403] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF0007795291000068.tif7139, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.
[0404] FIG. 12 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 12, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·). The network model NN3(·) generates an allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated, which is also the same network model NN w (·) output NN w (w k ) and mapped by a function f(·). Both outputs are combined to give the estimated presentation likelihood u k Generate.
[0405] In another implementation, the implicit per-allele presentation likelihood for MHC allele h can be calculated as: The likelihood of the proposed model is shown in Figure 1. Generated by TIFF0007795291000070.tif14128.
[0406] VIII.C.7. Example 4: Quadratic Model In one implementation, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k teeth, TIFF0007795291000071.tif14137, where the element u' k h is the implicit per-allele presentation likelihood for MHC allele h. A set of values for parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit per-allele presentation likelihood can be in any of the forms shown in equations (18), (20), and (22) above.
[0407] In one embodiment, the model of equation (23) is k However, there is a possibility that a given antigen may be simultaneously presented by two MHC alleles, which may imply that presentation by the two HLA alleles is statistically independent.
[0408] According to equation (23), one or more MHC alleles H bind to the peptide sequence p k The presentation likelihood is calculated by combining the implicit per-allele presentation likelihood and the MHC allele H k Each pair of MHC alleles is assigned to a peptide p such that it generates a presentation likelihood that p will be presented. k can be generated by subtracting from the sum the likelihood that
[0409] For example, the affine transformation function g h The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF0007795291000072.tif6128 can be generated byk , x3 k are the allele interaction variables identified for HLA alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for HLA alleles h=2, h=3.
[0410] As another example, the network transformation function g h (·), g w The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF0007795291000073.tif6137, where NN2(·) and NN3(·) are the network models specified for HLA alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for HLA alleles h=2 and h=3.
[0411] IX. Example 5: Prediction Module The prediction module 320 receives sequence data and selects candidate neoantigens in the sequence data using the proposed model. Specifically, the sequence data may be DNA sequences, RNA sequences, and / or protein sequences extracted from tumor tissue cells of a patient. The prediction module 320 converts the sequence data into a plurality of peptide sequences p having 8-15 amino acids for MHC-I or 6-30 amino acids for MHC-II. k For example, the prediction module 320 processes a predetermined sequence TIFF0007795291000074.tif4128, three peptide sequences with nine amino acids TIFF0007795291000075.tif9128. In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by comparing sequence data extracted from a patient's normal tissue cells with sequence data extracted from the patient's tumor tissue cells to identify segments that have one or more mutations.
[0412] The prediction module 320 applies one or more presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation models to the candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences with an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences with the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the selected candidate neoantigens for a given patient can be injected into the patient to induce an immune response.
[0413] X. Example 6: Patient Selection Module The patient selection module 324 selects a subset of patients for vaccine treatment based on whether the patient meets the inclusion criteria. In one embodiment, the inclusion criteria are determined based on the patient's likelihood of presenting the neoantigen candidates generated by the presentation model. By adjusting the inclusion criteria, the patient selection module 324 can adjust the number of patients who receive the vaccine based on the patient's likelihood of presenting the neoantigen candidates. Specifically, strict inclusion criteria may result in a smaller number of patients being treated with the vaccine, but a higher proportion of vaccine-treated patients receiving effective treatment (e.g., delivering one or more tumor-specific neoantigens (TSNAs)). In contrast, looser inclusion criteria may result in a larger number of patients being treated with the vaccine, but a lower proportion of vaccine-treated patients receiving effective treatment. The patient selection module 324 varies the inclusion criteria based on a desired balance between a target proportion of patients receiving the vaccine and the proportion of patients receiving effective treatment as a result of the vaccine treatment.
[0414] In one embodiment, a patient is associated with a corresponding therapeutic subset of v neoantigen candidates that can potentially be included in a personalized vaccine for the patient, having a vaccine volume v. In one embodiment, the therapeutic subset for a patient is the neoantigen candidate with the highest likelihood of presentation as determined by the presentation model. For example, if a vaccine can include v=20 epitopes, the vaccine can include a therapeutic subset for each patient with the highest likelihood of presentation as determined by the presentation model. However, it will be appreciated that in other embodiments, the therapeutic subset for a patient can be determined based on other methods. For example, the therapeutic subset for a patient can be randomly selected from the set of neoantigen candidates for that patient, or can be determined in part based on a combination of factors including prior art models that model the binding affinity or stability of peptide sequences, or presentation likelihoods obtained from presentation models and affinity or stability information for those peptide sequences.
[0415] In one embodiment, the patient selection module 324 determines that a patient meets the inclusion criteria if the patient's tumor mutation burden is equal to or higher than a minimum mutation burden. A patient's tumor mutation burden (TMB) indicates the total number of nonsynonymous mutations in the tumor exome. In one embodiment, the patient selection module 324 selects patients suitable for vaccine treatment if the patient's absolute TMB number is equal to or higher than a predetermined threshold. In another implementation, the patient selection module 324 selects patients suitable for vaccine treatment if the patient's TMB is within a threshold percentile among the TMBs determined for the set of patients.
[0416] In another embodiment, the patient selection module 324 determines that a patient meets the inclusion criteria if the patient's utility score based on the patient's therapeutic subset is equal to or greater than a minimum utility score. In one embodiment, the utility score is a measure of the estimated number of presented antigens from the therapeutic subset.
[0417] The expected number of presented neoantigens can be predicted by modeling the presentation of neoantigens as random variables with one or more probability distributions. In one implementation, the utility score for patient i is the expected number of presented neoantigen candidates from the treatment subset, or a specific function thereof. As an example, the presentation of each neoantigen can be modeled as a Bernoulli random variable, where the probability of presentation (success) is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation (success) of each neoantigen candidate is given by the presentation likelihood of the neoantigen candidate. Specifically, the probability of presentation of each neoantigen candidate is given by the likelihood of presentation of each neoantigen candidate, u i1 , u i2 , …, u iv V types of neoantigen candidates p i1 , p i2 , …, p iv Treatment subset S i Regarding neoantigen candidate p ij The presentation of the random variable A ij where: The expected number of neoantigens presented is given by the sum of the likelihoods of presentation of each neoantigen candidate. In other words, the utility score for patient i is expressed as: The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility for vaccine treatment.
[0418] In another implementation, the utility score for patient i is the probability that at least a threshold number of neoantigens k are presented. In one example, a therapeutic subset of neoantigen candidates S i The number of presented antigens in is modeled as a Poisson binomial random variable where the probability of presentation (success) is given by the likelihood of presentation of each of the epitopes. In particular, the number of presented antigens for patient i is determined by the random variable N i can be given by: TIFF0007795291000078.tif14128In the formula, PBD(·) denotes the Poisson binomial distribution. The probability that at least a threshold number of neoantigens k are presented is given by the number of presented antigens N. iis given by the operation of the probability that k is equal to or greater than k. In other words, the utility score for patient i is expressed as: The patient selection module 324 selects a subset of patients with a utility score equal to or greater than the minimum utility for vaccine treatment.
[0419] In another embodiment, the number of neoantigens in a therapeutic subset for a patient need not be limited to the vaccine dose v, and the patient selection module 324 can select patients using a utility score determined based on any set of candidate neoantigens for that patient. For example, the utility score can be determined based on all mutations or candidate neoantigens identified for that patient. The utility score can be generated, for example, using the method described in conjunction with equations (24)-(27), where v is a variable v(i) that depends on patient i, and indicates the total number of mutations or candidate neoantigens identified for that patient.
[0420] In another implementation, the utility score for patient i is calculated based on a therapeutic subset S of neoantigen candidates that have binding affinities or predicted binding affinities below a fixed threshold (e.g., 500 nM) for one or more of the patient's HLA alleles. i The threshold is the number of neoantigens in the target region. In one example, the fixed threshold ranges from 1000 nM to 10 nM. In some cases, the utility score may only count neoantigens detected as expressed by RNA-seq.
[0421] In another implementation, the utility score for patient i is calculated based on the therapeutic subset S of candidate neoantigens whose binding affinity to one or more HLA alleles for that patient is less than or equal to a threshold percentile of the binding affinity of a random peptide to that HLA allele. iThe threshold percentile is the number of neoantigens in the sample. In one example, the threshold percentile ranges from the 10th percentile to the 0.1th percentile. In some cases, the utility score may only count neoantigens detected as expressed by RNA-seq.
[0422] It will be appreciated that the example utility scores described with respect to equations (25) and (27) are for illustrative purposes only, and that the patient selection module 324 may use other statistics or probability distributions to generate utility scores.
[0423] XI. Example 7: Neoantigen Loading for Immune Checkpoint Inhibitor Therapy and Other Immunotherapies The patient selection module 324 can also use the utility score defined in Section X above to select patients for immune checkpoint inhibitor therapy (e.g., PD-1, CTLA4) or any other immunotherapy for which neoantigen load may be associated with efficacy, including immunostimulatory agents, immunostimulatory molecule agonists (e.g., CD40), oncolytic viruses (e.g., T-VEC), neoantigen- or other cancer antigen-containing therapeutic vaccines, neoantigen- or other cancer antigen-targeted adoptive cell therapy, tumor microenvironment modulators (e.g., TGFβ), or any combination of these with immune checkpoint inhibitors.
[0424] For example, in some embodiments, the immunostimulatory agent is an agent that blocks signaling of an inhibitory receptor or its ligand on an immune cell. In some embodiments, the inhibitory receptor or ligand is selected from CTLA-4, PD-1, PD-L1, LAG-3, Tim3, TIGIT, neuritin, BTLA, KIR, and combinations thereof. In some aspects, the agent is selected from an anti-PD-1 antibody (e.g., pembrolizumab or nivolumab), an anti-PD-L1 antibody (e.g., atezolizumab), an anti-CTLA-4 antibody (e.g., ipilimumab), and combinations thereof. In some aspects, the agent is pembrolizumab. In some aspects, the agent is nivolumab. In some aspects, the agent is atezolizumab.
[0425] In some embodiments, the therapeutic agent is an agent that inhibits the interaction of PD-1 and PD-L1. In some aspects, the additional agent that inhibits the interaction of PD-1 and PD-L1 is selected from an antibody, a peptidomimetic, and a small molecule. In some aspects, the additional agent that inhibits the interaction of PD-1 and PD-L1 is selected from pembrolizumab, nivolumab, atezolizumab, avelumab, durvalumab, BMS-936559, sulfamonomethoxine 1, and sulfamethizole 2. In some embodiments, the additional agent that inhibits the interaction of PD-1 and PD-L1 is any therapeutic agent known in the art having such activity, for example, as described in Weinmann et al., Chem Med Chem, 2016, 14:1576 (DOI: 10.1002 / cmdc.201500566), which is incorporated by reference in its entirety.
[0426] In some embodiments, the immunostimulatory agent is an agonist of a costimulatory receptor on an immune cell. In some aspects, the costimulatory receptor is selected from OX40, ICOS, CD27, CD28, 4-1BB, and CD40. In some embodiments, the agonist is an antibody.
[0427] In some embodiments, the immunostimulatory agent is a cytokine, hi some aspects, the cytokine is selected from IL-2, IL-5, IL-7, IL-12, IL-15, IL-21, and combinations thereof.
[0428] In some embodiments, the immunostimulant is an oncolytic virus. In some aspects, the oncolytic virus is selected from herpes simplex virus, vesicular stomatitis virus, adenovirus, Newcastle disease virus, vaccinia virus, and Maraba virus.
[0429] In some embodiments, the immunostimulatory agent is a chimeric antigen receptor-bearing T cell (CAR-T cell). In some embodiments, the immunostimulatory agent is a bispecific or multispecific T cell-directed antibody. In some embodiments, the immunostimulatory agent is an anti-TGF-β antibody. In some embodiments, the immunostimulatory agent is a TGF-β trap.
[0430] In some embodiments, the therapeutic agent is a vaccine against a tumor antigen. The vaccine can be targeted to any suitable antigen, provided that the antigen is present in the tumor treated by the methods provided herein. In some aspects, the tumor antigen is a tumor antigen that is overexpressed relative to its expression level in normal tissue. In some aspects, the tumor antigen is selected from a cancer testis antigen, a differentiation antigen, NY-ESO-1, MAGE-A1, MART, and combinations thereof. In some embodiments, the therapeutic agent is a vaccine against one or more neoantigens. The neoantigens in the vaccine can be identified by the methods provided herein.
[0431] In particular, the patient selection module 324 determines a neoantigen load, which indicates the total expected number of presented neoantigens for each patient. Patients with a neoantigen load that meets the inclusion criteria can be administered checkpoint inhibitor therapy. For example, such therapy can be administered to patients with a neoantigen load above a predetermined threshold. In one embodiment, the neoantigen load is a utility score as described in Section XI, where v is the total number of mutations or candidate neoantigens identified for the patient, rather than a subset of candidate antigens for that patient.
[0432] A high neoantigen burden relative to the median in a particular tumor may indicate that a subject with that tumor is more likely to benefit from treatment with a checkpoint inhibitor, such as anti-CTLA4, anti-PD1, and / or anti-PDL1. For example, neoantigen burden may be a better indicator of efficacy of checkpoint inhibitors compared to mutational burden, because neoantigens are generally presented on the surface of tumor cells and are more likely to be recognized by T cells with higher activity against tumors following checkpoint inhibitor therapy.
[0433] In another embodiment, the patient selection module 324 can use a utility score generated from a combination of one or more of the following characteristics: predicted HLA class I neoantigen load, predicted HLA class II neoantigen load, and tumor mutational load. The predicted HLA class I neoantigen load for a patient is the neoantigen load for that patient's set of class I HLA alleles, indicating the total expected number of neoantigens presented on that patient's class I HLA alleles. The predicted HLA class II neoantigen load for a patient is the neoantigen load for that patient's set of class II HLA alleles, indicating the total expected number of neoantigens presented on that patient's class II HLA alleles. For example, the utility score can be calculated as f(class I neoantigen load, class II neoantigen load, tumor mutational load; b), where f(·) is a function parameterized by the set of machine-learned parameters b. The set of machine-learned parameters b can depend on the tumor type (e.g., b can be different for melanoma and non-small cell lung cancer).
[0434] In another embodiment, the patient selection module 324 can use a utility score that incorporates information about immunogenic tumor antigens other than neoantigens. Examples of immunogenic tumor antigens other than neoantigens include cancer-germline antigens (CGAs, e.g., MAGEA3), differentiation antigens (e.g., tyrosinase), and antigens overexpressed in tumors (e.g., CEA). The expression levels of these antigens can be determined using at least tumor RNA sequencing data, and the expected number of HLA class I or class II epitopes from these genes presented by HLA alleles in the patient's tumor can be determined by applying a presentation model to each peptide from the set of tumor antigens using the RNA sequencing data for each gene. These presentation likelihoods can be incorporated into a utility score calculated as f(class I neoantigen load, class II neoantigen load, tumor mutation load, class I non-neoantigen tumor antigen load, class II non-neoantigen tumor antigen load; b), where f(·) is a function parameterized by the set of machine-learned parameters b. The set of machine-learned parameters b may depend on the tumor type (e.g., b may be different for melanoma and non-small cell lung cancer).
[0435] A higher efficacy score indicates that the tumor presents more HLA epitopes that are recognized as foreign or non-self by the patient's immune system. Patients with tumors that present more non-self HLA epitopes may be more likely to benefit from checkpoint inhibitors or other immunotherapy, as these tumors are more likely to be recognized by T cells with higher activity against the tumor after immunotherapy.
[0436] The utility score described in Section X above can also be adapted to select patients for treatment with adoptive cell therapy (e.g., expanded TILs, CAR-T, or engineered TCRs) by using f(Class I neoantigen burden, Class II neoantigen burden, tumor mutation burden, Class I non-neoantigen tumor antigen burden, Class II non-neoantigen tumor antigen burden) (wherein Class I and Class II neoantigens and non-neoantigens are only considered as present or predicted to be present in the adoptive immunotherapy). For example, in the case of engineering a TCR therapy against a single neoantigen epitope, f can be reduced to the likelihood of presentation of that single epitope.
[0437] XII. Example 8: Experimental Results Demonstrating Exemplary Patient Selection Performance The validity of the patient selection described in Section X is validated by selecting patients from a set of simulated patients, each associated with a test set of simulated neoantigen candidates, for which a subset of the simulated neoantigens is known to be represented in the mass spectrometry data. Specifically, each simulated neoantigen candidate in the test set is associated with a label indicating whether that neoantigen is represented in the mass spectrometry data set of the multi-allelic JY cell line HLA-A*02:01 and HLA-B*07:02 from the Bassani-Sternberg dataset (dataset "D1") (data available at www.ebi.ac.uk / pride / archive / projects / PXD0000394). As described in more detail below in conjunction with Figure 13A, a number of neoantigen candidates for the simulated patients are sampled from the human proteome based on known frequency distributions of mutation burden in non-small cell lung cancer (NSCLC) patients.
[0438] Allele-specific presentation models for the same HLA alleles are trained using a training set that is a subset of the mass spectrometry data for the single alleles HLA-A*02:01 and HLA-B*07:02 from the IEDB dataset (dataset "D2") (data can be found at http: / / www.iedb.org / doc / mhc_ligand_full.zip). Specifically, the presentation model for each allele is trained using the network dependency function g h (·) and g w The allele-specific models were modeled as shown in equation (8), incorporating the allele-specific expression (·) and exponent function f(·). The presentation model for the HLA-A*02:01 allele generates the presentation likelihood of a particular peptide being presented on the HLA-A*02:01 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables. The presentation model for the HLA-B*07:02 allele generates the presentation likelihood of a particular peptide being presented on the HLA-B*07:02 allele, given the peptide sequence as the allele interaction variable and the N- and C-terminal flanking sequences as the allele non-interaction variables.
[0439] As disclosed in the following examples with reference to Figures 13A-13G, different models, such as a presentation model trained for peptide binding prediction and a prior art model, are applied to a test set of neoantigen candidates for each simulated patient to identify different therapeutic subsets for the patient based on the predictions. Patients who meet the inclusion criteria for vaccine treatment are selected and associated with a personalized vaccine containing epitopes in the patient's therapeutic subset. The size of the therapeutic subsets varies depending on the different vaccine doses. No overlap is introduced between the training set used to train the presentation model and the test set of simulated neoantigen candidates.
[0440] In the following example, the proportion of selected patients with at least a certain number of presented neoantigens among the epitopes included in the vaccine is analyzed. This statistic indicates the effectiveness of the simulated vaccine in delivering potential neoantigens that will elicit an immune response in patients. Specifically, simulated neoantigens in a test set are presented if the neoantigen is presented in mass spectrometry dataset D2. A high proportion of patients with presented neoantigens indicates the likelihood of successful treatment with the neoantigen vaccine by inducing an immune response.
[0441] XII.A. Example 8A: Frequency Distribution of Tumor Mutation Burden in NSCLC Cancer Patients Figure 13A shows the sample frequency distribution of mutation burden in NSCLC patients. Mutation burden and mutations in different tumor types, including NSCLC, can be found, for example, in the Cancer Genome Atlas (TCGA) (https: / / cancergenome.nih.gov). The X-axis represents the number of nonsynonymous mutations for each patient, and the Y-axis represents the proportion of sample patients with a specific number of nonsynonymous mutations. The sample frequency distribution in Figure 13A shows a range of 3 to 1786 mutations, with 30% of patients having fewer than 100 mutations. Although not shown in Figure 13A, studies have shown that mutation burden is higher in smokers compared to nonsmokers, and that mutation burden can be a strong indicator of neoantigen burden in patients.
[0442] As introduced at the beginning of Section XI above, each of the simulated patient populations is associated with a test set of neoantigen candidates. Each patient's test set is determined by selecting the mutation load m from the frequency distribution shown in Figure 13A for each patient. iThe D1 dataset is generated by sampling the D1 sequence. For each mutation, a 21-mer peptide sequence from the human proteome is randomly selected to represent the mutant sequence to be simulated. A test set of candidate neoantigen sequences is generated for patient i by identifying each (8, 9, 10, 11)-mer peptide sequence across the mutations in the 21-mer. Each candidate neoantigen is associated with a label indicating whether the candidate neoantigen sequence is present in the mass spectrometry D1 dataset. For example, candidate neoantigen sequences present in dataset D1 can be associated with the label "1," and sequences not present in dataset D1 can be associated with the label "0." As described in more detail below, Figures 13B-13G show experimental results of patient selection based on the neoantigens presented by patients in the test set.
[0443] XII.B. Example 8B: Proportion of Selected Patients with Neoantigen Presentation Based on Inclusion Criteria for Tumor Mutational Burden Figure 13B shows the number of neoantigens presented in the simulated vaccine for patients selected based on the inclusion criteria of whether the patient met a minimum tumor mutation burden. The proportion of selected patients with at least a certain number of presented neoantigens in the corresponding study is identified.
[0444] In Figure 13B, the x-axis shows the proportion of patients excluded from vaccine treatment based on tumor mutation burden, as indicated by the label "Minimum Number of Mutations." For example, the data point at "Minimum Number of Mutations" 200 indicates that the patient selection module 324 selected only a subset of simulated patients with a tumor mutation burden of at least 200 mutations. As another example, the data point at "Minimum Number of Mutations" 300 indicates that the patient selection module 324 selected a lower proportion of simulated patients with at least 300 mutations. The y-axis shows the proportion of selected patients associated with at least a certain number of presented neoantigens in the test set without vaccine dose v. Specifically, the top plot shows the proportion of selected patients presenting at least one neoantigen, the middle plot shows the proportion of selected patients presenting at least two antigens, and the bottom plot shows the proportion of selected patients presenting at least three antigens.
[0445] As shown in Figure 13B, the proportion of selected patients with presented neoantigens significantly increased with increasing tumor mutation burden, indicating that tumor mutation burden as an inclusion criterion can be effective in selecting patients in whom neoantigen vaccines are likely to induce effective immune responses.
[0446] XII.C. Example 8C: Comparison of Neoantigen Presentation in Vaccines Identified by Presentation Models and Prior Art Models Figure 13C compares the number of presented neoantigens in simulated vaccines between selected patients associated with vaccines containing therapeutic subsets identified based on the presented model and selected patients associated with vaccines containing therapeutic subsets identified by prior art models. The plot on the left assumes a limiting vaccine volume of v=10, and the plot on the right assumes a limiting vaccine volume of v=20. Patients are selected based on a utility score indicating the expected number of presented neoantigens.
[0447] In Figure 13C, the solid lines indicate patients associated with a vaccine containing therapeutic subsets identified based on presentation models for the alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient is identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted lines indicate patients associated with a vaccine containing therapeutic subsets identified based on the prior art model NETMHCpan for the single allele HLA-A*02:01. Implementation details for NETMHCpan are provided at http: / / www.cbs.dtu.dk / services / NetMHCpan. The therapeutic subset for each patient is identified by applying the NETMHCpan model to the sequences in the test set and identifying the v neoantigen candidates with the highest predicted binding affinity. The x-axis of both graphs indicates the proportion of patients excluded from vaccine treatment based on the expected utility score, which indicates the expected number of presented neoantigens in the therapeutic subsets identified based on the presentation model. The expected utility score is determined as described in Section X in connection with Equation (25). The y-axis represents the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens) contained in the vaccine.
[0448] As shown in Figure 13C, a significantly higher proportion of patients associated with vaccines containing therapeutic subsets based on the presentation model receive vaccines containing presented neoantigens than patients associated with vaccines containing therapeutic subsets based on the prior art model. For example, as shown in the graph on the right, 80% of selected patients associated with vaccines based on the presentation model receive at least one presented neoantigen in the vaccine, compared to only 40% of selected patients associated with vaccines based on the prior art model. These results demonstrate that the presentation model described herein is effective in selecting neoantigen candidates for vaccines that are likely to elicit an immune response to treat tumors.
[0449] XII.D. Example 8D: Effect of HLA Coverage on Neoantigen Presentation of Vaccines Identified by Presentation Models Figure 13D compares the number of neoantigens presented in simulated vaccines between selected patients associated with a vaccine containing a therapeutic subset identified based on a single allele-per-presentation model for HLA-A*02:01 and selected patients associated with a vaccine containing a therapeutic subset identified based on both an allele-per-presentation model for HLA-A*02:01 and HLA-B*07:02. Vaccine volume is set to v = 20 epitopes. For each experiment, patients are selected based on expected utility scores determined based on different therapeutic subsets.
[0450] In Figure 13D, the solid line indicates patients associated with a vaccine containing therapeutic subsets based on both presentation models for HLA alleles HLA-A*02:01 and HLA-B*07:02. The therapeutic subset for each patient was identified by applying each of the presentation models to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. The dotted line indicates patients associated with a vaccine containing therapeutic subsets based on a single presentation model for HLA allele HLA-A*02:01. The therapeutic subset for each patient was identified by applying a presentation model for only a single HLA allele to the sequences in the test set and identifying the v neoantigen candidates with the highest presentation likelihood. In the solid plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score for the therapeutic subsets identified by both presentation models. In the dotted plot, the x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score for the therapeutic subset identified by a single presentation model. The y-axis shows the proportion of selected patients presenting at least a specified number of neoantigens (1, 2, or 3 neoantigens).
[0451] As shown in Figure 13D, patients associated with vaccines containing therapeutic subsets identified by presentation models for both HLA alleles presented neoantigens at a significantly higher rate than patients associated with vaccines containing therapeutic subsets identified by a single presentation model. These results demonstrate the importance of establishing presentation models with high HLA coverage.
[0452] XII.E. Example 8E: Comparison of Neoantigen Presentation in Patients Selected by Tumor Mutational Burden and Expected Number of Presented Neoantigens Figure 13E compares the number of neoantigens presented in simulated vaccines between patients selected based on tumor mutation burden and patients selected by expected utility score, which is determined based on the therapeutic subset identified by the presentation model with v = 20 different epitope sizes.
[0453] In Figure 13E, the solid lines indicate patients selected based on the expected utility score associated with a vaccine containing the therapeutic subset identified by the proposed model. The therapeutic subset for each patient was identified by applying each of the proposed models to the sequences in the test set and identifying the v = 20 neoantigen candidates with the highest likelihood of presentation. The therapeutic utility score was determined based on the likelihood of presentation of the therapeutic subset identified in Section X based on Equation (25). The dotted lines indicate patients selected based on the tumor mutation burden associated with a vaccine containing the therapeutic subset identified by the proposed model. The x-axis indicates the proportion of patients excluded from vaccine treatment based on the expected utility score in the solid plot and the proportion of patients excluded based on tumor mutation burden in the dotted plot. The y-axis indicates the proportion of selected patients who will receive a vaccine containing at least a specified number of presented neoantigens (1, 2, or 3 neoantigens).
[0454] As shown in Figure 13E, patients selected based on expected utility score receive a vaccine containing a higher proportion of presented neoantigens than patients selected based on tumor mutation burden. However, patients selected based on tumor mutation burden receive a vaccine containing a higher proportion of presented neoantigens than patients not selected. Thus, while tumor mutation burden is an effective patient selection criterion for effective neoantigen vaccine therapy, expected utility score is more effective.
[0455] XIII. Exemplary Computer Figure 14 illustrates an exemplary computer 1400 for implementing the entities shown in Figures 1 and 3. Computer 1400 includes at least one processor 1402 coupled to a chipset 1404. Chipset 1404 includes a memory controller hub 1420 and an input / output (I / O) controller hub 1422. Memory 1406 and a graphics adapter 1412 are coupled to memory controller hub 1420, and a display 1418 is coupled to graphics adapter 1412. Storage device 1408, input device 1414, and network adapter 1416 are coupled to I / O controller hub 1422. Other embodiments of computer 1400 have different architectures.
[0456] The storage device 1408 is a non-transitory computer-readable storage medium, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 1406 holds instructions and data used by the processor 1402. The input interface 1414 is a touchscreen interface, a mouse, a trackball, or other type of pointing device, a keyboard, or some combination thereof, used to input data into the computer 1400. In some embodiments, the computer 1400 may be configured to receive input (e.g., commands) from the input interface 1414 via gestures from a user. The graphics adapter 1412 displays images and other information on the display 1418. The network adapter 1416 couples the computer 1400 to one or more computer networks.
[0457] The computer 1400 is adapted to execute computer program modules to provide the functionality described herein. As used herein, the term "module" refers to computer program logic used to provide particular functionality. Thus, modules can be implemented in hardware, firmware, and / or software. In one embodiment, program modules are stored in the storage device 1408, loaded into the memory 1406, and executed by the processor 1402.
[0458] 1 can vary depending on the implementation and processing power required by the entity. For example, presentation specification system 160 can run on a single computer 1400 or on multiple computers 1400 communicating with each other over a network, such as in a server farm. Computer 1400 may lack some of the components described above, such as graphics adapter 1412 and display 1418.
[0459] References TIFF0007795291000080.tif223158TIFF0007795291000081.tif238159TIFF0007795291 000082.tif238159TIFF0007795291000083.tif233160TIFF0007795291000084.tif52153
Claims
1. 1. A method for identifying a subset of patients suitable for treatment, comprising: obtaining, for each patient, at least one of exome, transcriptome, or whole genome tumor nucleotide sequencing data from the patient's tumor cells and normal cells, wherein the tumor nucleotide sequencing data is used to obtain a peptide sequence for each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen for the patient includes at least one alteration that makes it different from a corresponding wild-type parent peptide sequence identified from the patient's normal cells; generating, for each patient, a set of numerical presentation likelihoods for the set of neoantigens for the patient by inputting the peptide sequences of each of the set of neoantigens into a machine-learned presentation model, each presentation likelihood representing the likelihood that the corresponding neoantigen will be presented by one or more MHC alleles on the surface of tumor cells in the patient, the set of presentation likelihoods being identified based at least on mass spectrometry data; identifying for each patient one or more neoantigens from said set of neoantigens for said patient; determining for each patient a utility score indicative of the estimated number of neoantigens presented on the surface of tumor cells of said patient, as determined by the corresponding presentation likelihood for said one or more neoantigens for said patient; selecting a subset of patients suitable for treatment, wherein each patient in said subset of patients is associated with a utility score that meets predetermined inclusion criteria; Including, The machine-learned presentation model is a label obtained by mass spectrometry to measure the presence of a peptide bound to at least one MHC allele identified as being present in at least one of the plurality of samples; a training peptide sequence containing information about a set of amino acids constituting the training peptide sequence and the positions of the amino acids within the training peptide sequence; at least one MHC allele associated with said training peptide sequence; a plurality of parameters determined based at least on the training data set, the plurality of parameters including: A function that represents the relationship between the peptide sequence and the likelihood of presentation based on the plurality of parameters. Including, The method.
2. 2. The method of claim 1, wherein identifying the one or more neoantigens for the patient comprises selecting a subset of neoantigens in the set of neoantigens for the patient.
3. 3. The method of claim 2, wherein the subset of neoantigens are neoantigens with the highest presentation likelihood among the set of presentation likelihoods for the patient.
4. 10. The method of claim 1, further comprising identifying, for each patient within the selected subset of patients, one or more T cells or T cell receptors that are antigen-specific for at least one of the one or more neoantigens identified for the patient.
5. 2. The method of claim 1, wherein identifying one or more neoantigens for the patient comprises selecting the entire set of identified neoantigens for the patient.
6. 2. The method of claim 1, wherein selecting a subset of patients suitable for treatment comprises selecting a subset of patients having a tumor mutational burden (TMB) higher than a minimum threshold, wherein the TMB of a patient indicates the number of neoantigens in a set of neoantigens associated with that patient.
7. Selecting a subset of patients suitable for treatment is crucial. Selecting a subset of patients with a utility score above a minimum threshold The method of claim 1 , comprising:
8. 2. The method of claim 1, wherein the utility score is the sum of the likelihoods of presentation for each neoantigen within the identified subset of neoantigens for the patient.
9. 2. The method of claim 1, wherein the utility score is the probability that the number of presented neoantigens among the one or more identified neoantigens for the patient is above a minimum threshold.
10. the training data set is (a) data relating to measurements of peptide-MHC binding affinity for at least one of the isolated peptides; and (b) data relating to measurements of peptide-MHC binding stability for at least one of the isolated peptides; The method of claim 1 , further comprising at least one of:
11. The set of numerical likelihoods is: (a) a C-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; and (b) an N-terminal sequence adjacent to the neoantigen-encoding peptide sequence within the source protein sequence; The method of claim 1 , further characterized by characteristics including at least one of:
12. 10. The method of claim 1, wherein the set of presentation likelihoods is further identified by at least the expression level of the one or more MHC alleles in the subject as measured by RNA-seq or mass spectrometry.
13. The set of presentation likelihoods is: (a) the predicted affinity between neoantigens within said set of neoantigens and said one or more MHC alleles; and (b) the predicted stability of the neoantigen-encoded peptide-MHC complex. The method of claim 1 , further characterized by characteristics including at least one of:
14. inputting the peptide sequence into the machine-learned presentation model; applying the machine-learned presentation model to the peptide sequence of each neoantigen to generate, for each of the one or more MHC alleles, a dependency score indicating whether the MHC allele will present the neoantigen based on a particular amino acid at a particular position in the peptide sequence. The method of claim 1 , comprising:
15. inputting the peptide sequence into the machine-learned presentation model; transforming the dependency scores to generate, for each MHC allele, a corresponding per-allele likelihood that indicates the likelihood that the corresponding MHC allele presents the corresponding neoantigen; and combining the per-allele likelihoods to generate a presentation likelihood for the neoantigen.
15. The method of claim 14, comprising:
16. 16. The method of claim 15, wherein transforming the dependency score models presentation of the neoantigen as mutually exclusive across the one or more class MHC alleles.
17. inputting the peptide sequence into the machine-learned presentation model; Transforming the combination of dependency scores to generate a presentation likelihood. and wherein transforming the combination of dependency scores models presentation of the neoantigen as interfering between the one or more MHC alleles.
Citation Information
Patent Citations
Compositions and Methods of Individualized Neoplastic Vaccines
JP2016518355A