Reducing junction epitope presentation for neoantigens

The use of next-generation sequencing and nonlinear deep learning models, along with cassette sequence design, addresses the low PPV issue in neoantigen-based vaccines, enhancing anti-tumor immunity and reducing undesirable immune responses.

JP2025175055APending Publication Date: 2025-11-28GRITSTONE BIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025147998
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-11-22
Filing Date
2025-09-08
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Current methods for designing neoantigen-based cancer vaccines suffer from low positive predictive value (PPV) due to incomplete modeling of the epitope generation process, leading to inefficient utilization of vaccine doses and potential autoimmunity, with junction epitopes often stimulating undesirable immune responses.

Method used

An optimized approach using next-generation sequencing and nonlinear deep learning models to identify and select neoantigens, combined with cassette sequence design to minimize junction epitope presentation, ensuring high-PPV and therapeutic efficacy.

Benefits of technology

Enhances the likelihood of eliciting anti-tumor immunity by selecting neoantigens with high presentation likelihood and reducing undesirable immune responses, thereby improving the effectiveness of personalized cancer vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175055000001_ABST
    Figure 2025175055000001_ABST
Patent Text Reader

Abstract

To provide an optimized approach for identifying and selecting neoantigens for personalized cancer vaccines.SOLUTION: Given a set of therapeutic epitopes, a cassette sequence is designed to reduce the likelihood that junction epitopes are presented in the patient. The cassette sequence is designed by taking into account presentation of junction epitopes that span the junction between a pair of therapeutic epitopes in the cassette. The cassette sequence may be designed based on a set of distance metrics each associated with a junction of the cassette. The distance metric may specify a likelihood that one or more of the junction epitopes spanning between a pair of adjacent epitopes will be presented.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Application No. 62 / 590,045, filed November 22, 2017, the entire contents of which are incorporated herein by reference. [Background technology]

[0002] background Therapeutic vaccines based on tumor-specific neoantigens hold great promise as the next generation of personalized cancer immunotherapy. 1~3 Cancers with high mutational burden, such as non-small cell lung cancer (NSCLC) and melanoma, are particularly promising targets for such therapies due to their relatively high likelihood of generating neoantigens. 4,5 Early evidence suggests that neoantigen-based vaccination induces T cell responses. 6 , Cell therapy targeting neoantigens can induce tumor regression in selected patients 7 Both MHC class I and MHC class II influence T cell responses. 70~71 .

[0003] One question regarding neoantigen vaccine design is which of the many coding mutations present in the target tumor can give rise to the "best" therapeutic neoantigen (e.g., an antigen capable of eliciting antitumor immunity and causing tumor regression).

[0004] Early methods have been proposed that incorporate mutation-based analysis using next-generation sequencing, RNA gene expression, and prediction of MHC binding affinity of neoantigen peptides. 8However, these proposed methods involve many steps in addition to gene expression and MHC binding (e.g., TAP transport, proteasomal cleavage, MHC binding, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-I; endocytosis or autophagy, cleavage by extracellular or lysosomal proteases (e.g., cathepsins), competition with CLIP peptides for HLA binding catalyzed by HLA-DM, transport of peptide-MHC complexes to the cell surface, and / or TCR recognition of MHC-II). 9 The entire epitope generation process cannot be modeled, and therefore existing methods tend to suffer from low positive predictive value (PPV) (Figure 1A).

[0005] Indeed, analyses of peptides presented by tumor cells conducted by multiple groups have shown that less than 5% of peptides predicted to be presented using gene expression and MHC binding affinity are found on tumor surface MHC. 10,11 (Figure 1B). This low correlation between binding prediction and MHC presentation is further supported by the lack of improved prediction accuracy for binding-restricted neoantigens in response to checkpoint inhibitors relative to the number of mutations alone. 12 .

[0006] Such low positive predictive values ​​(PPV) of existing methods for predicting presentation present a problem in the design of neoantigen-based vaccines. If vaccines are designed using low PPV predictions, most patients are unlikely to receive therapeutic neoantigens, and even fewer will receive multiple neoantigens (even assuming all presented peptides are immunogenic). Thus, neoantigen vaccination using current methods is unlikely to be successful in a significant number of tumor-bearing subjects (Figure 1C).

[0007] Furthermore, previous approaches have only used cis-acting mutations to generate candidate neoantigens, including splicing factor mutations that occur in multiple tumor types and lead to aberrant splicing of many genes. 13 , and additional sources of nascent ORFs, including mutations that create or remove protease cleavage sites, were not considered in most cases.

[0008] Standard approaches to tumor genome and transcriptome analysis may miss somatic mutations that give rise to candidate neoantigens due to suboptimal conditions in library construction, exome and transcriptome capture, sequencing, or data analysis. Similarly, standard tumor analysis approaches may falsely promote sequence artifacts or germline polymorphisms as neoantigens, which can lead to inefficient utilization of vaccine doses or the risk of autoimmunity, respectively.

[0009] Neoantigen vaccines are also commonly designed as vaccine cassettes, in which a series of therapeutic epitopes are linked together. The vaccine cassette sequence may or may not include a linker sequence between adjacent pairs of therapeutic epitopes. The cassette sequence can provide junction epitopes, which are novel, yet unrelated, epitope sequences spanning the junction between pairs of therapeutic epitopes. Junction epitopes can be presented by a patient's HLA class I or class II alleles and stimulate CD8 or CD4 T cell responses, respectively. Such responses are often undesirable because T cells that react to these junction epitopes lack therapeutic efficacy and can reduce the immune response to the selected therapeutic epitopes in the cassette due to antigenic competition. Summary of the Invention [Means for solving the problem]

[0010] overview Optimized approaches for identifying and selecting neoantigens for personalized cancer vaccines are disclosed herein. First, we address an optimized tumor exome and transcriptome analysis approach to identify neoantigen candidates using next-generation neoantigen (NGS). These methods build on standard approaches for tumor analysis by NGS so that the most sensitive and specific neoantigen candidates are developed across all classes of genomic alterations. Second, novel approaches for high-PPV neoantigen selection are provided to overcome specificity issues and ensure that neoantigens developed for vaccine administration are more likely to elicit anti-tumor immunity. Depending on the embodiment, these approaches include trained statistical regression or nonlinear deep learning models that jointly model peptide-allele mapping, as well as motifs for each allele of peptides of multiple lengths that share statistical efficacy across peptides of different lengths. In particular, nonlinear deep learning models can be designed and trained to treat different MHC alleles within the same cell as independent, thereby resolving the problem associated with linear models interfering with each other. Finally, additional concerns regarding the design and production of personalized neoantigen-based vaccines are addressed.

[0011] Given a set of therapeutic epitopes, cassette sequences are designed to reduce the likelihood of junction epitope presentation in a patient. The cassette sequences are designed with consideration given to the presentation of junction epitopes across the junction between a pair of therapeutic epitopes in the cassette. In one embodiment, the cassette sequences are designed based on a set of distance metrics, each associated with a junction in the cassette. The distance metrics may identify the likelihood that one or more of the junction epitopes across a pair of adjacent epitopes will be presented. In one embodiment, one or more candidate cassette sequences are generated by randomly permuting the arrangement linking the set of therapeutic epitopes, and a cassette sequence having a presentation score (e.g., sum of distance metrics) below a predetermined threshold is selected. In another embodiment, the therapeutic epitopes are modeled as nodes, and the distance metric for a pair of adjacent epitopes represents the distance between corresponding nodes. Cassette sequences are selected that result in a total distance to "visit" each therapeutic epitope once strictly below a predetermined threshold. [The present invention 1001] 1. A method for identifying a cassette sequence for a neoantigen vaccine, comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one alteration that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and the obtaining includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; using a computer processor to input peptide sequences of the neoantigens into a machine learning display model to generate a set of numerical display likelihoods for the set of neoantigens, wherein each display likelihood in the set represents the likelihood that the corresponding neoantigen will be displayed by one or more MHC alleles on the surface of tumor cells of the subject; for each sample in the set of samples, a label obtained by mass spectrometry that measures the presence of a peptide bound to at least one MHC allele in the set of MHC alleles identified as being present in said sample; a training peptide sequence containing information about a plurality of amino acids constituting the training peptide sequence and a set of positions of the amino acids within the training peptide sequence for each of the samples; and A function that represents the relationship between the neoantigen peptide sequence received as input and the likelihood of presentation generated as output. the inputting step includes a plurality of parameters identified based at least on a training dataset including: identifying for the subject a therapeutic subset of neoantigens from the set of neoantigens, the therapeutic subset of neoantigens corresponding to a predetermined number of neoantigens having a likelihood of presentation above a predetermined threshold; and identifying for the subject a cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in a therapeutic subset of neoantigens, wherein the cassette sequence is identified based on the representation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes. The method comprising: [The present invention 1002] 1001. The method of claim 1001, wherein the presentation of said one or more junction epitopes is determined based on a presentation likelihood generated by inputting the sequences of said one or more junction epitopes into said machine learning presentation model. [The present invention 1003] 1001. The method of claim 1001, wherein the presentation of said one or more junction epitopes is determined based on binding affinity predictions between one or more junction epitopes and said one or more MHC alleles of the subject. [The present invention 1004] 1001. The method of claim 1001, wherein the presentation of said one or more junction epitopes is determined based on a binding stability prediction of said one or more junction epitopes. [The present invention 1005] 1001. The method of claim 1001, wherein said one or more junction epitopes comprise a junction epitope that overlaps with the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after said first therapeutic epitope. [The present invention 1006] 1001. The method of claim 1001, wherein a linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after said first therapeutic epitope, and said one or more junction epitopes comprise a junction epitope that overlaps with said linker sequence. [The present invention 1007] identifying the cassette sequence, For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between said ordered pair of therapeutic epitopes; and For each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for the ordered pair on said one or more MHC alleles of the subject. The method of the present invention 1001, comprising: [The present invention 1008] identifying the cassette sequence, generating a set of candidate cassette sequences corresponding to different sequences of said therapeutic epitope; For each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine. The method of the present invention 1001, comprising: [The present invention 1009] 1008. The method of claim 10, wherein said set of candidate cassette sequences is randomly generated. [The present invention 1010] identifying the cassette sequence, The following optimization problem: x in TIFF2025175055000002.tif60128 km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the therapeutic epitope, and P is TIFF2025175055000003.tif8128, where D is a v×v matrix and element D(k, m) indicates the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of The method of the present invention 1007 further comprising: [The present invention 1011] 1001. The method of claim 1001, further comprising producing or having produced a tumor vaccine comprising said cassette sequence. [The present invention 1012] 1. A method for identifying a cassette sequence for a neoantigen vaccine, comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one alteration that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and the obtaining includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; identifying a therapeutic subset of neoantigens from the set of neoantigens for the subject; and identifying for the subject a cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in a therapeutic subset of neoantigens, wherein the cassette sequence is identified based on the representation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes. The method comprising: [The present invention 1013] The method of claim 1012, wherein presentation of the one or more junction epitopes is determined based on presentation likelihoods generated by inputting the sequences of the one or more junction epitopes into a machine learning presentation model, the presentation likelihoods indicating the likelihood that the one or more junction epitopes are presented by one or more MHC alleles on the surface of tumor cells of the patient, and the set of presentation likelihoods has been identified based at least on received mass spectrometry data. [The present invention 1014] 1013. The method of claim 1012, wherein presentation of said one or more junction epitopes is determined based on binding affinity prediction between one or more junction epitopes and one or more MHC alleles of the subject. [The present invention 1015] 1013. The method of claim 1012, wherein the presentation of said one or more junction epitopes is determined based on a binding stability prediction of said one or more junction epitopes. [The present invention 1016] 1012. The method of claim 1012, wherein said one or more junction epitopes comprise a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after said first therapeutic epitope. [The present invention 1017] 1012. The method of claim 1012, wherein a linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after said first therapeutic epitope, and said one or more junction epitopes comprise a junction epitope that overlaps with said linker sequence. [The present invention 1018] identifying the cassette sequence, For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between said ordered pair of therapeutic epitopes; and For each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for said ordered pair on said one or more MHC alleles of the subject. The method of the present invention 1012, comprising: [The present invention 1019] identifying the cassette sequence, generating a set of candidate cassette sequences corresponding to different sequences of said therapeutic epitope; For each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine. The method of the present invention 1012, comprising: [The present invention 1020] The method of claim 1019, wherein the set of candidate cassette sequences is randomly generated. [The present invention 1021] identifying the cassette sequence, The following optimization problem: x in TIFF2025175055000004.tif60128 km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the therapeutic epitope, and P is TIFF2025175055000005.tif8128, where D is a v×v matrix and element D(k, m) indicates the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of The method of the present invention 1018 further comprises: [The present invention 1022] The method of claim 1012, further comprising producing or having produced a tumor vaccine comprising said cassette sequence. [The present invention 1023] 1. A method for identifying a cassette sequence for a neoantigen vaccine, comprising: obtaining peptide sequences for a therapeutic subset of shared antigens or a therapeutic subset of shared neoantigens for treating a plurality of subjects, wherein the therapeutic subsets corresponding to a predetermined number of peptide sequences have a likelihood of presentation that exceeds a predetermined threshold; and identifying said cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a corresponding peptide sequence in a therapeutic subset of a shared antigen or a therapeutic subset of a shared neoantigen; Including, identifying the cassette sequence, For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between the ordered pair of therapeutic epitopes; and determining, for each ordered pair of therapeutic epitopes, a distance metric indicative of presentation of the set of junction epitopes for said ordered pair, said distance metric being determined as a combination of a set of weights each indicative of the prevalence of a corresponding MHC allele and corresponding sub-distance metrics indicative of the likelihood of presentation of the set of junction epitopes on said MHC allele. Including, The method. [The present invention 1024] 1. A tumor vaccine comprising a cassette sequence comprising a sequence of linked therapeutic epitopes, the cassette sequence comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one modification that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and the obtaining includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; identifying a therapeutic subset of neoantigens from the set of neoantigens for the subject; and identifying, for the subject, the cassette sequences comprising sequences of a plurality of linked therapeutic epitopes, each comprising a peptide sequence of a corresponding neoantigen in a therapeutic subset of neoantigens, wherein the cassette sequences are identified based on the representation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes. are identified by performing The tumor vaccine. [The present invention 1025] A tumor vaccine of the present invention 1024, wherein the presentation of the one or more junction epitopes is determined based on a presentation likelihood generated by inputting the sequences of the one or more junction epitopes into a machine learning presentation model, the presentation likelihood indicating the likelihood that the one or more junction epitopes are presented by one or more MHC alleles on the surface of the patient's tumor cells, and the set of presentation likelihoods is identified based at least on received mass spectrometry data. [The present invention 1026] The tumor vaccine of the present invention 1024, wherein the presentation of said one or more junction epitopes is determined based on predicted binding affinity between one or more junction epitopes and one or more MHC alleles of a subject. [The present invention 1027] The tumor vaccine of the present invention 1024, wherein the presentation of said one or more junction epitopes is determined based on predicted binding stability of said one or more junction epitopes. [The present invention 1028] The tumor vaccine of the present invention 1024, wherein the one or more junction epitopes comprise a junction epitope that overlaps with the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope. [The present invention 1029] The tumor vaccine of the present invention 1024, wherein a linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and the one or more junction epitopes include a junction epitope that overlaps with the linker sequence. [The present invention 1030] the step of identifying the cassette sequence comprises: For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between said ordered pair of therapeutic epitopes; and For each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for said ordered pair on one or more MHC alleles of the subject. The tumor vaccine of the present invention, comprising: [The present invention 1031] the step of identifying the cassette sequence comprises: generating a set of candidate cassette sequences corresponding to different sequences of said therapeutic epitope; For each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine. The tumor vaccine of the present invention, comprising: [The present invention 1032] The tumor vaccine of the present invention 1031, wherein the set of candidate cassette sequences is randomly generated. [The present invention 1033] the step of identifying the cassette sequence comprises: The following optimization problem: x in TIFF2025175055000006.tif60128 km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the first therapeutic epitope, and P is TIFF2025175055000007.tif8128, where D is a v×v matrix and element D(k, m) indicates the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of The tumor vaccine of the present invention further comprising: [The present invention 1034] The tumor vaccine of the present invention 1024, further comprising producing or having produced a tumor vaccine comprising the cassette sequence. [This invention 1035] A tumor vaccine comprising a cassette sequence comprising linked therapeutic epitope sequences, wherein the cassette sequences are ordered to each comprise a peptide sequence of a corresponding neoantigen within a therapeutic subset of neoantigens, and the therapeutic epitope sequences are identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes, and the junction epitopes of the cassette sequence have an HLA binding affinity below a threshold binding affinity. [The present invention 1036] The tumor vaccine of the present invention 1035, wherein the threshold binding affinity is 1000 nM or more. [This invention 1037] A tumor vaccine comprising a cassette sequence comprising linked therapeutic epitope sequences, wherein the cassette sequences are ordered so that each comprises a peptide sequence of a corresponding neoantigen within a therapeutic subset of neoantigens, and the therapeutic epitope sequences are identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes, and at least a threshold percentage of the junction epitopes of the cassette sequence have a presentation likelihood below a threshold presentation likelihood. [The present invention 1038] The tumor vaccine of the present invention 1037, wherein said threshold percentage is 50%. [Brief explanation of the drawings]

[0012] These and other features, aspects, and aspects of the present invention will become better understood with regard to the following description and accompanying drawings.

[0013] [Figure 1A] Current clinical approaches to neoantigen identification are presented. [Figure 1B] It shows that less than 5% of the predicted binding peptides are displayed on tumor cells. [Figure 1C] Illustrates the impact of specificity issues on neoantigen prediction. [Figure 1D] This shows that binding prediction is not sufficient to identify neoantigens. [Figure 1E] Probability of MHC-I presentation as a function of peptide length. [Figure 1F] Figure 1F shows an exemplary peptide spectrum generated from a dynamic range standard from Promega. Figure 1F discloses SEQ ID NO: 1. [Figure 1G] We show how adding features increases the positive predictive value of the model. [Figure 2A] 1 is a schematic of an environment for identifying the likelihood of peptide presentation in a patient, according to one embodiment. [Figure 2B]A method for obtaining presentation information according to one embodiment is described (SEQ ID NO:72). [Figure 2C] A method for obtaining presentation information according to one embodiment will be described (in order of appearance, SEQ ID NOs: 3 to 8, respectively). [Figure 3] FIG. 1 is a high-level block diagram illustrating computer logic components of a presentation specification system, according to one embodiment. [Figure 4] An exemplary set of training data according to one embodiment is illustrated (in order of appearance, SEQ ID NOs: 10-13, 15, 73-74, and 74, respectively). [Figure 5] 1 illustrates an exemplary network model related to MHC alleles. [Figure 6A] 1 illustrates an exemplary network model NNH(·) shared by MHC alleles, according to one embodiment. [Figure 6B] 1 illustrates an exemplary network model NNH(·) shared by MHC alleles, according to another embodiment. [Figure 7] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 8] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 9] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 10] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 11] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 12] 1 illustrates the generation of presentation likelihoods of peptides associated with MHC alleles using an exemplary network model. [Figure 13]Illustrated is the determination of distance metrics for two example cassette sequences (SEQ ID NOs: 75-76, respectively, in order of appearance). [Figure 14] An exemplary computer for implementing the entities shown in FIGS. 1 and 3 will now be described. DETAILED DESCRIPTION OF THE INVENTION

[0014] Detailed Description I. Definition In general, terms used in the claims and the specification shall be interpreted as having their ordinary meaning as understood by one of ordinary skill in the art. Certain terms are defined below to provide further clarity. If there is a conflict between the ordinary meaning and a given definition, the given definition shall control.

[0015] As used herein, the term "antigen" refers to a substance that induces an immune response.

[0016] As used herein, the term "neoantigen" refers to an antigen that has at least one change that makes it different from its corresponding wild-type parent antigen, for example, due to a tumor cell mutation or tumor cell-specific post-translational modification. Neoantigens may include polypeptide or nucleotide sequences. Mutations can include frameshift or non-frameshift insertion / deletions (indels), missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression change that results in a new ORF. Mutations can also include splice variants. Tumor cell-specific post-translational modifications can include aberrant phosphorylation. Tumor cell-specific post-translational modifications can also include splice antigens generated by the proteasome. See Liepe et al., "A large fraction of HLA class I ligands are proteasome-generated spliced ​​peptides"; Science. 2016 Oct 21;354(6310):354-358.

[0017] As used herein, the term "tumor neoantigen" refers to a neoantigen that is present in tumor cells or tissues of a subject, but is not present in the corresponding normal cells or tissues of the subject.

[0018] As used herein, the term "neoantigen-based vaccine" refers to a vaccine construct that is based on one or more neoantigens, eg, multiple neoantigens.

[0019] As used herein, the term "candidate neoantigen" refers to a mutation or other abnormality that gives rise to a new sequence that may represent a neoantigen.

[0020] As used herein, the term "coding region" refers to the portion of a gene that encodes a protein.

[0021] As used herein, the term "coding mutation" refers to a mutation that occurs in a coding region.

[0022] As used herein, the term "ORF" means open reading frame.

[0023] As used herein, the term "neo-ORF" refers to a tumor-specific ORF that arises due to mutation or other abnormalities such as splicing.

[0024] As used herein, the term "missense mutation" is a mutation that results in the substitution of one amino acid for another.

[0025] As used herein, the term "nonsense mutation" is a mutation that results in the substitution of an amino acid for a stop codon.

[0026] As used herein, the term "frameshift mutation" is a mutation that causes an alteration in the frame of a protein.

[0027] As used herein, the term "indel" refers to the insertion or deletion of one or more nucleic acids.

[0028] As used herein, the term "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences in which a certain percentage of nucleotides or amino acid residues are the same when compared and aligned for maximum correspondence, as determined using one of the sequence comparison algorithms described below (e.g., BLASTP and BLASTN, or other algorithms available to those of skill in the art), or by visual inspection. Depending on the application, the "percent identity" can exist over a region of the sequences being compared, e.g., over a functional domain, or over the full length of the two sequences being compared.

[0029] In sequence comparison, generally, one sequence serves as a reference sequence to which test sequences are compared.When using a sequence comparison algorithm, test sequences and reference sequences are input into a computer, subsequence coordinates are designated if necessary, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the percent sequence identity (%) of the test sequence to the reference sequence based on the designated program parameters.Alternatively, sequence similarity or difference can also be established by the combination of the presence or absence of a specific nucleotide at a selected sequence position (e.g., sequence motif) or an amino acid in a translated sequence.

[0030] Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by visual inspection (see generally Ausubel et al., infra).

[0031] One example of an algorithm that is suitable for determining percent sequence identity and percent sequence similarity is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.

[0032] As used herein, the term "non-stop or read-through" refers to a mutation that results in the removal of the natural stop codon.

[0033] As used herein, the term "epitope" refers to a specific portion of an antigen that is typically bound by an antibody or T-cell receptor.

[0034] As used herein, the term "immunogenic" refers to the ability to elicit an immune response, for example, via T cells, B cells, or both.

[0035] As used herein, the terms "HLA binding affinity" and "MHC binding affinity" refer to the affinity of binding between a specific antigen and a specific MHC allele.

[0036] As used herein, the term "bait" refers to a nucleic acid probe used to enrich a specific sequence of DNA or RNA from a sample.

[0037] As used herein, the term "mutation" is a difference between the nucleic acid of a subject and a reference human genome used as a control.

[0038] As used herein, the term "variant calling" is the algorithmic determination, typically from sequencing, of the presence of a mutation.

[0039] As used herein, the term "polymorphism" refers to a germline mutation, ie, a mutation found in all DNA-bearing cells of an individual.

[0040] As used herein, the term "somatic mutation" is a mutation that occurs in a non-germline cell of an individual.

[0041] As used herein, the term "allele" refers to one version of a gene or one version of a gene sequence or one version of a protein.

[0042] As used herein, the term "HLA type" refers to the complement of HLA gene alleles.

[0043] As used herein, the term "nonsense-mediated decay" or "NMD" refers to the degradation of mRNA by the cell due to a premature stop codon.

[0044] As used herein, the term "truncal mutation" is a mutation that occurs early in the development of a tumor and is present in the majority of the cells of the tumor.

[0045] As used herein, the term "subclonal mutation" is a mutation that occurs late in the development of a tumor and is present in only a portion of the cells of the tumor.

[0046] As used herein, the term "exome" refers to the subset of the genome that encodes proteins. The exome can be the collection of exons of the genome.

[0047] As used herein, the term "logistic regression" is a regression model for binary data from statistics in which the logit of the probability that the dependent variable is equal to 1 is modeled as a linear function of the dependent variable.

[0048] As used herein, the term "neural network" refers to a machine learning model for classification or regression that consists of multiple layers of linear transformations followed by element-wise nonlinear transformations typically trained by stochastic gradient descent and backpropagation.

[0049] As used herein, the term "proteome" refers to the set of all proteins expressed and / or translated by a cell, a group of cells, or an individual.

[0050] As used herein, the term "peptidome" refers to the set of all peptides presented by MHC-I or MHC-II on the cell surface. Peptidome can also refer to the properties of a cell or a collection of cells (e.g., a tumor peptidome refers to the union of the peptidomes of all cells that comprise a tumor).

[0051] As used herein, the term "ELISPOT" refers to enzyme-linked immunosorbent spot assay, a common method for monitoring immune responses in humans and animals.

[0052] As used herein, the term "dextramer" refers to a dextran-based peptide-MHC multimer used for antigen-specific T cell staining in flow cytometry.

[0053] As used herein, the term "tolerance or immune tolerance" refers to a state of immune unresponsiveness to one or more antigens, eg, self-antigens.

[0054] As used herein, the term "central tolerance" is tolerance conferred in the thymus by either deleting autoreactive T cell clones or promoting their differentiation into immunosuppressive regulatory T cells (Tregs).

[0055] As used herein, the term "peripheral tolerance" refers to tolerance conferred in the peripheral system by downregulating or anergizing autoreactive T cells that survive central tolerance or by promoting the differentiation of these T cells into Tregs.

[0056] The term "sample" can include a single cell, or multiple cells, or fragments of cells, or an aliquot of bodily fluid obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or intervention, or other means known in the art.

[0057] The term "subject" includes cells, tissues, or organisms, human or non-human, whether male or female, in vivo, ex vivo, or in vitro. The term subject includes mammals, including humans.

[0058] The term "mammal" encompasses both humans and non-humans, and includes, but is not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.

[0059] The term "clinical factor" refers to a measurement of a subject's condition, e.g., disease activity or severity. "Clinical factor" encompasses all markers of a subject's health status, including non-sample markers, and / or other characteristics of the subject, such as, but not limited to, age and sex. A clinical factor can be a score, value, or set of values ​​that can be obtained from assessing a subject or a sample (or a population of samples) from a subject under a given condition. A clinical factor can also be predicted by other parameters, such as markers and / or gene expression surrogates. Clinical factors can include tumor type, tumor subtype, and smoking history.

[0060] Abbreviations: MHC: major histocompatibility complex; HLA: human leukocyte antigen, or human MHC locus; NGS: next-generation sequencing; PPV: positive predictive value; TSNA: tumor-specific neoantigen; FFPE: formalin-fixed, paraffin-embedded; NMD: nonsense-mediated decay; NSCLC: non-small cell lung cancer; DC: dendritic cell.

[0061] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0062] Terms not directly defined herein should be understood to have the meanings generally associated with them as understood within the technical field of the present invention. Certain terms are discussed herein to provide further guidance to the practitioner in describing the compositions, devices, methods, etc. of embodiments of the present invention, as well as how to make or use them. It will be recognized that multiple ways of saying the same thing may be used. Accordingly, alternative terms and synonyms may be used for any one or more of the terms discussed herein. No weight should be placed on whether a term is detailed or discussed herein. Several synonyms or alternative methods, materials, etc. are provided. The recitation of one or more synonyms or equivalents does not exclude the use of other synonyms or equivalents, unless expressly stated. The use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the inventive embodiments herein.

[0063] All references, issued patents, and patent applications cited within the body of this specification are hereby incorporated by reference in their entirety for all purposes.

[0064] II. Methods for inhibiting the presentation of junction epitopes Disclosed herein are methods for identifying cassette sequences for neoantigen vaccines. As an example, one such method may include the following steps: obtaining at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of a patient, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one modification that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; and inputting the peptide sequences of the neoantigens into a machine learning presentation model using a computer processor to generate a set of numerical presentation likelihoods for the set of neoantigens, wherein each presentation likelihood in the set represents the likelihood that the corresponding neoantigen will be presented by one or more MHC alleles on the surface of the subject's tumor cells. The machine learning presentation model includes a plurality of parameters identified based at least on a training dataset including, for each sample in a set of samples, labels obtained by mass spectrometry that measure the presence of peptides bound to at least one MHC allele in a set of MHC alleles identified as present in the sample, a training peptide sequence for each of the samples that includes information about a plurality of amino acids that make up the training peptide sequence and a set of amino acid positions within the training peptide sequence, and a function that represents the relationship between the neoantigen peptide sequence received as input and the presentation likelihood generated as output.The method may further include the following steps: identifying a therapeutic subset of neoantigens from the set of neoantigens for the subject, wherein the therapeutic subset of neoantigens corresponds to a predetermined number of neoantigens having a presentation likelihood exceeding a predetermined threshold; and identifying a cassette sequence for the subject comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in the therapeutic subset of neoantigens, wherein the cassette sequence is identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes.

[0065] The presentation of the one or more junction epitopes may be determined based on a presentation likelihood generated by inputting the sequences of the one or more junction epitopes into the machine learning presentation model.

[0066] The presentation of the one or more junction epitopes may be determined based on predicted binding affinity between the one or more junction epitopes and one or more MHC alleles of the subject.

[0067] The presentation of the one or more junction epitopes may be determined based on predicted binding stability of the one or more junction epitopes.

[0068] The one or more junction epitopes may include a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope.

[0069] A linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and the one or more junction epitopes may include a junction epitope that overlaps with the linker sequence.

[0070] Identifying the cassette sequence may further include determining, for each ordered pair of therapeutic epitopes, a set of junction epitopes spanning the junction between the ordered pair of therapeutic epitopes; and, for each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for the ordered pair on one or more MHC alleles of the subject.

[0071] Identifying the cassette sequences may further include generating a set of candidate cassette sequences corresponding to different sequences of the therapeutic epitopes; for each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine.

[0072] The set of candidate cassette sequences may be generated randomly.

[0073] Identifying the cassette sequence includes: The following optimization problem: x in TIFF2025175055000008.tif60128 km A step of determining the value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the therapeutic epitope, and P is a path matrix given by TIFF2025175055000009.tif8128, where D is a v×v matrix and element D(k, m) denotes the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of It may further include:

[0074] The method may further comprise the step of producing, or having produced, a tumor vaccine comprising the cassette sequence.

[0075] Also disclosed herein is a method for identifying cassette sequences for a neoantigen vaccine, the method comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome from tumor cells and normal cells of the subject; the nucleotide sequencing data being used to obtain data representing peptide sequences for each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells; and wherein the peptide sequence for each neoantigen is determined by comparing the peptide sequence with a corresponding wild-type parent peptide sequence identified from the subject's normal cells. and the peptide sequence includes information about a plurality of amino acids comprising the peptide sequence and a set of amino acid positions within the peptide sequence; identifying a therapeutic subset of neoantigens from the set of neoantigens for the subject; and identifying a cassette sequence for the subject comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in the therapeutic subset of neoantigens, wherein the cassette sequence is identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes.

[0076] The presentation of the one or more junction epitopes can be determined based on presentation likelihoods generated by inputting the sequences of the one or more junction epitopes into a machine learning presentation model, the presentation likelihoods indicating the likelihood that the one or more junction epitopes are presented by one or more MHC alleles on the surface of tumor cells of the patient, and the set of presentation likelihoods has been identified based at least on the received mass spectrometry data.

[0077] The presentation of the one or more junction epitopes may be determined based on predicted binding affinity between the one or more junction epitopes and one or more MHC alleles of the subject.

[0078] The presentation of the one or more junction epitopes may be determined based on predicted binding stability of the one or more junction epitopes.

[0079] The one or more junction epitopes may include a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope.

[0080] A linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and the one or more junction epitopes may include a junction epitope that overlaps with the linker sequence.

[0081] Identifying the cassette sequence may further include determining, for each ordered pair of therapeutic epitopes, a set of junction epitopes spanning the junction between the ordered pair of therapeutic epitopes; and, for each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for the ordered pair on one or more MHC alleles of the subject.

[0082] Identifying the cassette sequences may further include generating a set of candidate cassette sequences corresponding to different sequences of the therapeutic epitopes; for each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine.

[0083] The set of candidate cassette sequences can be generated randomly.

[0084] Identifying the cassette sequence includes: The following optimization problem: x in TIFF2025175055000010.tif60128 km A step of determining the value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the therapeutic epitope, and P is a path matrix given by TIFF2025175055000011.tif8128, where D is a v×v matrix and element D(k, m) denotes the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of It may further include:

[0085] The method may further comprise the step of producing a tumor vaccine comprising the cassette sequence.

[0086] Also disclosed herein is a method for identifying a cassette sequence for a neoantigen vaccine, the method comprising the steps of obtaining peptide sequences for a therapeutic subset of a shared antigen or a therapeutic subset of a shared neoantigen for treating a plurality of subjects, the therapeutic subsets corresponding to a predetermined number of peptide sequences having a likelihood of presentation above a predetermined threshold; and identifying the cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a corresponding peptide sequence in the therapeutic subset of the shared antigen or the therapeutic subset of the shared neoantigen, and The identifying step includes: for each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between the ordered pair of therapeutic epitopes; and for each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for the ordered pair, wherein the distance metric is determined as a combination of a set of weights, each indicative of the prevalence of a corresponding MHC allele, and corresponding sub-distance metrics indicative of the likelihood of presentation of the set of junction epitopes on the MHC allele.

[0087] Also disclosed herein is a tumor vaccine comprising a cassette sequence comprising sequences of linked therapeutic epitopes, the cassette sequence being identified by performing the following steps: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, the nucleotide sequencing data being used to obtain data representing each peptide sequence of a set of neoantigens identified by comparing the nucleotide sequencing data derived from the tumor cells with the nucleotide sequencing data derived from the normal cells, and the peptide sequence of each neoantigen being identified by comparing the peptide sequence with the nucleotide sequencing data derived from the normal cells of the subject. the step of obtaining a set of neoantigens from the set of neoantigens for the subject, the set of neoantigens comprising at least one modification that makes the set of neoantigens different from the corresponding wild-type parent peptide sequence identified from the cell, and the set of neoantigens comprising information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; the step of identifying a therapeutic subset of neoantigens from the set of neoantigens for the subject; and the step of identifying a cassette sequence for the subject comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in the therapeutic subset of neoantigens, the cassette sequence being identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes.

[0088] The presentation of the one or more junction epitopes is determined based on presentation likelihoods generated by inputting the sequences of the one or more junction epitopes into a machine learning presentation model, the presentation likelihoods indicating the likelihood that the one or more junction epitopes are presented by one or more MHC alleles on the surface of tumor cells of the patient, and the set of presentation likelihoods has been identified based at least on the received mass spectrometry data.

[0089] The presentation of the one or more junction epitopes may be determined based on predicted binding affinity between the one or more junction epitopes and one or more MHC alleles of the subject.

[0090] The presentation of the one or more junction epitopes may be determined based on predicted binding stability of the one or more junction epitopes.

[0091] The one or more junction epitopes may include a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope.

[0092] A linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and the one or more junction epitopes may include a junction epitope that overlaps with the linker sequence.

[0093] Identifying the cassette sequence may further include determining, for each ordered pair of therapeutic epitopes, a set of junction epitopes spanning the junction between the ordered pair of therapeutic epitopes; and, for each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for the ordered pair on one or more MHC alleles of the subject.

[0094] Identifying the cassette sequences may further include generating a set of candidate cassette sequences corresponding to different sequences of the therapeutic epitopes; for each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine.

[0095] The set of candidate cassette sequences can be generated randomly.

[0096] Identifying the cassette sequence includes: The following optimization problem: x in TIFF2025175055000012.tif60128 km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the first therapeutic epitope, and P is a path matrix given by TIFF2025175055000013.tif8128, where D is a v×v matrix and element D(k, m) denotes the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of It may further include:

[0097] The tumor vaccine of claim 24, further comprising producing or having produced a tumor vaccine comprising the cassette sequence.

[0098] Also disclosed herein is a tumor vaccine comprising a cassette sequence comprising linked therapeutic epitope sequences, the cassette sequences being ordered to each comprise a peptide sequence of a corresponding neoantigen within a therapeutic subset of neoantigens, the therapeutic epitope sequences being identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes, the junction epitopes of the cassette sequence having an HLA binding affinity below a threshold binding affinity.

[0099] The threshold binding affinity may be 1000 nM or greater.

[0100] Also disclosed herein is a tumor vaccine comprising a cassette sequence comprising linked therapeutic epitope sequences, the cassette sequences being ordered to each comprise a peptide sequence of a corresponding neoantigen within a therapeutic subset of neoantigens, the therapeutic epitope sequences being identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes, and at least a threshold percentage of the junction epitopes of the cassette sequence having a presentation likelihood below a threshold presentation likelihood.

[0101] The threshold percentage may be 50%.

[0102] III. Identification of tumor-specific mutations in neoantigens Also disclosed herein are methods for identifying certain mutations (e.g., mutations or alleles present in cancer cells). In particular, these mutations may be present in the genome, transcriptome, proteome, or exome of cancer cells of a subject with cancer, but may not be present in normal tissues from the subject.

[0103] Genetic mutations in tumors can be considered useful for immunological targeting of tumors if they result in changes in the amino acid sequence of proteins exclusively in tumors. Useful mutations include: (1) non-synonymous mutations that result in different amino acids in proteins; (2) read-through mutations in which the stop codon is modified or deleted, resulting in the translation of a longer protein with a new tumor-specific sequence at the C-terminus; (3) splice site mutations that result in the inclusion of an intron in mature mRNA, thus resulting in a unique tumor-specific protein sequence; (4) chromosomal rearrangements (i.e., gene fusions) that result in chimeric proteins with tumor-specific sequences at the junction of two proteins; (5) frameshift mutations or deletions that result in new open reading frames with new tumor-specific protein sequences. Mutations can also include one or more of non-frameshift insertions / deletions, missense or nonsense substitutions, splice site changes, genomic rearrangements or gene fusions, or any genomic or expression changes that result in new ORFs.

[0104] For example, mutated peptides or mutated polypeptides resulting from splice site, frameshift, readthrough, or gene fusion mutations in tumor cells can be identified by sequencing DNA, RNA, or protein in tumor versus normal cells.

[0105] Mutations can also include previously identified tumor-specific mutations. Known tumor mutations can be found in the Catalogue of Somatic Mutations in Cancer (COSMIC) database.

[0106] Various methods are available for detecting the presence of specific mutations or alleles in an individual's DNA or RNA. Advances in this field have provided accurate, easy, and inexpensive large-scale SNP genotyping. Several techniques have been described, including dynamic allele-specific hybridization (DASH), microplate array diagonal gel electrophoresis (MADGE), pyrosequencing, oligonucleotide-specific ligation, the TaqMan system, and various DNA "chip" technologies such as the Affymetrix SNP chip. These methods utilize amplification of target gene regions, typically by PCR. Still other methods rely on the generation of small signal molecules by invasive cleavage followed by mass spectrometry, or on immobilized padlock probes and rolling circle amplification. Some of the methods known in the art for detecting specific mutations are summarized below.

[0107] PCR-based detection means can involve the multiplex amplification of multiple markers simultaneously.For example, it is well known in the art to select PCR primers so as to generate PCR products that do not overlap in size and can be analyzed simultaneously.Alternatively, it is possible to amplify different markers with primers that are differentially labeled and therefore can be differentially detected.Of course, hybridization-based detection means allows the differential detection of multiple PCR products in samples.Other techniques that allow multiplex analysis of multiple markers are known in the art.

[0108] Several methods have been developed to facilitate the analysis of single nucleotide polymorphisms in genomic DNA or cellular RNA.For example, single nucleotide polymorphisms can be detected by using special exonuclease-resistant nucleotides, as disclosed in Mundy, CR (US Patent No. 4,656,127).According to this method, a primer complementary to the allele sequence immediately 3' of the polymorphic site is hybridized to a target molecule obtained from a specific animal or human.If the polymorphic site on the target molecule contains a nucleotide that is complementary to the specific exonuclease-resistant nucleotide derivative present, this derivative will be incorporated onto the end of the hybridized primer.This incorporation makes the primer resistant to exonucleases, thereby enabling its detection.Since the identity of the exonuclease-resistant derivative of the sample is known, the knowledge that the primer has become resistant to exonucleases reveals that the nucleotide present at the polymorphic site of the target molecule is complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of exogenous sequence data.

[0109] To determine the identity of the nucleotide at a polymorphic site, a solution-based method can be used (Cohen, D. et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087)). As in the method of Mundy, U.S. Pat. No. 4,656,127, a primer is used that is complementary to the allelic sequence immediately 3' to the polymorphic site. This method uses a labeled dideoxynucleotide derivative that becomes incorporated onto the end of the primer if it is complementary to the nucleotide at the polymorphic site to determine the identity of the nucleotide at that site. An alternative method, known as Genetic Bit Analysis or GBA, has been described by Goelet, P. et al. (PCT Application No. 92 / 15712). The Goelet, P. et al. method uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' to the polymorphic site. Goelet, P. et al. The method of Goelet, P. et al. uses a mixture of labeled terminators and a primer that is complementary to the sequence 3' of the polymorphic site. In contrast to the method of Cohen et al. (French Patent No. 2,650,840; PCT Application No. WO 91 / 02087), the method of Goelet, P. et al. can be a heterogeneous phase assay in which either the primer or the target molecule is immobilized on a solid phase.

[0110] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. et al., Nucl. Acids. Res. 17:7779-7784 (1989); Sokolov, B. P., Nucl. Acids Res. 18:3671 (1990); Syvanen, A.-C., et al., Genomics 8:684-692 (1990); Kuppuswamy, M. et al., Proc. Natl. Acad. Sci. (USA) 88:1143-1147 (1991); Prezant, T. R. et al., Hum. Mutat. 1:159-164 (1992); Ugozzoli, L. et al., GATA 9:107-112 (1992); Nyren, P. et al. al., Anal. Biochem. 208:171-175 (1993)). These methods differ from GBA in that they utilize the incorporation of labeled deoxynucleotides to discriminate between bases at the polymorphic site. In such formats, signal is proportional to the number of incorporated deoxynucleotides, so that polymorphisms occurring in runs of the same nucleotide can result in a signal proportional to the length of the run (Syvanen, A.-C., et al., Amer. J. Hum. Genet. 52:46-59 (1993)).

[0111] Numerous initiatives obtain sequence information directly from millions of individual molecules of DNA or RNA in parallel. Real-time single-molecule sequencing by synthesis techniques rely on the detection of fluorescent nucleotides as they are incorporated into nascent strands of DNA complementary to the template to be sequenced. In one method, oligonucleotides 30–50 bases in length are covalently anchored at their 5′ ends to glass coverslips. These anchored strands serve two functions. First, they act as capture sites for the target template strands when the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also act as primers for template-directed primer extension, which forms the basis for sequence reading. The capture primers serve as fixed-location sites for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of the addition of a polymerase / labeled nucleotide mixture, rinsing, imaging, and dye cleavage. In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the γ-phosphate. As the nucleotide becomes incorporated into the new strand, the system detects the interaction between the fluorescently tagged polymerase and the fluorescently modified nucleotide. Other sequencing-by-synthesis techniques also exist.

[0112] Any suitable sequencing-by-synthesis platform can be used to identify mutations. As mentioned above, four major sequencing-by-synthesis platforms are currently available: the Genome Sequencer sold by Roche / 454 Life Sciences, the 1G Analyzer sold by Illumina / Solexa, the SOLiD system sold by Applied BioSystems, and the Heliscope system sold by Helicos Bioscience. Sequencing-by-synthesis platforms have also been described by Pacific BioSciences and VisiGen Biotechnologies. In some embodiments, the multiple nucleic acid molecules to be sequenced are bound to a support (e.g., a solid support). To immobilize the nucleic acid on the support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. The nucleic acid can be bound to the support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. A capture sequence (also called a universal capture sequence) is a nucleic acid sequence complementary to a sequence attached to a support that can double as a universal primer.

[0113] As an alternative to capture sequences, a member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair, e.g., as described in U.S. Patent Application Publication No. 2006 / 0252077) can be linked to each fragment and captured on a surface coated with the respective second member of the coupling pair.

[0114] Following capture, the sequence can be analyzed by single-molecule detection / sequencing, including, for example, template-dependent sequencing by synthesis, as described, for example, in the Examples and in U.S. Patent No. 7,283,337. In sequencing by synthesis, surface-bound molecules are exposed to a multitude of labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides incorporated into the 3' end of the growing strand. This can be done in real time, in a step-and-repeat mode. For real-time analysis, a different optical label can be incorporated for each nucleotide, and multiple lasers can be utilized for stimulation of the incorporated nucleotides.

[0115] Sequencing can also include other massively parallel sequencing or next-generation sequencing (NGS) techniques and platforms. Additional examples of massively parallel sequencing techniques and platforms are Illumina HiSeq or MiSeq, ThermoPGM or Proton, Pac Bio RS II or Sequel, Qiagen's Gene Reader, and Oxford Nanopore MinION. Additional similar current massively parallel sequencing technologies, and future generations of these technologies, can be used.

[0116] Any cell type or tissue can be used to obtain nucleic acid samples for use in the methods described herein.For example, DNA or RNA samples can be obtained from tumor or body fluids, for example, blood obtained by known techniques (for example, venipuncture) or saliva.Alternatively, nucleic acid testing can be performed on dry samples (for example, hair or skin).In addition, a sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal tissue is of the same tissue type as tumor.A sample can be obtained from tumor for sequencing, and another sample can be obtained from normal tissue for sequencing, if the normal sample is of a different tissue type from tumor.

[0117] The tumor may include one or more of lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, stomach cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.

[0118] Alternatively, protein mass spectrometry can be used to identify or demonstrate the presence of mutated peptides bound to MHC proteins on tumor cells. Peptides can be acid-eluted from tumor cells or from HLA molecules immunoprecipitated from tumors, and then identified using mass spectrometry.

[0119] IV. Neoantigens Neoantigens can comprise nucleotides or polynucleotides. For example, neoantigens can be RNA sequences that encode polypeptide sequences. Neoantigens useful in vaccines can therefore comprise nucleotide sequences or polypeptide sequences.

[0120] Disclosed herein are isolated peptides comprising tumor-specific mutations identified by the methods disclosed herein, peptides comprising known tumor-specific mutations, and mutant polypeptides or fragments thereof identified by the methods disclosed herein. Neoantigen peptides can be described in the context of their coding sequences when the neoantigen comprises nucleotide sequences (e.g., DNA or RNA) that encode the associated polypeptide sequence.

[0121] The one or more polypeptides encoded by the neoantigen nucleotide sequences can comprise at least one of the following: a binding affinity to MHC with an IC50 value of less than 1000 nM; a length of 8-15, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids for MHC class I peptides; the presence of a sequence motif within or near the peptide that promotes proteasomal cleavage; and the presence of a sequence motif within or near the peptide that promotes TAP transport; a length of 6-30, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids for MHC class II polypeptides; and the presence of a sequence motif within or near the peptide that promotes cleavage by extracellular or lysosomal proteases (e.g., cathepsins) or HLA binding catalyzed by HLA-DM.

[0122] One or more neoantigens can be present on the surface of a tumor.

[0123] The one or more neoantigens can be immunogenic in a tumor-bearing subject, for example, capable of eliciting a T cell or B cell response in the subject.

[0124] One or more neoantigens that induce an autoimmune response in a subject can be eliminated from consideration in the context of generating a vaccine for a tumor-bearing subject.

[0125] The size of the at least one neoantigenic peptide molecule is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35 , about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, or more amino acid residues, and any range derivable therein. In a specific embodiment, the neoantigenic peptide molecule is 50 amino acids or less.

[0126] Neoantigenic peptides and polypeptides can be 15 residues or less in length, typically between about 8 and about 11 residues, particularly 9 or 10 residues, for MHC class I; and 6 to 30 residues for MHC class II.

[0127] If desired, longer peptides can be designed in several ways. In one example, if the likelihood of peptide presentation on HLA alleles is predicted or known, the longer peptides can consist of either (1) individual presented peptides with extensions of 2-5 amino acids toward the N- and C-termini of each corresponding gene product; or (2) a concatenation of some or all of the presented peptides, each with its extended sequence. In another example, if sequencing reveals long (more than 10 residues) neo-epitope sequences present in the tumor (e.g., due to frameshifts, readthrough, or intron inclusion resulting in novel peptide sequences), the longer peptides would (3) consist of the entire novel tumor-specific stretch of amino acids, thus avoiding the need for computational or in vitro test-based selection of shorter peptides that are presented to the strongest HLA alleles. In either example, the use of longer peptides may allow for endogenous processing by patient cells, resulting in more effective antigen presentation and induction of T cell responses.

[0128] Neoantigenic peptides and polypeptides can be presented on HLA proteins. In some embodiments, the neoantigenic peptides and polypeptides are presented on HLA proteins with greater affinity than wild-type peptides. In some embodiments, the neoantigenic peptide or polypeptide can have an IC50 of at least 5000 nM or less, at least 1000 nM or less, at least 500 nM or less, at least 250 nM or less, at least 200 nM or less, at least 150 nM or less, at least 100 nM or less, at least 50 nM or less, or even less.

[0129] In some embodiments, the neoantigenic peptides and polypeptides do not induce an autoimmune response and / or do not cause immune tolerance when administered to a subject.

[0130] Also provided are compositions comprising at least two or more neoantigenic peptides. In some embodiments, the composition contains at least two different peptides. The at least two different peptides can be derived from the same polypeptide. Different polypeptides mean that the peptides differ in length, amino acid sequence, or both. The peptides can be derived from any polypeptide known or found to contain tumor-specific mutations. Suitable polypeptides from which neoantigenic peptides can be derived can be found, for example, in the COSMIC database. COSMIC manages comprehensive information on somatic mutations in human cancers. The peptides contain tumor-specific mutations. In some embodiments, the tumor-specific mutations are driver mutations for a particular cancer type.

[0131] Neoantigenic peptides and polypeptides with desired activities or properties can be modified to confer certain desirable attributes, e.g., improved pharmacological characteristics, while enhancing or at least retaining substantially all of the biological activity of the unmodified peptide, which binds to desired MHC molecules and activates appropriate T cells. For example, neoantigenic peptides and polypeptides can be further subjected to various modifications, such as conservative or non-conservative substitutions, which may provide certain advantages in their use, such as improved MHC binding, stability, or presentation. Conservative substitutions refer to the replacement of an amino acid residue with another that is biologically and / or chemically similar, e.g., one hydrophobic residue with another hydrophobic residue, or one polar residue with another polar residue. Substitutions include combinations such as Gly, Ala; Val, Ile, Leu, Met; Asp, Glu; Asn, Gln; Ser, Thr; Lys, Arg; and Phe, Tyr. The effects of single amino acid substitutions can also be explored using D-amino acids. Such modifications can be carried out using well-known peptide synthesis procedures, as described, for example, in Merrifield, Science 232:341-347 (1986), Barany & Merrifield, The Peptides, Gross & Meienhofer, eds. (NY, Academic Press), pp. 1-284 (1979); and Stewart & Young, Solid Phase Peptide Synthesis, (Rockford, Ill., Pierce), 2nd Ed. (1984).

[0132] Modification of peptides and polypeptides with various amino acid mimetics or unnatural amino acids can be particularly useful for increasing peptide and polypeptide stability in vivo. Stability can be assayed in a number of ways. For example, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See, e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11:291-302 (1986). Peptide half-life can be conveniently determined using a 25% human serum (v / v) assay. The protocol generally follows: Pooled human serum (type AB, non-heat-inactivated) is defatted by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture medium and used to test peptide stability. At predetermined time intervals, small aliquots of the reaction solution are removed and added to either 6% aqueous trichloroacetic acid or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun to precipitate the precipitated serum proteins. The presence of the peptide is then determined by reverse-phase HPLC using stability-specific chromatographic conditions.

[0133] Peptides and polypeptides can be modified to provide desirable attributes other than improved serum half-life. For example, the ability of a peptide to induce CTL activity can be enhanced by linking it to a sequence containing at least one epitope capable of inducing a T helper cell response. The immunogenic peptide / T helper conjugate can be linked by a spacer molecule. The spacer is typically composed of relatively small, neutral molecules, such as amino acids or amino acid mimetics, that are substantially uncharged under physiological conditions. The spacer is typically selected from, for example, Ala, Gly, or other neutral spacers of nonpolar or neutral polar amino acids. It will be understood that the optional spacer need not be composed of the same residues and can therefore be a hetero- or homo-oligomer. If present, the spacer will usually be at least one or two residues, more usually three to six residues. Alternatively, the peptide can be linked to the T helper peptide without a spacer.

[0134] The neoantigenic peptide can be linked to a T helper peptide at either the amino or carboxy terminus of the peptide, either directly or via a spacer. The amino terminus of either the neoantigenic peptide or the T helper peptide can be acylated. Exemplary T helper peptides include tetanus toxin 830-843, influenza 307-319, and malaria sporozoite peritoneal sites 382-398 and 378-389.

[0135] Proteins or peptides can be produced by any technique known to those of skill in the art, including expressing proteins, polypeptides, or peptides through standard molecular biology techniques, isolating proteins or peptides from natural sources, or chemically synthesizing proteins or peptides. Nucleotide and protein, polypeptide, and peptide sequences corresponding to various genes have been previously disclosed and can be found in computerized databases known to those of skill in the art. One such database is the Genbank and GenPept databases of the National Center for Biotechnology Information, located on the National Institutes of Health website. The coding regions of known genes can be amplified and / or expressed using the techniques disclosed herein or as known to those of skill in the art. Alternatively, various commercial preparations of proteins, polypeptides, and peptides are known to those of skill in the art.

[0136] In a further embodiment, the neoantigen comprises a nucleic acid (e.g., a polynucleotide) encoding a neoantigenic peptide or a portion thereof. The polynucleotide can be, for example, a single-stranded and / or double-stranded polynucleotide, such as DNA, cDNA, PNA, CNA, RNA (e.g., mRNA), or a polynucleotide having a phosphorothioate backbone, either in a natural or stabilized form, or a combination thereof, and may or may not contain introns. Yet a further embodiment provides an expression vector capable of expressing the polypeptide or a portion thereof. Expression vectors for various cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in the proper orientation and correct reading frame for expression. If necessary, the DNA can be linked to appropriate transcriptional and translational regulatory control nucleotide sequences recognized by the desired host; such controls are generally available in the expression vector. The vector is then introduced into the host through standard techniques. Guidance can be found, for example, in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY.

[0137] V. Vaccine Compositions Also disclosed herein are immunogenic compositions, e.g., vaccine compositions, that can generate a specific immune response, e.g., a tumor-specific immune response. Vaccine compositions typically include multiple neoantigens selected, e.g., using the methods described herein. Vaccine compositions may also be referred to as vaccines.

[0138] The vaccine can contain 1 to 30 different peptides, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 different peptides, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different peptides, or 12, 13, or 14 different peptides. The peptides can include post-translational modifications. The vaccine may contain 1 to 100 or more nucleotide sequences, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, It may contain 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different nucleotide sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different nucleotide sequences, or 12, 13, or 14 different nucleotide sequences.The vaccine contains 1-30 neoantigen sequences: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 It can contain 6, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more different neoantigen sequences, 6, 7, 8, 9, 10, 11, 12, 13, or 14 different neoantigen sequences, or 12, 13, or 14 different neoantigen sequences.

[0139] In one embodiment, the different peptides and / or polypeptides, or the nucleotide sequences encoding them, are selected such that the peptides and / or polypeptides are capable of binding to different MHC molecules, such as different MHC class I molecules and / or different MHC class II molecules. In some embodiments, a vaccine composition comprises coding sequences for peptides and / or polypeptides capable of binding to the most frequently occurring MHC class I molecules and / or MHC class II molecules. Thus, the vaccine composition can comprise different fragments capable of binding to at least two preferred, at least three preferred, or at least four preferred MHC class I molecules and / or MHC class II molecules.

[0140] The vaccine composition may generate a specific cytotoxic T cell response and / or a specific helper T cell response.

[0141] The vaccine composition can further comprise an adjuvant and / or a carrier. Examples of useful adjuvants and carriers are provided herein below. The composition can be combined with a carrier, such as a protein, or an antigen-presenting cell, such as a dendritic cell (DC), which can present peptides to T cells.

[0142] An adjuvant is any substance whose incorporation into a vaccine composition enhances or otherwise modifies the immune response to a neoantigen. The carrier can be a scaffold, such as a polypeptide or polysaccharide, to which the neoantigen can be bound. Optionally, the adjuvant is covalently or non-covalently conjugated.

[0143] The ability of adjuvant to increase the immune response to antigen is typically manifested by a significant or substantial increase in immune-mediated reaction or a reduction in disease symptoms.For example, the increase in humoral immunity is typically manifested by a significant increase in the titer of antibody produced against antigen, and the increase in T cell activity is typically manifested in an increase in cell proliferation, or cellular cytotoxicity, or cytokine secretion.Adjuvant can also change immune response, for example, by changing mainly humoral or Th response to mainly cellular or Th response.

[0144] Suitable adjuvants include 1018 ISS, alum, aluminum salts, Amplivax, AS15, BCG, CP-870,893, CpG7909, CyaA, dSLIM, GM-CSF, IC30, IC31, imiquimod, ImuFact IMP321, IS Patch, ISS, ISCOMATRIX, JuvImmune, LipoVac, MF59, monophosphoryl lipid A, Montanide IMS 1312, Montanide ISA206, Montanide ISA 50V, and Montanide. Adjuvants include, but are not limited to, ISA-51, OK-432, OM-174, OM-197-MP-EC, ONTAK, PepTel vector system, PLG microparticles, resiquimod, SRL172, virosomes and other virus-like particles, YF-17D, VEGF trap, R848, β-glucan, Pam3Cys, Aquila's QS21 stimulon (Aquila Biotech, Worcester, Mass., USA) derived from saponins, mycobacterial extracts and synthetic bacterial cell wall mimics, and other proprietary adjuvants such as Ribi's Detox.Quil or Superfos. Adjuvants such as incomplete Freund's or GM-CSF are also useful. Several immunological adjuvants specific for dendritic cells (e.g., MF59) and their preparations have been previously described (Dupuis M, et al., Cell Immunol. 1998;186(1):18-27; Allison AC; Dev Biol Stand. 1998;92:3-11). Cytokines can also be used. Several cytokines have been directly linked to influencing dendritic cell migration to lymphoid tissues (e.g., TNF-α), accelerating dendritic cell maturation into efficient antigen-presenting cells for T lymphocytes (e.g., GM-CSF, IL-1, and IL-4) (U.S. Pat. No. 5,849,589, specifically incorporated herein by reference in its entirety), and acting as immune adjuvants (e.g., IL-12) (Gabrilovich DI, et al., J Immunother Emphasis Tumor Immunol. 1996(6):414-418).

[0145] CpG immunostimulatory oligonucleotides have also been reported to enhance the effects of adjuvants in vaccine settings. Other TLR-binding molecules, such as RNA that binds to TLR 7, TLR 8, and / or TLR 9, may also be used.

[0146] Other examples of useful adjuvants include, but are not limited to, chemically modified CpG (e.g., CpR, Idera), Poly(I:C) (e.g., polyi:CI2U), non-CpG bacterial DNA or RNA, and immunologically active small molecules and antibodies, such as cyclophosphamide, sunitinib, bevacizumab, Celebrex, NCX-4016, sildenafil, tadalafil, vardenafil, sorafinib, XL-999, CP-547632, pazopanib, ZD2171, AZD2171, ipilimumab, tremelimumab, and SC58175, which may act therapeutically and / or as adjuvants. The amounts and concentrations of adjuvants and additives can be readily determined by one of ordinary skill in the art without undue experimentation. Additional adjuvants include colony-stimulating factors such as granulocyte-macrophage colony-stimulating factor (GM-CSF, sargramostim).

[0147] A vaccine composition can include more than one different adjuvant. Additionally, a therapeutic composition can include any adjuvant material, including any of the above or combinations thereof. It is also contemplated that the vaccine and adjuvant can be administered together or separately in any suitable sequence.

[0148] The carrier (or excipient) can exist independently of the adjuvant. The function of the carrier can be, for example, to increase activity or immunogenicity, to provide stability, to increase biological activity, or to increase serum half-life, particularly to increase the molecular weight of the variant. Furthermore, the carrier can help present the peptide to T cells. The carrier can be any suitable carrier known to those skilled in the art, such as a protein or an antigen-presenting cell. The carrier protein can be, but is not limited to, a serum protein such as keyhole limpet hemocyanin, transferrin, bovine serum albumin, human serum albumin, thyroglobulin or ovalbumin, an immunoglobulin, or a hormone such as insulin, or palmitic acid. For human immunization, the carrier is generally a physiologically acceptable carrier that is tolerated and safe for humans. However, tetanus toxoid and / or diphtheria toxoid are suitable carriers. Alternatively, the carrier can be a dextran, such as Sepharose.

[0149] Cytotoxic T cells (CTLs) recognize antigens in the form of peptides bound to MHC molecules rather than the intact foreign antigen itself. MHC molecules themselves are located on the cell surface of antigen-presenting cells. Therefore, CTL activation is possible when a trimeric complex of peptide antigen, MHC molecule, and APC is present. Correspondingly, not only when peptides are used to activate CTLs, but also when APCs bearing the respective MHC molecules are added, it can enhance the immune response. Therefore, in some embodiments, the vaccine composition additionally contains at least one antigen-presenting cell.

[0150] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba virus, adenovirus (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1):45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3):603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter, Nucl. Acids Res. (2015) 43(1):682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72(12):9873-9880), etc. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22(4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352(6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20( 13):3401-10). Upon introduction into the host, the infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.

[0151] VA neoantigen cassette The methods used for the selection of one or more neoantigens, cloning and construction of "cassettes," and their insertion into viral vectors are within the skill of the art, given the teachings provided herein. A "neoantigen cassette" refers to the combination of a selected neoantigen or neoantigens with other regulatory elements necessary to transcribe the neoantigen(s) and express the transcripts. The neoantigen or neoantigens can be operably linked to regulatory elements in a manner that allows transcription. Such elements include conventional regulatory elements capable of driving expression of the neoantigen(s) in cells transfected with the viral vector. Thus, the neoantigen cassette can also include a selected promoter linked to the neoantigen(s) and located within the selected viral sequence of the recombinant vector, along with any other regulatory elements.

[0152] Useful promoters include constitutive promoters or regulated (inducible) promoters, which allow for control of the amount of neoantigen(s) expressed. For example, a desirable promoter is the cytomegalovirus immediate-early promoter / enhancer promoter [see, e.g., Boshart et al., Cell, 41:521-530 (1985)]. Another desirable promoter is the Rous sarcoma virus LTR promoter / enhancer. Yet another promoter / enhancer sequence is the chicken cytoplasmic beta-actin promoter [TAKost et al., Nucl. Acids Res., 11(23):8287 (1983)]. Other suitable or desirable promoters can be selected by those skilled in the art.

[0153] The neoantigen cassette can also include nucleic acid sequences heterologous to the viral vector sequence, including sequences providing signals for efficient polyadenylation of the transcript (poly-A, or pA), and an intron with functional splice donor and acceptor sites. A common polyA sequence for use in exemplary vectors of the invention is derived from the papovavirus SV-40. The polyA sequence can generally be inserted into the cassette after the neoantigen-based sequence and before the viral vector sequence. A common intron sequence can be derived from SV-40, also referred to as the SV-40 T intron sequence. The neoantigen cassette can also include such an intron located between the promoter / enhancer sequence and the neoantigen(s). Selection of these and other common vector elements is conventional (see, e.g., Sambrook et al., "Molecular Cloning. A Laboratory Manual," 2d ed., Cold Spring Harbor Laboratory, New York (1989), and references cited therein), and many such sequences are available from commercial and industrial sources, as well as from Genbank.

[0154] A neoantigen cassette can have one or more neoantigens. For example, a given cassette can include 1-10, 1-20, 1-30, 10-20, 15-25, 15-20, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more neoantigens. Neoantigens can be directly linked to each other. Neoantigens can also be linked to each other using a linker. Neoantigens can be in any relative geometric configuration to each other, such as NC or CN.

[0155] As noted above, the neoantigen cassette can be located at any deletion site of choice in the viral vector, such as at the site of a selectable E1 gene region deletion or E3 gene region deletion, among others.

[0156] VB immune checkpoint A vector described herein, such as a C68 vector described herein or an alphavirus vector described herein, can contain a nucleic acid encoding at least one neoantigen, and the same or a separate vector contains a nucleic acid encoding at least one immune modulator (e.g., an antibody such as an scFv) that binds to and blocks the activity of an immune checkpoint molecule. The vector can contain a neoantigen cassette and one or more nucleic acid molecules encoding a checkpoint inhibitor.

[0157] Examples of immune checkpoint molecules that can be targeted for blocking or inhibition include, but are not limited to, CTLA-4, 4-1BB (CD137), 4-1BBL (CD137L), PDL1, PDL2, PD1, B7-H3, B7-H4, BTLA, HVEM, TIM3, GAL9, LAG3, TIM3, B7H3, B7H4, VISTA, KIR, 2B4 (which belongs to the CD2 family of molecules and is expressed on all NK, gamma delta, and memory CD8+ (alpha beta) T cells), CD160 (also known as BY55), and CGEN-15049. Immune checkpoint inhibitors include antibodies, antigen-binding fragments thereof, or other binding proteins that bind to and block or inhibit the activity of one or more of CTLA-4, PDL1, PDL2, PD1, B7-H3, B7-H4, BTLA, HVEM, TIM3, GAL9, LAG3, TIM3, B7H3, B7H4, VISTA, KIR, 2B4, CD160, and GEN-15049. Examples of immune checkpoint inhibitors include tremelimumab (a CTLA-4 blocking antibody), anti-OX40, PD-L1 monoclonal antibody (anti-B7-H1; MEDI4736), ipilimumab, MK-3475 (a PD-1 blocker), nivolumab (an anti-PD1 antibody), CT-011 (an anti-PD1 antibody), BY55 monoclonal antibody, AMP224 (an anti-PDL1 antibody), BMS-936559 (an anti-PDL1 antibody), MPLDL3280A (an anti-PDL1 antibody), MSB0010718C (an anti-PDL1 antibody), and yervoy / ipilimumab (an anti-CTLA-4 checkpoint inhibitor). The antibody-encoding sequence can be engineered into a vector such as C68 using techniques routine in the art. Exemplary methods are described in Fang et al., Stable antibody expression at therapeutic levels using the 2A peptide. Nat Biotechnol. 2005 May;23(5):584-90. Epub 2005 Apr 17, the contents of which are incorporated by reference herein for all purposes.

[0158] Further considerations for VA vaccine design and manufacturing VA1. Determination of a set of peptides covering all tumor subclones Truncal peptides, meaning those presented by all or most tumor subclones, are prioritized for inclusion in the vaccine. 53 Optionally, if there are no truncal peptides that are predicted to be highly likely to be presented and immunogenic, or if the number of truncal peptides that are predicted to be highly likely to be presented and immunogenic is small enough that additional non-truncal peptides can be included in the vaccine, additional peptides can be prioritized by estimating the number and identity of tumor subclones and selecting peptides to maximize the number of tumor subclones covered by the vaccine. 54 .

[0159] VA2. Neoantigen prioritization After applying all of the above neoantigen filters, it is possible that more candidate neoantigens remain available for vaccine inclusion than vaccine technology can accommodate. Additionally, uncertainty about various aspects of neoantigen analysis may remain, and trade-offs may exist between various attributes of candidate vaccine neoantigens. Therefore, instead of predetermined filters at each stage of the selection process, an integral multidimensional model can be considered, in which candidate neoantigens are placed in a space with at least the following axes, and selection is optimized using an integral approach: 1. Risk of autoimmunity or tolerance (germline risk) (lower autoimmune risk is typically preferred) 2. Probability of sequencing artifacts (lower artifact probabilities are typically preferred) 3. Probability of immunogenicity (higher probability of immunogenicity is typically preferred) 4. Probability of presentation (higher probability of presentation is typically preferable) 5. Gene Expression (higher expression is typically preferred) 6. HLA gene coverage (a greater number of HLA molecules involved in presenting a set of neoantigens may decrease the probability that tumors will evade immune attack through downregulation or mutation of HLA molecules) 7. HLA class coverage (covering both HLA-I and HLA-II may increase the probability of therapeutic response and decrease the probability of tumor immune evasion)

[0160] VI. Methods of Treatment and Manufacturing Also provided are methods for inducing a tumor-specific immune response in a subject, vaccinating against the tumor, and treating and / or alleviating symptoms of cancer in a subject by administering to the subject one or more neoantigens, such as multiple neoantigens identified using the methods disclosed herein.

[0161] In some embodiments, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject can be a human, dog, cat, horse, or any animal in which a tumor-specific immune response is desired. The tumor can be any solid tumor, such as breast, ovarian, prostate, lung, kidney, stomach, colon, testicular, head and neck, pancreas, brain, melanoma, and other tissue organ tumors, as well as hematological tumors, such as lymphomas and leukemias, including acute myeloid leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, T-cell lymphocytic leukemia, and B-cell lymphoma.

[0162] The neoantigen can be administered in an amount sufficient to induce a CTL response.

[0163] The neoantigen can be administered alone or in combination with other therapeutic agents, such as chemotherapeutic agents, radiation, or immunotherapy. Any suitable therapeutic treatment for the particular cancer can be administered.

[0164] In addition, the subject can be further administered with an anti-immunosuppressive / immunostimulatory substance, such as a checkpoint inhibitor.For example, the subject can be further administered with an anti-CTLA antibody, or anti-PD-1 or anti-PD-L1.Blocking CTLA-4 or PD-L1 with an antibody can enhance the immune response against cancerous cells in patients.In particular, blocking CTLA-4 has been shown to be effective when used in vaccination protocols.

[0165] The optimal amount of each neoantigen to be included in the vaccine composition and the optimal dosing regimen can be determined. For example, the neoantigen or its variants can be formulated for intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. Methods of injection include sc, id, ip, im, and iv. Methods of DNA or RNA injection include id, im, sc, ip, and iv. Other methods of administering vaccine compositions are known to those skilled in the art.

[0166] Vaccines can be edited so that the selection, number, and / or amount of neoantigens present in the composition are tissue-, cancer-, and / or patient-specific. For example, the exact selection of peptides can be guided by the expression pattern of the parent protein in a given tissue. Selection can depend on the specific type of cancer, the state of the disease, earlier treatment regimens, the patient's immune status, and, of course, the patient's HLA haplotype. Furthermore, vaccines can contain components that are personalized according to the individual needs of a particular patient. Examples include altering the selection of neoantigens according to the expression of neoantigens in a particular patient, or adjusting for secondary treatments after a first round or scheme of treatment.

[0167] For compositions to be used as vaccines for cancer, neoantigens with similar normal self-peptides that are abundantly expressed in normal tissues can be avoided or present in low amounts in the compositions described herein. On the other hand, if a patient's tumor is known to express a high amount of a particular neoantigen, the respective pharmaceutical composition for treating that cancer can be present in high amount and / or can include more than one neoantigen specific to that particular neoantigen or pathway of that neoantigen.

[0168] Compositions containing neoantigens can be administered to individuals already suffering from cancer. In therapeutic applications, the compositions are administered to patients in an amount sufficient to elicit an effective CTL response against the tumor antigen and cure or at least partially halt symptoms and / or complications. An amount adequate to achieve this is defined as a "therapeutically effective dose." An amount effective for this use will depend, for example, on the composition, the mode of administration, the stage and severity of the disease being treated, the patient's weight and general health, and the judgment of the prescribing physician. It should be kept in mind that compositions can generally be used in severe disease states, i.e., life-threatening or potentially life-threatening situations, particularly when cancer has metastasized. In such instances, it is possible, and the treating physician may find it desirable, to administer substantial excesses of these compositions, taking into account the minimization of adventitious substances and the relatively non-toxic nature of the neoantigens.

[0169] For therapeutic use, administration can begin at the time of detection or surgical removal of a tumor, followed by boosting doses until at least symptoms are substantially abated, and for a period thereafter.

[0170] Pharmaceutical compositions for therapeutic treatment (e.g., vaccine compositions) are intended for parenteral, topical, nasal, oral, or local administration. Pharmaceutical compositions can be administered parenterally, for example, intravenously, subcutaneously, intradermally, or intramuscularly. The compositions can be administered at the site of surgical resection to induce a local immune response against tumors. Disclosed herein are compositions for parenteral administration that contain a solution of a neoantigen, where the vaccine composition is dissolved or suspended in an acceptable carrier, e.g., an aqueous carrier. Various aqueous carriers can be used, such as water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid, and the like. These compositions can be sterilized by conventional, well-known sterilization techniques or sterile filtered. The resulting aqueous solutions can be packaged for use as is or lyophilized, with the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may contain pharmaceutically acceptable auxiliary substances required to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, and the like, for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and the like.

[0171] Neoantigens can also be administered via liposomes, which target them to specific cellular tissues, such as lymphoid tissues. Liposomes are also useful for increasing half-life. Liposomes include emulsions, foams, micelles, insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers, and the like. In these preparations, the neoantigen to be delivered is incorporated as part of the liposome, either alone or in combination with a molecule that binds to a receptor dominant among lymphoid cells, such as a monoclonal antibody that binds to the CD45 antigen, or with other therapeutic or immunogenic compositions. Liposomes filled with the desired neoantigen can thus be directed to the site of lymphoid cells, where they then deliver the selected therapeutic / immunogenic composition. Liposomes can generally be formed from standard vesicle-forming lipids, including neutral and negatively charged phospholipids and sterols, such as cholesterol. The choice of lipid is generally guided by considerations, for example, of liposome size, acid lability, and stability of the liposomes in the bloodstream. Various methods are available for preparing liposomes, as described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980), U.S. Pat. Nos. 4,235,871, 4,501,728, 4,501,728, 4,837,028, and 5,019,369.

[0172] For targeting to immune cells, the ligand to be incorporated into the liposome can include, for example, an antibody or fragment thereof specific for a cell surface determinant of the desired immune system cell. The liposome suspension can be administered intravenously, topically, locally, etc., at doses that vary according to, inter alia, the mode of administration, the peptide being delivered, and the stage of the disease being treated.

[0173] The peptides described herein, and optionally nucleic acids encoding one or more of the peptides, can also be administered to a patient for therapeutic or immunization purposes. Numerous methods are conveniently used to deliver nucleic acids to a patient. For example, nucleic acids can be delivered directly as "naked DNA." This approach is described, for example, in Wolff et al., Science 247:1465-1468 (1990), and U.S. Pat. Nos. 5,580,859 and 5,589,466. Nucleic acids can also be administered using ballistic delivery, as described, for example, in U.S. Pat. No. 5,204,253. Particles consisting solely of DNA can be administered. Alternatively, DNA can be attached to particles, such as gold particles. Approaches for delivering nucleic acid sequences include viral vectors, mRNA vectors, and DNA vectors, with or without electroporation.

[0174] Nucleic acids can also be delivered by complexing them with cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for example, in 9618372 WOAWO 96 / 18372; 9324640 WOAWO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Patent No. 5,279,833 Rose, U.S. Patent No. 5,279,833; 9106309 WOAWO 91 / 06309; and Felgner et al., Proc.Natl.Acad.Sci.USA 84: 7413-7414 (1987).

[0175] Neoantigens can also be derived from viruses such as vaccinia, fowlpox, self-replicating alphaviruses, Maraba viruses, adenoviruses (see, e.g., Tatsis et al., Adenoviruses, Molecular Therapy (2004) 10, 616-629), or lentiviruses, including, but not limited to, second, third, or hybrid second / third generation lentiviruses, and any generation of recombinant lentiviruses designed to target specific cell types or receptors (see, e.g., Hu et al., Immunization Delivered by Lentiviral Vectors for Cancer and Infectious Diseases, Immunol Rev. (2011) 239(1): 45-61; Sakuma et al., Lentiviral vectors: basic to translational, Biochem J. (2012) 443(3): 603-18; Cooper et al., Rescue of splicing-mediated intron loss maximizes expression in lentiviral vectors containing the human The ubiquitin C promoter can also be included in a viral vector-based vaccine platform, such as the ubiquitin C promoter, Nucl. Acids Res. (2015) 43 (1): 682-690; Zufferey et al., Self-Inactivating Lentivirus Vector for Safe and Efficient In Vivo Gene Delivery, J. Virol. (1998) 72 (12): 9873-9880. Depending on the packaging capacity of the viral vector-based vaccine platform described above, this approach can deliver one or more nucleotide sequences encoding one or more neoantigen peptides.The sequence may be flanked by non-mutated sequences, separated by linkers, or preceded by one or more sequences that target intracellular compartments (see, e.g., Gros et al., Prospective identification of neoantigen-specific lymphocytes in the peripheral blood of melanoma patients, Nat Med. (2016) 22 (4):433-8; Stronen et al., Targeting of cancer neoantigens with donor-derived T cell receptor repertoires, Science. (2016) 352 (6291):1337-41; Lu et al., Efficient identification of mutated cancer antigens recognized by T cells associated with durable tumor regressions, Clin Cancer Res. (2014) 20(13):3401-10). Upon introduction into the host, infected cells express the neoantigen, thereby eliciting a host immune (e.g., CTL) response against the peptide. Vaccinia vectors and methods useful in immunization protocols are described, for example, in U.S. Patent No. 4,722,848. Another vector is Bacillus Calmette-Guerin (BCG). BCG vectors are described by Stover et al. (Nature 351:456-460 (1991)). A wide variety of other vaccine vectors useful for therapeutic administration or immunization of neoantigens, such as Salmonella typhi vectors, will be apparent to those skilled in the art from the description herein.

[0176] A means of administering nucleic acids uses minigene constructs encoding one or more epitopes. To generate DNA sequences (minigenes) encoding selected CTL epitopes for expression in human cells, the amino acid sequences of the epitopes are reverse-translated. A human codon usage table is used to guide codon selection for each amino acid. The DNA sequences encoding these epitopes are then directly adjacent to generate a continuous polypeptide sequence. Additional elements can be incorporated into the minigene design to optimize expression and / or immunogenicity. Examples of amino acid sequences that can be reverse-translated and included in the minigene sequence include helper T lymphocyte epitopes, leader (signal) sequences, and endoplasmic reticulum retention signals. In addition, MHC presentation of CTL epitopes can be improved by including synthetic (e.g., polyalanine) or naturally occurring flanking sequences adjacent to the CTL epitopes. The minigene sequence is converted to DNA by assembling oligonucleotides encoding the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified, and annealed under appropriate conditions using well-known techniques. The ends of the oligonucleotides are ligated using T4 DNA ligase. This synthetic minigene encoding the CTL epitope polypeptide can then be cloned into a desired expression vector.

[0177] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is reconstitution of lyophilized DNA in sterile phosphate-buffered saline (PBS). Various methods have been described, and new techniques may become available. As mentioned above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides, and compounds collectively referred to as protective, interactive, non-condensing (PINC) compounds can also be complexed with purified plasmid DNA to affect variables such as stability, intramuscular distribution, or transport to specific organs or cell types.

[0178] Also disclosed herein is a method of producing a tumor vaccine, comprising performing the steps of the methods disclosed herein; and producing a tumor vaccine comprising multiple neoantigens or a subset of multiple neoantigens.

[0179] The neoantigens disclosed herein can be produced using methods known in the art. For example, a method for producing a neoantigen or vector (e.g., a vector comprising at least one sequence encoding one or more neoantigens) disclosed herein can include culturing host cells under conditions suitable for expression of the neoantigen or vector, wherein the host cells comprise at least one polynucleotide encoding the neoantigen or vector, and purifying the neoantigen or vector. Standard purification methods include chromatographic, electrophoretic, immunological, precipitation, dialysis, filtration, concentration, and chromatofocusing techniques.

[0180] The host cell can comprise a Chinese hamster ovary (CHO) cell, an NS0 cell, yeast, or an HEK293 cell. The host cell can be transformed with one or more polynucleotides comprising at least one nucleic acid sequence encoding a neoantigen or vector disclosed herein, and optionally, the isolated polynucleotide further comprises a promoter sequence operably linked to the at least one nucleic acid sequence encoding the neoantigen or vector. In certain embodiments, the isolated polynucleotide can be a cDNA.

[0181] VII. Identification of neoantigens VII.A. Identification of Candidate Neoantigens A research method for NGS analysis of tumor and normal exomes and transcriptomes is described and applied in the specific space of neoantigens. 6,14,15The examples below consider certain optimizations for greater sensitivity and specificity for identifying neoantigens in a clinical setting. These optimizations can be grouped into two areas: those related to laboratory processes and those related to NGS data analysis.

[0182] VII.A.1. Laboratory Process Optimization The process improvements presented herein build on the concepts developed for reliable assessment of cancer driver genes in targeted cancer panels. 16 This addresses the challenges in high-precision neoantigen discovery from clinical specimens with low tumor content and small volumes by expanding the method to the whole-exome and whole-transcriptome settings required for neoantigen identification. Specifically, these improvements include: 1. Targeting deep (greater than 500x) unique average coverage across the tumor exome to detect mutations present at low mutant allele frequency due to either low tumor content or subclonal status. 2. Fewer than 5% of bases are covered at less than 100x to minimize missed potential neoantigens, e.g. a. Use of DNA-based capture probes with individual probe QC 17 b. Inclusion of additional baits for areas that are not sufficiently covered 3. Targeting uniform coverage across the normal exome, with less than 5% of bases covered below 20x, to minimize the chance of potential neoantigens remaining unclassified for somatic / germline status (and therefore unusable as TSNAs). 4. To minimize the total amount of sequencing required, sequence capture probes are designed only for the coding regions of the gene, since non-coding RNAs cannot give rise to neoantigens. Additional optimizations include: a. Supplementary probes for HLA genes that are GC-rich and not well captured by standard exome sequencing18 . b. Elimination of genes predicted to produce few or no candidate neoantigens due to factors such as poor expression, suboptimal digestion by the proteasome, or atypical sequence characteristics. 5. Tumor RNA is also sequenced at high depth (greater than 100M reads) to enable mutation detection, quantification of gene and splice variant ("isoform") expression, and fusion detection. RNA from FFPE samples can be subjected to probe-based enrichment with the same or similar probes used to capture the exome in DNA. 19 It is extracted using

[0183] VII.A.2. Optimizing NGS Data Analysis Analytical method improvements address the suboptimal sensitivity and specificity of common research variant calling approaches and specifically allow for customization relevant for identifying neoantigens in the clinical setting. These include: 1. Use of the HG38 reference human genome or later versions for alignment, as it contains multiple MHC region assemblies that better reflect population polymorphism, as opposed to earlier genome releases. 2. Various programs 5 Overcoming the limitations of single mutation callers 20 by merging results from a. Single nucleotide mutations and indels are detected in tumor DNA, tumor RNA, and normal DNA with a range of tools including: Strelka 21 and Mutect 22 and programs based on comparison of tumor and normal DNA, such as; and 23 , UNCeqR, and other programs that incorporate tumor DNA, tumor RNA, and normal DNA. b. Indels are found in Strelka and ABRA 24 This is determined by a program that performs local reassembly, such as c. Structural rearrangements are 25or Breakseq 26 It is determined using specialized tools such as 3. To detect and prevent sample swapping, mutation calls from samples for the same patient are compared at a selected number of polymorphic sites. 4. Extensive filtering of artificial calls is performed, for example, by: a. Removal of mutations found in normal DNA, potentially with relaxed detection parameters in the case of low coverage and permissive proximity criteria in the case of indels. b. Removal of mutations due to poor mapping quality or poor base quality 27 . c. Elimination of mutations resulting from re-emerging sequencing artifacts, even if not observed in the corresponding normal 27 Examples include mutations that are detected primarily on one strand. d. Removal of mutations detected in a set of unrelated controls 27 . 5.seq2HLA 28 , ATHLATES 29 or Optitype, and also combine exome and RNA sequencing data 28 , accurate HLA calling from normal exomes. Additional potential optimizations include the adoption of dedicated assays for HLA typing, such as long-read DNA sequencing. 30 or adaptation of methods for linking RNA fragments to maintain continuity. 31 Includes. 6. Robust Detection of Nascent ORFs Arising from Tumor-specific Splice Variants in CLASS 32 , Bayesembler 33 , StringTie 34 This is done by assembling transcripts from RNA-seq data using Cufflinks, or a similar program in its reference-guided mode (i.e., using known transcript structures rather than attempting to recreate the entire transcripts from each experiment). 35Although commonly used for this purpose, it frequently produces an incredibly large number of splice variants, many of which are much shorter than the full-length gene, and may not be able to recover a simple positive control. The coding sequence and potential nonsense-mediated decay mechanisms reintroduced the mutant sequence, SpliceR. 36 and MAMBA 37 Gene expression is determined using tools such as Cufflinks 35 or Express (Roberts and Pachter, 2013). Wild-type and mutant-specific expression counts and / or relative levels are determined using ASE. 38 or HTSeq 39 Potential filtering steps include: a. Removal of candidate nascent ORFs that are thought to be poorly expressed. b. Removal of candidate nascent ORFs predicted to trigger nonsense-mediated decay (NMD). 7. Candidate neoantigens observed only in RNA (e.g., neo-ORFs) that cannot be directly validated as tumor-specific are classified as likely to be tumor-specific according to additional parameters, for example, by considering the following: a. Presence of supporting cis-acting frameshift or splice site mutations in tumor DNA only. b. The presence of confirmed trans-acting mutations in splicing factors in tumor DNA only. As an example, the gene that exhibited the most differential splicing in three independently published experiments with R625 mutant SF3B1 was 1. One experiment examined patients with uveal melanoma. 40 The second experiment examined uveal melanoma cell lines. 41 , and a third study looked at breast cancer patients. 42 Nevertheless, there was agreement. c. For novel splicing isoforms, the presence of confirmatory "novel" splice-junction reads in the RNASeq data. d. For de novo rearrangements, the presence of confirmatory exon-proximal reads in tumor DNA that are not present in normal DNA. e.GTEx 43 and absence from the gene expression compendium (i.e., making germline origin less likely). 8. Complementing reference genome alignment-based analyses by comparing tumor and normal reads (or k-mers derived from such reads) of assembled DNA to directly avoid alignment- and annotation-based errors and artifacts (e.g., for somatic mutations occurring near germline mutations or repeat-context indels).

[0184] In samples with polyadenylated RNA, the presence of viral and microbial RNA in the RNA-seq data will be assessed using RNA CoMPASS44 or similar methods to identify additional factors that may predict patient response.

[0185] VII.B. HLA Peptide Isolation and Detection Isolation of HLA peptide molecules was performed using classical immunoprecipitation (IP) techniques after lysis and solubilization of tissue samples. 55~58 The clarified lysates were used for HLA-specific IP.

[0186] Immunoprecipitation was performed using antibodies coupled to beads, where the antibodies are specific for HLA molecules. For pan-class I HLA immunoprecipitation, a pan-class I CR antibody is used, and for class II HLA-DR, an HLA-DR antibody is used. The antibodies are covalently attached to NHS-Sepharose beads during overnight incubation. After covalent attachment, the beads were washed and aliquoted for IP. 59、60Immunoprecipitation can also be performed using antibodies that are not covalently attached to beads. Typically, this is done using Sepharose or magnetic beads coated with Protein A and / or Protein G to retain the antibody on the column. Some antibodies that can be used to selectively enrich MHC / peptide complexes are listed below. TIFF2025175055000014.tif42146

[0187] The clarified tissue lysate is added to antibody beads for immunoprecipitation. After immunoprecipitation, the beads are removed from the lysate and the lysate is saved for further experiments, including additional IPs. The IP beads are washed to remove nonspecific binding, and the HLA / peptide complexes are eluted from the beads using standard techniques. Protein components are removed from the peptides using molecular weight spin columns or C18 fractionation. The resulting peptides are dried by SpeedVac evaporation and, in some cases, stored at -20°C prior to MS analysis.

[0188] The dried peptides were reconstituted in an HPLC buffer suitable for reversed-phase chromatography and loaded onto a C-18 microcapillary HPLC column for gradient elution on a Fusion Lumos mass spectrometer (Thermo). MS1 spectra of peptide mass / charge (m / z) were collected at high resolution on an Orbitrap detector, followed by MS2 low-resolution scans on an ion trap detector after HCD fragmentation of selected ions. Additionally, MS2 spectra can be acquired using either CID or ETD fragmentation methods, or any combination of the three techniques to obtain greater amino acid coverage of the peptide. MS2 spectra can also be measured with high-resolution mass accuracy on an Orbitrap detector.

[0189] The MS2 spectra from each analysis were analyzed using Comet 61、62 and peptide identifications were analyzed using Percolator63~65 Further sequencing is performed using PEAKS studio (Bioinformatics Solutions Inc.) and other search engines, or spectral matching and de novo sequencing are performed. 75 Sequencing methods including:

[0190] VII.B.1. Investigation of MS detection limits for comprehensive HLA peptide sequencing Using the peptide YVYVADVAAK (SEQ ID NO:1), the limit of detection was determined using various amounts of peptide loaded onto the LC column. The amounts of peptide tested were 1 pmol, 100 fmol, 10 fmol, 1 fmol, and 100 amol (Table 1). The results are shown in Figure 1F. These results indicate that the lowest limit of detection (LoD) was in the attomolar range (10 -18 ), a dynamic range spanning five orders of magnitude, and a signal-to-noise ratio in the low femtomole range (10 -15 ) appears to be sufficient for sequencing. TIFF2025175055000015.tif62128

[0191] VIII. Presented Model VIII.A. System Overview 2A is an overview of an environment 100 for identifying the likelihood of peptide presentation in a patient, according to one embodiment. The environment 100 provides a context for implementing a presentation identification system 160, which itself includes a presentation information store 165.

[0192] The presentation identification system 160 is a computer model, embodied in a computational system such as that discussed below with respect to FIG. 14, that receives a peptide sequence associated with a set of MHC alleles and determines the likelihood that the peptide sequence will be presented by one or more of the set of associated MHC alleles. The presentation identification system 160 can be applied to both class I and class II MHC alleles, making it useful in a variety of contexts. One specific example application of the presentation identification system 160 is to receive the nucleotide sequence of a candidate neoantigen associated with a set of MHC alleles derived from tumor cells in a patient 110 and determine the likelihood that the candidate neoantigen will be presented by one or more of the tumor's associated MHC alleles and / or will induce an immunogenic response in the patient's 110 immune system. Those candidate neoantigens with a high likelihood, as determined by the system 160, can be selected for inclusion in a vaccine 118, such that an anti-tumor immune response can be elicited from the immune system of the patient 110 that provided the tumor cells.

[0193] The presentation identification system 160 determines presentation likelihood through one or more presentation models. Specifically, a presentation model generates a likelihood that a given peptide sequence will be presented for a set of associated MHC alleles, the likelihood being generated based on the presentation information stored in the storage device 165. For example, a presentation model may determine whether the peptide sequence "YVYVADVAAK" (SEQ ID NO:1) will be presented for a set of alleles, HLA-A, on the cell surface of a sample. * 02:01, HLA-A * 03:01, HLA-B * 07:02, HLA-B * 08:03, HLA-C *01:04。 Presentation information 165 contains information about whether these peptides bind to various types of MHC alleles so that the peptides are presented by the MHC alleles, which is determined in the model according to the position of the amino acid in the peptide sequence. Based on the presentation information 165, the presentation model can predict whether an unrecognized peptide sequence will be presented in association with a related set of MHC alleles. As mentioned above, the presentation model can be applied to both class I and class II MHC alleles.

[0194] VIII.B. Presentation information 2 illustrates a method for obtaining presentation information according to one embodiment. Presentation information 165 includes two general categories of information: allele interaction information and allele non-interaction information. Allele interaction information includes information that affects presentation of peptide sequences that is dependent on the type of MHC allele. Allele non-interaction information includes information that affects presentation of peptide sequences that is independent of the type of MHC allele.

[0195] VIII.B.1. Allelic Interaction Information The allele interaction information primarily includes identified peptide sequences known to be presented by one or more identified MHC molecules from humans, mice, etc. Of note, this may or may not include data obtained from tumor samples. Presented peptide sequences may be identified from cells expressing a single MHC allele. In this example, presented peptide sequences are generally collected from a single-allelic cell line engineered to express a predetermined MHC allele and then exposed to a synthetic protein. Peptides presented on the MHC allele are isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2B shows the peptides presented on a cell line expressing a predetermined MHC allele, HLA-DRB1. * 12:01 Exemplary peptides presented above An example of this is shown in TIFF2025175055000016.tif4128, where a peptide is isolated and identified by mass spectrometry. In this situation, the direct association between the presented peptide and the MHC protein to which it binds is definitively known, since the peptide is identified through cells engineered to express a single, predetermined MHC protein.

[0196] Presented peptide sequences may also be collected from cells expressing multiple MHC alleles. Typically, in humans, six different types of MHC1 molecules and up to 12 different types of MHCII molecules are expressed by cells. Such presented peptide sequences may be identified from multi-allelic cell lines engineered to express multiple predetermined MHC alleles. Such presented peptide sequences may also be identified from tissue samples, either normal or tumor tissue samples. In this particular example, MHC molecules can be immunoprecipitated from normal or tumor tissue. Peptides presented on multiple MHC alleles can similarly be isolated by techniques such as acid elution and identified by mass spectrometry. Figure 2C shows six exemplary peptides. TIFF2025175055000017.tif18170 identifies class I MHC alleles HLA-A * 01:01, HLA-A * 02:01, HLA-B * 07:02, HLA-B * 08:01, and class II MHC allele HLA-DRB1 * An example of this is shown in Figure 1, where a peptide presented on HLA-DRB1:10:01, HLA-DRB1:11:01, was isolated and identified by mass spectrometry. In contrast to monoallelic cell lines, the bound peptide is isolated from the MHC molecule before it is identified, so the direct association between the presented peptide and the MHC protein to which it is bound may be unknown.

[0197] Allele interaction information can also include mass spectrometry ion currents, which depend on both the concentration of peptide-MHC molecule complexes and the ionization efficiency of the peptides. Ionization efficiency varies from peptide to peptide in a sequence-dependent manner. Generally, ionization efficiency varies from peptide to peptide over approximately two orders of magnitude, while the concentration of peptide-MHC complexes varies over an even larger range.

[0198] Allele interaction information can also include measured or predicted binding affinities between a given MHC allele and a given peptide. One or more affinity models can generate such predictions (72, 73, 74). For example, returning to the example shown in Figure 1D, representation 165 may represent a binding affinity between the peptide YEMFNDKSF (SEQ ID NO:3) and the class I allele HLA-A * The presentation information 165 may include a predicted binding affinity value of 1000 nM between the peptide KNFLENFIESOFI (SEQ ID NO:8) and the class II allele HLA-DRB1:11:01. Few peptides with IC50>1000 nM are presented by the MHC, with lower IC50 values ​​increasing the probability of presentation. The presentation information 165 may include a predicted binding affinity value between the peptide KNFLENFIESOFI (SEQ ID NO:8) and the class II allele HLA-DRB1:11:01.

[0199] The allele interaction information can also include measured or predicted stability values ​​for MHC complexes. One or more stability models can generate such predictions. More stable peptide-MHC complexes (i.e., complexes with longer half-lives) are more likely to be presented in high copy number on tumor cells and on antigen-presenting cells that encounter vaccine antigens. For example, returning to the example shown in FIG. 2C, the presentation information 165 can include a predicted stability value for a half-life of 1 hour for the class I molecule HLA-A*01:01. The presentation information 165 can also include a predicted stability value for the half-life of the class II molecule HLA-DRB1:11:01.

[0200] Allele interaction information can also include measured or predicted rates of peptide-MHC complex formation. Complexes that form at a faster rate are more likely to be presented at high concentrations on the cell surface.

[0201] Allele interaction information can also include peptide sequence and length. MHC class I molecules typically prefer to present peptides with a length of 8-15 peptides. 60-80% of presented peptides have a length of 9 peptides. MHC class II molecules generally tend to present peptides with a length of 6-30 peptides.

[0202] Allele interaction information can also include the presence of kinase sequence motifs on the neoantigen-encoded peptide and the presence or absence of specific post-translational modifications on the neoantigen-encoded peptide. The presence of kinase motifs influences the probability of post-translational modifications that may enhance or interfere with MHC binding.

[0203] Allelic interaction information can also include expression or activity levels of proteins involved in post-translational modification processes, such as kinases (as measured or predicted by RNA-seq, mass spectrometry, or other methods).

[0204] Allelic interaction information can also include the probability of presentation of peptides with similar sequences in cells from other individuals expressing particular MHC alleles, as assessed by mass spectrometry proteomics or other means.

[0205] Allelic interaction information can also include the expression levels of particular MHC alleles in the individual in question (e.g., as measured by RNA-seq or mass spectrometry): peptides that bind most strongly to MHC alleles expressed at high levels are more likely to be presented than peptides that bind most strongly to MHC alleles expressed at low levels.

[0206] Allelic interaction information can also include the overall neoantigen-encoded peptide sequence-independent probability of presentation by a particular MHC allele in other individuals that express that particular MHC allele.

[0207] Allele interaction information can also include the overall peptide sequence-independent probability of presentation by MHC alleles of the same family of molecules (e.g., HLA-A, HLA-B, HLA-C, HLA-DQ, HLA-DR, HLA-DP) in other individuals. For example, HLA-C molecules are typically expressed at lower levels than HLA-A or HLA-B molecules, and therefore, presentation of peptides by HLA-C is a priori less likely than presentation by HLA-A or HLA-B II. As another example, because HLA-DP is generally expressed at lower levels than HLA-DR or HLA-DQ, presentation of peptides by HLA-DP is predicted to be less likely than presentation by HLA-DR or HLA-DQ.

[0208] The allele interaction information can also include the protein sequence of a particular MHC allele.

[0209] Any of the MHC allele non-interacting information listed in the section below can also be modeled as MHC allele interacting information.

[0210] VIII.B.2. Allelic Non-Interaction Information The allele-non-interacting information can include the C-terminal sequence adjacent to the neoantigen-encoded peptide within its source protein sequence. In MHC-I, the C-terminal flanking sequence can affect proteasomal processing of the peptide. However, the C-terminal flanking sequence is cleaved from the peptide by the proteasome before the peptide is transported to the endoplasmic reticulum and encounters an MHC allele on the cell surface. As a result, the MHC molecule does not receive any information about the C-terminal flanking sequence, and therefore, the effect of the C-terminal flanking sequence cannot vary depending on the MHC allele type. For example, returning to the example shown in Figure 2C, the presentation information 165 can include the C-terminal flanking sequence FOEIFNDKSLDKFJI (SEQ ID NO:9) of the presented peptide FJIEJFOESS (SEQ ID NO:5), identified from the peptide's source protein.

[0211] Allele non-interaction information can also include mRNA quantification measurements. For example, mRNA quantification data can be obtained for the same samples that provide mass spectrometry training data. As described below with respect to Figure 13H, RNA expression has been identified as a strong predictor of peptide presentation. In one embodiment, mRNA quantification measurements are determined from the software tool RSEM. A detailed implementation of the RSEM software tool can be found in Bo Li and Colin N. Dewey. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics, 12:323, August 2011. In one embodiment, mRNA quantification is measured in units of fragments per kilobase of transcript per million mapped reads (FPKM).

[0212] The allele non-interacting information can also include N-terminal sequences adjacent to the peptide within its source protein sequence.

[0213] The allelic non-interaction information can also include a source gene for the peptide sequence. The source gene can be defined as an Ensembl protein family for the peptide sequence. In another example, the source gene can be defined as a source DNA or source RNA for the peptide sequence. The source gene can be represented, for example, as a string of nucleotides that encodes a protein, or alternatively, in a more categorized form based on a named set of known DNA or RNA sequences known to encode specific proteins. In another example, the allelic non-interaction information can also include a source transcript or isoform or a set of potential source transcripts or isoforms for the peptide sequence extracted from a database such as Ensembl or RefSeq.

[0214] The allelic non-interaction information can also include the tissue type, cell type, or tumor type of the cell from which the peptide sequence is derived.

[0215] The allele non-interaction information can also include the presence of protease cleavage motifs in peptides, optionally weighted according to the expression of the corresponding proteases in tumor cells (as measured by RNA-seq or mass spectrometry). Peptides containing protease cleavage motifs are more easily degraded by proteases and therefore less stable in cells, and therefore less likely to be presented.

[0216] Allelic non-interaction information can also include the turnover rate of the source protein when measured in the appropriate cell type. A faster turnover rate (i.e., a lower half-life) increases the probability of presentation, but this characteristic has low predictive power when measured in dissimilar cell types.

[0217] The allelic non-interaction information can also include the length of the source protein, optionally taking into account the specific splice variants ("isoforms") that are most highly expressed in tumor cells, as measured by RNA-seq or proteome mass spectrometry, or as predicted from annotation of germline or somatic splicing mutations detected in DNA or RNA sequence data.

[0218] Allele-free interaction information can also include the expression level of proteasomes, immunoproteasomes, thymoproteasomes, or other proteases in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Different proteasomes have different cleavage site preferences. More weight is given to the cleavage preference of each type of proteasome in proportion to its expression level.

[0219] Allele-free interaction information can also include the expression of the peptide's source gene (e.g., as measured by RNA-seq or mass spectrometry). Possible optimizations include adjusting the measured expression to account for the presence of stromal cells and tumor-infiltrating lymphocytes in the tumor sample. Peptides from genes with higher expression are more likely to be presented. Peptides from genes with undetectable levels of expression can be eliminated from consideration.

[0220] Allelic non-interaction information can also include the probability that the source mRNA of the neoantigen-encoding peptide will be subject to nonsense-mediated decay as predicted by a model of nonsense-mediated decay, e.g., the model from Rivas et al., Science 2015.

[0221] Allelic non-interaction information can also include typical tissue-specific expression of the peptide source gene during various stages of the cell cycle. Genes that are expressed at low levels overall (as measured by RNA-seq or sample analysis proteomics) but are known to be expressed at high levels during specific stages of the cell cycle are more likely to produce peptides that are displayed than genes that are stably expressed at very low levels.

[0222] Allelic non-interaction information can also include a comprehensive catalog of source protein properties, such as those provided in uniProt or the PDB (http: / / www.rcsb.org / pdb / home / home.do). These properties can include, among others, protein secondary and tertiary structure, subcellular localization, and Gene Ontology (GO) terms. Specifically, this information can include annotations operating at the protein level, e.g., 5'UTR length, and annotations operating at the level of specific residues, e.g., a helix motif between residues 300 and 310. These properties can also include turn motifs, sheet motifs, and disordered residues.

[0223] Allelic non-interacting information can also include features that describe the nature of the domain of the source protein containing the peptide, such as secondary or tertiary structure (e.g., alpha-helix vs. beta-sheet); alternative splicing.

[0224] The allelic non-interaction information can also include properties that describe the presence or absence of presentation hotspots at the peptide's position in its source protein.

[0225] Allelic non-interaction information can also include the probability of presentation of peptides derived from the source protein of the peptide in question in other individuals (after adjusting for the expression level of the source protein in those individuals and the influence of the various HLA types of those individuals).

[0226] Allelic non-interaction information can also include the probability that a peptide will be undetected or over-represented by mass spectrometry due to technical bias.

[0227] Expression of various gene modules / pathways (not necessarily containing the source protein of the peptides) as measured by gene expression assays such as RNASeq, microarray(s), targeted panel(s) such as Nanostring, or single / multiple genes representing gene modules measured by assays such as RT-PCR, which give information about the status of tumor cells, stroma, or tumor infiltrating lymphocytes (TILs).

[0228] Allele non-interaction information can also include the copy number of the peptide's source gene in the tumor cell. For example, a peptide derived from a gene that is subject to homozygous deletion in the tumor cell can be assigned a presentation probability of zero.

[0229] The allele-non-interaction information can also include the probability that the peptide will bind to TAP, or the measured or predicted binding affinity of the peptide to TAP. Peptides that are more likely to bind to TAP or that bind with higher affinity to TAP are more likely to be presented by MHC-I.

[0230] Allelic non-interaction information can also include the expression level of TAP in tumor cells (which can be measured by RNA-seq, proteome mass spectrometry, or immunohistochemistry). Higher TAP expression levels at MHC-I increase the probability of presentation of all peptides.

[0231] Allelic non-interaction information can also include the presence or absence of tumor mutations, including but not limited to: i. Driver mutations in known cancer driver genes such as EGFR, KRAS, ALK, RET, ROS1, TP53, CDKN2A, CDKN2B, NTRK1, NTRK2, and NTRK3. ii. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOB, HLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome). Peptides whose presentation relies on components of the antigen presentation machinery affected by loss-of-function mutations in the tumor have a reduced probability of presentation.

[0232] Presence or absence of functional germline polymorphisms, including but not limited to: i. In genes encoding proteins involved in antigen presentation machinery (e.g., B2M, HLA-A, HLA-B, HLA-C, TAP-1, TAP-2, TAPBP, CALR, CNX, ERP57, HLA-DM, HLA-DMA, HLA-DMB, HLA-DO, HLA-DOA, HLA-DOBHLA-DP, HLA-DPA1, HLA-DPB1, HLA-DQ, HLA-DQA1, HLA-DQA2, HLA-DQB1, HLA-DQB2, HLA-DR, HLA-DRA, HLA-DRB1, HLA-DRB3, HLA-DRB4, HLA-DRB5, or any of the genes encoding components of the proteasome or immunoproteasome).

[0233] The allelic non-interaction information can also include tumor type (eg, NSCLC, melanoma).

[0234] Allele non-interaction information can also include the known functionality of the HLA allele, e.g., as reflected by the HLA allele suffix. For example, the N suffix in the allele name HLA-A*24:09N indicates a null allele that is not expressed and therefore unlikely to present an epitope; the complete HLA allele suffix nomenclature is described at https: / / www.ebi.ac.uk / ipd / imgt / hla / nomenclature / suffixes.html.

[0235] Allelic non-interaction information can also include clinical tumor subtype (eg, squamous cell lung cancer vs. non-squamous).

[0236] The allele non-interaction information can also include smoking history.

[0237] Allele non-interaction information can also include a history of sunburn, sun exposure, or exposure to other mutagens.

[0238] The allelic non-interaction information can also include regional expression of the peptide's source gene in relevant tumor types or clinical subtypes, optionally stratified by driver mutations. Genes that are typically expressed at high levels in relevant tumor types are more likely to be represented.

[0239] The allelic non-interaction information can also include the frequency of the mutation in all tumors, or in tumors of the same type, or in tumors from individuals with at least one shared MHC allele, or in tumors of the same type in individuals with at least one shared MHC allele.

[0240] In the example of a mutated tumor-specific peptide, the list of characteristics used to predict the probability of presentation can also include the mutation's annotation (e.g., missense, readthrough, frameshift, fusion, etc.) or whether the mutation is predicted to result in nonsense-mediated decay (NMD). For example, a peptide derived from a protein segment that is not translated in tumor cells due to a homozygous premature termination mutation can be assigned a presentation probability of zero. NMD results in reduced mRNA translation, which reduces the probability of presentation.

[0241] VIII.C. Presentation Specific Systems 3 is a high-level block diagram illustrating the computer logic components of presentation specific system 160, according to one embodiment. In this exemplary embodiment, presentation specific system 160 includes a data management module 312, an encoding module 314, a training module 316, and a prediction module 320. Presentation specific system 160 also comprises a training data store 170 and a presentation model store 175. Some embodiments of model management system 160 have different modules than those described herein. Likewise, functionality may be distributed among the modules in a manner different from that described herein.

[0242] VIII.C.1. Data Management Module The data management module 312 generates sets of training data 170 from the representation information 165. Each training data set contains a number of data examples, each of which contains at least one of the represented or unrepresented peptide sequences p i and the peptide sequence p i one or more relevant MHC alleles combined with i and the dependent variable y, which represents information that the presentation specification system 160 is interested in predicting new values ​​of the independent variables. i and the independent variable z i Contains a set of

[0243] In one particular implementation described later in this specification, the dependent variable y i is the peptide p i but one or more associated MHC alleles a i However, in other implementations, the dependent variable y i is the presentation specific system 160 determines the independent variable z i It will be appreciated that the dependent variable y may represent any other type of information that one is interested in predicting. For example, in another implementation, the dependent variable y i σ may also be a numerical value indicating the mass analysis ion current determined for the example data.

[0244] Peptide sequence p for data example i i is k i is a sequence of k amino acids, i can vary within a range among data instances i. For example, the range can be 8 to 15 for MHC class I, or 6 to 30 for MHC class II. In one specific implementation of system 160, all peptide sequences p in the training data set are i may have the same length, e.g., 9. The number of amino acids in a peptide sequence may vary depending on the type of MHC allele (e.g., MHC allele in humans). MHC allele a for data example i i is the peptide sequence p i indicates whether it existed in combination with

[0245] The data management module 312 also manages the peptide sequences p contained in the training data 170. i and bound MHC allele a i Together with the binding affinity b i and stability i For example, the training data 170 may include a predictor of the peptide p i and a i The predicted binding affinity b between each of the bound MHC molecules shown in iAs another example, the training data 170 may contain a i The predicted stability value s for each of the MHC alleles shown in i may contain

[0246] The data management module 312 also receives the peptide sequence p i along with non-allele interacting variables such as C-terminal flanking sequences and mRNA quantification measurements. i It may also include.

[0247] The data management module 312 also identifies peptide sequences that are not presented by MHC alleles to generate the training data 170. Generally, this involves identifying a "longer" sequence of the source protein that contains the peptide sequence to be presented prior to presentation. If the presentation information contains a genetically engineered cell line, the data management module 312 identifies a set of peptide sequences in the synthetic protein to which the cells were exposed that were not presented on the MHC alleles of the cells. If the presentation information contains a tissue sample, the data management module 312 identifies the source protein from which the presented peptide sequences originated and identifies a set of peptide sequences in the source protein that were not presented on the MHC alleles of the tissue sample cells.

[0248] The data management module 312 also artificially generates peptides with random sequences of amino acids and identifies the generated sequences as peptides that are not presented on MHC alleles. This can be achieved by randomly generating peptide sequences, allowing the data management module 312 to easily generate large amounts of synthetic data for peptides that are not presented on MHC alleles. In practice, because a small percentage of peptide sequences are presented by MHC alleles, synthetically generated peptide sequences are very likely not presented by MHC alleles, even if they are included in proteins processed by cells.

[0249] 4 illustrates an exemplary set of training data 170A, according to one embodiment. Specifically, the first three data examples in training data 170A are alleles HLA-C * Monoallelic cell lines containing 01:03 and three peptide sequences The peptide presentation information from TIFF2025175055000018.tif11128 is shown. The fourth data example in training data 170A shows the allele HLA-B * 07:02, HLA-C * 01:03, HLA-A * The training data 170A shows peptide information from a multi-allelic cell line containing HLA-DRB3:01:01 and peptide sequence QIEJOEIJE (SEQ ID NO:13). The first data example shows that peptide sequence QCEIOWARE (SEQ ID NO:14) was not presented by the allele HLA-DRB3:01:01. As discussed in the previous two paragraphs, negatively labeled peptide sequences may be randomly generated by the data management module 312 or may be identified from the source protein of the presented peptide. The training data 170A also includes a predicted binding affinity of 1000 nM and a predicted stability value of 1 hour half-life for the peptide sequence-allele pair. The training data 170A also shows that peptides were not presented by the allele HLA-DRB3:01:01. The C-terminal flanking sequence of TIFF2025175055000019.tif4128, and 10 2 It also includes non-allele-interacting variables, such as mRNA quantification measurements of TPM. A fourth data example shows that the peptide sequence QIEJOEIJE (SEQ ID NO: 13) interacts with the allele HLA-B * 07:02, HLA-C * 01:03, or HLA-A * 01:01. Training data 170A also includes predicted binding affinity and stability values ​​for each of the alleles, as well as C-terminal flanking sequences of the peptides and mRNA quantification measurements for the peptides.

[0250] VIII.C.2. Coding Module The encoding module 314 encodes the information contained in the training data 170 into a numerical representation that can be used to generate one or more representation models. In one implementation, the encoding module 314 one-hot encodes sequences (e.g., peptide sequences or C-terminal flanking sequences) for a predetermined 20-letter amino acid alphabet. Specifically, k i Peptide sequence p having amino acids i is 20·k i p, which is represented as a row vector of elements, corresponding to the alphabet of the amino acid at the jth position of the peptide sequence. i 20·(j-1)+1 ,p i 20·(j-1)+2 ,...,p i 20·j A single element in has a value of 1. The remaining elements have a value of 0. As an example, for a given alphabet {A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y}, the three amino acid peptide sequence EAF of data example i is a 60-element row vector The C-terminal flanking sequence c can be represented by TIFF2025175055000020.tif11128. i , and the protein sequence for the MHC allele d h , and other sequence data in the presentation information can be similarly coded as above.

[0251] If the training data 170 contains sequences of amino acids of different lengths, the encoding module 314 may further encode the peptides into vectors of equivalent length by adding PAD characters to extend the predetermined alphabet. For example, this may be done by left-padding the peptide sequence with PAD characters until the length of the peptide sequence reaches the peptide sequence with the longest length in the training data 170. Thus, if the peptide sequence with the longest length is k 最大 amino acids, the encoding module 314 encodes each sequence as (20+1) k 最大It is represented numerically as a row vector of elements. For example, consider the extended alphabet {PAD,A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y} and k 最大 For a maximum amino acid length of ≡5, the same exemplary peptide sequence EAF of 3 amino acids is represented as a 105-element row vector The C-terminal flanking sequence c can be represented by TIFF2025175055000021.tif18142. i or other sequence data can be similarly encoded as above. Thus, the peptide sequence p i or c i Each argument or row in represents the occurrence of a particular amino acid at a particular position in the sequence.

[0252] Although the above method for encoding sequence data has been described with respect to sequences having amino acid sequences, the method can be similarly extended to other types of sequence data, such as, for example, DNA or RNA sequence data.

[0253] The encoding module 314 also encodes one or more MHC alleles a for data instance i. i is encoded into a row vector of m elements, where each element h=1,2,...,m corresponds to a uniquely identified MHC allele. The element corresponding to the identified MHC allele for data instance i has a value of 1. The remaining elements have a value of 0. As an example, let * 01:01, HLA-C * 01:08, HLA-B * 07:02, HLA-DRB1 * 10:01}, allele HLA-B for data example i corresponding to a multi-allele cell line * 07:02 and HLA-DRB1 * 10:01 is a four-element row vector a i = [0 0 1 1], and a3 i =1 and a4 i= 1. An example with four identified MHC allele types is described herein, but the number of MHC allele types can actually be hundreds or thousands. As noted above, each data instance i typically contains a peptide sequence p i It contains up to six different MHC allele types associated with

[0254] The encoding module 314 also generates a label y for each data instance i. i We code, as a binary variable with values ​​from the set {0,1}, where a value of 1 indicates that the peptide x i However, the associated MHC allele a i a value of 0 indicates that the peptide was presented by one of the peptides x i However, the associated MHC allele a i The dependent variable y i If represents the mass analysis ion current, the encoding module 314 may additionally scale the value using various functions, such as a log function with a range of [-∞,∞] for ion current values ​​between [0,∞].

[0255] The coding module 314 encodes the peptide p i and the allele interaction variable x for the associated MHC allele h h i pairs as row vectors in which the numerical representations of the allele interaction variables are concatenated one after the other. For example, the encoding module 314 may h i [p i ], [p i b h i ], [p i s h i ], or [p i b h i s h i ], where b h iis the predicted binding affinity for peptide p and associated MHC allele h, and similarly, s h i is for stability. Alternatively, one or more combinations of allele interaction variables may be stored individually (e.g., as individual vectors or matrices).

[0256] In one example, the encoding module 314 encodes the measured or predicted values ​​for binding affinity as a function of the allele interaction variable x h i The binding affinity information is expressed by incorporating the

[0257] In one example, the encoding module 314 encodes the measured or predicted values ​​for binding stability as allele interaction variables x h i By incorporating the information into the

[0258] In one example, the encoding module 314 encodes the measured or predicted values ​​for the binding on-rate as a function of the allele interaction variable x h i The combined on-rate information is expressed by incorporating

[0259] In one example, for peptides presented by class I MHC molecules, the encoding module 314 encodes the peptide length in the vector TIFF2025175055000022.tif11128 (in the formula, TIFF2025175055000023.tif3128 is an index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i In another example, for peptides presented by class II MHC molecules, the encoding module 314 may include the peptide length in the vector TIFF2025175055000024.tif24131 (in the formula, TIFF2025175055000025.tif3128 is the index function, L k is peptide p k (meaning the length of the vector T) k the allele interaction variable x h i can be included in

[0260] In one example, the encoding module 314 represents the RNA expression information of MHC alleles by incorporating the RNA-seq-based expression levels of the MHC alleles into an allele interaction variable xhi.

[0261] Similarly, the encoding module 314 encodes the allele non-interacting variable w i can be represented as a row vector in which the numerical representations of the allele-non-interacting variables are concatenated one after the other. For example, w i is [c i ] or [c i m i w i ], and w i is the C-terminal flanking sequence of peptide pi and the mRNA quantification measurement m associated with the peptide. i Alternatively, one or more combinations of allele non-interacting variables may be stored individually (e.g., as individual vectors or matrices).

[0262] In one example, the encoding module 314 encodes the turnover rate or half-life as a function of the allele non-interacting variable w i represents the turnover rate of the source protein for the peptide sequence.

[0263] In one example, the encoding module 314 encodes the protein length as a function of the allele non-interacting variable w i represents the length of the source protein or isoform by incorporating

[0264] In one example, the encoding module 314 generates β1 i , β2 i , β5 i The mean expression of immunoproteasome-specific proteasome subunits, including the subunits, was calculated using the allele-noninteracting variable w i Incorporation into the IL-1 protein results in activation of the immunoproteasome.

[0265] In one example, the encoding module 314 encodes the RNA-seq abundance of a peptide (quantified in units of FPKM, TPM by techniques such as RSEM) or a source protein of a gene or transcript of the peptide, by correlating the source protein abundance with an allele-non-interacting variable w i This is expressed by incorporating it into

[0266] In one example, the encoding module 314 calculates the probability that the transcript of the peptide's origin will undergo nonsense-mediated decay (NMD), for example, as estimated by the model in Rivas et al. Science, 2015, and calculates this probability as a function of the allele non-interaction variable w i This is expressed by incorporating it into

[0267] In one example, encoding module 314 represents the activation status of a gene module or pathway assessed via RNA-seq, for example, by quantifying the expression of genes in the pathway in units of TPM using, for example, RSEM, for each of the genes in the pathway, and then computing a summary statistic, such as a mean, across the genes in the pathway. The mean is calculated using the allele-non-interaction variable w i can be incorporated into.

[0268] In one example, the encoding module 314 encodes the copy number of the source gene by dividing the copy number by the allele non-interacting variable w i This is expressed by incorporating it into

[0269] In one example, the encoding module 314 encodes the measured or predicted TAP binding affinity (e.g., in nanomolar units) relative to the allele-non-interacting variable w i The TAP binding affinity is expressed by including

[0270] In one example, the encoding module 314 encodes TAP expression levels measured by RNA-seq (and quantified, for example, by RSEM in units of TPM) into an allele-non-interacting variable w i The expression level of TAP is represented by the inclusion of

[0271] In one example, the encoding module 314 encodes the tumor mutations as allele-non-interacting variables w i vector of indicator variables in (i.e., peptide p k is derived from a sample with a KRAS G12D mutation, d k = 1, otherwise 0).

[0272] In one example, the encoding module 314 encodes germline polymorphisms in antigen-presenting genes as a vector of indicator variables (i.e., peptide p k If is derived from a sample with a specific germline polymorphism in TAP, then d k = 1). These indicator variables are expressed as the allele non-interaction variables w i can be included in

[0273] In one example, the encoding module 314 represents tumor types as one-hot coded vectors of length 1 for an alphabet of tumor types (e.g., NSCLC, melanoma, colon cancer, etc.). These one-hot coded variables are combined into an allelic non-interaction variable w i can be included in

[0274] In one example, the encoding module 314 represents MHC allele suffixes by processing four-digit HLA alleles with various suffixes. For example, HLA-A * 24:09N is a model for HLA-A* Alternatively, since HLA alleles ending in an N suffix are not expressed, the probability of presentation by an MHC allele with an N suffix can be set to zero for all peptides.

[0275] In one example, the encoding module 314 represents tumor subtypes as one-hot coded vectors of length 1 for an alphabet of tumor subtypes (e.g., lung adenocarcinoma, lung squamous cell carcinoma, etc.). These one-hot coded variables are combined into an allelic non-interaction variable w i can be included in

[0276] In one example, the encoding module 314 encodes smoking history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a smoking history) can be included in k = 1 otherwise 0). Alternatively, smoking history can be coded as a one-hot coded variable of length 1 for the alphabet of smoking severity. For example, smoking status can be assessed on a 1-5 scale, with 1 indicating non-smoker and 5 indicating current heavy smoker. Because smoking history is primarily associated with lung tumors, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a smoking history and the tumor type is lung tumor, and zero otherwise.

[0277] In one example, the encoding module 314 encodes sunburn history as a function of the allele-non-interacting variable w i A binary indicator variable (d if the patient has a history of severe sunburn) can be included in k = 1 otherwise 0). Because severe sunburn is primarily associated with melanoma, when training models for multiple tumor types, this variable can also be defined as equal to 1 if the patient has a history of severe sunburn and the tumor type is melanoma, and zero otherwise.

[0278] In one example, the encoding module 314 represents the distribution of expression levels of a particular gene or transcript for each gene or transcript in the human genome as a summary statistic (e.g., mean, median) of the distribution of expression levels by using a reference database such as TCGA. Specifically, the expression of peptide p in samples with the tumor type melanoma is expressed as a summary statistic (e.g., mean, median) of the distribution of expression levels. k Regarding peptide p k The measured gene or transcript expression levels of the genes or transcripts of origin are compared with the allele-non-interacting variable w i Not only can it be included in the peptide p in melanoma as measured by TCGA, k The mean and / or median gene or transcript expression of the genes or transcripts of a given source may also be included.

[0279] In one example, the encoding module 314 represents the variant types as one-hot coded variables of length 1 for an alphabet of variant types (e.g., missense, frameshift, NMD-induced, etc.). These one-hot coded variables are grouped together into allelic non-interaction variables w i can be included in

[0280] In one example, the encoding module 314 encodes the protein-level characteristics of the protein as values ​​of the source protein annotation (e.g., 5′ UTR length) and the allele-non-interacting variable w i In another example, the encoding module 314 encodes the peptide p i The residue-level annotation of the source protein for peptide p i is equal to 1 if overlaps with the helical motif, otherwise it is equal to 0, or i The allele-non-interaction variable wi represents the allele-non-interaction variable wi by including an indicator variable equal to 1 if p is completely contained within the helix motif. i The property that represents the proportion of residues in ican be included in

[0281] In one example, the encoding module 314 encodes the types of proteins or isoforms in the human proteome into an index vector o having a length equivalent to the number of proteins or isoforms in the human proteome. k and the corresponding element o k i is the peptide p k is 1 if comes from protein i, and 0 otherwise.

[0282] In one example, the encoding module 314 encodes the peptide p i Source gene G = gene(p i ) as a categorical variable with L possible categories (where L denotes an upper bound on the number of subscripted source genes 1, 2, ..., L).

[0283] In one example, the encoding module 314 encodes the peptide p i T = tissue type, cell type, tumor type, or tumor histology type of T = tissue (p i ) as a categorical variable with M possible categories (where M denotes an upper limit on the number of subscripted types 1, 2, ..., M). Tissue types can include, for example, lung tissue, cardiac tissue, intestinal tissue, and neural tissue. Cell types can include, for example, dendritic cells, macrophages, and CD4 T cells. Cancers can include, for example, lung adenocarcinoma, lung squamous cell carcinoma, melanoma, and non-Hodgkin's lymphoma.

[0284] The encoding module 314 also encodes the peptide p i and the variable z for the associated MHC allele h i The entire set of alleles is expressed as the allele interaction variable x i and the allele non-interaction variable w i For example, the encoding module 314 may represent z h i [xh i w i ] or [w i x h i ] can be represented as a row vector equivalent to

[0285] IX. Training Module The training module 316 constructs one or more presentation models that generate a likelihood of whether a peptide sequence will be presented by an MHC allele associated with the peptide sequence. k and peptide sequence p k MHC alleles associated with a k Given a set of p k However, the associated MHC allele a k the estimate u, which indicates the likelihood that one or more of k Generate.

[0286] IX.A. Overview The training module 316 constructs one or more representation models based on a training data set stored in storage 170, which is generated from the representation information stored in 165. Generally, regardless of the specific type of representation model, all representation models capture the dependencies between independent and dependent variables in the training data 170 such that a loss function is minimized. Specifically, the loss function TIFF2025175055000026.tif4128 is a graph of the dependent variable y for one or more data examples S in the training data 170. i∈S and the estimated likelihood u for the data example S generated by the proposed model. i∈S In one particular implementation described later in this specification, the loss function (y i∈S ,u i∈S ;θ) is the negative log likelihood function given by equation (1a) as follows: However, in practice, another loss function may be used. For example, if a prediction is made for mass spectrometry ion current, the loss function is the mean square loss given by Equation 1b as follows: TIFF2025175055000028.tif10128

[0287] The proposed model can be a parametric model, where one or more parameters θ mathematically specify the dependency between the independent and dependent variables. Typically, the loss function (y i∈S ,u i∈S The various parameters of the proposed parametric model that minimizes θ ( ; θ ) are determined through a gradient-based numerical optimization algorithm, such as a batch gradient algorithm, a stochastic gradient algorithm, etc. Alternatively, the proposed model may be a non-parametric model, in which the model structure is determined from training data 170 and is not strictly based on a fixed set of parameters.

[0288] IX.B. Allele-specific model The training module 316 may build a presentation model to predict the presentation likelihood of a peptide on an allele-by-allele basis. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele.

[0289] In one implementation, the training module 316: TIFF2025175055000029.tif7128 estimates the likelihood of peptide pk being presented for a particular allele h. k where the peptide sequence x h k is the peptide p k and the coded allele interaction variable for the corresponding MHC allele h, where f(·) is an arbitrary function, which for convenience of description will be referred to as a transformation function throughout this specification. h(·) is an arbitrary function, which for convenience of description will be referred to as the dependence function throughout this specification, and the parameter θ determined for the MHC allele h h Based on the set of allele interaction variables x h k Generate a dependency score for the parameter θ for each MHC allele h. h The set of values ​​is θ h where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele h.

[0290] Dependence function g h (x h k ;θ h ) output is the MHC allele h with at least the allele interaction characteristic x h k and in particular the peptide p k The dependency score for MHC allele h indicates whether the corresponding neoantigen is presented based on the amino acid position of the peptide sequence of p. For example, the dependency score for MHC allele h is determined by the relationship between the MHC allele h and the peptide p. k The transformation function f(·) transforms the input, more specifically, g in this example. h (x h k ;θ h ) is used to calculate the dependency score for peptide p k is converted to an appropriate value indicating the likelihood that it will be presented by the MHC allele.

[0291] In one particular implementation described later in this specification, f(·) is a function with a range in [0,1] for the appropriate domain range. In one example, f(·) is This is the expit function given by TIFF2025175055000030.tif10128. As another example, f(·) also has the following meaning: for values ​​in the domain z greater than or equal to 0, f(z)=tanh(z) (5) Alternatively, if the prediction is made for mass analysis ion currents with values ​​outside the range [0,1], f(·) can be any function, for example, the identity function, the exponential function, the log function, etc.

[0292] Therefore, the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is given by the dependency function g h (·) is the peptide sequence p k to generate a corresponding dependency score. k may be transformed by a transformation function f(·) to generate the allele-specific likelihood that h will be presented by MHC allele h.

[0293] IX.B.1 Dependence Functions for Allelic Interaction Variables In one particular implementation mentioned throughout this specification, the dependency function g h (·) is x h k Each allele interaction variable in is compared with the parameter θ determined for the relevant MHC allele h. h with the corresponding parameters in the set This is an affine function given by TIFF2025175055000031.tif5128.

[0294] In another specific implementation mentioned throughout this specification, the dependency function g h (·) is a network model NN with a set of nodes arranged in one or more layers. h (·), The network function is given by TIFF2025175055000032.tif5128. The nodes are connected by parameters θh A node may be connected to other nodes through connections, each having an associated parameter in a set of . The value at one particular node may be represented as the sum of the values ​​of the nodes connected to the particular node, weighted by the associated parameter mapped by the activation function associated with the particular node. In contrast to affine functions, network models are advantageous because the presentation model can incorporate nonlinearity and process data having amino acid sequences of different lengths. Specifically, through nonlinear modeling, the network model can capture the interactions between amino acids at different positions in a peptide sequence and how these interactions affect peptide presentation.

[0295] Generally speaking, the network model NN h (·) can be structured as feedforward networks such as artificial neural networks (ANNs), convolutional neural networks (CNNs), deep neural networks (DNNs), and / or recurrent networks such as long short-term memory networks (LSTMs), bidirectional recurrent networks, and deep bidirectional recurrent networks.

[0296] In one example described later in this specification, each MHC allele in h=1, 2,..., m is associated with a separate network model, NN h (·) denotes the output from the network model related to MHC allele h.

[0297] FIG. 5 illustrates an exemplary network model NN3(·) associated with an arbitrary MHC allele h=3. As shown in FIG. 5, the network model NN3(·) for MHC allele h=3 includes three input nodes at layer l=1, four nodes at layer l=2, two nodes at layer l=3, and one output node at layer l=4. The network model NN3(·) is associated with a set of ten parameters θ3(1), θ3(2), ..., θ3(10). The network model NN3(·) includes three allele interaction variables x3 for MHC allele h=3.k (1), x3 k (2) and x3 k (3) receives input values ​​(individual data examples, including encoded polypeptide sequence data and any other training data used) and generates the value NN3(x3 k ) The network function may include one or more network models, each taking a different allele interaction variable as input.

[0298] In another example, the identified MHC alleles h=1,2,...,m are used to model the MHC alleles in a single network NN H (·) and NN h (·) denotes one or more outputs of a single network model associated with MHC allele h. In such an example, the parameters θ h may correspond to the set of parameters for a single network model, and thus the parameters θ h The set of can be shared by all MHC alleles.

[0299] Figure 6A shows an exemplary network model NN shared by MHC alleles h = 1, 2, ..., m. H (·). As shown in Figure 6A, the network model NN H (·) contains m output nodes, each corresponding to an MHC allele. The network model NN3(·) has an allele interaction variable x3 for MHC allele h=3. k , and the value NN3(x3 k ) and outputs m values.

[0300] In yet another example, a single network model NN H (·) is the allele interaction variable x for MHC allele h h k and the encoded protein sequence d h In such an example, the parameter θ hmay again correspond to the set of parameters for a single network model, and thus the parameters θ h The set of MHC alleles can be shared by all MHC alleles. Therefore, in such an example, NNh(·) is a single network model with inputs [x h k d h ], a single network model NN H Such a network model is advantageous because it can correctly predict peptide presentation probabilities for MHC alleles that were unknown in the training data simply by identifying their protein sequences.

[0301] Figure 6B shows an exemplary network model NN shared by MHC alleles. H (·). As shown in Figure 6B, the network model NN H (·) takes as input the allele interaction variables and protein sequence of MHC allele h=3 and calculates the dependency score NN3 (x3 k ) is output.

[0302] In yet another example, the dependency function g h (·)teeth, TIFF2025175055000033.tif5128, where g' h (x h k ;θ' h ) is the parameter θ' h a bias parameter θ in the set of parameters for the allele interaction variables of the MHC alleles, which represents the baseline probability of presentation for the MHC allele h. h 0 accompanied by.

[0303] In another implementation, the bias parameter θ h 0may be shared according to the gene family of the MHC allele h. That is, the bias parameter θ for the MHC allele h h 0 is θ 遺伝子(h) 0 and gene (h) is the gene family for the MHC allele h. For example, the class I MHC allele HLA-A * 02:01, HLA-A * 02:02, and HLA-A * 02:03 may be assigned to the "HLA-A" gene family, and the bias parameters θ for each of these MHC alleles h 0 As another example, if the class II MHC alleles HLA-DRB1:10:01, HLA-DRB1:11:01, and HLA-DRB3:01:01 are assigned to the "HLA-DRB" gene family, and the bias parameters θ for each of these MHC alleles are h 0 can be shared.

[0304] As an example, returning to equation (2), the affine dependency function g h Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000034.tif5128 can be generated by k is the allele interaction variable identified for MHC allele h=3, and θ3 is the set of parameters determined for MHC allele h=3 through loss function minimization.

[0305] As another example, separate network transformation functions g h Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000035.tif5128 can be generated, where x3k is the allele interaction variable identified for MHC allele h = 3, and θ3 is the set of parameters determined for the network model NN3(·) associated with MHC allele h = 3.

[0306] Figure 7 shows the correlation coefficients of peptide p associated with MHC allele h=3 using the exemplary network model NN3(·). k As shown in Figure 7, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The output is then mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0307] IX.B.2. Per allele with non-allele interaction variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF2025175055000036.tif7128, peptide p k We model the estimated presentation likelihood uk, where w k is the peptide p k means the coded allele non-interaction variable for g w (·) is the parameter θ determined for the allele non-interacting variable w Based on the set of allele-non-interacting variables w k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values ​​of θ h and θ w where i is each example in a subset S of training data 170 generated from cells expressing a single MHC allele.

[0308] Dependence function g w (w k ;θ w) output is a measure of the peptide p expression by one or more MHC alleles based on the influence of allele-non-interacting variables. k represents a dependency score for an allele-non-interacting variable, indicating whether peptide p k The C-terminal flanking sequences and peptide p, which are known to positively influence the presentation of k If peptide p is bound, it may have a high value k The C-terminal flanking sequences and peptide p k If bound, it may have a low value.

[0309] According to equation (8), the peptide sequence p k The allele-specific likelihood that is presented by MHC allele h is the function g h (·) is the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. w (·) is also applied to the coded version of the allele-non-interacting variable to generate a dependency score for the allele-non-interacting variable. Both scores are combined, and the combined score is used to estimate the association of peptide sequence p with MHC allele h. k is transformed by a transformation function f(·) to produce the allele-specific likelihoods that will be presented.

[0310] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (2) k the allele interaction variable x h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF2025175055000037.tif7128.

[0311] IX.B.3 Dependence Functions for Allelic Non-Interacting Variables Dependence function g for allelic interaction variables h Similarly to (·), the dependence function g for allelic non-interacting variables w (·) is an affine function, or a separate network model for the allelic non-interaction variables w k It can be a network function related to

[0312] Specifically, the dependency function g w (·) is w k The allele non-interaction variables in w with the corresponding parameters in the set This is an affine function given by TIFF2025175055000038.tif5128.

[0313] Dependence function g w (·) also corresponds to the parameter θ w the network model NN with relevant parameters in the set w (·), TIFF2025175055000039.tif5128. The network function may include one or more network models, each taking different allele-non-interacting variables as input.

[0314] In another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF2025175055000040.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of m k is the peptide p k is the mRNA quantitative measurement for , h(·) is a function that transforms the quantitative measurement, and θ w mis a parameter in the set of parameters for the allele-non-interacting variables that is combined with the mRNA quantification measurement to generate a dependency score for the mRNA quantification measurement. In one particular embodiment described herein below, h(·) is a log function, although in practice h(·) can be any one of a variety of different functions.

[0315] In yet another example, the dependence function g for the allelic non-interacting variables w (·)teeth, TIFF2025175055000041.tif5128, where g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc., with a set of k is the peptide p k is the indicator vector described in Section VII.C.2, which represents proteins and isoforms in the human proteome, and θ w o is the set of parameters in the set of parameters for the allele non-interacting variables that are combined with the indicator vector. k and parameter θ w o If the dimension of the set is significantly higher, TIFF2025175055000042.tif4128(in the formula, Parameter regularization terms such as λ (representing L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined through an appropriate method.

[0316] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025175055000044.tif13128In formula, g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF2025175055000045.tif5128 is peptide p k is an indicator function that is equal to 1 if θ is derived from source gene 1 as described above for allele-non-interacting variables, and θ w l is a parameter indicating the "antigenicity" of source gene 1. In one variation, L is sufficiently large, and therefore the number of parameters θ w l=1, 2,...,L If is large enough, TIFF2025175055000046.tif5128, where TIFF2025175055000047.tif4128 represents the L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter value. The optimal value of the hyperparameter λ can be determined by an appropriate method.

[0317] In yet another example, the dependence function g on the allelic non-interacting variables w (·) is given by the following formula: TIFF2025175055000048.tif21128In formula, g' w (w k ;θ' w ) is the allele non-interaction parameter θ' w are affine functions, network functions, etc. with a set of TIFF2025175055000049.tif5128 is a peptide p k is derived from source gene 1, and peptide p k is an indicator function that is equal to 1 if originates from tissue type m, and θ wlm is a parameter indicating the antigenicity of the combination of source gene 1 and tissue type m. Specifically, the antigenicity of gene 1 in tissue type m may indicate the residual tendency of cells of tissue type m to present peptides derived from gene 1 after adjustment for RNA expression and peptide sequence context.

[0318] In one variation, L or M is sufficiently large, so that the number of parameters θ w lm=1, 2,...,LM If is large enough, TIFF2025175055000050.tif5128, where TIFF2025175055000051.tif4128 represents L1 norm, L2 norm, combination, etc.) can be added to the loss function when determining the parameter values. The optimal value of the hyperparameter λ can be determined by an appropriate method. In another variation, a parameter regularization term can be added to the loss function when determining the parameter values ​​so that the coefficients for the same source gene do not differ significantly between tissue types. For example, a penalty term such as: TIFF2025175055000052.tif18128 (in the formula, TIFF2025175055000053.tif5128 is the average antigenicity across tissue types for source gene 1), one can add a penalty to the standard deviation of antigenicity across different tissue types in the loss function.

[0319] In practice, the dependence function g on the allelic non-interacting variables can be calculated by combining any of the additional terms in equations (10), (11), (12a) and (12b). w For example, the term h(·) representing the mRNA quantification measurement in equation (10) and the term representing the antigenicity of the source gene in equation (12) can be added together along with any other affine or network functions to generate a dependency function for the allele-non-interacting variables.

[0320] As an example, returning to equation (8), the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000054.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0321] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC allele h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000055.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0322] FIG. 8 shows exemplary network models NN3(·) and NN w Peptide p associated with MHC allele h=3 using (·) k As shown in Figure 8, the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood uk.

[0323] IX.C. Multi-allele models The training module 316 may also build a presentation model to predict the presentation likelihood of a peptide in a multi-allelic setting where two or more MHC alleles are present. In this example, the training module 316 may train the presentation model based on example data S in the training data 170 generated from cells expressing a single MHC allele, cells expressing multiple MHC alleles, or a combination thereof. [Example]

[0324] IX.C.1. Example 1: Maximum Values ​​for Per-Allele Models In one implementation, the training module 316 trains peptides p associated with a set of MHC alleles H. k The estimated presentation likelihood u k is the presentation likelihood u determined for each of the MHC alleles h in the set H determined based on cells expressing a single allele, as explained above in conjunction with equations (2)-(11). k h∈H Specifically, the presentation likelihood u k u k h∈H In one implementation, the function is a maximum function, as shown in equation (12), where the proposed likelihood u k can be determined as the maximum of the presentation likelihood for each MHC allele h in set H. TIFF2025175055000056.tif5128

[0325] IX.C.2. Example 2.1: Sum Function Model In one implementation, the training module 316 trains peptides p k The estimated presentation likelihood u k of, Modeled by TIFF2025175055000057.tif13128, where element ah k is the peptide sequence p k 1 for multiple MHC alleles H associated with x h k is the peptide p k and the coded allele interaction variables for the corresponding MHC alleles. The parameter θ for each MHC allele h h The set of values ​​is θ h The dependence function g can be determined by minimizing a loss function for i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. h is the dependency function g introduced above in Section VIII.B.1. h It can be in any of the following forms:

[0326] According to equation (13), the peptide sequence p k The likelihood that a given allele will be presented by one or more MHC alleles h is given by the dependency function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate a corresponding score for the allele interaction variable. The scores for each MHC allele h are combined to generate a corresponding score for the peptide sequence p k is transformed by a transformation function f(·) to produce the presentation likelihood that MHC allele H will be presented by the set of MHC alleles H.

[0327] The model presented in equation (13) is that for each peptide p k It differs from the allele-by-allele model of equation (2) in that the number of relevant alleles for a can be greater than 1. In other words, h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with

[0328] For example, the affine transformation function gh Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000058.tif5128 can be generated by k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.

[0329] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000059.tif5128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.

[0330] FIG. 9 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0331] IX.C.3. Example 2.2: Sum Function Model with Allelic Non-Interacting Variables In one implementation, the training module 316 incorporates allelic non-interacting variables to: TIFF2025175055000060.tif13130, peptide p k The estimated presentation likelihood u k where w k is the peptide p k Specifically, the parameter θ for each MHC allele h h and the parameters θ for the allele-non-interacting variables w The set of values ​​of θ h and θ w The dependence function g can be determined by minimizing a loss function for i, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. w is the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:

[0332] Therefore, according to equation (14), one or more MHC alleles H can bind to a peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele-non-interacting variables to generate a dependency score for the allele-non-interacting variables. The scores are combined and the combined score is used to estimate the association of the peptide sequence p with the MHC allele H. k is transformed by a transformation function f(·) to produce the presentation likelihood that

[0333] In the model presented in equation (14), each peptide p k The number of relevant alleles for a can be greater than 1. h k More than one element in the peptide sequence p k can have a value of 1 for multiple MHC alleles H associated with

[0334] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000061.tif5128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0335] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000062.tif5128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0336] FIG. 10 shows exemplary network models NN2(·), NN3(·), and NN w Using (·), peptide p associated with MHC alleles h=2 and h=3 kAs shown in Figure 10, the network model NN2(·) generates a representation likelihood of the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. The network model NN3(·) generates the allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w for k receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·) to produce an estimated presentation likelihood u k Generate.

[0337] Alternatively, the training module 316 may use the allelic non-interacting variable w in equation (15) k the allele interaction variable x h k By adding to the allele non-interaction variable w k Thus, the presentation likelihood may include The image can be given by TIFF2025175055000063.tif13128.

[0338] IX.C.4. Example 3.1: Model with Implicit Allele-by-Allele Likelihood In another implementation, the training module 316 k The estimated presentation likelihood u k of, TIFF2025175055000064.tif7130, where element a h k is the peptide sequence p k 1 for multiple MHC alleles h∈H associated with u' k h is the implicit allele-specific presentation likelihood for MHC allele h, and vector v has elements v h But, a hk ···u' k h where s(·) is a vector corresponding to v, s(·) is a function that maps the elements of v, and r(·) is a clipping function that clips the values ​​of the input within a predetermined range. As described in more detail below, s(·) may be a summation function or a quadratic function, although it will be recognized that in other embodiments, s(·) may be any function, such as a maximum function. A set of values ​​for the parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in the subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles.

[0339] The presentation likelihood in the presentation model of equation (17) is the likelihood that each peptide p is presented by an individual MHC allele h. k The implicit allele-specific presentation likelihood u' corresponds to the likelihood that k h The implicit per-allele likelihood differs from the per-allele presentation likelihood of Section VIII.B in that the parameters for the implicit per-allele likelihood can be learned from a multi-allelic setting, in addition to a single-allelic setting, where the direct association between the presented peptide and the corresponding MHC allele is unknown. Thus, in a multi-allelic setting, the presentation model is based on the likelihood of the peptide p k Not only can we estimate whether peptide p is presented by the set of MHC alleles H as a whole, but also which MHC alleles h are present in peptide p k The individual likelihood u' indicates which person is most likely to have presented k h∈H The advantage of this is that the presented model can generate implicit likelihoods without training data for cells expressing a single MHC allele.

[0340] In one particular implementation described later in this specification, r(·) is a function with a range [0,1]. For example, r(·) is a clip function: r(z)=min(max(z,0),1) and the minimum value between z and 1 is chosen as the proposed likelihood uk. In another implementation, r(·) is r(z)=tanh(z) where the domain z is greater than or equal to 0.

[0341] IX.C.5. Example 3.2: Function Sum Model In one particular implementation, s(·) is a summation function, and the presentation likelihood is given by summing the implicit per-allele presentation likelihoods. TIFF2025175055000065.tif15128

[0342] In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: Generated by TIFF2025175055000066.tif7128, the presented likelihood is Let it be estimated by TIFF2025175055000067.tif13128.

[0343] According to equation (19), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables. Each dependency score can be generated by first applying the implicit per-allele presentation likelihood u' k h The allele likelihood u' is transformed by the function f(·) to generate k h are combined and a clipping function is applied to the combined likelihood to clip the values ​​into the range [0,1] to produce a peptide sequence p k A presentation likelihood can be generated that g will be presented by a set of MHC alleles H. The dependency function g h is the dependency function g introduced above in Section VIII.B.1.h It can be in any of the following forms:

[0344] For example, the affine transformation function g h Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000068.tif7128 can be generated by k , x3 k are the allele interaction variables identified for MHC alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for MHC alleles h=2, h=3.

[0345] As another example, the network transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000069.tif7128, where NN2(·) and NN3(·) are the network models specified for MHC alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for MHC alleles h=2 and h=3.

[0346] FIG. 11 shows the correlation of peptide p associated with MHC alleles h=2 and h=3 using exemplary network models NN2(·) and NN3(·). k As shown in Figure 9, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) and the network model NN3(·) generates the allele interaction variable x3 for the MHC allele h=3. k receives and outputs NN3(x3 k) are then mapped by a function f(·) and combined to produce an estimated presentation likelihood u k Generate.

[0347] In another implementation, if the prediction is made in terms of the log of the mass analysis ion current, then r(·) is the log function and f(·) is the exponential function.

[0348] IX.C.6. Example 3.3: Sum of Functions Model with Allelic Non-Interacting Variables In one implementation, the implicit per-allele presentation likelihood for MHC allele h is defined as: Generated by TIFF2025175055000070.tif7128, the presented likelihood is As generated by TIFF2025175055000071.tif13129, the effects of allelic non-interacting variables are incorporated into peptide presentation.

[0349] According to equation (21), one or more MHC alleles H bind to the peptide sequence p k The likelihood of being presented is given by the function g h (·) for each of the MHC alleles H, the peptide sequence p k to generate the corresponding dependency scores for the allele interaction variables for each MHC allele h. w (·) is also applied to the coded versions of the allele non-interaction variables to generate dependency scores for the allele non-interaction variables. The scores of the allele non-interaction variables are combined with each of the dependency scores of the allele interaction variables. Each of the combined scores is transformed by the function f(·) to generate an implicit per-allele presentation likelihood. The implicit likelihoods are combined, and a clipping function is applied to the combined output to clip values ​​into the range [0,1] to determine the likelihood of presentation of the peptide sequence p by the MHC allele H. k A likelihood of being presented can be generated. wis the dependency function g introduced above in Section VIII.B.3. w It can be in any of the following forms:

[0350] For example, the affine transformation function g h (·), g w Using (·), peptide p was identified by MHC alleles h=2 and h=3 among m=4 different identified MHC alleles. k The likelihood that will be presented is TIFF2025175055000072.tif7128, where w k is the peptide p k are allele-non-interacting variables identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0351] As another example, let us consider the network transformation functions gh(·) and gw(·) to determine the peptide p k The likelihood that will be presented is TIFF2025175055000073.tif7128, where w k is the peptide p k is the allele interaction variable identified for θ w is the set of parameters determined for the allele non-interacting variables.

[0352] FIG. 12 shows exemplary network models NN2(·), NN3(·), and NN w Peptide p associated with MHC alleles h = 2 and h = 3 using (·) k As shown in Figure 12, the network model NN2(·) generates the allele interaction variable x2 for the MHC allele h=2. k receives and outputs NN2(x2 k ) is generated. w (·) indicates peptide p k Allele non-interaction variable w fork receives and outputs NN w (w k ) The outputs are combined and mapped by a function f(·). The network model NN3(·) generates an allele interaction variable x3 for MHC allele h=3. k receives and outputs NN3(x3 k ) is generated, which is also the same network model NN w (·) output NN w (w k ) and mapped by a function f(·). Both outputs are combined to produce the estimated presentation likelihood uk.

[0353] In another implementation, the implicit per-allele presentation likelihood for MHC allele h can be calculated as: Generated by TIFF2025175055000074.tif7128, the presented likelihood is Generated by TIFF2025175055000075.tif13128.

[0354] IX.C.7. Example 4: Quadratic Model In one implementation, s(·) is a quadratic function, and the peptide p k The estimated presentation likelihood u k teeth, TIFF2025175055000076.tif13135, where element u' k h is the implicit per-allele presentation likelihood for MHC allele h. A set of values ​​for parameters θ for the implicit per-allele likelihood can be determined by minimizing a loss function with respect to θ, where i is each example in subset S of training data 170 generated from cells expressing a single MHC allele and / or cells expressing multiple MHC alleles. The implicit per-allele presentation likelihood can be in any of the forms shown in equations (18), (20), and (22) above.

[0355] In one embodiment, the model of equation (23) is kHowever, there is a possibility that a given antigen may be simultaneously presented by two MHC alleles, which may imply that presentation by the two HLA alleles is statistically independent.

[0356] According to equation (23), one or more MHC alleles H bind to the peptide sequence p k The presentation likelihood is calculated by combining the implicit per-allele presentation likelihood and the MHC allele H k Each pair of MHC alleles is assigned to a peptide p such that it generates a presentation likelihood that p will be presented. k can be generated by subtracting from the sum the likelihood that

[0357] For example, the affine transformation function g h The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF2025175055000077.tif5128 can be generated by k , x3 k are the allele interaction variables identified for HLA alleles h=2, h=3, and θ2, θ3 are the set of parameters determined for HLA alleles h=2, h=3.

[0358] As another example, the network transformation function g h (·), g w The peptide p was identified by HLA alleles h = 2 and h = 3 among m = 4 different identified HLA alleles using (·). k The likelihood that will be presented is TIFF2025175055000078.tif5128, where NN2(·) and NN3(·) are the network models specified for HLA alleles h=2 and h=3, and θ2 and θ3 are the sets of parameters determined for HLA alleles h=2 and h=3.

[0359] X. Example 5: Prediction Module The prediction module 320 receives sequence data and selects candidate neoantigens in the sequence data using the proposed model. Specifically, the sequence data may be DNA sequences, RNA sequences, and / or protein sequences extracted from tumor tissue cells of a patient. The prediction module 320 converts the sequence data into a plurality of peptide sequences p having 8-15 amino acids for MHC-I or 6-30 amino acids for MHC-II. k For example, the prediction module 320 can process a given sequence "IEFROEIFJEF (SEQ ID NO:16)" into three nine amino acid peptide sequences "IEFROEIFJ (SEQ ID NO:17)," "EFROEIFJE (SEQ ID NO:18)," and "FROEIFJEF (SEQ ID NO:19)." In one embodiment, the prediction module 320 can identify candidate neoantigens that are mutated peptide sequences by comparing sequence data extracted from a patient's normal tissue cells with sequence data extracted from the patient's tumor tissue cells to identify segments containing one or more mutations.

[0360] The prediction module 320 applies one or more presentation models to the processed peptide sequences to estimate the presentation likelihood of the peptide sequences. Specifically, the prediction module 320 can select one or more candidate neoantigen peptide sequences that are likely to be presented on tumor HLA molecules by applying the presentation models to the candidate neoantigens. In one implementation, the prediction module 320 selects candidate neoantigen sequences with an estimated presentation likelihood above a predetermined threshold. In another implementation, the presentation model selects v candidate neoantigen sequences with the highest estimated presentation likelihood (v is generally the maximum number of epitopes that can be delivered in a vaccine). A vaccine containing the selected candidate neoantigens for a given patient can be injected into the patient to induce an immune response.

[0361] XI. Example 6: Cassette Design Module XI.A. Overview The cassette design module 324 generates a vaccine cassette sequence based on the v candidate peptides selected for injection into a patient. Specifically, the peptides p selected for incorporation into a volume v of a vaccine are: k , k=1, 2, ..., v, the cassette sequence is a set of therapeutic epitope sequences p ‘k , k=1,2,...,v, each of which corresponds to a peptide p k The cassette design module 324 can link epitopes that are directly adjacent to each other. For example, vaccine cassette C can be represented as follows: TIFF2025175055000079.tif5128, p' ti indicates the i-th epitope of the cassette. i where k corresponds to the index k=1, 2, ..., v of the selected peptide at the i position of the cassette. The cassette design module 324 may link epitopes with one or more optional linker sequences between adjacent epitopes. For example, a vaccine cassette C may be represented as follows: TIFF2025175055000080.tif5128, l (ti,tj) is the i-th epitope p' of the cassette ti and the j=i+1th epitope p' j=i+1 The cassette design module 324 designates the selected epitope p' k , k=1, 2, ..., v, which are placed at different positions in the cassette, as well as linker sequences placed between the epitopes. Cassette sequence C can be loaded as a vaccine using any of the methods described herein.

[0362] The set of therapeutic epitopes may be generated based on selected peptides determined by the prediction module 320 to be associated with a presentation likelihood above a predetermined threshold, where the presentation likelihood is determined by a presentation model. However, in other embodiments, the set of therapeutic epitopes may be generated based on any one or more of a number of methods (single or in combination), such as, for example, based on binding affinity or predicted binding affinity to the patient's HLA class I or class II alleles, binding stability or predicted binding stability to the patient's HLA class I or class II alleles, random sampling, etc.

[0363] In one embodiment, the therapeutic epitope p ‘k is the selected peptide p k The therapeutic epitope p ‘k In addition to the selected peptide, the cassette may also include C-terminal and / or N-terminal flanking sequences. For example, the epitope p ‘k Haha, the array [n k p k c k ], wherein c k is the selected peptide p k and a C-terminal flanking sequence attached to the C-terminus of k is the selected peptide p k In one example described herein below, the N- and C-terminal flanking sequences are the native N- and C-terminal flanking sequences of the therapeutic vaccine epitope associated with its source protein. In one example described herein below, the therapeutic epitope p ‘k represents a certain length of the epitope. In another example, the therapeutic epitope p ‘k can represent epitopes of varying lengths, and the length of the epitope can vary, for example, depending on the length of the C- or N-flanking sequence. For example, the C-terminal flanking sequence c k , and N-terminal flanking sequence n kcan each have a variable length of 2 to 5 residues, so that the epitope p ‘k So, there are 16 options available.

[0364] The cassette design module 324 generates cassette sequences taking into consideration the representation of junction epitopes across the junction between a pair of therapeutic epitopes in the cassette. A junction epitope is a novel, non-self, but unrelated epitope sequence that arises in the cassette due to the process of joining a therapeutic epitope and a linker sequence in the cassette. The novel sequence of the junction epitope is different from the therapeutic epitope of the cassette itself. Epitope p' ti and p' tj The junction epitope spanning the therapeutic epitope p' ti and p' tj p', which is different from its own sequence ti or p' tj In particular, any linker sequence l may be used. (ti,tj) The epitope p' of the cassette, with or without ti and the adjacent epitope p' tj Each junction between (ti,tj) Junction epitope e n (ti,tj) , n=1,2,…,n (ti,tj) The junction epitope may be related to both epitopes p'. ti and p' tj or epitope p', ti and p' tj The junction epitope may be presented by MHC class I, MHC class II, or both.

[0365] Figure 13 shows two exemplary cassette sequences, Cassette 1 (C1) and Cassette 2 (C2). Each cassette has a vaccine capacity of v=2 and contains the therapeutic epitope p' t1 =p 1 = SINFEKL (SEQ ID NO: 20), and p' t2 =p 2 =LLLLLVVVV (SEQ ID NO: 21), and a linker sequence between these two epitopes (t1,t2) Specifically, the sequence of cassette C1 is [p 1 l (t1,t2) p 2 ], while the sequence of cassette C2 is given by [p 2 l (t1,t2) p 1 The junction epitope of cassette C1 is given by n (1,2) As an example of the epitope p' in the cassette, 1 and p' 2 Examples of sequences that can be used include sequences such as EKLAAYLLL (SEQ ID NO:22), KLAAYLLLLL (SEQ ID NO:23), and FEKLAAYL (SEQ ID NO:24), which span both the linker sequence and a single selected epitope within the cassette, and sequences such as AAYLLLLL (SEQ ID NO:25) and YLLLLLVVV (SEQ ID NO:26), which span the linker sequence and a single selected epitope within the cassette. Similarly, exemplary junction epitopes of cassette C2 include: m (2,1) can be sequences such as VVVVAAYSIN (SEQ ID NO: 27), VVVVAAY (SEQ ID NO: 28), and AYSINFEK (SEQ ID NO: 29). Both cassettes contain the same set of sequences p 1 , l (c1,c2) , and p 2 However, the set of junction epitopes identified will vary depending on the ordered sequence of therapeutic epitopes within the cassette.

[0366] The cassette design module 324 generates cassette sequences that suppress the potential presentation of junction epitopes to the patient. Specifically, when the cassette is injected into the patient, the junction epitopes are presented by the patient's HLA class I or HLA class II alleles and stimulate CD8 or CD4 T cell responses, respectively. Such responses are often undesirable because T cells that react to the junction epitopes have no therapeutic effect and can reduce the immune response to the therapeutic epitope selected in the cassette due to antigen competition. 76

[0367] In one embodiment, the cassette design module 324 iterates through one or more candidate cassettes and determines a cassette sequence for which the junction epitope presentation score associated with the cassette sequence is below a numerical threshold. The junction epitope presentation score is a quantity related to the likelihood of presentation of the junction epitope in the cassette, and a higher junction epitope presentation score indicates a greater likelihood that the junction epitope of the cassette will be presented by HLA class I, HLA class II, or both.

[0368] The cassette design module 324 may determine the cassette sequence associated with the smallest junction epitope presentation score among the candidate cassette sequences, or may select cassette sequences having a presentation score below a predetermined threshold. In one example, the presentation score of a given cassette sequence C is calculated using a distance metric d(e n (ti,tj) ,n=1,2,…,n (ti,tj) )=d (ti,tj) , each associated with a junction of the cassette C. Specifically, the distance metric d (ti,tj) is the adjacent therapeutic epitope p' ti and p' tjThe likelihood of presenting one or more junction epitopes spanning between and is determined. The junction epitope presentation score of cassette C can then be determined by applying a function (e.g., sum, statistical function) to the set of distance metrics for that cassette C. Mathematically, the presentation score is: TIFF2025175055000081.tif5128, where h(·) is some function that maps the distance metric of each junction to a score. In one particular case described later in this specification, the function h(·) is the sum over the distance metrics of the cassette.

[0369] The cassette design module 324 may iterate through one or more candidate cassette sequences, determine the junction epitope presentation scores of the candidate cassettes, and identify optimal cassette sequences associated with junction epitope presentation scores below a threshold. In certain embodiments described below, the distance metric d(·) for a given junction may be determined by the presentation likelihood or the sum of the predicted presentation junction epitopes as determined by the presentation models described in Sections VII and VIII of this specification. However, in other embodiments, the distance metric may be determined from other factors alone or in combination with models such as those described above, which may include determining the distance metric from one or more (alone or in combination) of: measured or predicted HLA binding affinity or stability for HLA class I or HLA class II; and HLA mass spectrometry for HLA class I or HLA class II; or presentation or immunogenicity models trained on T-cell epitope data. For example, the distance metric can combine information about HLA class I and HLA class II presentation. For example, the distance metric can be the number of junction epitopes predicted to bind to either the patient's HLA class I or HLA class II alleles with a binding affinity below a threshold. In another example, the distance metric can be a prediction of the number of junction epitopes predicted to be presented by either the patient's HLA class I or HLA class II alleles.

[0370] The cassette design module 324 may further check one or more candidate cassette sequences to identify whether any of the junction epitopes in the candidate cassette sequences are self-epitopes for a given patient for whom a vaccine is being designed. To accomplish this, the cassette design module 324 checks the junction epitopes against a known database, such as BLAST. In one embodiment, the cassette design module checks the epitopes against the known database, such as BLAST. i The epitopej A pair of epitopes that bind to the N-terminus of t to form a junction self-epitope i ,t j For the distance metric d (ti,tj) Setting Λ to a very large value (e.g., 100) can be configured to design cassettes that avoid junction self-epitopes.

[0371] Returning to the example of FIG. 13, the cassette design module 324 may select all possible junction epitopes e, e.g., having a length of 8 to 15 amino acids for MHC class I or 9 to 30 amino acids for MHC class II. n (t1,t2) =e n (1,2) The distance metric d for a single junction (t1, t2) in cassette C1 (for example) is obtained by summing the likelihoods of (t1,t2) =d (1,2) = 0.39. Because there are no other junctions in cassette C1, the junction epitope presentation score, summed over all distance metrics for cassette C1, is also 0.39. The cassette design module 324 determines all possible junction epitopes e having a length of 8 to 15 amino acids for MHC class I or 9 to 30 amino acids for MHC class II. n (t1,t2) =e n (1,2) The distance metric d for a single junction (t1, t2) in cassette C2 is obtained by summing the likelihoods of (t1,t2) =d (2,1) = 0.068. In this example, the junction epitope presentation score for cassette C2 also gives a single-junction distance metric of 0.068. The cassette design module 324 outputs the cassette sequence of C2 as the optimal cassette because the junction epitope presentation score is lower than the cassette sequence of C1.

[0372] The cassette design module 324 can perform a brute force approach and iterate through all or most likely candidate cassette sequences to select the sequence with the lowest junction epitope presentation score. However, the number of such candidate cassettes becomes prohibitively large as the vaccine volume v increases. For example, for a vaccine volume of v=20 epitopes, the cassette design module 324 may iterate through approximately 10 candidate cassettes to determine the cassette with the lowest junction epitope presentation score. 18 To achieve this, the cassette design module 324 must iterate through all possible candidate cassettes. This determination can be computationally burdensome (in terms of the computing resources required) and may be unmanageable for the cassette design module 324 to complete the generation of a vaccine for the patient within a reasonable time. Furthermore, it is even more cumbersome to consider the possible junction epitopes of each candidate cassette. Therefore, the cassette design module 324 may select a cassette sequence based on a method that iterates through a significantly smaller number of candidate cassette sequences than the number of candidate cassette sequences in a brute force approach.

[0373] In one embodiment, the cassette design module 324 generates a subset of randomly, or at least pseudo-randomly, generated candidate cassettes and selects as cassette sequences those candidate cassettes associated with junction epitope presentation scores below a predetermined threshold. Additionally, the cassette design module 324 may select as cassette sequences those candidate cassettes from the subset with the lowest junction epitope presentation scores. For example, the cassette design module 324 may generate a subset of approximately 1 million candidate cassettes for a set of v=20 selected epitopes and select the candidate cassette with the lowest junction epitope presentation score. While generating a subset of random cassette sequences and selecting those with low junction epitope presentation scores from the subset is suboptimal compared to a brute-force approach, it requires significantly fewer computational resources and is therefore technically feasible. Furthermore, in contrast to this more efficient approach, brute force methods provide only small or negligible improvements in junction epitope presentation scores and are therefore not worthwhile from a resource allocation perspective.

[0374] In another embodiment, the cassette design module 324 determines an improved cassette configuration by formulating the epitope sequences for the cassette as an asymmetric traveling salesperson problem (TSP). Given a list of nodes and the distances between each pair of nodes, the TSP determines the sequence of nodes associated with the shortest total distance to visit each node exactly once and return to the original node. For example, given cities A, B, and C with known distances from each other, the solution to the TSP generates a tightly packed sequence of cities such that the total distance traveled to visit each city exactly once is the smallest possible route. An asymmetric version of the TSP determines the optimal sequence of nodes when the distances between pairs of nodes are asymmetric. For example, the "distance" to travel from node A to node B may be different from the "distance" to travel from node B to node A.

[0375] The cassette design module 324 determines whether each node contains a therapeutic epitope p' k The improved cassette sequence is determined by solving the asymmetric TSP corresponding to the epitope p'. k From the node corresponding to epitope p' m The distance to another node corresponding to (k,m) Although the epitope p' is given by m From the node corresponding to epitope p' k The distance to the node corresponding to (m,k) which is given by the distance metric d (k,m) By using the asymmetric TSP to find an improved optimal cassette solution, the cassette design module 324 can recognize cassette sequences that result in a decrease in the overall junction presentation score between the epitopes of the cassette. The asymmetric TSP solution indicates a sequence of therapeutic epitopes and corresponds to the order in which the epitopes are linked within a cassette that minimizes the overall junction junction epitope presentation score of the cassette. Specifically, given a set of therapeutic epitopes k=1, 2, ..., v, the cassette design module 324 calculates the distance metric d for each ordered pair of potential therapeutic epitopes in the cassette. (k、m) , k, m = 1, 2, ..., v. In other words, for a given pair of epitopes k, m, these distance metrics may differ from each other, so that the distance metrics for the epitope p' k a therapeutic epitope p' linked after m and the distance metric for the epitope p' m a therapeutic epitope p' linked after k Determine both the distance metric for

[0376] The cassette design module 324 solves the asymmetric TSP through an integer linear programming problem. Specifically, the cassette design module 324 solves the following: We generate a (v+1) × (v+1) path matrix P given by TIFF2025175055000082.tif8128. The v × v matrix D is an asymmetric distance matrix, and each element D(k, m), k = 1, 2, ..., v; m = 1, 2, ..., v, represents the distance between the epitopes p ’k From epitope p ’m The matrix k=2,...,v corresponds to the distance metric for junctions up to the last epitope in P. The rows k=2,...,v of P correspond to the original epitope nodes, and row 1 and column 1 correspond to "ghost nodes" that are at zero distance from all other nodes. Adding "ghost nodes" to the matrix encodes the notion that the vaccine cassette is linear rather than circular, so there are no junctions between the first and last epitope. In other words, the sequence is not circular, and it is not expected that the first epitope will be concatenated after the last epitope in the sequence. Epitope p ’k , epitope p ’m A binary variable x is defined as a value of 1 if there is a specified path (i.e., an epitope-epitope junction in the cassette) that is linked to the N-terminus of km In addition, let E denote the set of all v therapeutic vaccine epitopes, and S ⊂ E denote a subset of epitopes. For such a subset S, let out(S) denote the number of epitope-epitope junctions x km =1, where k is an epitope in S and m is an epitope in E\S. Given a known path matrix P, the cassette design module 324 solves the following integer linear programming problem: TIFF2025175055000083.tif13128 is solved to derive the path matrix X, where P km denotes an element P(k,m) in the path matrix P, subject to the following constraints: TIFF2025175055000084.tif44128The first two constraints ensure that each epitope appears only once in the cassette. The last constraint ensures that the cassettes are connected. In other words, the cassettes encoded by x are connected linear protein sequences.

[0377] x in the integer linear programming problem of equation (27) km , k, m=1, 2,..., v+1, the solution indicates a close alignment of nodes and ghost nodes that can be used to infer one or more sequences of therapeutic epitopes for that cassette that reduce the presentation score of the junction epitope. Specifically, x km A value of =1 indicates that there is a "path" from node k to node m, or in other words, that therapeutic epitope p ’m The improved cassette sequence is used to identify therapeutic epitopes. ’k It must be concatenated after x km = 0 means that no such path exists, or in other words, the therapeutic epitope p ’m However, the improved cassette sequence contains the therapeutic epitope p ’k In summary, the integer programming problem in equation (27) km The value of represents an array of nodes and an array of ghost nodes, and the path is introduced and exists only once for each node. For example, x ghost,1 =1, x 13 =1, X 32 = 1 and X 2.ghost A value of =1 (otherwise 0) may indicate the ordering of nodes ghost->1->3->2->ghost and ghost node.

[0378] Once the sequence is solved, ghost nodes are removed from the sequence to generate a refined sequence with only the original nodes corresponding to the therapeutic epitopes in the cassette. This refined sequence shows the arrangement in which selected epitopes are linked to the cassette to improve the presentation score. For example, continuing from the example in the previous paragraph, ghost nodes can be removed to generate the refined sequence 1→3→2. This refined sequence illustrates one method that can be used to link epitopes in the cassette, namely, p 1 →p 3 →p 2 Shows.

[0379] The therapeutic epitope p ’k is a variable length epitope, the cassette design module 324 may design therapeutic epitopes of different lengths, p ’k and p ’m and determine the candidate distance metric corresponding to the distance metric d as the minimum candidate distance metric. (k、m) For example, epitope p 、k =[n k p k c k ], and p ‘m =[n m p m c m ] may each include corresponding N- and C-terminal flanking sequences which may (in some embodiments) differ by 2 to 5 amino acids. ’k and p ‘m The junction between the k The four length numbers and c m The cassette design module 324 determines candidate distance metrics for each set of junction epitopes and calculates the distance metric d (k、m) The cassette design module 324 can then construct the path matrix P and solve the integer linear programming problem of equation (27) to determine the cassette arrangement.

[0380] Compared to the random sampling approach, solving cassette sequences using an integer programming problem requires the determination of v × (v-1) distance metrics, each corresponding to a pair of therapeutic epitopes for the vaccine. Cassette sequences determined by this approach exhibit significantly fewer junction epitopes, while still requiring fewer computational resources than the random sampling approach, especially when a large number of candidate cassette sequences are generated.

[0381] XI.B. Comparison of Junction Epitope Presentation for Cassette Sequences Generated by Random Sampling and Asymmetric TSP Two cassette sequences containing v=20 therapeutic epitopes were generated by randomly sampling 1,000,000 permutations (cassette sequence C1) and by solving the integer linear programming problem of Equation (27) (cassette sequence C2). The distance metric, i.e., presentation score, was determined based on the presentation model described in Equation (14), where f is a sigmoid function and x h i is the peptide p i is the sequence of , gh(·) is the neural network function, w is the flanking sequence, log transcript / peptide p i of kilobases million (TPM), peptide p i The antigenicity of the protein and peptide p i The sample ID of the origin of the gene and the flanking sequence of the gene w (·) and log TPM are neural network functions, respectively. hEach neural network function (·) had input dimensions of 231 (11 residues × 21 letters / residue, including padding), width 256, rectified linear unit (ReLU) activation in the hidden layer, linear activation in the output layer, and one output node for each HLA allele in the training dataset. The neural network function for the flanking sequence was a first-order hidden MLP with input dimensions of 210 (5 residues of N-terminal flanking sequence + 5 residues of C-terminal flanking sequence × 21 letters / residue, including padding), width 32, ReLU activation in the hidden layer, and linear activation in the output layer. The neural network function for RNA log TPM was a first-order hidden MLP with input dimensions of 1, width 16, ReLU activation in the hidden layer, and linear activation in the output layer. The presented model was used to train the HLA allele HLA-A. * 02:04, HLA-A * 02:07, HLA-B * 40:01, HLA-B * 40:02, HLA-C * 16:02 and HLA-C * The system was constructed for 16:04. The presentation scores, which indicate the predicted value of junction epitopes presented by the two cassette sequences, were compared. The results showed that the presentation score of the cassette sequence generated by solving equation (27) was associated with an approximately four-fold improvement over the presentation score of the cassette sequence generated by random sampling.

[0382] Specifically, the epitope with v=20 is Give it as TIFF2025175055000085.tif95128. In a first example, 1,000,000 different candidate cassette sequences were randomly generated with 20 therapeutic epitopes. A presentation score was generated for each of the candidate cassette sequences. The candidate cassette sequence identified as having the lowest presentation score was: TIFF2025175055000086.tif24170, and had a presentation score of 6.1, which predicted the proposed junction epitope. The median presentation score for 1,000,000 random sequences was 18.3. This experiment demonstrates that identifying cassette sequences from randomly sampled cassettes can significantly reduce the predictive value of proposed junction epitopes.

[0383] In a second example, the cassette sequence C2 was identified by solving the integer linear programming problem of equation (27). Specifically, the distance metric for each potential junction between a pair of therapeutic epitopes was determined. This distance metric was used to solve the integer programming problem. The cassette sequence identified by this approach was: TIFF2025175055000087.tif24170, which yielded a presentation score of 1.7. The presentation score of cassette array C2 was approximately four times better than that of cassette array C1 and approximately 11 times better than the median presentation score of 1,000,000 randomly generated candidate cassettes. The execution time to generate cassette C1 was 20 seconds on a single-threaded 2.30 GHz Intel Xeon E5-2650 CPU. The execution time to generate cassette C2 was 1 second on the same CPU, single-threaded. Thus, in this example, the cassette array identified by solving the integer programming problem in Equation (27) yields a solution that is approximately four times better at one-twentieth the computational cost.

[0384] These results indicate that integer programming can potentially provide cassette sequences with a smaller number of presented junction epitopes than those identified from random sampling, potentially with fewer computational resources.

[0385] XI.C. Comparison of junction epitope presentation for selection of cassette sequences generated by MHC flurry and presentation models In this example, cassette sequences containing v = 20 therapeutic epitopes were selected based on tumor / normal exome sequencing. Tumor transcriptome sequencing and HLA typing of lung cancer samples were generated by randomly sampling 1,000,000 permutations and solving the integer programming problem of Equation (27). The distance metric, or presentation score, was determined based on the number of junction epitopes that bind to the patient's HLA with affinity below various thresholds (e.g., 50-1000 nM, above, or below) as predicted by MHCflurry, an HLA peptide binding affinity predictor. In this example, 20 nonsynonymous somatic mutations selected as therapeutic epitopes were selected from 98 somatic mutations identified in tumor samples by ranking the mutations according to the presentation model described in Section XI.B above. However, it is understood that in other embodiments, the therapeutic epitopes may be selected based on other criteria; stability, or criteria based on a combination of presentation scores, affinity, etc. Additionally, it is understood that the criteria used to prioritize therapeutic epitopes for inclusion in a vaccine need not be the same as the criteria used to determine the distance metric D(k, m) used in the cassette design module 324.

[0386] The patient's HLA class I allele was HLA-A * 01:01, HLA-A * 03:01, HLA-B * 07:02, HLA-B * 35:03, HLA-C * 07:02, HLA-C * It was 14:02.

[0387] Specifically, in this example, the therapeutic epitope for v=20 is: The file was TIFF2025175055000088.tif89128.

[0388] The results of this example in the table below compare the number of junction epitopes predicted by MHC flurry that bind to the patient's HLA with an affinity below the threshold value (nM stands for nanomolar) found via the three example methods. In the first method, the optimal cassette was found via the Traveling Salesperson Problem (ATSP) formulation described above using a 1 s execution time. In the second method, the best cassette was obtained and determined after 1 million random samples. In the third method, the median number of junction epitopes was found in 1 million random samples. TIFF2025175055000089.tif37146

[0389] The results of this example demonstrate that any of several criteria can be used to identify whether a particular cassette design meets the design requirements. Specifically, as demonstrated in the previous examples, cassette sequences selected from a large number of candidates can be identified by the lowest junction epitope presentation score, or at least the cassette sequence with a score below an identified threshold. This example demonstrates that other criteria, such as binding affinity, can be used to identify whether a given cassette design meets the design requirements. In this case, a threshold binding affinity (e.g., 50-1000, or higher or lower) can be set to identify the cassette design sequence as having fewer than some threshold number of junction epitopes above a threshold (e.g., 0), and any one of a number of methods (e.g., methods 1-3 shown in the table) can be used to identify whether a given candidate cassette sequence meets these requirements. The methods of these examples further demonstrate that the threshold may need to be set differently depending on the method used. Other criteria, such as stability or a combination of criteria such as presentation score and affinity, can be envisioned.

[0390] In another example, the same cassette was generated using the same HLA types and the 20 therapeutic epitopes previously described in this section (XI.C), but instead of using a distance metric based on binding affinity predictions, the distance metric for epitopes m and k was the predicted number of peptides spanning the junction of m through k that would be presented by the patient's HLA class I alleles with presentation probabilities above a set threshold (0.005 to 0.5, or higher or lower), as determined by the presentation model described above in Section XI.B. This example further illustrates broad criteria that can be considered when identifying whether a given candidate cassette sequence meets the design requirements for use in a vaccine. TIFF2025175055000090.tif37170

[0391] The above examples identified criteria for determining whether a candidate cassette sequence has varied implementations. Each of these examples demonstrated that a count of the number of junction epitopes above or below the criteria can be used to determine whether a candidate cassette sequence satisfies the criteria. For example, if the criteria is the number of epitopes that meet or exceed a threshold HLA binding affinity, whether the candidate cassette sequence has more or less than that number can determine whether the candidate cassette sequence meets the criteria for use as a selected cassette for a vaccine. Similarly, if the criteria is the number of junction epitopes that exceed a threshold likelihood of presentation,

[0392] However, in other embodiments, calculations other than counting can be performed to determine whether a candidate cassette sequence meets the design criteria. For example, rather than determining whether the number of epitopes exceeds / below a particular threshold, one can instead determine what proportion of junction epitopes is above or below a threshold, such as whether the top X% of junction epitopes have a likelihood of presentation above some threshold Y, or whether X% of junction epitopes have an HLA binding affinity less than or greater than ZnM. These are merely examples, and the criteria can generally be based on statistics derived from any attribute of the individual junction epitopes or aggregation of some or all of the junction epitopes. Here, X can generally be any number between 0 and 100% (e.g., 75% or less), Y can be any number between 0 and 1, and Z can be any number appropriate to the criteria in question. These values ​​can be determined empirically and will vary depending on the model and criteria used and the quality of the training data used.

[0393] Thus, in certain embodiments, junction epitopes with a high probability of presentation can be removed; junction epitopes with a low probability of presentation can be retained; tightly binding junction epitopes, i.e., junction epitopes with a binding affinity below 1000 nM, or 500 nM, or some other threshold, can be removed; and / or loosely binding junction epitopes, i.e., junction epitopes with a binding affinity above 1000 nM, or 500 nM, or some other threshold, can be retained.

[0394] Although the examples described above identify candidate sequences using implementations of the suggested models described above, these principles apply equally to implementations in which epitopes to be placed in cassette sequences are identified based on other types of models, such as those based on affinity, stability, etc.

[0395] XI.D. Cassette Selection of Shared Antigens and Shared Neoantigens Rather than selecting a subset of therapeutic epitopes for a personalized vaccine for an individual patient, a set of therapeutic epitopes is selected. ‘k , k=1, 2, ..., v can be a set of epitopes associated with high likelihood of presentation in a population of cancer patients. For example, the set of therapeutic epitopes can be shared antigen sequences that are sequences from genes identified as overexpressed in cancer patients and are associated with high likelihood of presentation in a population of cancer patients. As another example, the set of therapeutic epitopes can be shared neoantigen sequences that are associated with common driver mutations in a population of cancer patients and are associated with high likelihood of presentation. Thus, instead of customizing the therapeutic epitope sequences of a cassette based on each individual patient's sequencing data and HLA allele type, the therapeutic epitope sequences can be shared among multiple patients.

[0396] When the cassette sequence is shared, a pair of epitopes i and t j the distance metric d between (ti,tj) may be determined as a weighted sum of sub-distance metrics, each associated with a corresponding HLA allele. (ti,tj) teeth: TIFF2025175055000091.tif13128, where d h,(ti,tj) is one or more junction epitopes spanning between a pair of adjacent therapeutic epitopes e n (ti,tj) ,n=1,2,...,n (ti,tj) is a sub-distance metric that specifies the likelihood of HLA allele h being present, and w his a weight indicating the prevalence of HLA allele h in a given patient population. By setting a distance metric as in equation (28) or other similar methods that use the prevalence of HLA alleles to weight the presentation of junction epitopes, cassette sequences can be selected that suppress the presentation of junction epitopes of HLA alleles that are predicted to be more prevalent in the patient population.

[0397] The sub-distance metric associated with HLA allele h is given by the total likelihood of presentation or the predicted number of junction epitopes presented for HLA allele h, as determined by the presentation models described in Sections VII and VIII of this specification. However, in other embodiments, the sub-distance metric may be derived from other factors alone or in combination with models such as those exemplified above, including sub-distance metrics derived from one or more of: HLA binding affinity or stability measurements, HLA class I or HLA class II predictions, and HLA mass spectrometry of HLA class I or HLA, presentation models trained on T-cell epitope data, or immunogenicity models (alone or in combination). The sub-distance metric may combine information about HLA class I and HLA class II presentation. For example, the sub-distance metric may be the number of junction epitopes predicted to bind to either the patient's HLA class I or HLA class II allele with a binding affinity below a threshold. In another example, the sub-distance metric can be a predictor of junction epitopes predicted to be presented on either the patient's HLA class I or HLA class II alleles.

[0398] Based on the distance metric defined in equation (28), the cassette design module 324 may iterate through one or more candidate cassette sequences using any of the methods introduced in Section XI.A above, determine the junction epitope presentation scores of the candidate cassettes, and identify the optimal cassette sequence associated with a junction epitope presentation score below a threshold.

[0399] XI.E. Comparison of Junction Epitope Presentation for Cassette Sequences Generated by Random Sampling and Asymmetric TSP for Shared Antigens and Shared Neoantigens In this example, the cassettes were generated using the same 20 therapeutic epitopes in Section XI.C, and the predicted junction epitopes of the cassette sequences observed by the methods of the three examples were compared. Unlike in Section XI.C, the distance metric and distance matrix were determined using Equation (28). In Equation (28), w h Allele frequencies, denoted as , were calculated across 28 HLA-A, 43 HLA-B, and 23 HLA-C alleles using the model training samples in Section XI.B. These were the alleles supported by the model. The frequencies were calculated separately for each HLA-A, HLA-B, and HLA-C gene. Each distance metric was determined based on the predicted value of the proposed junction epitope exceeding a threshold likelihood of proposal weighted by the corresponding allele frequency at a different threshold probability. As in Section XI.B, the first method identified the optimal cassette via the Traveling Salesperson Problem (ATSP) formulation described above. In the second method, the optimal cassette was determined using the best cassette identified after 1 million random samples. In the third method, the median junction epitope was identified in 1 million random samples. Specifically, the distance matrix of the ATSP method is a weighted sum of single-allele distance sub-matrices weighted by allele frequency. TIFF2025175055000092.tif42146

[0400] Because the distance metric for each method is an expectation weighted to the junction epitopes based on allele frequency, the distance matrices are not integer-valued, and therefore the results in Section XI.C are not integer-valued, as shown in the table above. These results demonstrate that the integer programming problem can significantly reduce the number of junction epitopes present for shared (neo)antigen vaccine cassette packaging compared to those identified from random sampling, and can also provide shared antigen or shared neoantigen cassette sequences that potentially require less computational resources.

[0401] In another example, the cassettes were generated using the same 20 therapeutic epitopes in Section XI.C, and the predicted junction epitopes of the cassette sequences observed in the three example methods were compared using MHCflurry. The distance metric and distance matrix were determined using Equation (28). In Equation (28), w h Allele frequencies, denoted as , were calculated across 22 HLA-A, 27 HLA-B, and 9 HLA-C alleles using the model training samples. The frequencies were calculated separately for each HLA-A, HLA-B, and HLA-C gene. Each distance metric was determined based on the predicted value of proposed junction epitopes below a threshold binding affinity weighted by the corresponding allele frequency at a different threshold probability. As in Section XI.B, in the first method, the optimal cassette was identified via the Traveling Salesperson Problem (ATSP) formulation described above. In the second method, the optimal cassette was determined using the best cassette identified after 1 million random samples. In the third method, the median junction epitope was identified in 1 million random samples. Specifically, the distance matrix for the ATSP method is a weighted sum of single-allele distance submatrices weighted by allele frequency. TIFF2025175055000093.tif42146

[0402] The results of this example demonstrate that any of several criteria can be used to identify whether a given cassette design meets the design requirements. Specifically, this example demonstrates that other criteria, such as binding affinity, can be used to identify whether a given cassette design meets the design requirements for shared antigen and neoantigen vaccine cassettes. For this criterion, a threshold binding affinity (e.g., 50-1000, or greater or less) can be set to specify that the cassette design sequence must have less than a threshold number of junction epitopes that exceed a threshold (e.g., 0). One of a number of methods (e.g., methods 1-3 shown in the table) can then be used to identify whether a given candidate cassette sequence meets these requirements. These example methods further demonstrate that thresholds may need to be set differently depending on the method used. Other criteria, such as stability-based criteria or a combination of criteria such as presentation score and affinity, are also contemplated.

[0403] XII. Exemplary Computer Figure 14 illustrates an exemplary computer 1400 for executing the entities shown in Figures 1 and 3. Computer 1400 includes at least one processor 1402 coupled to a chipset 1404. Chipset 1404 includes a memory controller hub 1420 and an input / output (I / O) controller hub 1422. Memory 1406 and a graphics adapter 1412 are coupled to memory controller hub 1420, and a display 1418 is coupled to graphics adapter 1412. Storage device 1408, input device 1414, and network adapter 1416 are coupled to I / O controller hub 1422. Other embodiments of computer 1400 have different architectures.

[0404] The storage device 1408 is a non-transitory computer-readable storage medium, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 1406 holds instructions and data used by the processor 1402. The input interface 1414 is a touchscreen interface, a mouse, a trackball, or other type of pointing device, a keyboard, or some combination thereof, used to input data into the computer 1400. In some embodiments, the computer 1400 may be configured to receive input (e.g., commands) from the input interface 1414 via gestures from a user. The graphics adapter 1412 displays images and other information on the display 1418. The network adapter 1416 couples the computer 1400 to one or more computer networks.

[0405] The computer 1400 is adapted to execute computer program modules to provide the functionality described herein. As used herein, the term "module" refers to computer program logic used to provide particular functionality. Thus, a module may be implemented in hardware, firmware, and / or software. In one embodiment, the program modules are stored in the storage device 1408, loaded into the memory 1406, and executed by the processor 1402.

[0406] 1 can vary depending on the implementation and processing power required by the entity. For example, presentation specification system 160 can run on a single computer 1400 or on multiple computers 1400 communicating with each other over a network, such as in a server farm. Computer 1400 may lack some of the components described above, such as graphics adapter 1412 and display 1418.

[0407] References TIFF2025175055000094.tif209142TIFF2025175055000095.tif218143TIFF2025175055 000096.tif218143TIFF2025175055000097.tif218143TIFF2025175055000098.tif62141

[0408] Sequence information SEQUENCE LISTING <110> GRITSTONE BIO, INC. <120> REDUCING JUNCTION EPITOPE PRESENTATION FOR NEOANTIGENS <150> US 62 / 590,045 <151> 2017-11-22 <160> 76 <170> PatentIn version 3.5 <210> 1 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 1 Tyr Val Tyr Val Ala Asp Val Ala Ala Lys 1 5 10 <210> 2 <211> 17 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 2 Tyr Glu Met Phe Asn Asp Lys Ser Gln Arg Ala Pro Asp Asp Lys Met 1 5 10 15 Phe <210> 3 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 3 Tyr Glu Met Phe Asn Asp Lys Ser Phe 1 5 <210> 4 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (3)..(3) <223> Pyrrolysine <220> <221> MOD_RES <222> (11)..(11) <223> Leu or Ile <400> 4 His Arg Xaa Glu Ile Phe Ser His Asp Phe Xaa 1 5 10 <210> 5 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Leu or Ile <220> <221> MOD_RES <222> (5)..(5) <223> Leu or Ile <220> <221> MOD_RES <222> (7)..(7) <223> Pyrrolysine <400> 5 Phe Xaa Ile Glu Xaa Phe Xaa Glu Ser Ser 1 5 10 <210> 6 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Pyrrolysine <400> 6 Asn Glu Ile Xaa Arg Glu Ile Arg Glu Ile 1 5 10 <210> 7 <211> 27 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (1)..(1) <223> Leu or Ile <220> <221> MOD_RES <222> (11)..(11) <223> Leu or Ile <220> <221> MOD_RES <222> (15)..(15) <223> Selenocysteine <220> <221> MOD_RES <222> (21)..(21) <223> Leu or Ile <220> <221> MOD_RES <222> (27)..(27) <223> Leu or Ile <400> 7 Xaa Phe Lys Ser Ile Phe Glu Met Met Ser Xaa Asp Ser Xaa Island 1 5 10 15 Phe Leu Lys Ser Xaa Phe Ile Glu Ile Phe Xaa 20 25 <210> 8 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (11)..(11) <223> Pyrrolysine <400> 8 Lys Asn Phe Leu Glu Asn Phe Ile Glu Ser Xaa Phe Ile 1 5 10 <210> 9 <211> 15 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)…(2) <223> Pyrrolysine <220> <221> MOD_RES <222> (14)..(14) <223> Leu or Ile <400> 9 Phe Xaa Glu Ile Phe Asn Asp Lys Ser Leu Asp Lys Phe Xaa Ile 1 5 10 15 <210> 10 <211> 16 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (16)..(16) <223> Leu or Ile <400> 10 Gln Cys Glu Ile Xaa Trp Ala Arg Glu Phe Leu Lys Glu Ile Gly Xaa 1 5 10 15 <210> 11 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Selenocysteine <400> 11 Phe Ile Glu Xaa His Phe Trp Ile 1 5 <210> 12 <211> 12 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (7)..(7) <223> Leu or Ile <220> <221> MOD_RES <222> (10)..(10) <223> Selenocysteine <220> <221> MOD_RES <222> (11)..(11) <223> Leu or Ile <400> 12 Phe Glu Trp Arg His Arg Xaa Thr Arg Xaa Xaa Arg 1 5 10 <210> 13 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Leu or Ile <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (8)..(8) <223> Leu or Ile <400> 13 Gln Ile Glu Xaa Xaa Glu Ile Xaa Glu 1 5 <210> 14 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <400> 14 Gln Cys Glu Ile Xaa Trp Ala Arg Glu 1 5 <210> 15 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)..(2) <223> Leu or Ile <220> <221> MOD_RES <222> (9)..(9) <223> Pyrrolysine <220> <221> MOD_RES <222> (11)..(11) <223> Leu or Ile <400> 15 Phe Xaa Glu Leu Phe Ile Ser Asx Xaa Ser Xaa Phe Ile Glu 1 5 10 <210> 16 <211> 11 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (9)..(9) <223> Leu or Ile <400> 16 Ile Glu Phe Arg Xaa Glu Ile Phe Xaa Glu Phe 1 5 10 <210> 17 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (5)..(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (9)..(9) <223> Leu or Ile <400> 17 Ile Glu Phe Arg Xaa Glu Ile Phe Xaa 1 5 <210> 18 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (4)..(4) <223> Pyrrolysine <220> <221> MOD_RES <222> (8)..(8) <223> Leu or Ile <400> 18 Glu Phe Arg Xaa Glu Ile Phe Xaa Glu 1 5 <210> 19 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (3)..(3) <223> Pyrrolysine <220> <221> MOD_RES <222> (7)..(7) <223> Leu or Ile <400> 19 Phe Arg Xaa Glu Ile Phe Xaa Glu Phe 1 5 <210> 20 <211> 7 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 20 Ser Ile Asn Phe Glu Lys Leu 1 5 <210> 21 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 21 Leu Leu Leu Leu Leu Val Val Val Val 1 5 <210> 22 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 22 Glu Lys Leu Ala Ala Tyr Leu Leu Leu 1 5 <210> 23 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 23 Lys Leu Ala Ala Tyr Leu Leu Leu Leu Leu 1 5 10 <210> 24 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 24 Phe Glu Lys Leu Ala Ala Tyr Leu 1 5 <210> 25 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 25 Ala Ala Tyr Leu Leu Leu Leu Leu 1 5 <210> 26 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 26 Tyr Leu Leu Leu Leu Leu Val Val Val 1 5 <210> 27 <211> 10 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 27 Val Val Val Val Ala Ala Tyr Ser Ile Asn 1 5 10 <210> 28 <211> 7 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 28 Val Val Val Val Ala Ala Tyr 1 5 <210> 29 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 29 Ala Tyr Ser Ile Asn Phe Glu Lys 1 5 <210> 30 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 30 Tyr Asn Tyr Ser Tyr Trp Ile Ser Ile Phe Ala His Thr Met Trp Tyr 1 5 10 15 Asn Ile Trp His Val Gln Trp Asn Lys 20 25 <210> 31 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 31 Ile Glu Ala Leu Pro Tyr Val Phe Leu Gln Asp Gln Phe Glu Leu Arg 1 5 10 15 Leu Leu Lys Gly Glu Gln Gly Asn Asn 20 25 <210> 32 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 32 Asp Ser Glu Glu Thr Asn Thr Asn Tyr Leu His Tyr Cys His Phe His 1 5 10 15 Trp Thr Trp Ala Gln Gln Thr Thr Val 20 25 <210> 33 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 33 Gly Met Leu Ser Gln Tyr Glu Leu Lys Asp Cys Ser Leu Gly Phe Ser 1 5 10 15 Trp Asn Asp Pro Ala Lys Tyr Leu Arg 20 25 <210> 34 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 34 Val Arg Ile Asp Lys Phe Leu Met Tyr Val Trp Tyr Ser Ala Pro Phe 1 5 10 15 Ser Ala Tyr Pro Leu Tyr Gln Asp Ala 20 25 <210> 35 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 35 Cys Val His Ile Tyr Asn Asn Tyr Pro Arg Met Leu Gly Ile Pro Phe 1 5 10 15 Ser Val Met Val Ser Gly Phe Ala Met 20 25 <210> 36 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 36 Phe Thr Phe Lys Gly Asn Ile Trp Ile Glu Met Ala Gly Gln Phe Glu 1 5 10 15 Arg Thr Trp Asn Tyr Pro Leu Ser Leu 20 25 <210> 37 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 37 Ala Asn Asp Asp Thr Pro Asp Phe Arg Lys Cys Tyr Ile Glu Asp His 1 5 10 15 Ser Phe Arg Phe Ser Gln Thr Met Asn 20 25 <210> 38 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 38 Ala Ala Gln Tyr Ile Ala Cys Met Val Asn Arg Gln Met Thr Ile Val 1 5 10 15 Tyr His Leu Thr Arg Trp Gly Met Lys 20 25 <210> 39 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 39 Lys Tyr Leu Lys Glu Phe Thr Gln Leu Leu Thr Phe Val Asp Cys Tyr 1 5 10 15 Met Trp Ile Thr Phe Cys Gly Pro Asp 20 25 <210> 40 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 40 Ala Met His Tyr Arg Thr Asp Ile His Gly Tyr Trp Ile Glu Tyr Arg 1 5 10 15 Gln Val Asp Asn Gln Met Trp Asn Thr 20 25 <210> 41 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 41 Thr His Val Asn Glu His Gln Leu Glu Ala Val Tyr Arg Phe His Gln 1 5 10 15 Val His Cys Arg Phe Pro Tyr Glu Asn 20 25 <210> 42 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 42 Gln Thr Phe Ser Glu Cys Leu Phe Phe His Cys Leu Lys Val Trp Asn 1 5 10 15 Asn Val Lys Tyr Ala Lys Ser Leu Lys 20 25 <210> 43 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 43 Ser Phe Ser Ser Trp His Tyr Lys Glu Ser His Ile Ala Leu Leu Met 1 5 10 15 Ser Pro Lys Lys Asn His Asn Asn Thr 20 25 <210> 44 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 44 Ile Leu Asp Gly Ile Met Ser Arg Trp Glu Lys Val Cys Thr Arg Gln 1 5 10 15 Thr Arg Tyr Ser Tyr Cys Gln Cys Ala 20 25 <210> 45 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 45 Tyr Arg Ala Ala Gln Met Ser Lys Trp Pro Asn Lys Tyr Phe Asp Phe 1 5 10 15 Pro Glu Phe Met Ala Tyr Met Pro Ile 20 25 <210> 46 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 46 Pro Arg Pro Gly Met Pro Cys Gln His His Asn Thr His Gly Leu Asn 1 5 10 15 Asp Arg Gln Ala Phe Asp Asp Phe Val 20 25 <210> 47 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 47 His Asn Ile Ile Ser Asp Glu Thr Glu Val Trp Glu Gln Ala Pro His 1 5 10 15 Ile Thr Trp Val Tyr Met Trp Cys Arg 20 25 <210> 48 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 48 Ala Tyr Ser Trp Pro Val Val Pro Met Lys Trp Ile Pro Tyr Arg Ala 1 5 10 15 Leu Cys Ala Asn His Pro Pro Gly Thr 20 25 <210> 49 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 49 His Val Met Pro His Val Ala Met Asn Ile Cys Asn Trp Tyr Glu Phe 1 5 10 15 Leu Tyr Arg Ile Ser His Ile Gly Arg 20 25 <210> 50 <211> 484 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 50 Thr His Val Asn Glu His Gln Leu Glu Ala Val Tyr Arg Phe His Gln 1 5 10 15 Val His Cys Arg Phe Pro Tyr Glu Asn Ala Met His Tyr Gln Met Trp 20 25 30 Asn Thr Tyr Arg Ala Ala Gln Met Ser Lys Trp Pro Asn Lys Tyr Phe 35 40 45 Asp Phe Pro Glu Phe Met Ala Tyr Met Pro Ile Cys Val His Ile Tyr 50 55 60 Asn Asn Tyr Pro Arg Met Leu Gly Ile Pro Phe Ser Val Met Val Ser 65 70 75 80 Gly Phe Ala Met Ala Tyr Ser Trp Pro Val Val Pro Met Lys Trp Ile 85 90 95 Pro Tyr Arg Ala Leu Cys Ala Asn His Pro Pro Gly Thr Ala Asn Asp 100 105 110 Asp Thr Pro Asp Phe Arg Lys Cys Tyr Ile Glu Asp His Ser Phe Arg 115 120 125 Phe Ser Gln Thr Met Asn Ile Glu Ala Leu Pro Tyr Val Phe Leu Gln 130 135 140 Asp Gln Phe Glu Leu Arg Leu Leu Lys Gly Glu Gln Gly Asn Asn Asp 145 150 155 160 Ser Glu Glu Thr Asn Thr Asn Tyr Leu His Tyr Cys His Phe His Trp 165 170 175 Thr Trp Ala Gln Gln Thr Thr Val Ile Leu Asp Gly Ile Met Ser Arg 180 185 190 Trp Glu Lys Val Cys Thr Arg Gln Thr Arg Tyr Ser Tyr Cys Gln Cys 195 200 205 Ala Phe Thr Phe Lys Gly Asn Ile Trp Ile Glu Met Ala Gly Gln Phe 210 215 220 Glu Arg Thr Trp Asn Tyr Pro Leu Ser Leu Ser Phe Ser Ser Trp His 225 230 235 240 Tyr Lys Glu Ser His Ile Ala Leu Leu Met Ser Pro Lys Lys Asn His 245 250 255 Asn Asn Thr Gln Thr Phe Ser Glu Cys Leu Phe Phe His Cys Leu Lys 260 265 270 Val Trp Asn Asn Val Lys Tyr Ala Lys Ser Leu Lys His Val Met Pro 275 280 285 His Val Ala Met Asn Ile Cys Asn Trp Tyr Glu Phe Leu Tyr Arg Ile 290 295 300 Ser His Ile Gly Arg His Asn Ile Ile Ser Asp Glu Thr Glu Val Trp 305 310 315 320 Glu Gln Ala Pro His Ile Thr Trp Val Tyr Met Trp Cys Arg Val Arg 325 330 335 Ile Asp Lys Phe Leu Met Tyr Val Trp Tyr Ser Ala Pro Phe Ser Ala 340 345 350 Tyr Pro Leu Tyr Gln Asp Ala Lys Tyr Leu Lys Glu Phe Thr Gln Leu 355 360 365 Leu Thr Phe Val Asp Cys Tyr Met Trp Ile Thr Phe Cys Gly Pro Asp 370 375 380 Ala Ala Gln Tyr Ile Ala Cys Met Val Asn Arg Gln Met Thr Ile Val 385 390 395 400 Tyr His Leu Thr Arg Trp Gly Met Lys Tyr Asn Tyr Ser Tyr Trp Ile 405 410 415 Ser Ile Phe Ala His Thr Met Trp Tyr Asn Ile Trp His Val Gln Trp 420 425 430 Asn Lys Gly Met Leu Ser Gln Tyr Glu Leu Lys Asp Cys Ser Leu Gly 435 440 445 Phe Ser Trp Asn Asp Pro Ala Lys Tyr Leu Arg Pro Arg Pro Gly Met 450 455 460 Pro Cys Gln His His Asn Thr His Gly Leu Asn Asp Arg Gln Ala Phe 465 470 475 480 Asp Asp Phe Val <210> 51 <211> 484 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 51 Ile Glu Ala Leu Pro Tyr Val Phe Leu Gln Asp Gln Phe Glu Leu Arg 1 5 10 15 Leu Leu Lys Gly Glu Gln Gly Asn Asn Ile Leu Asp Gly Ile Met Ser 20 25 30 Arg Trp Glu Lys Val Cys Thr Arg Gln Thr Arg Tyr Ser Tyr Cys Gln 35 40 45 Cys Ala His Val Met Pro His Val Ala Met Asn Ile Cys Asn Trp Tyr 50 55 60 Glu Phe Leu Tyr Arg Ile Ser His Ile Gly Arg Thr His Val Asn Glu 65 70 75 80 His Gln Leu Glu Ala Val Tyr Arg Phe His Gln Val His Cys Arg Phe 85 90 95 Pro Tyr Glu Asn Phe Thr Phe Lys Gly Asn Ile Trp Ile Glu Met Ala 100 105 110 Gly Gln Phe Glu Arg Thr Trp Asn Tyr Pro Leu Ser Leu Ala Met His 115 120 125 Tyr Gln Met Trp Asn Thr Ser Phe Ser Ser Trp His Tyr Lys Glu Ser 130 135 140 His Ile Ala Leu Leu Met Ser Pro Lys Lys Asn His Asn Asn Thr Val 145 150 155 160 Arg Ile Asp Lys Phe Leu Met Tyr Val Trp Tyr Ser Ala Pro Phe Ser 165 170 175 Ala Tyr Pro Leu Tyr Gln Asp Ala Gln Thr Phe Ser Glu Cys Leu Phe 180 185 190 Phe His Cys Leu Lys Val Trp Asn Asn Val Lys Tyr Ala Lys Ser Leu 195 200 205 Lys Tyr Arg Ala Ala Gln Met Ser Lys Trp Pro Asn Lys Tyr Phe Asp 210 215 220 Phe Pro Glu Phe Met Ala Tyr Met Pro Ile Ala Tyr Ser Trp Pro Val 225 230 235 240 Val Pro Met Lys Trp Ile Pro Tyr Arg Ala Leu Cys Ala Asn His Pro 245 250 255 Pro Gly Thr Cys Val His Ile Tyr Asn Asn Tyr Pro Arg Met Leu Gly 260 265 270 Ile Pro Phe Ser Val Met Val Ser Gly Phe Ala Met His Asn Ile Ile 275 280 285 Ser Asp Glu Thr Glu Val Higher than Glu Gln Ala Pro His Ile Thr Higher Val 290,295,300 Tyr Met Trp Cys Arg Ala Ala Gln Tyr Ile Ala Cys Met Val Asn Arg 305 310 315 320 Gln Met Thr Ile Val Tyr His Leu Thr Arg Trp Gly Met Lys Tyr Asn 325 330 335 Tyr Ser Tyr Trp Ile Ser Ile Phe Ala His Thr Met Trp Tyr Asn Ile 340 345 350 Trp His Val Gln Trp Asn Lys Gly Met Leu Ser Gln Tyr Glu Leu Lys 355 360 365 Asp Cys Ser Leu Gly Phe Ser Trp Asn Asp Pro Ala Lys Tyr Leu Arg 370 375 380 Lys Tyr Leu Lys Glu Phe Thr Gln Leu Leu Thr Phe Val Asp Cys Tyr 385 390 395 400 Met Trp Ile Thr Phe Cys Gly Pro Asp Ala Asn Asp Asp Thr Pro Asp 405 410 415 Phe Arg Lys Cys Tyr Ile Glu Asp His Ser Phe Arg Phe Ser Gln Thr 420 425 430 Met Asn Asp Ser Glu Glu Thr Asn Thr Asn Tyr Leu His Tyr Cys His 435 440 445 Phe His Trp Thr Trp Ala Gln Gln Thr Thr Val Pro Arg Pro Gly Met 450 455 460 Pro Cys Gln His His Asn Thr His Gly Leu Asn Asp Arg Gln Ala Phe 465 470 475 480 Asp Asp Phe Val <210> 52 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 52 Ser Ser Thr Pro Tyr Leu Tyr Tyr Gly Thr Ser Ser Val Ser Tyr Gln 1 5 10 15 Phe Pro Met Val Pro Gly Gly Asp Arg 20 25 <210> 53 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 53 Glu Met Ala Gly Lys Ile Asp Leu Leu Arg Asp Ser Tyr Ile Phe Gln 1 5 10 15 Leu Phe Trp Arg Glu Ala Ala Glu Pro 20 25 <210> 54 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 54 Ala Leu Lys Gln Arg Thr Trp Gln Ala Leu Ala His Lys Tyr Asn Ser 1 5 10 15 Gln Pro Ser Val Ser Leu Arg Asp Phe 20 25 <210> 55 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 55 Val Ser Ser His Ser Ser Gln Ala Thr Lys Asp Ser Ala Val Gly Leu 1 5 10 15 Lys Tyr Ser Ala Ser Thr Pro Val Arg 20 25 <210> 56 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 56 Lys Glu Ala Ile Asp Ala Trp Ala Pro Tyr Leu Pro Glu Tyr Ile Asp 1 5 10 15 His Val Ile Ser Pro Gly Val Thr Ser 20 25 <210> 57 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 57 Ser Pro Val Ile Thr Ala Pro Pro Ser Ser Pro Val Phe Asp Thr Ser 1 5 10 15 Asp Ile Arg Lys Glu Pro Met Asn Ile 20 25 <210> 58 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 58 Pro Ala Glu Val Ala Glu Gln Tyr Ser Glu Lys Leu Val Tyr Met Pro 1 5 10 15 His Thr Phe Phe Ile Gly Asp His Ala 20 25 <210> 59 <211> 22 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 59 Met Ala Asp Leu Asp Lys Leu Asn Ile His Ser Ile Ile Gln Arg Leu 1 5 10 15 Leu Glu Val Arg Gly Ser 20 <210> 60 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 60 Ala Ala Ala Tyr Asn Glu Lys Ser Gly Arg Ile Thr Leu Leu Ser Leu 1 5 10 15 Leu Phe Gln Lys Val Phe Ala Gln Ile 20 25 <210> 61 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 61 Lys Ile Glu Glu Val Arg Asp Ala Met Glu Asn Glu Ile Arg Thr Gln 1 5 10 15 Leu Arg Arg Gln Ala Ala Ala His Thr 20 25 <210> 62 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 62 Asp Arg Gly His Tyr Val Leu Cys Asp Phe Gly Ser Thr Thr Asn Lys 1 5 10 15 Phe Gln Asn Pro Gln Thr Glu Gly Val 20 25 <210> 63 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 63 Gln Val Asp Asn Arg Lys Ala Glu Ala Glu Glu Ala Ile Lys Arg Leu 1 5 10 15 Ser Tyr Ile Ser Gln Lys Val Ser Asp 20 25 <210> 64 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 64 Cys Leu Ser Asp Ala Gly Val Arg Lys Met Thr Ala Ala Val Arg Val 1 5 10 15 Met Lys Arg Gly Leu Glu Asn Leu Thr 20 25 <210> 65 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 65 Leu Pro Pro Arg Ser Leu Pro Ser Asp Pro Phe Ser Gln Val Pro Ala 1 5 10 15 Ser Pro Gln Ser Gln Ser Ser Ser Gln 20 25 <210> 66 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 66 Glu Leu Val Leu Glu Asp Leu Gln Asp Gly Asp Val Lys Met Gly Gly 1 5 10 15 Ser Phe Arg Gly Ala Phe Ser Asn Ser 20 25 <210> 67 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 67 Val Thr Met Asp Gly Val Arg Glu Glu Asp Leu Ala Ser Phe Ser Leu 1 5 10 15 Arg Lys Arg Trp Glu Ser Glu Pro His 20 25 <210> 68 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 68 Ile Val Gly Val Met Phe Phe Glu Arg Ala Phe Asp Glu Gly Ala Asp 1 5 10 15 Ala Ile Tyr Asp His Ile Asn Glu Gly 20 25 <210> 69 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 69 Thr Val Thr Pro Thr Pro Thr Pro Thr Gly Thr Gln Ser Pro Thr Pro 1 5 10 15 Thr Pro Ile Thr Thr Thr Thr Thr Val 20 25 <210> 70 <211> 25 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 70 Gln Glu Glu Met Pro Pro Arg Pro Cys Gly Gly His Thr Ser Ser Ser 1 5 10 15 Leu Pro Lys Ser His Leu Glu Pro Ser 20 25 <210> 71 <211> 21 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 71 Pro Asn Ile Gln Ala Val Leu Leu Pro Lys Lys Thr Asp Ser His His 1 5 10 15 Lys Ala Lys Gly Lys 20 <210> 72 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 72 Tyr Glu Met Phe Asn Asp Lys Ser Phe Gln Arg Ala Pro Asp Asp Lys 1 5 10 15 Met Phe <210> 73 <211> 9 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (6)..(6) <223> Selenocysteine <220> <221> MOD_RES <222> (7)..(8) <223> Pyrrolysine <400> 73 Phe Glu Gly Arg Lys Molecule Island 1 5 <210> 74 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <220> <221> MOD_RES <222> (2)…(2) <223> Lion or Island <220> <221> MOD_RES <222> (5)…(5) <223> Pyrrolysine <220> <221> MOD_RES <222> (7)…(7) <223> Lion or Island <220> <221> MOD_RES <222> (8)…(8) <223> Pyrrolysine <220> <221> MOD_RES <222> (10)..(10) <223> Lion or Island <220> <221> MOD_RES <222> (14)..(14) <223> Pyrrolysine <400> 74 Pro Half Phe Island Half Glu Half Island Half Gly Glu Island 1 5 10 <210> 75 <211> 19 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 75 Ser Ile Asn Phe Glu Lys Leu Ala Ala Tyr Leu Leu Leu Leu Leu Val 1 5 10 15 Val Val Val <210> 76 <211> 19 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic peptide <400> 76 Leu Leu Leu Leu Leu Val Val Val Val Ala Ala Tyr Ser Ile Asn Phe 1 5 10 15 Glu Lys Leu

Claims

1. 1. A method for identifying a cassette sequence for a neoantigen vaccine, comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one alteration that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and the obtaining includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; using a computer processor to input peptide sequences of the neoantigens into a machine learning display model to generate a set of numerical display likelihoods for the set of neoantigens, wherein each display likelihood in the set represents the likelihood that the corresponding neoantigen will be displayed by one or more MHC alleles on the surface of tumor cells of the subject; for each sample in the set of samples, a label obtained by mass spectrometry that determines the presence of a peptide bound to at least one MHC allele in the set of MHC alleles identified as being present in said sample; a training peptide sequence containing information about a plurality of amino acids constituting the training peptide sequence and a set of positions of the amino acids within the training peptide sequence for each of the samples; and A function that represents the relationship between the neoantigen peptide sequence received as input and the likelihood of presentation generated as output. the inputting step includes a plurality of parameters identified based at least on a training dataset including: identifying for the subject a therapeutic subset of neoantigens from the set of neoantigens, the therapeutic subset of neoantigens corresponding to a predetermined number of neoantigens having a likelihood of presentation above a predetermined threshold; and identifying for the subject a cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in a therapeutic subset of neoantigens, wherein the cassette sequence is identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes. The method comprising:

2. 2. The method of claim 1, wherein presentation of the one or more junction epitopes is determined based on a presentation likelihood generated by inputting sequences of the one or more junction epitopes into the machine learning presentation model.

3. 2. The method of claim 1, wherein the presentation of the one or more junction epitopes is determined based on binding affinity prediction between the one or more junction epitopes and the one or more MHC alleles of the subject.

4. The method of claim 1, wherein the presentation of the one or more junction epitopes is determined based on a binding stability prediction of the one or more junction epitopes.

5. 2. The method of claim 1, wherein the one or more junction epitopes comprise a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope.

6. 2. The method of claim 1, wherein a linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and wherein the one or more junction epitopes comprise a junction epitope that overlaps with the linker sequence.

7. identifying the cassette sequence, For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between said ordered pair of therapeutic epitopes; and For each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for the ordered pair on the one or more MHC alleles of the subject. The method of claim 1 , comprising:

8. identifying the cassette sequence, generating a set of candidate cassette sequences corresponding to different sequences of said therapeutic epitope; for each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine. The method of claim 1 , comprising:

9. The method of claim 8 , wherein the set of candidate cassette sequences is randomly generated.

10. identifying the cassette sequence, The following optimization problem: x in km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the therapeutic epitope, and P is where D is a v×v matrix, and element D(k, m) indicates the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of The method of claim 7 further comprising:

11. The method of claim 1, further comprising producing or having produced a tumor vaccine comprising said cassette sequence.

12. 1. A method for identifying a cassette sequence for a neoantigen vaccine, comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one alteration that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and the obtaining includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; identifying a therapeutic subset of neoantigens from the set of neoantigens for said subject; and identifying for the subject a cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a peptide sequence of a corresponding neoantigen in a therapeutic subset of neoantigens, wherein the cassette sequence is identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes. The method comprising:

13. 13. The method of claim 12, wherein presentation of the one or more junction epitopes is determined based on presentation likelihoods generated by inputting sequences of the one or more junction epitopes into a machine learning presentation model, the presentation likelihoods indicating the likelihood that the one or more junction epitopes are presented by one or more MHC alleles on the surface of tumor cells of the patient, and the set of presentation likelihoods has been identified based at least on received mass spectrometry data.

14. 13. The method of claim 12, wherein the presentation of the one or more junction epitopes is determined based on binding affinity prediction between the one or more junction epitopes and one or more MHC alleles of the subject.

15. The method of claim 12, wherein the presentation of the one or more junction epitopes is determined based on a binding stability prediction of the one or more junction epitopes.

16. 13. The method of claim 12, wherein the one or more junction epitopes comprise a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope.

17. 13. The method of claim 12, wherein a linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and wherein the one or more junction epitopes comprise a junction epitope that overlaps with the linker sequence.

18. identifying the cassette sequence, For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between said ordered pair of therapeutic epitopes; and For each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for said ordered pair on said one or more MHC alleles of the subject.

13. The method of claim 12, comprising:

19. identifying the cassette sequence, generating a set of candidate cassette sequences corresponding to different sequences of said therapeutic epitope; for each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine.

13. The method of claim 12, comprising:

20. 20. The method of claim 19, wherein the set of candidate cassette sequences is randomly generated.

21. identifying the cassette sequence, The following optimization problem: x in km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the therapeutic epitope, and P is where D is a v×v matrix, and element D(k, m) indicates the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of 20. The method of claim 18, further comprising:

22. 13. The method of claim 12, further comprising producing or having produced a tumor vaccine comprising said cassette sequence.

23. 1. A method for identifying a cassette sequence for a neoantigen vaccine, comprising: Obtaining peptide sequences for a therapeutic subset of shared antigens or a therapeutic subset of shared neoantigens for treating a plurality of subjects, wherein the therapeutic subsets corresponding to a predetermined number of peptide sequences have a likelihood of presentation that exceeds a predetermined threshold; and identifying said cassette sequence comprising a plurality of linked therapeutic epitope sequences, each comprising a corresponding peptide sequence in a therapeutic subset of a shared antigen or a therapeutic subset of a shared neoantigen; Including, identifying the cassette sequence, For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between the ordered pair of therapeutic epitopes; and determining, for each ordered pair of therapeutic epitopes, a distance metric indicative of presentation of a set of junction epitopes for said ordered pair, said distance metric being determined as a combination of a set of weights each indicative of the prevalence of a corresponding MHC allele and corresponding sub-distance metrics indicative of the likelihood of presentation of the set of junction epitopes on said MHC allele. Including, The method.

24. 1. A tumor vaccine comprising a cassette sequence comprising a sequence of linked therapeutic epitopes, the cassette sequence comprising: obtaining, for a patient, at least one of tumor nucleotide sequencing data of the exome, transcriptome, or whole genome derived from tumor cells and normal cells of the subject, wherein the nucleotide sequencing data is used to obtain data representing the peptide sequences of each of a set of neoantigens identified by comparing the nucleotide sequencing data from the tumor cells with the nucleotide sequencing data from the normal cells, wherein the peptide sequence of each neoantigen contains at least one alteration that makes the peptide sequence different from a corresponding wild-type parent peptide sequence identified from the subject's normal cells, and the obtaining includes information about a plurality of amino acids that make up the peptide sequence and a set of amino acid positions within the peptide sequence; identifying a therapeutic subset of neoantigens from the set of neoantigens for the subject; and identifying, for the subject, the cassette sequences comprising sequences of a plurality of linked therapeutic epitopes, each comprising a peptide sequence of a corresponding neoantigen in a therapeutic subset of neoantigens, wherein the cassette sequences are identified based on the representation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes. are identified by performing The tumor vaccine.

25. 25. The tumor vaccine of claim 24, wherein the presentation of the one or more junction epitopes is determined based on presentation likelihoods generated by inputting the sequences of the one or more junction epitopes into a machine learning presentation model, the presentation likelihoods indicating the likelihood that the one or more junction epitopes are presented by one or more MHC alleles on the surface of tumor cells of the patient, and the set of presentation likelihoods has been identified based at least on received mass spectrometry data.

26. 25. The tumor vaccine of claim 24, wherein the presentation of the one or more junction epitopes is determined based on predicted binding affinity between the one or more junction epitopes and one or more MHC alleles of the subject.

27. The tumor vaccine of claim 24, wherein the presentation of the one or more junction epitopes is determined based on predicted binding stability of the one or more junction epitopes.

28. 25. The tumor vaccine of claim 24, wherein the one or more junction epitopes comprise a junction epitope that overlaps the sequence of a first therapeutic epitope and the sequence of a second therapeutic epitope linked after the first therapeutic epitope.

29. 25. The tumor vaccine of claim 24, wherein a linker sequence is disposed between a first therapeutic epitope and a second therapeutic epitope linked after the first therapeutic epitope, and the one or more junction epitopes comprise a junction epitope that overlaps with the linker sequence.

30. the step of identifying the cassette sequence comprises: For each ordered pair of therapeutic epitopes, determining a set of junction epitopes spanning the junction between said ordered pair of therapeutic epitopes; and For each ordered pair of therapeutic epitopes, determining a distance metric indicative of the presentation of the set of junction epitopes for said ordered pair on one or more MHC alleles of interest. The tumor vaccine of claim 24, comprising:

31. the step of identifying the cassette sequence comprises: generating a set of candidate cassette sequences corresponding to different sequences of said therapeutic epitope; for each candidate cassette sequence, determining a presentation score for the candidate cassette sequence based on a distance metric for each ordered pair of therapeutic epitopes in the candidate cassette sequence; and selecting candidate cassette sequences associated with presentation scores below a predetermined threshold as cassette sequences for the neoantigen vaccine. The tumor vaccine of claim 24, comprising:

32. The tumor vaccine of claim 31 , wherein the set of candidate cassette sequences is randomly generated.

33. the step of identifying the cassette sequence comprises: The following optimization problem: x in km To find the numerical value of where v corresponds to a predetermined number of neoantigens, k corresponds to a therapeutic epitope, and m corresponds to an adjacent therapeutic epitope linked after the first therapeutic epitope, and P is where D is a v×v matrix, and element D(k, m) indicates the distance metric for an ordered pair of therapeutic epitopes k, m; and x km selecting the cassette arrangement based on the numerical value of the solution of The tumor vaccine of claim 30, further comprising:

34. The tumor vaccine of claim 24, further comprising producing or having produced a tumor vaccine comprising said cassette sequence.

35. A tumor vaccine comprising a cassette sequence comprising linked therapeutic epitope sequences, wherein the cassette sequences are ordered so that each comprises a peptide sequence of a corresponding neoantigen within a therapeutic subset of neoantigens, and the therapeutic epitope sequences are identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes, and the junction epitopes of the cassette sequence have an HLA binding affinity below a threshold binding affinity.

36. 36. The tumor vaccine of claim 35, wherein the threshold binding affinity is 1000 nM or greater.

37. A tumor vaccine comprising a cassette sequence comprising linked therapeutic epitope sequences, wherein the cassette sequences are ordered so that each comprises a peptide sequence of a corresponding neoantigen within a therapeutic subset of neoantigens, and the therapeutic epitope sequences are identified based on the presentation of one or more junction epitopes across corresponding junctions between one or more adjacent pairs of therapeutic epitopes, and at least a threshold percentage of the junction epitopes of the cassette sequence have a presentation likelihood below a threshold presentation likelihood.

38. 38. The tumor vaccine of claim 37, wherein the threshold percentage is 50%.