Predicting the immunogenicity of T cell epitopes

By evaluating the binding scores of modified peptides with MHC molecules and T cell receptors, immunogenic amino acid modifications are screened out, which solves the problem of insufficient prediction of immunogenicity of tumor-associated neoantigens in existing technologies and improves the targeting and immune response effects of personalized cancer vaccines.

CN113219179BActive Publication Date: 2025-09-30BIONTECH SE +1

Patent Information

Application Number
CN202110057195.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2013-05-10
Filing Date
2014-05-07
Publication Date
2025-09-30
Estimated Expiration
2034-06-06

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately predict the immunogenicity of tumor-associated neoantigens, resulting in insufficient targeting and effectiveness of personalized cancer vaccines.

Method used

By evaluating the binding scores of modified peptides to MHC molecules and T cell receptors, combined with chemical and physical similarity scores, amino acid modifications that may be immunogenic can be screened for the preparation of personalized cancer vaccines.

Benefits of technology

It improves the targeting and immune response effects of cancer vaccines, reduces the number of peptides tested in experiments, and enhances the effectiveness of cancer treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113219179B_ABST
    Figure CN113219179B_ABST
Patent Text Reader

Abstract

The present invention relates to methods for predicting T cell epitopes. In particular, the present invention relates to methods for predicting whether a modification in a peptide or polypeptide, such as a tumor-associated neoantigen, is immunogenic. The methods of the present invention are particularly suitable for providing a vaccine specific to a patient's tumor and, therefore, are suitable for use in the context of personalized cancer vaccines.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical field of the invention

[0002] The present invention relates to methods for predicting T cell epitopes. In particular, the present invention relates to methods for predicting whether modifications in peptides or polypeptides, such as tumor-associated neoantigens, are immunogenic. The methods of the present invention are particularly suitable for providing vaccines specific to a patient's tumor, and are therefore suitable for use in the context of personalized cancer vaccines. Background of the Invention

[0003] Personalized cancer vaccines are therapeutic vaccines tailored to the unique target tumor-specific mutations of a given patient. Since such treatments do not harm healthy cells and have the potential to provide lifelong relief, they provide great hope for cancer patients. However, not every mutation expressed by cancer can be used as a target for a vaccine. In fact, when targeted by vaccines, the somatic mutations of most cancers do not result in an immune response (JC Castle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012)). Since tumors can encode up to 100,000 somatic mutations (MR Stratton, Science Signalling 331, 1553 (2011)), and vaccines only target a small amount of epitopes, it is obvious that the key goal of cancer immunotherapy is to identify which mutation may be immunogenic.

[0004] From a biological perspective, in order for a somatic mutation to generate an immune response, several criteria must be met: the cell must express the allele containing the mutation, the mutation must be located in a protein-coding region and be non-synonymous, the translated protein must be cleaved by the proteasome, and the epitope containing the mutation must be presented by the MHC complex, the presented epitope must be recognized by the T cell receptor (TCR), and finally, the TCR-pMHC complex must initiate the signaling cascade that activates T cells (S. Whelan, N. Goldman, Molecular biology and evolution 18, 691 (2001)). To date, no algorithm has been proposed that can predict with high certainty which mutations are likely to meet all of these criteria. In this report, we consider several factors that may contribute to immunogenicity, compare these factors with experimental data, and propose a simple model for identifying immunogenic mutations.

[0005] MHC binding prediction: current state of the art

[0006] More than 20 years ago, it was determined that there are sites on MHC-binding peptides that contribute more to binding capacity than other sites (e.g., (A. Sette et al., Proceedings of the National Academy of Sciences 86, 3296 (1989))). The identification and description of those anchor sites made it possible to find patterns in MHC binding to peptides and thus formed the basis for the development of prediction methods. In recent years, significant progress has been made in the field of in silico models of antigen processing mechanisms. Two pioneering methods developed in the late 1990s were BIMAS (K. C. Parker, M. A. Bednarek, J. E. Coligan, The Journal of Immunology 152, 163 (1994)) and SYFPEITHI (H.-G. Rammensee, J. Bachmann, N. P. Emmerich, O. A. Bachor, S. Immunogenetics 50, 213 (1999)) is based on the knowledge of anchor sites and derived allele-specific motifs. As more and more experimental MHC peptide-binding data become available, more tools have been developed using a variety of statistical and computational techniques (see Figure 1 As reviewed). So-called matrix-based methods use site-specific scoring matrices to determine whether a peptide sequence matches a binding motif for a specific MHC allele. Another class of MHC binding prediction methods utilizes machine learning techniques such as artificial neural networks or support vector machines (see Figure 1 The performance of these algorithms strongly depends on the quantity and quality of the training datasets available for each allele model (e.g., "HLA-A*02:01," "H2-Db," etc.) to "learn" the underlying patterns / features that have predictive power for binding. Recently, structure-based methods have emerged that rely solely on peptide-MHC crystal structures and scoring functions (e.g., different energy functions) to predict peptide-MHC interactions, for example, by energy minimization (see Figure 1), thus avoiding the bottleneck of having a large training set. However, the accuracy of those methods is still far lower than that of sequence-based methods. Benchmark studies have shown that the artificial neural network-based tool NetMHC (C. Lundegaard et al., Nucleic Acids Research 36, W509 (2008)) and the matrix-based algorithm SMM (B. Peters, A. Sette, BMC bioinformatics 6, 132 (2005)) perform best on test evaluation data (B. Peters, A. Sette, BMC bioinformatics 6, 132 (2005); HHLin, S. Ray, S. Tongchusak, EL Reinherz, V. Brusic, BMC immunology 9, 8 (2008)). Both methods have been integrated into the so-called IEDB consensus method, which is available in the Immune Epitope Database (Y. Kim et al., Nucleic Acids Research 40, W525 (2012)). Because MHC II molecules have open-ended binding grooves on either side, allowing them to bind peptides of varying lengths, the interaction model for peptide-MHC II binding is far more complex than for MHC I. However, peptides that bind to MHC I are primarily limited to 8-12 amino acids, a length that can differ significantly from MHC II peptides (9-30 amino acids). Recent benchmark studies have shown that available MHC II prediction methods offer limited accuracy compared to MHC I predictions (HHLin, S. Ray, S. Tongchusak, EL Reinherz, V. Brusic, BMC immunology 9, 8 (2008)).

[0007] The first large-scale systematic application of those algorithms for finding T cell epitopes was performed by Moutaftsi et al. (M. Moutaftsi et al., Nature Biotechnology 24, 817 (2006)). They combined different tools to predict possible vaccine candidates in C57BL76 mice infected with vaccinia virus, extracted spleen cells and measured the CD8+ T cell response to the top 1% of predicted peptides. They identified 49 peptides (out of 2256 peptides) that induced T cell responses. Since then, many studies have been published using different MHC binding prediction tools to find T cell epitopes for candidate vaccines, mainly for pathogens such as Leishmania major (C. Herrera-Najera, R. F. Xacur-Garcia, MJ Ramirez-Sierra, E. Dumonteil, Proteomics 9, 1293 (2009). However, using MHC I binding prediction tools alone to predict immunogenicity is misleading because those tools are trained to predict whether a given peptide has the potential to bind to a given MHC allele. The basic principle of using MHC binding prediction to predict immunogenicity is to assume that peptides that bind with high affinity to each MHC allele are more likely to be immunogenic (A. Sette et al., The Journal of Immunology 153, 5586 (1994)). However, many studies have shown that low MHC binding affinity can also lead to high immunogenicity (MC Feltkamp, ​​MP Vierboom, WM Kast, CJ Melief, Molecular Immunology 31, 1391 (1994)), and peptide-MHC stability may be a better predictor of immunogenicity than peptide affinity (M. Harndahl et al., European Journal of Immunology 42, 1405 (2012)). Therefore, immunogenicity prediction has not been very accurate to date, which is reflected in the low success rate of predicting immunogenicity. However, peptide binding is a necessary but not sufficient condition for T cell epitope recognition, and effective prediction will significantly reduce the number of peptides tested experimentally.

[0008] Obviously, the development of models that predict immunogenicity also needs to take into account T cell receptor (TCR) recognition and central tolerance, that is, negative and positive selection of T cells during thymic development.

[0009] Predictive models that can simulate all the aspects mentioned above to accurately predict the immunogenicity of epitopes, not just binding, are needed.

[0010] Description of the Invention SUMMARY OF THE INVENTION

[0011] In one aspect, the present invention relates to a method for predicting immunogenic amino acid modifications, the method comprising the steps of:

[0012] a) determining a score for binding of the modified peptide to one or more MHC molecules, and

[0013] b) determining a score for binding of the non-modified peptide to one or more MHC molecules, and / or

[0014] c) determining a score for binding of the modified peptide to one or more T cell receptors when presented in an MHC-peptide complex.

[0015] In one embodiment, the modified peptide comprises a fragment of the modified protein, said fragment containing the modification present in said protein.In one embodiment, the non-modified peptide or protein contains germline amino acids at a site corresponding to the modification site in the modified peptide or protein.

[0016] In one embodiment, the non-modified peptide or protein is identical to the modified peptide or protein except for the modification.Preferably, the non-modified peptide or protein has the same length and / or sequence (except for the modification) as the modified peptide or protein.

[0017] In one embodiment, the length of the non-modified peptide and the modified peptide is 8-15 amino acids, preferably 8-12 amino acids.

[0018] In one embodiment, the one or more MHC molecules comprise different MHC molecule types, in particular different MHC alleles. In one embodiment, the one or more MHC molecules are MHC class I molecules and / or MHC class II molecules. In one embodiment, the one or more MHC molecules comprise a set of MHC alleles, such as a set of MHC alleles or a subset thereof for an individual.

[0019] In one embodiment, the binding score to one or more MHC molecules is determined by a method comprising sequence comparison to a database of MHC-binding motifs.

[0020] In one embodiment, step a) comprises determining whether the score meets a predetermined threshold for binding to one or more MHC molecules and / or step b) comprises determining whether the score meets a predetermined threshold for binding to one or more MHC molecules. In one embodiment, the threshold applied in step a) is different from the threshold applied in step b). In one embodiment, the predetermined threshold for binding to one or more MHC molecules reflects the likelihood of binding to one or more MHC molecules.

[0021] In one embodiment, the one or more T cell receptors comprise a set of T cell receptors, such as a set of T cell receptors of an individual or a subset thereof. In one embodiment, step c) comprises assuming that the set of T cell receptors does not comprise a T cell receptor that binds to the non-modified peptide when present in an MHC-peptide complex and / or does not comprise a T cell receptor that binds with high affinity to the non-modified peptide when present in an MHC-peptide complex.

[0022] In one embodiment, step c) comprises determining a score for the chemical and physical similarity between the non-modified amino acid and the modified amino acid. In one embodiment, step c) comprises determining whether the score meets a predetermined threshold for the chemical and physical similarity between the amino acids. In one embodiment, the predetermined threshold for the chemical and physical similarity between the amino acids reflects the likelihood that the amino acids are chemically and physically similar. In one embodiment, the chemical and physical similarity score is determined based on the likelihood that the amino acids are interchangeable in their natural state. In one embodiment, the more frequently the amino acids are interchanged in their natural state, the more similar the amino acids are considered to be, and vice versa. In one embodiment, the chemical and physical similarity is determined using an evolution-based log-odds matrix.

[0023] In one embodiment, if the binding score of the non-modified peptide to one or more MHC molecules meets a threshold value indicative of binding to one or more MHC molecules, and the binding score of the modified peptide to one or more MHC molecules meets a threshold value indicative of binding to one or more MHC molecules, then the modification or modified peptide is predicted to be immunogenic, provided that the chemical and physical similarity scores of the non-modified amino acid and the modified amino acid meet a threshold value indicative of chemical and physical dissimilarity.

[0024] In one embodiment, if the non-modified peptide binds or has the potential to bind to one or more MHC molecules, and the modified peptide binds or has the potential to bind to one or more MHC molecules, then the modification or modified peptide is predicted to be immunogenic, provided that the non-modified amino acid and the modified amino acid are not chemically and physically similar or have the potential to be chemically and physically similar.

[0025] In one embodiment, the modification is not at the anchor site for binding to one or more MHC molecules.

[0026] In one embodiment, a modification or modified peptide is predicted to be immunogenic if the binding score of the non-modified peptide to one or more MHC molecules meets a threshold value indicating no binding to one or more MHC molecules, and the binding score of the modified peptide to one or more MHC molecules meets a threshold value indicating binding to one or more MHC molecules.

[0027] In one embodiment, if the non-modified peptide does not bind or has the potential not to bind to one or more MHC molecules, and the modified peptide does bind or has the potential to bind to one or more MHC molecules, then the modification or modified peptide is predicted to be immunogenic.

[0028] In one embodiment, the modification is at an anchor site for binding to one or more MHC molecules.

[0029] In one embodiment, the method of the present invention comprises performing step a) on two or more differently modified peptides, wherein the two or more differently modified peptides comprise the same modification. In one embodiment, the two or more differently modified peptides comprising the same modification comprise different fragments of a modified protein, wherein the different fragments comprise the same modification present in the protein. In one embodiment, the two or more differently modified peptides comprising the same modification comprise all possible MHC binding fragments of the modified protein, wherein the fragments comprise the same modification present in the protein. In one embodiment, the method of the present invention further comprises selecting a modified peptide from the two or more differently modified peptides comprising the same modification, wherein the two or more differently modified peptides have a probability or maximum probability of binding to one or more MHC molecules. In one embodiment, the two or more differently modified peptides comprising the same modification differ in modification length and / or site.

[0030] In one embodiment, method of the present invention comprises implementing step a) to two or more differently modified peptides, and optionally implementing one or both of step b) and c).In one embodiment, described two or more differently modified peptides comprise identical modification and / or comprise different modifications.In one embodiment, different modifications are present in the same protein and / or different proteins.Step a) and optionally step b) and c) the group of two or more differently modified peptides used in one or both can be identical or different.In one embodiment, step b) and / or step c) the group of two or more differently modified peptides used is the subset of the group of two or more differently modified peptides used in step a).Preferably, described subset comprises the peptide with the highest score in step a).

[0031] In one embodiment, the method of the present invention comprises comparing the scores of two or more of the differently modified peptides. In one embodiment, the method of the present invention comprises ranking the two or more differently modified peptides. In one embodiment, the weight of the modified peptide binding to one or more MHC molecules is higher than the weight of the modified peptide binding to one or more T cell receptors when present in an MHC-peptide complex, preferably higher than the score of the chemical and physical similarity between the non-modified amino acids and the modified amino acids, and the score of the modified peptide binding to one or more T cell receptors when present in an MHC-peptide complex, preferably the chemical and physical similarity score between the non-modified amino acids and the modified amino acids, is higher than the score of the non-modified peptide binding to one or more MHC molecules.

[0032] In one embodiment, the methods of the invention further comprise identifying non-synonymous mutations in one or more protein coding regions.

[0033] In one embodiment, according to the present invention, modifications are identified by partially or completely sequencing the genome or transcriptome of one or more cells, such as one or more cancer cells and, optionally, one or more non-cancerous cells, and identifying mutations in one or more protein coding regions.

[0034] In one embodiment, the mutation is a somatic mutation. In one embodiment, the mutation is a cancer mutation.

[0035] In one embodiment, the method of the invention is used for the preparation of a vaccine.In one embodiment, the vaccine is derived from one or more modifications or one or more modified peptides predicted to be immunogenic by the method of the invention.

[0036] In another aspect, the present invention provides a method for providing a vaccine comprising the step of identifying one or more mutations or one or more modified peptides predicted to be immunogenic by the method of the present invention.

[0037] In one embodiment, the method further comprises the step of providing a vaccine comprising a peptide or polypeptide, or a vaccine comprising a nucleic acid encoding the peptide or polypeptide, containing a modification or a modified peptide predicted to be immunogenic.

[0038] In another aspect, the present invention provides a vaccine obtainable using a method according to the present invention. Preferred embodiments of such a vaccine are described herein.

[0039] The vaccine provided according to the present invention may comprise a pharmaceutically acceptable carrier, and may optionally comprise one or more adjuvants, stabilizers, etc. The vaccine may be in the form of a therapeutic or prophylactic vaccine.

[0040] Another aspect relates to a method of inducing an immune response in a patient comprising administering to the patient a vaccine provided according to the present invention.

[0041] Another aspect relates to a method of treating a cancer patient comprising the steps of:

[0042] (a) providing a vaccine by a method according to the present invention; and

[0043] (b) administering the vaccine to a patient.

[0044] Another aspect relates to a method of treating a patient with cancer comprising administering to the patient a vaccine according to the invention.

[0045] In other aspects, the present invention provides vaccines as described herein for use in the methods of treatment as described herein, particularly for the treatment or prevention of cancer.

[0046] Treatment of cancers described herein can be combined with surgery and / or radiation and / or traditional chemotherapy.

[0047] Other features and advantages of the invention will be apparent from the following detailed description, and from the claims. Detailed Description of the Invention

[0048] Although the present invention is described in detail below, it should be understood that the invention is not limited to the specific methods, protocols, and reagents described herein, as they may vary. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the scope of the invention, which is limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art.

[0049] Hereinafter, elements of the present invention will be described. These elements are listed together with specific embodiments, however, it should be understood that they can be combined in any manner, in any quantity to produce other embodiments. Various description examples and preferred embodiments should not be interpreted as limiting the present invention to only the embodiments clearly described. This description should be understood to support and include embodiments that combine the embodiments clearly described with any number of disclosed and / or preferred elements. In addition, any arrangement and combination of all the elements described in this application should be considered to be disclosed by the application's description, unless context otherwise indicates.

[0050] Preferably, as in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", HGW Leuenberger, B. Nagel, and H. The terms used herein are defined as described in Helvetica Chimica Acta, CH-4010 Basel, Switzerland.

[0051] Unless otherwise indicated, the practice of the present invention utilizes conventional methods of biochemistry, molecular biology, immunology and recombinant DNA technology as explained in the literature in the field (cf., e.g., Molecular Cloning: A Laboratory Manual, 2 ndEdition, J. Sambrook et al. eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989).

[0052] Unless the context requires otherwise, throughout this specification and the following claims, the term "comprise" and variations such as "comprises" and "comprising" should be understood to mean the inclusion of the stated member, integer or step or group of members, integers or steps, but not the exclusion of any other member, integer or step or group of members, integers or steps, although in some embodiments, such other members, integers or steps or groups of members, integers or steps may be excluded, i.e., the subject matter consists of the stated member, integer or step or group of members, integers or steps. The terms "a", "an", and "the", and similar references used in the text of the present invention (particularly in the text of the claims) should be interpreted to cover the singular and the plural, unless otherwise indicated herein or clearly contradicted by the context. The recitation of numerical ranges herein is intended merely to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into this specification as if it were individually recited herein.

[0053] All methods described herein can be implemented in any suitable order, unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the present invention and does not limit the scope of the claimed invention. The language of this description should not be construed as indicating any unrequired elements necessary for the implementation of the present invention.

[0054] Throughout the text of this specification, several documents are cited. Each document cited herein (including all patents, patent applications, specific publications, manufacturer's specifications, instructions, etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein should be construed as an admission that the present invention is not entitled to antedate such disclosure by virtue of prior invention.

[0055] According to the present invention, the term "peptide" refers to a substance containing two or more, preferably three or more, preferably four or more, preferably six or more, preferably eight or more, preferably ten or more, preferably 13 or more, preferably 16 or more, preferably 21 or more and up to preferably 8, 10, 20, 30, 40 or 50, in particular 100 amino acids covalently linked by peptide bonds. The term "polypeptide" or "protein" refers to large peptides, preferably peptides containing more than 100 amino acid residues, but in general the terms "peptide", "polypeptide" and "protein" are synonymous and are used interchangeably herein.

[0056] According to the present invention, the term "modification" with respect to a peptide, polypeptide or protein refers to a sequence change in the peptide, polypeptide or protein compared to the sequence of the parent sequence, such as the wild-type peptide, polypeptide or protein. The term includes amino acid insertion variants, amino acid addition variants, amino acid deletion variants and amino acid substitution variants, preferably amino acid substitution variants. All of these sequence changes according to the present invention may produce new epitopes.

[0057] Amino acid insertion variants include insertions of single or two or more amino acids into the specified amino acid sequence.

[0058] Amino acid addition variants include fusion of one or more amino acids, such as 1, 2, 3, 4 or 5, or more amino acids, to the amino terminus and / or carboxyl terminus.

[0059] Amino acid deletion variants are characterized by the removal of one or more amino acids from the sequence, such as the removal of 1, 2, 3, 4 or 5, or more amino acids.

[0060] Amino acid substitution variants are characterized by at least one residue in the sequence being removed and another residue inserted in its place.

[0061] According to the present invention, the modification or modified peptide to be tested in the method of the present invention may be derived from a protein containing the modification.

[0062] The term "derived" means that according to the present invention, a specific entity, in particular a specific peptide sequence, is present in the object from which it is derived. With respect to amino acid sequences, in particular specific sequence regions, "derived" means in particular that the relevant amino acid sequence is derived from the amino acid sequence from which it is present.

[0063] A protein containing a modification or a potential source of a modified peptide for testing in the methods of the invention may be a neoantigen.

[0064] According to the present invention, the term "neoantigen" relates to a peptide or protein that contains one or more amino acid modifications compared to a parent peptide or protein. For example, a neoantigen may be a tumor-associated neoantigen, wherein the term "tumor-associated neoantigen" encompasses a peptide or protein that contains amino acid modifications resulting from tumor-specific mutations.

[0065] According to the present invention, the term "tumor-specific mutation" or "cancer-specific mutation" refers to a somatic mutation that is present in the nucleic acid of a tumor cell or cancer cell but not in the nucleic acid of a corresponding normal cell (i.e., a non-tumor cell or non-cancer cell). The terms "tumor-specific mutation" and "tumor mutation" and the terms "cancer-specific mutation" and "cancer mutation" are used interchangeably herein.

[0066] The term "immune response" refers to the overall body response to a target such as an antigen, and preferably refers to a cellular immune response or a cellular and humoral immune response. The immune response can be protective / preventative / prophylactic and / or therapeutic.

[0067] "Inducing an immune response" may mean that there is no immune response before induction, but it may also mean that there is a certain level of immune response before induction and that the immune response is enhanced after induction. Therefore, "inducing an immune response" also includes "enhancing an immune response". Preferably, after inducing an immune response in a subject, the subject is protected from developing a disease such as cancer, or the disease condition is improved by inducing an immune response. For example, an immune response against a tumor-expressed antigen can be induced in a patient with cancer or a subject at risk of developing cancer. Inducing an immune response in this case may mean that the subject's disease condition is improved, the subject does not develop metastasis, or a subject at risk of developing cancer does not develop cancer.

[0068] The terms "cellular immune response" and "cellular response" or similar terms refer to an immune response directed against cells characterized by antigen presentation by class I or class II MHC, including T cells or T-lymphocytes that act as "helper" or "killer". Helper T cells (also known as CD4 + T cells) play a major role by regulating immune responses, and killer cells (also known as cytotoxic T cells, cytolytic T cells, CD8 + In a preferred embodiment, the present invention comprises stimulating an anti-tumor CTL response against tumor cells that express one or more tumor-expressed antigens, and preferably presents such tumor-expressed antigens using MHC class I.

[0069] "Antigen" according to the present invention encompasses any substance, preferably a peptide or protein, which is the target of an immune response and / or induces an immune response, such as a specific reaction with an antibody or T-lymphocyte (T cell). Preferably, the antigen comprises at least one epitope, such as a T cell epitope. Preferably, the antigen in the context of the present invention may be a molecule, optionally after processing, which induces an immune response that is preferably specific for the antigen (including cells expressing the antigen). The antigen or its T cell epitope is preferably presented by a cell, preferably an antigen presenting cell, which, in the case of an MHC molecule, includes a diseased cell, particularly a cancer cell, which results in an immune response against the antigen (including cells expressing the antigen).

[0070] In one embodiment, the antigen is a tumor antigen (also referred to herein as a tumor-expressed antigen), i.e., a portion of a tumor cell, such as a protein or peptide expressed in a tumor cell, which may originate from the cytoplasm, cell surface, or cell nucleus, particularly those that are primarily present in the cell or surface antigens of tumor cells. For example, tumor antigens include carcinoembryonic antigen, α1-fetoprotein, isoferritin, and embryonic thioglycoprotein, α2-H-ferritin, and γ-fetoprotein. According to the present invention, tumor antigens preferably include any antigen that is expressed in a tumor or cancer and tumor cells or cancer cells, and optionally characterized by its type and / or expression level for the tumor or cancer and tumor cells or cancer cells, i.e., tumor-associated antigens. In one embodiment, the term "tumor-associated antigen" refers to a protein that is specifically expressed in a limited number of tissues and / or organs or at a specific stage of development under normal conditions. For example, a tumor-associated antigen may be specifically expressed in gastric tissue, preferably in the gastric mucosa, in reproductive organs, such as the testis, in trophoblastic tissue, such as the placenta or germline cells, under normal conditions, and expressed or abnormally expressed in one or more tumor tissues or cancer tissues. In this regard, "limited number" preferably means no more than 3, more preferably no more than 2. Tumor antigens in the context of the present invention include, for example, differentiation antigens, preferably cell type-specific differentiation antigens, i.e., proteins that are specifically expressed at a certain differentiation stage of a certain cell type under normal conditions, cancer / testis antigens, i.e., proteins that are specifically expressed in the testis and sometimes in the placenta under normal conditions, and germ cell-specific antigens. Preferably, cancer cells are identified based on tumor antigens or abnormal expression of tumor antigens. In the context of the present invention, tumor antigens expressed by cancer cells of an object (e.g., a patient with cancer) are preferably self-proteins of the object. In a preferred embodiment, tumor antigens in the context of the present invention are specifically expressed in non-essential tissues or organs under normal conditions, i.e., tissues or organs that do not lead to the death of the object when damaged by the immune system, or are expressed in organs or structures of the body that cannot or can hardly be accessed by the immune system.

[0071] According to the present invention, the terms "tumor antigen", "tumor-expressed antigen", "cancer antigen" and "cancer-expressed antigen" are equivalent and are used interchangeably herein.

[0072] The term "immunogenicity" relates to the relative effectiveness of inducing an immune response, preferably in connection with a treatment, such as a treatment for cancer. As used herein, the term "immunogenic" relates to the property of being immunogenic. For example, the term "immunogenic modification" as used in relation to a peptide, polypeptide or protein relates to the effectiveness of the peptide, polypeptide or protein in inducing an immune response caused by and / or directed against the modification. Preferably, the non-modified peptide, polypeptide or protein does not induce an immune response, induces a different immune response or induces an immune response of a different, preferably lower, level.

[0073] The term "major histocompatibility complex" and the abbreviation "MHC" encompass both MHC class I and MHC class II molecules and refer to a set of genes present in all vertebrates. MHC proteins or molecules are important for signaling between lymphocytes and antigen-presenting cells or diseased cells during immune responses, where they bind to peptides and present them for recognition by T-cell receptors. Proteins encoded by MHC are expressed on the cell surface and display self-antigens (peptide fragments from the cell itself) as well as non-self antigens (e.g., fragments of invading microorganisms) to T cells.

[0074] The MHC family is divided into three subgroups: class I, class II, and class III. MHC class I proteins contain α-chains and β2-microglobulin (not part of the MHC encoded by chromosome 15). They present antigen fragments to cytotoxic T cells. On most immune system cells, particularly antigen-presenting cells, MHC class II proteins contain α-chains and β-chains, and they present antigen fragments to helper T cells. MHC class III regions encode other immune components such as complement components, and some encode cytokines.

[0075] The MHC is polygenic (there are several MHC class I and MHC class II genes) and polymorphic (there are multiple alleles for each gene).

[0076] As used herein, the term "haplotype" refers to the HLA alleles and the proteins they encode found on a chromosome. A haplotype also refers to the alleles present at any locus within the MHC. Several loci are used to represent each type of MHC: for example, for class I, HLA-A (human leukocyte antigen-A), HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-P, and HLA-V; and for class II, HLA-DRA, HLA-DRB1-9, HLA-, HLA-DQA1, HLA-DQB1, HLA-DPAl, HLA-DPBl, HLA-DMA, HLA-DMB, HLA-DOA, and HLA-DOB. The terms "HLA allele" and "MHC allele" are used interchangeably herein.

[0077] The MHC exhibits extreme polymorphism: in humans, a large number of haplotypes containing different alleles exist at each locus. Different polymorphic class I and class II MHC alleles have different peptide specificities: each allele encodes a protein that binds to peptides displaying a specific sequence type.

[0078] In a preferred embodiment of all aspects of the invention the MHC molecule is an HLA molecule.

[0079] In the context of the present invention, the term "MHC binding peptide" includes MHC class I and / or class II binding peptides or peptides that can be processed to produce MHC class I and / or class II binding peptides. In the case of MHC class I / peptide complexes, the length of the binding peptide is generally 8-12, preferably 8-10 amino acids, although longer or shorter peptides may be effective. In the case of MHC class II / peptide complexes, the length of the binding peptide is generally 9-30, preferably 10-25 amino acids, and in particular 13-18 amino acids, although longer or shorter peptides may be effective.

[0080] If the peptide is to be presented directly, i.e. without processing, in particular without cleavage, it has a length suitable for binding to an MHC molecule, in particular to an MHC class I molecule, and is preferably 7-30 amino acids in length, such as 7-20 amino acids in length, more preferably 7-12 amino acids in length, more preferably 8-11 amino acids in length, in particular 9 or 10 amino acids in length.

[0081] If the peptide is part of a larger entity containing other sequences, such as a vaccine sequence or polypeptide, and is to be presented after processing, in particular after shearing, the treatment results in a peptide of a length suitable for binding to MHC molecules, in particular MHC class I molecules, and is preferably 7-30 amino acids in length, such as 7-20 amino acids in length, more preferably 7-12 amino acids in length, more preferably 8-11 amino acids in length, in particular 9 or 10 amino acids in length. Preferably, the sequence of the peptide to be presented after processing is derived from the amino acid sequence of the antigen or polypeptide used for vaccination, i.e. its sequence corresponds substantially to, and preferably is identical to, the fragment of the antigen or polypeptide.

[0082] Thus, in one embodiment, the MHC binding peptide comprises a sequence that substantially corresponds to, and preferably is identical to, the antigen fragment.

[0083] The term "epitope" refers to an antigenic determinant in a molecule, such as an antigen, i.e., a portion or fragment of a molecule that is recognized by the immune system, for example, by T cells, particularly when presented by MHC molecules. An epitope of a protein, such as a tumor antigen, preferably comprises a continuous or discontinuous portion of the protein and is preferably 5-100, preferably 5-50, more preferably 8-30, and most preferably 10-25 amino acids in length. For example, an epitope may preferably be 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length. In the context of the present invention, an epitope is particularly preferably a T cell epitope.

[0084] According to the present invention, an epitope may be bound to an MHC molecule, such as an MHC molecule on the surface of a cell, and may therefore be an "MHC binding peptide."

[0085] As used herein, the term "neo-epitope" refers to an epitope that is not present in a reference cell, such as a normal non-cancerous cell or germ cell, but is present in a cancer cell. In particular, this includes situations where the corresponding epitope is present in a normal non-cancerous cell or germ cell, but the sequence of the epitope has been altered due to one or more mutations in the cancer cell, thereby generating a neo-epitope.

[0086] As used herein, the term "T cell epitope" refers to a peptide that binds to an MHC molecule of a structure recognized by a T cell receptor. Typically, a T cell epitope is present on the surface of an antigen presenting cell.

[0087] As used herein, the term "predicting T cell epitopes" refers to predicting whether a peptide binds to an MHC molecule and is recognized by a T cell receptor. The term "predicting T cell epitopes" is essentially synonymous with the phrase "predicting whether a peptide is immunogenic."

[0088] According to the present invention, the T cell epitope may be present in the vaccine as part of a larger entity such as a vaccine sequence and / or a polypeptide comprising multiple T cell epitopes.The presented peptide or T cell epitope is generated after appropriate processing.

[0089] A T cell epitope may be modified at one or more residues that are not essential for TCR recognition or binding to MHC. Such modified T cell epitopes may be considered immunoequivalent.

[0090] Preferably, when presented by MHC and recognized by a T cell receptor, in the presence of an appropriate co-stimulatory signal, the T cell epitope is capable of inducing clonal expansion of T cells bearing a T cell receptor that specifically recognizes the peptide / MHC complex.

[0091] Preferably, the T cell epitope comprises an amino acid sequence that substantially corresponds to the amino acid sequence of the antigen fragment.Preferably, the antigen fragment is a peptide presented by MHC class I and / or class II.

[0092] The T cell epitope according to the present invention preferably relates to a part or fragment of an antigen that is capable of stimulating an immune response, preferably a cellular immune response to the antigen or a cell characterized by expression of the antigen, and preferably achieved by presentation of the antigen (e.g., a diseased cell, in particular a cancer cell). Preferably, the T cell epitope is capable of stimulating a cellular response to a cell characterized by presentation of the antigen by MHC class I, and is preferably capable of stimulating cytotoxic T lymphocytes (CTLs) that respond to the antigen.

[0093] "Antigen processing" or "processing" refers to the degradation of a peptide, polypeptide or protein into processing products, which are fragments of the peptide, polypeptide or protein (e.g., degradation of a polypeptide into peptides), and one or more of these fragments are associated with an MHC molecule (e.g., by binding) for presentation to a specific T cell by a cell, preferably an antigen-presenting cell.

[0094] "Antigen presenting cells" (APCs) are cells that present peptide fragments of protein antigens linked to MHC molecules on their cell surface. Some APCs can activate antigen-specific T cells.

[0095] Professional antigen-presenting cells are very efficient at internalizing antigens, either through phagocytosis or receptor-mediated endocytosis, and then display the antigen fragments bound to class II MHC molecules on their cell membrane. T cells recognize and interact with the antigen-class II MHC molecule complex on the surface of antigen-presenting cells. The antigen-presenting cells then produce additional costimulatory signals, leading to T cell activation. The expression of costimulatory molecules is a defining characteristic of professional antigen-presenting cells.

[0096] The major types of professional antigen-presenting cells are dendritic cells, macrophages, B cells, and certain activated epithelial cells, of which dendritic cells have the widest range of antigen presentation and are perhaps the most important antigen-presenting cells. Dendritic cells (DCs) are a population of white blood cells that present antigens captured in peripheral tissues to T cells via the MHC class II and class I antigen presentation pathways. Dendritic cells are well known to be potent inducers of immune responses, and activation of these cells is a key step in the induction of anti-tumor immunity. Dendritic cells are conveniently classified as "immature" and "mature" cells, which can be used as a simple way to distinguish between two well-studied phenotypes. However, this nomenclature should not be interpreted as excluding all possible intermediate stages of differentiation. Immature dendritic cells are characterized as antigen-presenting cells with a high capacity for antigen uptake and processing, which is associated with high expression of Fcγ receptors and mannose receptors. The typical characteristics of the mature phenotype are low expression of these markers and high expression of cell surface molecules such as class I and class II MHC, adhesion molecules (such as CD54 and CD11) and co-stimulatory molecules (such as CD40, CD80, CD86 and 4-1BB) responsible for T cell activation. Dendritic cell maturation refers to the state of dendritic cell activation, in which such antigen-presenting dendritic cells cause T cell activation, while the presentation of immature dendritic cells causes tolerance. Dendritic cell maturation is mainly caused by the following reasons: detection of biomolecules with microbial characteristics (bacterial DNA, viral RNA, endotoxins, etc.) by genetic receptors, pro-inflammatory cytokines (TNF, IL-1, IFN), substances released by cells connecting CD40 and undergoing stress cell death on the surface of dendritic cells via CD40L. Dendritic cells can be obtained by culturing bone marrow cells in vitro with cytokines, such as granulocyte-macrophage colony-stimulating factor (GM-CSF) and tumor necrosis factor α.

[0097] Non-professional APCs do not continuously express the MHC class II proteins required for interaction with naive T cells; these proteins are expressed only when the non-professional APCs are stimulated by certain cytokines such as IFNγ.

[0098] Antigen presenting cells can be loaded with MHC class I-presented peptides by transducing the cells with a nucleic acid, preferably RNA, encoding a peptide or polypeptide comprising the peptide being presented, such as a nucleic acid encoding an antigen or polypeptide for vaccination.

[0099] In some embodiments, a pharmaceutical composition or vaccine comprising a nucleic acid delivery vehicle targeted to dendritic cells or other antigen-presenting cells can be administered to a patient, resulting in in vivo transfection. In vivo transfection of dendritic cells can generally be performed using any method known in the art, such as those described in WO 97 / 24447 or the gene gun method described in Mahvi et al. Immunology and Cell Biology 75:456-460, 1997.

[0100] According to the present invention, the term "antigen presenting cells" also includes target cells.

[0101] "Target cell" means a cell that is the target of an immune response, such as a cellular immune response. Target cells include cells that present antigens (i.e., peptide fragments derived from antigens), and include any undesirable cells, such as cancer cells. In preferred embodiments, target cells are cells that express an antigen as described herein, and preferably present the antigen by MHC class I.

[0102] The term "portion" refers to a fraction. With respect to a specific structure, such as an amino acid sequence or a protein, the term "portion thereof" may refer to a continuous portion or a discontinuous portion of said structure. Preferably, a portion of an amino acid sequence comprises at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, preferably at least 40%, preferably at least 50%, more preferably at least 60%, more preferably at least 70%, even more preferably at least 80% and most preferably at least 90% of the amino acids of said amino acid sequence. Preferably, if said portion is a discontinuous portion, said discontinuous portion consists of 2, 3, 4, 5, 6, 7, 8 or more portions of the structure, each portion being a continuous element of said structure. For example, a discontinuous portion of an amino acid sequence may consist of 2, 3, 4, 5, 6, 7, 8 or more, preferably no more than 4, portions of said amino acid sequence, wherein each portion preferably comprises at least 5 contiguous amino acids, at least 10 contiguous amino acids, preferably at least 20 contiguous amino acids, preferably at least 30 contiguous amino acids of said amino acid sequence.

[0103] The terms "part" and "fragment" are used interchangeably herein and refer to a continuous element. For example, a portion of a structure, such as an amino acid sequence or a protein, refers to a continuous element of the structure. An a portion of, a part of, or a fragment of a structure preferably comprises one or more functional properties of the structure. For example, a portion of an epitope, peptide or protein is preferably immunoequivalent to the epitope, peptide or protein from which it is derived. In the context of the present invention, a "portion" of a structure (such as an amino acid sequence) preferably comprises, preferably consists of, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98%, at least 99% of the entire structure or amino acid sequence.

[0104] The term "immunoreactive cell" in the context of the present invention relates to a cell that performs effector functions during an immune response. An "immunoreactive cell" is preferably capable of binding to an antigen or a cell characterized by presenting an antigen or a peptide fragment thereof (e.g., a T cell epitope) and mediating an immune response. For example, such cells secrete cytokines and / or chemokines, secrete antibodies, recognize cancer cells, and optionally eliminate such cells. For example, immune reactive cells include T cells (cytotoxic T cells, helper T cells, tumor infiltrating T cells), B cells, natural killer cells, neutrophils, macrophages, and dendritic cells. Preferably, in the context of the present invention, an "immunoreactive cell" is a T cell, preferably a CD4 + and / or CD8 + T cells.

[0105] Preferably, an "immune responsive cell" recognizes an antigen or a peptide fragment thereof, with a degree of specificity, particularly if presented by an MHC molecule, such as on the surface of an antigen presenting cell or a diseased cell such as a cancer cell. Preferably, the recognition is such that the cell that recognizes the antigen or a peptide fragment thereof is responsive or reactive. If the cell is a helper T cell (CD4 T cell) that contains a receptor that recognizes the antigen or a peptide fragment thereof, + T cells), in the case of MHC class II molecules, such responses or reactions may include the release of cytokines and / or CD8 +Activation of lymphocytes (CTLs) and / or B cells. If the cell is a CTL, such response or reaction may include the clearance of cells presented by MHC class I molecules, i.e., cells characterized by antigen presentation by class I MHC, such as by apoptosis or perforin-mediated cell lysis. According to the present invention, CTL responses may include sustained calcium efflux, cell division, the production of cytokines such as IFN-γ and TNF-α, the upregulation of activation markers such as CD44 and CD69, and specific cell lysis and killing of antigen-expressing target cells. CTL responses can also be determined using artificial reporter genes that accurately display CTL responses. Such CTLs that recognize antigens or antigen fragments and are responsive or reactive are also referred to herein as "antigen-responsive CTLs." If the cell is a B cell, such response may include the release of immunoglobulins.

[0106] The term "T cell" is used interchangeably with "T lymphocyte" herein and includes helper T cells (CD4+ T cells) and cytotoxic T cells (CTL, CD8+ T cells) including cytolytic T cells.

[0107] T cells belong to a group of white blood cells called lymphocytes and play a central role in cell-mediated immunity. They can be distinguished from other lymphocyte types, such as B cells and natural killer cells, by the presence of specialized receptors on their cell surface called T cell receptors (TCRs). The thymus is the primary organ responsible for T cell maturation. Several different T cell subsets have been identified, each with distinct functions.

[0108] In the immune process, helper T cells help other white blood cells, including the maturation of B cells into plasma cells and the activation of cytotoxic T cells and macrophages, among other functions. These cells are also called CD4+ T cells because they express the CD4 protein on their surface. When helper T cells are faced with peptide antigens presented by MHC class II molecules expressed on the surface of antigen presenting cells (APCs), they are activated. Once activated, they rapidly divide and secrete small proteins called cytokines that regulate or help the active immune response.

[0109] Cytotoxic T cells destroy virus-infected cells and tumor cells and are also involved in transplant rejection. Because these cells express the CD8 glycoprotein on their surface, they are also called CD8+ T cells. These cells recognize their targets by binding to antigens associated with MHC class I, which is present on the surface of nearly every cell in the body.

[0110] Most T cells have a T cell receptor (TCR) that exists as a complex of several proteins. The true T cell receptor is composed of two separate peptide chains, which are produced by separate T cell receptor α and β (TCRα and TCRβ) genes and are called the α-TCR chain and the β-TCR chain. γδ T cells (γδ T cells) represent a small T cell subpopulation with different T cell receptors (TCRs) on their surface. However, in γδ T cells, the TCR is composed of one γ-chain and one δ-chain. This T cell population is less common than the αβ T cell population (2% of total T cells).

[0111] The first signal in T cell activation is provided by the binding of the T cell receptor to a short peptide presented by the MHC on another cell. This ensures that only T cells with a TCR specific for that peptide are activated. The companion cell is usually an antigen-presenting cell such as a professional antigen-presenting cell, which in the case of an initial response is usually a dendritic cell, although B cells and macrophages are important APCs.

[0112] According to the present invention, a molecule is capable of binding to a predetermined target if it has a significant affinity for the target and binds to the target in a standard assay. D A molecule cannot (substantially) bind to a target if it does not have a significant affinity for the target and does not significantly bind to the target in a standard assay.

[0113] Cytotoxic T lymphocytes can be generated in vivo by incorporating antigens or peptide fragments thereof into antigen-presenting cells. Antigens or peptide fragments thereof can be expressed as proteins, DNA (e.g., in vectors), or RNA. Antigens can be processed to produce peptide partners for MHC molecules, and the fragments can be presented without further processing. The latter is a special case if they are able to bind to MHC molecules. Generally, administration to patients is possible via intradermal injection. However, injection can also be performed via intranodal injection into lymph nodes (Maloy et al. (2001), Proc Natl Acad Sci USA 98:3299-303). The generated cells present the target complex and are recognized by autologous cytotoxic T lymphocytes, which then proliferate.

[0114] Specific activation of CD4+ or CD8+ T cells can be detected by a variety of methods. Methods for detecting specific T cell activation include detecting T cell proliferation, cytokine (e.g., lymphokine) production, or cytolytic activity. For CD4+ T cells, a preferred method for detecting specific T cell activation is to detect T cell proliferation. For CD8+ T cells, a preferred method for detecting specific T cell activation is to detect cytolytic activity.

[0115] By "cell characterized by antigen presentation" or "cell presenting an antigen" or similar expressions, it is meant that a cell, such as a diseased cell (e.g., a cancer cell) or an antigen-presenting cell, presents an antigen it expresses or a fragment derived from said antigen in the context of an MHC molecule, particularly an MHC class I molecule, e.g., by processing the antigen. Similarly, the term "disease characterized by antigen presentation" refers to a disease involving cells characterized by antigen presentation, particularly presentation by MHC class I. Antigen presentation by a cell may be effected by transfecting the cell with a nucleic acid, such as an RNA encoding said antigen.

[0116] By "fragment of a presented antigen" or similar expressions, it is meant that the fragment can be presented by MHC class I or class II, preferably by MHC class I, e.g. when added directly to an antigen presenting cell. In one embodiment, the fragment is a fragment that is naturally presented by cells expressing the antigen.

[0117] The term "immune equivalent" means that immunoequivalent molecules, such as immunoequivalent amino acid sequences, exhibit the same or substantially the same immunological properties and / or exhibit the same or substantially the same immunological effect, for example with respect to the type of immunological effect, such as the induction of humoral and / or cellular immune responses, the intensity and / or duration of the induced immune response, or the specificity of the induced immune response. In the context of the present invention, the term "immune equivalent" is preferably used with respect to the immunological effects or properties of the peptides used for immunization. For example, an amino acid sequence is immunoequivalent to a reference amino acid sequence if, when exposed to the immune system of a subject, it induces an immune response that is specific for the reaction with the reference amino acid sequence.

[0118] The term "immune effector function" in the context of the present invention includes any function mediated by components of the immune system that results in, for example, killing of tumor cells, or inhibition of tumor growth and / or inhibition of tumor progression, including inhibition of tumor spread and metastasis. Preferably, the immune effector function in the context of the present invention is a T cell mediated effector function. Such functions include, when for helper T cells (CD4 + T cells), in the case of MHC class II molecules, antigen or antigen fragment recognition by T cell receptors, cytokine release and / or CD8 +Activation of lymphocytes (CTLs) and / or B cells and, in the case of CTLs, recognition of antigens or antigen fragments by T cell receptors in the case of MHC class I molecules, clearance of cells presented by MHC class I molecules (i.e., cells characterized by presentation of antigen by MHC class I, e.g., by apoptosis or perforin-mediated cytolysis), production of cytokines such as IFN-γ and TNF-α, and specific cytolytic killing of target cells expressing the antigen.

[0119] According to the present invention, the term "score" relates to the result of a test or examination, usually expressed as a number. Terms such as "score is better" or "score is best" relate to the better result or the best result of a test or examination.

[0120] Terms such as "predict," "predicting," or "prediction" relate to the determination of a likelihood.

[0121] According to the present invention, determining a score for binding of a peptide to one or more MHC molecules comprises determining the likelihood that the peptide will bind to the one or more MHC molecules.

[0122] The score of a peptide binding to one or more MHC molecules can be determined by utilizing any peptide:MHC binding prediction tool. For example, the Immune Epitope Database Analysis Resource (IEDB-AR) can be utilized. http: / / tools.iedb.org ).

[0123] Predictions are typically made for a panel of MHC molecules, such as a panel of different MHC alleles, such as all possible MHC alleles or a panel of MHC alleles or a subset thereof present in a patient, preferably with modifications, and to be determined for immunogenicity according to the invention.

[0124] According to the present invention, determining a score for binding of a modified peptide to one or more T cell receptors when present in an MHC-peptide complex comprises determining the likelihood of the peptide binding to a T cell receptor when present in a complex with an MHC molecule.

[0125] Predictions can be made for one T cell receptor, such as a T cell receptor present in a patient, or preferably for a panel of T cell receptors, such as an unknown panel of different T cell receptors or a panel of T cell receptors present in a patient, preferably with modifications, which will be assayed for immunogenicity according to the invention, or a subset thereof.

[0126] Furthermore, predictions are typically made for a panel of MHC molecules, such as a panel of different MHC alleles, such as all possible MHC alleles or a panel of MHC alleles or a subset thereof present in a patient, preferably with a modification, which immunogenicity is to be determined according to the invention.

[0127] The score for binding of a modified peptide to one or more T cell receptors when present in an MHC-peptide complex can be determined by evaluating the effect of the modification on T cell receptor-peptide-MHC complex binding for a given (unknown) T cell receptor library. The score for binding of a modified peptide to one or more T cell receptors when present in an MHC-peptide complex is generally defined as a representation of recognition of a given peptide-MHC molecule by a matching T cell receptor.

[0128] The score of binding of the modified peptide to one or more T cell receptors when present in an MHC-peptide complex can be determined by the physicochemical differences between the modified amino acid and the non-modified amino acid. For example, a substitution matrix can be used. Such a matrix describes the rate at which one amino acid in the sequence changes to another amino acid state over time.

[0129] For example, a log-odds matrix, such as an evolution-based log-odds matrix, can be used: substitutions with low log-odds scores have a greater chance of finding a matching T cell receptor from the pool of (unknown) T cell receptor molecules than substitutions with high log-odds scores (due to negative selection of T cell receptors that match non-modified peptides). However, there are other methods for determining this score. For example, taking into account the site of the mutation in the peptide (some sites may have less impact on binding than other sites), taking into account the nearest neighbors of the substituted amino acid (which may affect the secondary structure of the substituted amino acid), taking into account the entire peptide sequence, taking into account the entire structural information of the peptide in the MHC molecule, etc. Determining the score also includes determining the T cell receptor repertoire (such as the patient's T cell receptor repertoire or a subset thereof), for example by NGS and performing docking simulations of T cell receptor-peptide-MHC complexes.

[0130] The present invention may also encompass performing the methods of the present invention on different peptides comprising the same modification and / or different modifications.

[0131] In one embodiment, the term "different peptides comprising the same modification" relates to different fragments of a protein comprising a modification or peptides consisting thereof, wherein the different fragments comprise the same modification present in the protein, but differ in length and / or modification site. If a protein is modified at site x, then two or more fragments of the protein, each of which comprises a different sequence window of the protein encompassing site x, are considered to be different peptides comprising the same modification.

[0132] In one embodiment, the term "different peptides comprising different modifications" relates to peptides of the same and / or different lengths comprising different modifications of the same and / or different proteins. If a protein is modified at positions x and y, then two fragments of the protein comprising a sequence window encompassing the protein at position x or position y are considered to be different peptides comprising different modifications.

[0133] The present invention may also include fragmenting protein sequences containing modifications (whose immunogenicity is determined according to the present invention) into peptide lengths suitable for MHC binding, and determining the binding scores of different modified peptides containing the same and / or different modifications of the same and / or different proteins to one or more MHC molecules. The output may be ranked, and the results may consist of a list of peptides and a predicted score indicating their binding likelihood.

[0134] The step of determining the binding score of the non-modified peptide to one or more MHC molecules and / or the step of determining the binding score of the modified peptide present in the MHC-peptide complex to one or more T cell receptors can then be performed, and can be performed on all different modified peptides comprising the same and / or different modifications, a subset thereof, such as those modified peptides comprising the same and / or different modifications that have the highest binding score to one or more MHC molecules, or only the one modified peptide that has the highest binding score to one or more MHC molecules.

[0135] Following these additional steps, the results may be ranked and may consist of a list of peptides and a prediction score indicating their likelihood of being immunogenic.

[0136] Preferably, in such an arrangement, the score for binding of the modified peptide to one or more MHC molecules is weighted higher than the score for binding of the modified peptide in an MHC-peptide complex to one or more T-cell receptors, preferably higher than the score for chemical and physical similarity between the non-modified amino acid and the modified amino acid, and the score for binding of the modified peptide in an MHC-peptide complex to one or more T-cell receptors, preferably the score for chemical and physical similarity between the non-modified amino acid and the modified amino acid, is weighted higher than the score for binding of the non-modified peptide to one or more MHC molecules.

[0137] The amino acid modifications (whose immunogenicity is determined according to the present invention) may result from mutations in the nucleic acid in the cell. Such mutations can be identified by known sequencing techniques.

[0138] In one embodiment, the mutation is a cancer-specific somatic mutation in a tumor sample from a cancer patient, which can be determined by determining sequence differences between the genome, exome, and / or transcriptome of the tumor sample and the genome, exome, and / or transcriptome of a non-tumor sample.

[0139] According to the present invention, tumor sample relates to any sample, such as a body sample derived from a patient containing or possibly containing tumor cells or cancer cells. The body sample can be any tissue sample such as blood, a tissue sample obtained from a primary tumor or from tumor metastasis, or any other sample containing tumor cells or cancer cells. Preferably, the body sample is blood, and the cancer-specific somatic mutations or sequence differences in one or more circulating tumor cells (CTCs) contained in the blood are measured. In another embodiment, the tumor sample relates to one or more isolated tumor cells or cancer cells such as circulating tumor cells (CTCs), or comprises a sample of one or more isolated tumor cells or cancer cells such as circulating tumor cells (CTCs).

[0140] Non-tumor samples refer to any sample such as a body sample derived from a patient or other individual, the other individual is preferably an individual of the same type as the patient, preferably a healthy individual that does not contain or is unlikely to contain tumor cells or cancer cells. The body sample can be any tissue sample such as blood or a sample derived from non-tumor tissue.

[0141] The present invention may include the determination of a patient's cancer mutation signature. The term "cancer mutation signature" may refer to all cancer mutations present in one or more cancer cells of a patient or it may refer to only a portion of cancer mutations present in one or more cancer cells of a patient. Thus, the present invention may include the identification of all cancer-specific mutations present in one or more cancer cells of a patient or may include the identification of only a portion of cancer-specific mutations present in one or more cancer cells of a patient. Typically, the methods of the present invention provide for the identification of multiple mutations, which provide a sufficient number of modifications or modified peptides for inclusion in the methods of the present invention.

[0142] Preferably, the mutations identified according to the present invention are non-synonymous mutations, preferably non-synonymous mutations of a protein expressed in a tumor cell or cancer cell.

[0143] In one embodiment, cancer-specific somatic mutations or sequence differences are determined in a genome, preferably the entire genome of a tumor sample. Thus, the present invention can include identifying a cancer mutation signature of a genome, preferably the entire genome of one or more cancer cells. In one embodiment, the step of identifying cancer-specific somatic mutations in a tumor sample from a cancer patient comprises identifying a genome-wide cancer mutation profile.

[0144] In one embodiment, cancer-specific somatic mutations or sequence differences are determined in an exome, preferably the entire exome of a tumor sample. Thus, the present invention may include identifying a cancer mutation signature in an exome, preferably the entire exome of one or more cancer cells. In one embodiment, the step of identifying cancer-specific somatic mutations in a tumor sample from a cancer patient comprises identifying a whole-exome cancer mutation profile.

[0145] In one embodiment, cancer-specific somatic mutations or sequence differences are determined in a transcriptome, preferably the entire transcriptome of a tumor sample. Thus, the present invention can include identifying a cancer mutation signature of a transcriptome, preferably the entire transcriptome of one or more cancer cells. In one embodiment, the step of identifying cancer-specific somatic mutations in a tumor sample from a cancer patient comprises identifying a transcriptome-wide cancer mutation profile.

[0146] In one embodiment, the step of identifying cancer-specific somatic mutations or identifying sequence differences comprises single cell sequencing of one or more, preferably 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or even more cancer cells. Thus, the present invention can include identifying cancer mutation signatures of one or more cancer cells. In one embodiment, the cancer cells are circulating tumor cells. Cancer cells such as circulating tumor cells can be isolated prior to single cell sequencing.

[0147] In one embodiment, the step of identifying cancer-specific somatic mutations or identifying sequence differences comprises the use of next generation sequencing (NGS).

[0148] In one embodiment, the step of identifying cancer-specific somatic mutations or identifying sequence differences comprises sequencing genomic DNA and / or RNA from a tumor sample.

[0149] To reveal cancer-specific somatic mutations or sequence differences, sequence information obtained from a tumor sample is preferably compared with sequence information obtained from a reference, such as nucleic acid (e.g., DNA or RNA) sequencing of normal non-cancerous cells, such as germ cells, which can be obtained from the patient or a different individual. In one embodiment, normal genomic germline DNA is obtained from peripheral blood mononuclear cells (PBMCs).

[0150] The term "genome" refers to the total amount of genetic information in the chromosomes of an organism or cell.

[0151] The term "exome" refers to the portion of an organism's genome formed by exons, which are the coding portions of expressed genes. The exome provides the genetic blueprint for the synthesis of proteins and other functional gene products. It is the part of the genome most relevant to function and therefore most likely to influence the phenotype of an organism. It has been estimated that the exome of the human genome comprises 1.5% of the total genome (Ng, PC et al., PLoS Gen., 4(8):1-15, 2008).

[0152] The term "transcriptome" refers to the set of all RNA molecules, including mRNA, rRNA, tRNA and other non-coding RNAs, produced in a cell or cell population. In the context of the present invention, the transcriptome means the set of all RNA molecules produced by a cell, a cell population, preferably a cancer cell population, or all cells of a given individual at a certain point in time.

[0153] According to the present invention, "nucleic acid" is preferably deoxyribonucleic acid (DNA or ribonucleic acid (RNA), more preferably RNA, most preferably in vitro transcribed RNA (IVT RNA) or synthetic RNA. According to the present invention, nucleic acids include genomic DNA, cDNA, mRNA, recombinantly produced and chemically synthesized molecules. According to the present invention, nucleic acids can be present in the form of single-stranded or double-stranded molecules and linear or covalently circular closed molecules. According to the present invention, nucleic acids can be isolated. According to the present invention, the term "isolated nucleic acid" means that the nucleic acid (i) is amplified in vitro, for example by polymerase chain reaction (PCR); (ii) is produced by cloning and recombination; (iii) is purified, for example by shearing and separation by gel electrophoresis; or (iv) is synthesized, for example by chemical synthesis. The nucleic acid can be used for introduction into cells, i.e. transfection, in particular in the form of RNA that can be prepared by in vitro transcription from a DNA template. In addition, the RNA can be modified by stabilizing sequences, capping and polyadenylation before use.

[0154] The term "genetic material" refers to an isolated nucleic acid, DNA or RNA, a segment of a double helix, a segment of a chromosome, or the entire genome of an organism or cell, particularly its exome or transcriptome.

[0155] The term "mutation" refers to a change or difference (nucleotide substitution, addition or deletion) in a nucleic acid sequence compared to a reference. "Somatic mutations" can occur in any cell of the body except germ cells (sperm and eggs) and are therefore not passed on to children. These changes can (but not always) cause cancer or other diseases. Preferably, the mutation is a non-synonymous mutation. The term "non-synonymous mutation" refers to a mutation that results in an amino acid change (such as an amino acid substitution) in the translation product, preferably a nucleotide substitution.

[0156] According to the present invention, the term "mutation" includes point mutations, insertions and deletions (Indels), fusions, chromothripsis, and RNA editing.

[0157] According to the present invention, the term "indel" describes a specific class of mutations defined as mutations that result in the coexistence of insertions and deletions of nucleotides, as well as net gains or losses. In coding regions of the genome, frameshift mutations occur unless the length of the indel is a multiple of three. Indels can be contrasted with point mutations; indels insert and delete nucleotides from a sequence, while point mutations are a form of substitution that replaces a single nucleotide.

[0158] Fusion can produce a hybrid gene formed by two previously separate genes. It can be caused by translocation, interstitial deletion or chromosomal inversion. Typically, fusion genes are oncogenes. Oncogenic fusion genes can produce gene products with new functions or functions that differ from the two fused genes from which they originated. Alternatively, the proto-oncogene is fused to a strong promoter, so that carcinogenesis begins to play a role through upregulation caused by the strong promoter of the upstream fusion part. Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events.

[0159] According to the present invention, the term "chromothripsis" refers to a genetic phenomenon by which specific regions of the genome are fragmented and then joined together by a single destructive event.

[0160] According to the present invention, the term "RNA editing" or "RNA editing" refers to a molecular process in which the information content in an RNA molecule is changed by chemical changes in base composition. RNA editing includes nucleoside modifications such as deamination of cytosine (C) to uridine (U) and adenine (A) to inosine (I), as well as the addition and insertion of non-templated nucleotides. RNA editing in mRNA effectively changes the amino acid sequence of the encoded protein, making it different from that predicted by the genomic DNA sequence.

[0161] The term "cancer mutation signature" refers to a set of mutations present in cancer cells when compared to non-cancerous reference cells.

[0162] According to the present invention, a "reference" can be used to correlate and compare results obtained from tumor samples using the methods of the present invention. Typically, a "reference" can be obtained based on one or more normal samples, particularly samples obtained from the patient or one or more different individuals, preferably healthy individuals, particularly individuals of the same species, that are not affected by cancer. The "reference" can be determined empirically by testing a sufficiently large number of normal samples.

[0163] According to the present invention, any suitable sequencing method can be used to measure mutation, preferably second generation sequencing (NGS) technology. Future third generation sequencing may replace NGS technology, to accelerate the speed of sequencing step in the method. For the purpose of illustration: term " second generation sequencing " or " NGS " in the context of the present invention means all high-throughput sequencing technologies, compared with " tradition " sequencing methodology known as Sanger chemistry (Sanger chemistry), it is by dividing the whole genome into small fragments, along the whole genome parallel random read nucleic acid template. Such NGS technology (also known as large-scale parallel sequencing technology) can deliver the nucleic acid sequence information of whole genome, exon group, transcriptome (whole transcription sequence of genome) or methylation group (whole methylation sequence of genome) in very short time period, for example, in 1-2 week, preferably in 1-7 days or most preferably in less than 24 hours, and in principle allow single cell sequencing method. In the context of the present invention, a variety of commercially available NGS platforms or NGS platforms mentioned in the literature can be used, such as those described in detail in Zhang et al. 2011: The impact of next-generation sequencing on genomics. J. Genet Genomics 38(3), 95-109; or Voelkerding et al. 2009: Next generation sequencing: From basic research to diagnostics. Clinical chemistry 55, 641-658. Non-limiting examples of such NGS technologies / platforms are:

[0164] 1) For example, the GS-FLX 454 Genome Sequencer from Roche's 454 Life Sciences TM (Branford, Connecticut) is a synthesis sequencing technology called pyrophosphate sequencing, which was first described in Ronaghi et al. 1998: A sequencing method based on real-time pyrophosphate. Science 281(5375), 363-365. This technology utilizes emulsion PCR (emulsion PCR amplification) in which single-stranded DNA-bound beads are encapsulated in aqueous micelles containing PCR reactants surrounded by oil by vigorous vortexing to perform emulsion PCR amplification. During pyrophosphate sequencing, the light emitted by the phosphate molecules upon nucleotide incorporation is recorded as the DNA chain synthesized by the polymerase.

[0165] 2) Synthesis sequencing method developed by Solexa (now part of Illumina Inc., San Diego, California), based on reversible dye terminators and used in, for example, the Illumina / Solexa Genome Analyzer TM and Illumina HiSeq 2000 Genome Analyzer TM In this technique, all four nucleotides and DNA polymerase are added simultaneously to oligonucleotide primer cluster fragments within a flow cell channel. Bridge amplification extends the cluster chain with all four fluorescently labeled nucleotides for sequencing.

[0166] 3) Sequencing by ligation, such as the SOLid sequencing method used by Applied Biosystems (now Life Technologies Corporation, Carlsbad, California) TM The ligation sequencing method of the platform. In this technology, all possible oligonucleotides of fixed length are labeled according to the sequencing site. The oligonucleotides are annealed and ligated; the preferential ligation of matching sequences by DNA ligase generates signal information of the nucleotide at that site. Before sequencing, the DNA is amplified by emulsion PCR. The resulting beads, each containing only copies of the same DNA molecule, are placed on a slide. As a second example, the Polonator TM The G.007 platform also applies a ligation sequencing approach, using random array, bead-based emulsion PCR amplification to amplify DNA fragments for parallel sequencing.

[0167] 4) Single-molecule sequencing technologies, such as the PacBio RS system used by Pacific Biosciences (Menlo Park, California) or the HeliScope system used by Helicos Biosciences (Cambridge, Massachusetts) TM Single-molecule sequencing technology in the platform. The remarkable feature of this technology is its ability to sequence single DNA or RNA molecules without amplification, which is defined as single-molecule real-time (SMRT) DNA sequencing. For example, HeliScope uses a highly sensitive fluorescence detection system to directly detect each nucleotide as it is synthesized. Visigen Biotechnology (Houston, Texas) has developed a similar technology based on fluorescence resonance energy transfer (FRET). Other fluorescence-based single-molecule technologies come from US Genomics (GeneEngineTM ) and Genovoxx (AnyGene TM ).

[0168] 5) Nanotechnology for single molecule sequencing, where different nanostructures are used, for example arranged on a chip, to monitor the movement of polymerase molecules on a single strand during replication. A non-limiting example of a nanotechnology-based approach is GridON from Oxford Nanopore Technologies (Oxford, UK). TM platform, hybridization-assisted nanopore sequencing (HANS) developed by Nabsys (Providence, Rhode Island) TM ) platform, and a technique called combined probe anchor ligation (cPAL TM )'s proprietary ligase-based DNA sequencing platform with DNA nanoball (DNB) technology.

[0169] 6) Electron microscopy-based single-molecule sequencing technologies, such as those developed by LightSpeed ​​Genomics (Sunnyvale, California) and Halcyon Molecular (Redwood City, California).

[0170] 7) Ionic semiconductor sequencing based on the detection of hydrogen ions released during DNA polymerization. For example, IonTorrent Systems (San Francisco, California) implements this biochemical process in a massively parallel manner using a high-density array of micromachined wells. Each well contains a different DNA template. Below the wells is an ion-sensitive layer, and below that is a proprietary ion sensor.

[0171] Preferably, DNA and RNA preparations are used as the starting material of NGS. Such nucleic acids can be easily obtained from samples such as biological materials, for example, from fresh, quick-frozen or formalin-fixed paraffin-embedded tumor tissue (FFPE) or from freshly separated cells or from the CTC present in patient peripheral blood. Normal non-mutated genomic DNA or RNA can be extracted from normal somatic tissue, but preferably germ cells in the case of the present invention. Germline DNA or RNA can be extracted from the peripheral blood mononuclear cells (PBMC) of patients suffering from non-hematological malignancies. Although the nucleic acids extracted from FFPE tissues or freshly separated unicellular are highly fragmented, they are suitable for NGS applications.

[0172] Several targeted NGS methods for exome sequencing are described in the literature (for review, see, e.g., Teer and Mullikin 2010: Human Mol Genet 19(2), R145-51), all of which can be used in conjunction with the present invention. Many of these methods (described as, e.g., genome capture, genome partitioning, genome enrichment, etc.) utilize hybridization techniques and include array-based hybridization methods (e.g., Hodges et al. 2007: Nat. Genet. 39, 1522-1527) and liquid-based hybridization methods (e.g., Choi et al. 2009: Proc. Natl. Acad. Sci USA 106, 19096-19101). Commercial kits for DNA sample preparation and subsequent exome capture are also available: for example, Illumina Inc. (San Diego, California) offers the TruSeq TM TruSeq DNA sample preparation kit and exome enrichment kit TM Exome enrichment kit.

[0173] In order to reduce the number of false positive results in the detection of cancer-specific somatic mutations or sequence differences, when, for example, a tumor sample sequence is compared to a reference sample sequence such as a germ cell sample sequence, it is preferred to determine the replicate sequences of one or both of these sample types. Therefore, it is preferred that the sequence of the reference sample, such as the sequence of the germ line sample, be determined twice, three times or more. Alternatively or in addition, the sequence of the tumor sample is determined twice, three times or more. It is also possible to repeatedly determine the sequence of the reference sample (such as the sequence of the germ line sample) and / or the sequence of the tumor sample by at least determining the genomic DNA sequence and at least determining the RNA sequence in the reference sample and / or the tumor sample once. For example, by determining the variation between replicates of a reference sample such as a germ line sample, the expected rate of false positive (FDR) somatic mutations as a statistic can be estimated. The technical repetitions of the sample should produce the same results, and any variation of the detection in this "same to same comparison" is a false positive. In particular, in order to determine the false discovery rate of somatic mutation detection of a tumor sample relative to a reference sample, the technical repetitions of the reference sample can be used as a reference to estimate the number of false positives. Furthermore, machine learning methods can be used to combine different quality-related metrics (e.g., coverage or SNP quality) into a single quality score. For a given somatic variant, all other variants with a superior quality score can be calculated, enabling the ranking of all variants in a dataset.

[0174] In the context of the present invention, the term "RNA" refers to a molecule comprising at least one ribonucleotide residue, and is preferably composed entirely of or substantially composed of ribonucleotide residues." Ribonucleotides" refer to nucleotides containing a hydroxyl group at the 2'-position of a β-D-ribofuranosyl group. The term "RNA" includes double-stranded RNA, single-stranded RNA, isolated RNA (such as partially or completely purified RNA), substantially purified RNA, synthetic RNA, and recombinantly produced RNA, such as modified RNA, which is different from naturally occurring RNA by the addition, deletion, substitution, and / or change of one or more nucleotides. Such changes can include the addition of non-nucleotide substances, such as addition to the RNA ends or interior, for example, at one or more nucleotides of the RNA. The nucleotides in the RNA molecule can also include non-standard nucleotides, such as non-naturally occurring nucleotides or chemically synthesized nucleotides or deoxynucleotides. The RNAs of these changes can be referred to as analogs or analogs of naturally occurring RNA.

[0175] According to the present invention, the term "RNA" includes and preferably refers to "mRNA." The term "mRNA" means "messenger-RNA" and refers to a "transcript" produced using a DNA template and encoding a peptide or polypeptide. Typically, mRNA comprises a 5'-UTR, a protein-coding region, and a 3'-UTR. mRNA has only a limited half-life in cells and in vitro. In the context of the present invention, mRNA can be produced by in vitro transcription from a DNA template. In vitro transcription methodologies are known to those skilled in the art. For example, various commercially available in vitro transcription kits are available.

[0176] According to the present invention, the stability and translation efficiency of RNA can be optionally improved.For example, RNA can be modified by one or more RNAs with stabilizing effect and / or increasing translation efficiency to stabilize RNA and enhance translation.This type of modification is described in, for example, PCT / EP2006 / 009448, which is incorporated herein by reference. In order to increase the expression of the RNA used according to the present invention, the coding region can be modified, i.e., the sequence of the peptide or protein encoded for expression, preferably the sequence of the peptide or protein expressed is not changed, to increase GC-content, thereby enhancing the stability of mRNA and implementing codon optimization, and therefore enhancing the translation in the cell.

[0177] In the context of the RNA of the invention, the term "modification" includes any modification of the RNA that does not naturally occur in the RNA in question.

[0178] In one embodiment of the invention, the RNA used according to the invention does not contain uncapped 5'-triphosphates. Removal of such uncapped 5'-triphosphates can be achieved by treating the RNA with a phosphatase.

[0179] The RNA according to the present invention may contain modified ribonucleotides to increase its stability and / or reduce cytotoxicity. For example, in one embodiment, in the RNA used according to the present invention, 5-methylcytosine is partially or completely substituted, preferably completely substituted, for cytosine. Alternatively or in addition, in one embodiment, in the RNA used according to the present invention, pseudouridine is partially or completely substituted, preferably completely substituted for uridine.

[0180] In one embodiment, the term "modification" refers to providing a 5'-cap or a 5'-cap analog to the RNA. The term "5'-cap" refers to the cap structure found at the 5'-end of an mRNA molecule and is typically composed of a guanine nucleotide attached to the mRNA via an unusual 5' to 5' triphosphate linkage. In one embodiment, the guanosine nucleoside is methylated at the 7-position. The term "conventional 5'-cap" refers to a naturally occurring RNA 5'-cap, preferably a 7-methylguanosine nucleoside cap (m 7 G) In the context of the present invention, the term "5'-cap" includes 5'-cap analogs that are structurally similar to the RNA cap and that have been modified to have the ability to stabilize RNA and / or enhance RNA translation (if attached thereto), preferably in vivo and / or in cells.

[0181] Providing RNA with a 5'-cap or a 5'-cap analog can be achieved by in vitro transcription of a DNA template in the presence of the 5'-cap or 5'-cap analog, wherein the 5'-cap is co-transcriptionally incorporated into the resulting RNA chain, or the RNA can be produced, for example, by in vitro transcription and the 5'-cap can be attached to the RNA post-transcriptionally using a capping enzyme, such as that of vaccinia virus.

[0182] The RNA may comprise other modifications. For example, other modifications of the RNA used in the present invention may be extensions or truncations of naturally occurring poly(A) tails or alterations of the 5'- or 3'-untranslated regions (UTRs), such as the introduction of UTRs that are not associated with the coding region of the RNA, for example alterations of existing 3'-UTRs or the insertion of one or more, preferably two copies of, 3'-UTRs derived from globin genes (e.g., α2-globin, α1-globin, β-globin, preferably β-globin, more preferably human β-globin).

[0183] RNA with unmasked poly-A sequence is more efficiently transcribed than RNA with masked poly-A sequence. The term "poly (A) tail" or "poly-A sequence" refers to a sequence of adenine (A) residues typically located at the 3'-end of an RNA molecule, and "unmasked poly-A sequence" means that the end of the poly-A sequence at the 3' end of the RNA molecule is the A of the poly-A sequence and, except for the A at the 3' end, thereafter, i.e., downstream of the poly-A sequence, does not contain nucleotides. In addition, a long poly-A sequence of about 120 base pairs produces optimal RNA transcript stability and translation efficiency.

[0184] Therefore, in order to enhance the stability and / or expression of the RNA used according to the present invention, it can be modified so as to be present in a form connected to a poly-A sequence, the poly-A sequence preferably having a length of 10-500, more preferably 30-300, even more preferably 65-200, and in particular 100-150 adenosine residues. In a particularly preferred embodiment, the poly-A sequence has a length of about 120 adenosine residues. In order to further enhance the stability and / or expression of the RNA used according to the present invention, the poly-A sequence may be unmasked.

[0185] In addition, incorporating a 3'-untranslated region (UTR) into the 3'-untranslated region of an RNA molecule can lead to an increase in translation efficiency. A synergistic effect can be achieved by incorporating two or more such 3'-untranslated regions. The 3'-untranslated region and the RNA into which it is introduced can be autologous or heterologous. In a specific embodiment, the 3'-untranslated region is derived from the human β-globin gene.

[0186] The combination of the modifications described above, ie, the incorporation of a poly-A sequence, the unmasking of a poly-A sequence, and the incorporation of one or more 3'-untranslated regions, has a synergistic effect on the stability of the RNA and the enhancement of translation efficiency.

[0187] The term "stability" of RNA relates to the half-life of the RNA. "Half-life" refers to the period of time required to clear half of the activity, amount, or number of molecules. In the context of the present invention, the half-life of an RNA indicates the stability of the RNA. The half-life of an RNA can affect the "duration of expression" of the RNA. It is expected that RNA with a long half-life will be expressed for an extended period of time.

[0188] Of course, if it is desired to reduce the stability and / or translation efficiency of RNA according to the present invention, the RNA can be modified to interfere with the function of the elements that enhance the stability and / or translation efficiency of RNA as described above.

[0189] The term "expression" as used in accordance with the present invention is used in its most general sense and includes the production of RNA and / or peptides, polypeptides or proteins, for example, by transcription and / or translation. With respect to RNA, the term "expression" or "translation" particularly relates to the production of peptides, polypeptides or proteins. It also includes partial expression of nucleic acids. Furthermore, expression can be transient or stable.

[0190] According to the present invention, the term expression also includes "aberrant expression" or "abnormal expression". "Abnormal expression" or "abnormal expression" means that according to the present invention, the expression is altered, preferably the expression is increased compared to a reference, for example compared to the state of a subject who does not suffer from a disease associated with aberrant or abnormal expression of a certain protein (e.g. a tumor antigen). Increased expression refers to an increase of at least 10%, in particular an increase of at least 20%, at least 50% or at least 100% or more. In one embodiment, the expression is only present in diseased tissue, while the expression in healthy tissue is suppressed.

[0191] The term "specifically expressed" means that the protein is basically only expressed in a specific tissue or organ. For example, a tumor antigen that is specifically expressed in the gastric mucosa means that the protein is mainly expressed in the gastric mucosa, and is not expressed in other tissues or can not be expressed to a significant degree in other tissues or organ types. Therefore, a protein that is only expressed in gastric mucosal cells and has a significantly lower expression level in any other tissue such as the testis is a protein that is specifically expressed in gastric mucosal cells. In some embodiments, under normal conditions, tumor antigens may also be specifically expressed in multiple tissue types or organs, such as in 2 or 3 tissue types or organs, but preferably in no more than 3 different tissues or organ types. In this case, then the tumor antigen is specifically expressed in these organs. For example, if under normal conditions, the tumor antigen is preferably expressed to an approximately equal degree in the lungs and stomach, then the antigen is specifically expressed in the lungs and stomach.

[0192] In the context of the present invention, the term "transcription" refers to the process by which the genetic code in a DNA sequence is transcribed into RNA. Subsequently, the RNA can be translated into protein. According to the present invention, the term "transcription" includes "in vitro transcription," wherein the term "in vitro transcription" refers to the process of synthesizing RNA, particularly mRNA, in vitro in a cell-free system, preferably using a suitable cell extract. Preferably, a cloning vector is used to generate the transcript. These cloning vectors are generally designated as transcription vectors and, according to the present invention, are encompassed by the term "vector." According to the present invention, the RNA used in the present invention is preferably in vitro transcribed RNA (IVT-RNA) and can be obtained by in vitro transcription from a suitable DNA template. The promoter controlling transcription can be any promoter of any RNA polymerase. Specific examples of RNA polymerases are T7, T3, and SP6 RNA polymerases. Preferably, according to the present invention, in vitro transcription is controlled by a T7 or SP6 promoter. The DNA template used for in vitro transcription can be obtained by cloning nucleic acids, particularly cDNA, and introducing it into a suitable vector for in vitro transcription. cDNA can be obtained by reverse transcription of RNA.

[0193] The term "translation" according to the present invention relates to the process by which a messenger RNA chain directs the assembly of an amino acid sequence in cellular ribosomes to produce a peptide, polypeptide or protein.

[0194] According to the present invention, an expression control sequence or regulatory sequence that can be functionally linked to a nucleic acid can be homologous or heterologous to the nucleic acid. A coding sequence and a regulatory sequence are "functionally" linked together if the coding sequence is controlled by or influenced by the regulatory sequence. If the coding sequence is to be translated into a functional protein, with the regulatory sequence being functionally linked to the coding sequence, induction of the regulatory sequence results in transcription of the coding sequence without causing a shift in the reading frame of the coding sequence or without causing the coding sequence to be unable to be translated into the desired protein or peptide.

[0195] According to the present invention, the term "expression control sequence" includes promoters, ribosome binding sequences, and other control elements that control nucleic acid transcription or translation of the obtained RNA. In certain embodiments of the present invention, the regulatory sequence can be controlled. The precise structure of the regulatory sequence can vary depending on the species or cell type, but generally includes 5'-untranscribed sequences and 5'- and 3'-untranslated sequences involved in the initiation of transcription or translation, such as TATA-boxes, capping sequences, CAAT sequences, etc. In particular, the 5'-untranscribed regulatory sequence includes a promoter region containing a promoter sequence that functions to control gene transcription. The regulatory sequence may also include an enhancer sequence or an upstream activator sequence.

[0196] Preferably, according to the present invention, the RNA expressed in the cell is introduced into said cell.In one embodiment of the method according to the present invention, the RNA to be introduced into the cell is obtained by in vitro transcription of a suitable DNA template.

[0197] According to the present invention, terms such as "RNA capable of expressing..." and "RNA encoding..." are used interchangeably herein and, with respect to a particular peptide or polypeptide, mean that the RNA is capable of expression to produce said peptide or polypeptide if present in a suitable environment, preferably in a cell. Preferably, according to the present invention, the RNA is capable of interacting with the cellular translation machinery to provide the peptide or polypeptide that it is capable of expressing.

[0198] Terms such as "transferring," "introducing," or "transfecting" are used interchangeably herein and relate to the introduction of a nucleic acid, particularly an exogenous or heterologous nucleic acid, particularly RNA, into a cell. According to the present invention, the cell may form part of an organ, tissue, and / or organism. According to the present invention, administration of the nucleic acid is achieved by naked nucleic acid or in combination with an administration agent. Preferably, the nucleic acid is administered in the form of naked nucleic acid. Preferably, the RNA is administered in combination with a stabilizing agent, such as an RNase inhibitor. The present invention also envisions repeated introduction of nucleic acids into cells to allow for sustained expression for extended periods of time.

[0199] Any carrier transfection cell that available RNA can be connected thereto, for example, by forming complex or forming vehicle with RNA, wherein RNA is sealing or is encapsulated, and this causes comparing with naked RNA, and the stability of RNA increases. Available carrier comprises according to the present invention, for example contains the carrier of lipid, such as cationic lipid (cationic lipid), liposome, cationic liposome and micelle and nanoparticle.Cationic lipid can form complex with negatively charged nucleic acid.Any cationic lipid can be used according to the present invention.

[0200] Preferably, the RNA encoding the peptide or polypeptide is introduced into the cell, particularly introduced into the cell present in the body, causing the peptide or polypeptide to be expressed in the cell. In a specific embodiment, the nucleic acid of the specific cell is preferably targeted. In this type of embodiment, the carrier (e.g., retrovirus or liposome) applied to the cell administration nucleic acid displays the target molecule. For example, the molecule (e.g., antibody) specific to the surface membrane protein on the target cell or the part of the receptor on the target cell can be incorporated into the nucleic acid vector or can be combined with it. In the case of administering nucleic acid by liposomes, the protein combined with the surface membrane protein related to endocytosis can be incorporated into the liposome formulation, thereby enabling targeting and / or uptake. This type of protein comprises the capsid protein fragment specific to a particular cell type, antibodies for internalized protein, proteins in the target cell intracellular site, etc.

[0201] The term "cell" or "host cell" is preferably a complete cell, i.e., a cell with an intact membrane that has not yet released its normal intracellular components such as enzymes, organelles or genetic material. The complete cell is preferably a living cell, i.e., a living cell that can perform its normal metabolic functions. According to the present invention, the term preferably relates to any cell that can be transformed or transfected with an exogenous nucleic acid. According to the present invention, the term "cell" includes prokaryotic cells (e.g., Escherichia coli (E. coli)) or eukaryotic cells (e.g., dendritic cells, B cells, CHO cells, COS cells, K562 cells, HEK293 cells, HELA cells, yeast cells and insect cells). The exogenous nucleic acid may exist in the cell in the following forms: (i) freely dispersed, (ii) incorporated into a recombinant vector, or (iii) integrated into the host cell genome or mitochondrial DNA. Mammalian cells are particularly preferred, such as cells from humans, mice, hamsters, pigs, sheep and primates. Cells can be derived from a variety of tissue types and include primary cells and cell lines. Specific examples include keratinocytes, peripheral blood lymphocytes, bone marrow stem cells and embryonic stem cells. In other embodiments, the cell is an antigen presenting cell, particularly a dendritic cell, a monocyte or a macrophage.

[0202] A cell comprising the nucleic acid molecule preferably expresses the peptide or polypeptide encoded by the nucleic acid.

[0203] The term "clonal expansion" refers to the process by which a specific entity increases. In the context of the present invention, the term is preferably used in the context of an immune response in which lymphocytes are stimulated by an antigen, proliferate, and specific lymphocytes that recognize the antigen expand. Preferably, clonal expansion leads to lymphocyte differentiation.

[0204] Terms such as "reduce" or "inhibit" relate to the ability to cause an overall decrease in levels, preferably a decrease of 5% or more, 10% or more, 20% or more, more preferably 50% or more and most preferably 75% or more. The term "inhibit" or similar phrases includes complete or substantially complete inhibition, i.e., a decrease to zero or substantially to zero.

[0205] Terms such as "increase", "enhance", "promote" or "extend" preferably relate to an increase, enhancement, promotion or extension of about at least 10%, preferably at least 20%, preferably at least 30%, preferably at least 40%, preferably at least 50%, preferably at least 80%, preferably at least 100%, preferably at least 200%, and in particular at least 300%. These terms can also relate to an increase, enhancement, promotion or extension from zero or an unmeasurable or undetectable level to a level greater than zero or to a measurable or detectable level.

[0206] The present invention provides vaccines, such as cancer vaccines, designed based on amino acid modifications or modified peptides predicted to be immunogenic by the methods of the present invention.

[0207] According to the present invention, the term "vaccine" refers to a pharmaceutical preparation (pharmaceutical composition) or product that, when administered, induces an immune response, particularly a cellular immune response, that recognizes and attacks antigens or diseased cells, such as cancer cells. Vaccines can be used for the prevention or treatment of diseases. The term "personalized cancer vaccine" or "individualized cancer vaccine" refers to a specific cancer patient and means a cancer vaccine that is adapted to the needs or special circumstances of an individual cancer patient.

[0208] In one embodiment, the vaccine provided according to the present invention may comprise a peptide or polypeptide or a nucleic acid encoding the peptide or polypeptide, preferably RNA, wherein the modification or modified peptide contains one or more amino acid modifications or one or more modified peptides predicted to be immunogenic by the method of the present invention.

[0209] The cancer vaccines provided herein, when administered to a patient, provide one or more T cell epitopes suitable for stimulating, priming, and / or expanding T cells specific to the patient's tumor. The T cells are preferably directed against cells expressing the antigen from which the T cell epitopes are derived. Thus, the vaccines described herein are preferably capable of eliciting or promoting a cellular response, preferably cytotoxic T cell activity, against cancers characterized by presentation of one or more tumor-associated neoantigens by MHC class I. Because the vaccines provided herein target cancer-specific mutations, they are specific for the patient's tumor.

[0210] The vaccines provided according to the present invention relate to vaccines that, when administered to a patient, preferably provide one or more T cell epitopes, such as 2 or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, and preferably up to 60, up to 55, up to 50, up to 45, up to 40, up to 35 or up to 30 T cell epitopes that incorporate amino acid modifications or modified peptides predicted to be immunogenic by the methods of the present invention. Such cellular epitopes are also referred to herein as "neo-epitopes". These epitopes are presented by the patient's cells, particularly antigen-presenting cells, and when bound to MHC, preferably result in T cells that target the epitopes, so that the patient's tumor, preferably a primary tumor as well as tumor metastasis, expresses the antigen from which the T cell epitope is derived and presents the same epitope on the surface of the tumor cell.

[0211] The methods of the invention may include further steps of determining the usability of the identified amino acid modifications or modified peptides in cancer vaccines. Thus, the further steps may include one or more of the following: (i) assessing whether the modification is located within a known or predicted MHC-presented epitope; (ii) in vitro and / or in silico testing whether the modification is located within an MHC-presented epitope, for example, testing whether the modification is part of a peptide sequence that is processed and / or presented as an MHC-presented epitope; and (iii) in vitro testing whether the proposed modified epitope, particularly when present in its native sequence, for example, when flanked by amino acid sequences of an epitope that is also flanked in a naturally occurring protein, and when expressed in an antigen-presenting cell, is able to stimulate T cells, such as T cells of a patient, having the desired specificity. Each such flanking sequence may comprise 3 or more, 5 or more, 10 or more, 15 or more, 20 or more, and preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids, and may be flanked at the N-terminus and / or C-terminus by an epitope sequence.

[0212] The modified peptides identified according to the present invention can be ranked according to their suitability as cancer vaccine epitopes. Thus, in one aspect, the methods of the present invention include a manual or computer-based analysis process in which the identified modified peptides are analyzed and selected for suitability in respective provided vaccines. In a preferred embodiment, the analysis process is based on a computational algorithm. Preferably, the analysis process includes identifying and / or ranking epitopes based on predictions of their immunogenic properties.

[0213] The neo-epitopes identified according to the present invention and provided by the vaccines of the present invention are preferably present in the form of polypeptides comprising said neo-epitopes, such as polyepitope polypeptides or nucleic acids encoding said polypeptides, in particular RNA. In addition, the neo-epitopes may be present in the form of vaccine sequences in the polypeptide, i.e., in its natural sequence environment, for example flanked by amino acid sequences that flank said epitopes in naturally occurring proteins. Each such flanking sequence may comprise 5 or more, 10 or more, 15 or more, 20 or more, and preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids, and may be flanked by epitope sequences at the N-terminus and / or C-terminus. Thus, the vaccine sequence may comprise 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, and preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids. In one embodiment, the neo-epitope and / or vaccine sequences are arranged head-to-tail in the polypeptide.

[0214] In one embodiment, the new epitopes and / or vaccine sequences are separated by a linker, in particular a neutral linker. According to the present invention, the term "linker" relates to a peptide added between two peptide domains, such as an epitope or vaccine sequence, for connecting the peptide domains. There are no particular restrictions on the linker sequence. However, it is preferred that the linker sequence reduces the steric hindrance between the two peptide domains, is well translated, and supports or allows the processing of the epitope. In addition, the linker should not contain immunogenic sequence elements or only contain sequence elements with lower immunogenicity. The linker should preferably not produce non-endogenous new epitopes that may produce an undesirable immune response (such as those produced by the connection stitching between adjacent new epitopes). Therefore, the multi-epitope vaccine should preferably contain a linker sequence that can reduce the number of undesirable NHC-bound junction epitopes. Hoyt et al. (EMBO J. 25 (8), 1720-9, 2006) and Zhang et al. (J. Biol. Chem., 279 (10), 8635-41, 2004) have shown that glycine-rich sequences impair proteasomal processing, and therefore the use of glycine-rich linker sequences minimizes the number of linker-containing peptides that can be processed by the proteasome. In addition, glycine has been observed to inhibit strong binding within the MHC binding groove site (Abastado et al., J. Immunol. 151 (7), 3569-75, 1993). Schlessinger et al. (Proteins, 61 (1), 115-26, 2005) have found that the inclusion of glycine and serine in the amino acid sequence results in more efficient translation and more flexible proteins processed by the proteasome, allowing better access to encoded neo-epitopes. Each linker may comprise 3 or more, 6 or more, 9 or more, 10 or more, 15 or more, 20 or more, and preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids. Preferably, the linker is rich in glycine and / or serine. Preferably, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or at least 95% of the amino acids in the linker are glycine and / or serine. In a preferred embodiment, the linker consists essentially of glycine and serine. In one embodiment, the linker comprises the amino acid sequence (GGS) a (GSS) b (GGG) c (SSG) d (GSG) e, wherein a, b, c, d and e are numbers independently selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, and wherein a+b+c+d+e is not 0 and is preferably 2 or more, 3 or more, 4 or more, or 5 or more. In one embodiment, the linker comprises a sequence as described herein including a linker sequence as described in the examples, such as the sequence GGSGGGGSG.

[0215] In a particularly preferred embodiment, according to the present invention, a polypeptide incorporating one or more neo-epitopes, such as a multi-epitope polypeptide, is administered to a patient in the form of a nucleic acid, preferably RNA, such as in vitro transcribed or synthesized RNA, which can be expressed in a patient's cells, such as antigen-presenting cells, to produce the polypeptide. The present invention also envisions the administration of one or more multi-epitope polypeptides, which for the purposes of the present invention encompass the term "multi-epitope polypeptide," preferably administered in the form of a nucleic acid, preferably RNA, such as in vitro transcribed or synthesized RNA, which can be expressed in a patient's cells, such as antigen-presenting cells, to produce one or more polypeptides. In the case of administration of multiple multi-epitope polypeptides, the neo-epitopes provided by the different multi-epitope polypeptides may be different or partially overlapping. Once present in a patient's cells, such as antigen-presenting cells, the polypeptides are processed according to the present invention to produce the neo-epitopes identified according to the present invention. Administration of a vaccine provided according to the present invention can provide an MHC class II presented epitope capable of eliciting a CD4+ helper T cell response against a cell expressing the antigen from which the MHC presented epitope is derived. Alternatively or in addition, administration of a vaccine provided according to the present invention can provide an MHC class I presented epitope capable of eliciting a CD8+ T cell response against a cell expressing the antigen from which the MHC presented epitope is derived. In addition, the administration of the vaccine provided according to the present invention can provide one or more new epitopes (including known new epitopes and new epitopes identified according to the present invention) and one or more epitopes that do not contain cancer-specific somatic mutations but are expressed by cancer cells and preferably induce an immune response (preferably a cancer-specific immune response) against cancer cells. In one embodiment, the administration of the vaccine provided according to the present invention provides a new epitope, which is an epitope presented by MHC class II and / or can cause a CD4+ helper T cell response to a cell expressing a source antigen that presents an MHC epitope, and provides an epitope that does not contain a cancer-specific somatic mutation, which is an epitope presented by MHC class I and / or can cause a CD8+ T cell response to a cell expressing a source antigen that presents an MHC epitope. In one embodiment, the epitope that does not contain a cancer-specific somatic mutation is derived from a tumor antigen. In one embodiment, the new epitope and the epitope that does not contain a cancer-specific somatic mutation have a synergistic effect in the treatment of cancer. Preferably, the vaccine provided according to the present invention is suitable for multi-epitope stimulation of cytotoxic T cells and / or helper T cell responses.

[0216] The vaccine provided according to the present invention may be a recombinant vaccine.

[0217] In the context of the present invention, the term "recombinant" means "produced by genetic engineering". Preferably, in the context of the present invention, a "recombinant entity", such as a recombinant polypeptide, is not naturally occurring and is preferably the result of combining entities, such as amino acids or nucleic acids, that are not combined in nature. For example, in the context of the present invention, a recombinant polypeptide may comprise several amino acid sequences, such as neo-epitopes or vaccine sequences derived from different proteins or different parts of the same protein fused together (e.g., via peptide bonds or suitable linkers).

[0218] As used herein, the term "naturally occurring" refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that is present in an organism (including viruses), can be isolated from a natural source, and has not been intentionally modified by an experimenter is naturally occurring.

[0219] The agents, compositions, and methods described herein can be used to treat subjects suffering from diseases, such as diseases characterized by the presence of diseased cells expressing an antigen and the presentation of antigenic fragments. A particularly preferred disease is cancer. The agents, compositions, and methods described herein can also be used for immunization or vaccination to prevent the diseases described herein.

[0220] According to the present invention, the term "disease" refers to any pathological state, including cancer, in particular those forms of cancer described herein.

[0221] The term "normal" refers to the healthy state or condition of a healthy subject or tissue, ie, a non-pathological condition, wherein "healthy" preferably refers to non-cancerous.

[0222] According to the present invention, "disease involving cells expressing the antigen" means that expression of the antigen has been detected in the cells of the pathological tissue or organ. Expression in the cells of the pathological tissue or organ may be increased compared to the state in healthy tissue or organ. Increase refers to an increase of at least 10%, in particular at least 20%, at least 50%, at least 100%, at least 200%, at least 500%, at least 1000%, at least 10000% or even more. In one embodiment, expression is only found in pathological tissue, while expression in healthy tissue is suppressed. According to the present invention, diseases comprising cells expressing the antigen or associated with cells expressing the antigen include cancer.

[0223] According to the present invention, the term "tumor" or "neoplastic disease" refers to the abnormal growth of cells (called neoplastic cells, tumor-producing cells or tumor cells), preferably forming a swelling or lesion. "Tumor cell" means an abnormal cell that grows by rapid, uncontrolled cell proliferation and continues to grow after the stimulus that initiated new growth stops. Tumors show partial or complete loss of the structural organization and functional coordination of normal tissue and usually form isolated tissue masses, which can be benign, pre-malignant or malignant.

[0224] Cancer (medical term: malignant tumor) is a group of cells that exhibit uncontrolled growth (dividing beyond normal limits), invasion (invasion and destruction of neighboring tissues), and sometimes metastasis (spreading to other parts of the body via the lymph or blood). These three malignant characteristics of cancer distinguish them from benign tumors, which are self-limited and do not invade or metastasize. Most cancers form tumors, but some, such as leukemias, do not. Malignancy, malignant neoplasm, and malignant tumor are essentially synonymous with cancer.

[0225] A neoplasm is an abnormal mass of tissue resulting from neoplasia. Neoplasia (Greek for new growth) is an abnormal proliferation of cells. The cells grow beyond and out of proportion to the surrounding normal tissue. Even after the stimulus ceases, the growth continues in the same excessive manner. It typically results in a lump or tumor. Tumors can be benign, pre-malignant, or malignant.

[0226] "Growth of a tumor" or "tumor growth" according to the present invention relates to the tendency of a tumor to increase its size and / or the tendency of tumor cells to proliferate.

[0227] For the purposes of the present invention, the terms "cancer" and "cancerous disease" are used interchangeably with the terms "tumor" and "tumor disease."

[0228] Cancers are classified by the type of cells that resemble the tumor and the tissue from which the tumor is presumed to have originated. These are histology and site.

[0229] According to the present invention, term " cancer " comprises carcinomas, adenocarcinoma, blastoma, leukemia, seminoma, melanoma, teratoma, lymphoma, neuroblastoma, glioma, rectal cancer, endometrial cancer, kidney cancer, adrenal cancer, thyroid cancer, blood cancer, skin cancer, brain cancer, cervical cancer, intestinal cancer, liver cancer, colon cancer, stomach cancer, intestinal cancer, head and neck cancer, gastrointestinal cancer, lymph node cancer (lymph node cancer), esophageal cancer, colorectal cancer, pancreatic cancer, ear, nose and throat (ENT) cancer, breast cancer, prostate cancer, uterine cancer, ovarian cancer and lung cancer and their metastasis.Their example is the metastasis of lung cancer, breast cancer, prostate cancer, colon cancer, renal cell carcinoma, cervical cancer or above-mentioned cancer type or tumor. According to the present invention, term cancer also comprises cancer metastasis and cancer recurrence.

[0230] " transfer " means the spread of cancer cells from its original position to other parts of the body. The formation of transfer is a very complicated process, and depends on the separation of malignant cells from primary tumors, the invasion of extracellular matrix, infiltration of endothelial basement membrane into body cavity and blood vessel, then after being transported by blood, the infiltration of target organs. Finally, the growth of new tumors, i.e., secondary tumors or metastatic tumors, depends on angiogenesis at the target site. Tumor metastasis often occurs, even occurs after removing the primary tumor, because tumor cells or components can remain and develop metastatic potential. In one embodiment, according to the present invention, term " transfer " relates to " distant metastasis ", which relates to the metastasis away from primary tumor and regional lymph node system.

[0231] The cells of a secondary or metastatic tumor are the same as those in the original tumor. This means, for example, that if ovarian cancer metastasizes to the liver, the secondary tumor is made up of abnormal ovarian cells, not abnormal liver cells. The tumor in the liver is then called metastatic ovarian cancer, not liver cancer.

[0232] The term "circulating tumor cell" or "CTC" refers to cells that are isolated from a primary tumor or tumor metastasis and circulate in the blood. CTCs can serve as seed cells for the subsequent growth of other tumors (metastases) in various tissues. Circulating tumor cells are found at a frequency of approximately 1-10 CTCs per mL of whole blood in patients with metastatic disease. Research methods for isolating CTCs have been developed. Several research methods have been described in the art to isolate CTCs, such as techniques that exploit the fact that epithelial cells often express the cell adhesion protein EpCAM, which is absent in normal blood cells. Immunomagnetic bead-based capture involves treating a blood sample with an EpCAM antibody linked to magnetic particles, followed by separation of the labeled cells in a magnetic field. The isolated cells are then stained with antibodies to cytokeratin, another epithelial marker, and CD45, a common leukocyte marker, to distinguish small numbers of CTCs from contaminating white blood cells. This robust and semi-automated method identifies CTCs with an average yield of approximately 1 CTC / mL and a purity of 0.1% (Allard et al., 2004: Clin Cancer Res 10, 6897-6904). A second method for isolating CTCs utilizes a microfluidics-based CTC capture device that involves passing whole blood through a chamber embedded with 80,000 microspheres that have been functionalized by coating the microspheres with EpCAM antibodies. The CTCs are then stained with a secondary antibody against cytokeratin or a tissue-specific marker (such as PSA in prostate cancer or HER2 in breast cancer) and visualized by automatically scanning the microspheres along three-dimensional coordinates in multiple planes. The CTC-chip is able to identify cytokeratin-positive circulating tumor cells in patients with an average yield of 50 cells / ml and a purity range of 1–80% (Nagrath et al., 2007: Nature 450, 1235-1239). Another possibility for isolating CTCs is to utilize the CellSearch™ from Veridex, LLC (Raritan, NJ). TM Circulating tumor cell (CTC) detection, which captures, identifies and counts CTCs in blood in a tube. CellSearch TM The TRANSPORT® system is a method for enumerating CTCs in whole blood approved by the U.S. Food and Drug Administration (FDA) and is based on a combination of immunomagnetic labeling and automated digital microscopy. Other methods for isolating CTCs are described in the literature, all of which can be used in conjunction with the present invention.

[0233] Relapse or recurrence occurs when a person is again affected by a situation that affected them in the past. For example, if a patient has suffered from a neoplastic disease, received successful treatment for the disease, and suffers from the disease again, the newly suffered disease can be considered a relapse or recurrence. However, according to the present invention, the relapse or recurrence of a neoplastic disease may, but need not, occur at the original neoplastic disease site. Thus, for example, if a patient has suffered from an ovarian tumor and has received successful treatment, a relapse or recurrence may be the occurrence of an ovarian tumor or a tumor at a site other than the ovary. The relapse or recurrence of a tumor also includes the occurrence of a tumor at a site other than the original tumor site and at the original tumor site. Preferably, the original tumor for which the patient has been treated is a primary tumor, and the tumor at a site other than the original tumor site is a secondary tumor or a metastatic tumor.

[0234] By "treating" is meant administering a compound or composition as described herein to a subject to prevent or eliminate a disease, including reducing the size or number of tumors in a subject; preventing or slowing the disease in a subject; inhibiting or slowing the development of a new disease in a subject; reducing the frequency and severity of symptoms and / or recurrences in a subject currently suffering from the disease or having previously suffered from the disease; and / or prolonging, i.e., increasing, the lifespan of a subject. In particular, the term "treatment of a disease" includes curing, shortening the duration, ameliorating, preventing, slowing or inhibiting the progression or worsening, or preventing or delaying the onset of a disease or its symptoms.

[0235] By "at risk of..." is meant a subject, i.e., a patient, identified as having a higher than normal likelihood of developing a disease, particularly cancer, compared to the general population. Furthermore, a subject who has previously had or currently has a disease, particularly cancer, is considered to have an increased likelihood of developing the disease because such a subject is likely to continue to have the disease. A subject who currently has or has previously had cancer also has an increased likelihood of the cancer metastasizing.

[0236] The term "immunotherapy" refers to a treatment that involves activation of a specific immune response. In the context of the present invention, terms such as "protection," "prevention," "prophylactic," "preventative," or "protective" refer to the prevention or treatment of a disease in a subject, or both the occurrence and / or spread of the disease, in particular, to minimize the likelihood of the patient developing the disease or to slow the progression of the disease. For example, a person at risk of developing cancer, as described above, is a candidate for treatment to prevent cancer.

[0237] Prophylactic administration of immunotherapy, such as the vaccines of the present invention, preferably protects the recipient from developing the disease. Therapeutic administration of immunotherapy, such as the vaccines of the present invention, can result in inhibition of disease progression / growth. This includes slowing the progression / growth of the disease, and particularly disrupting the progression of the disease, which preferably results in clearance of the disease.

[0238] Immunotherapy can be performed using any of a variety of techniques, wherein the function of the reagents provided herein is to remove diseased cells from the patient. Such removal can occur due to enhancing or inducing a patient's immune response specific for the antigen or cells expressing the antigen.

[0239] In certain embodiments, immunotherapy may be active immunotherapy, wherein treatment relies on in vivo stimulation of the endogenous host immune system by administration of immune response-modifying agents (such as the polypeptides and nucleic acids described herein) to act against diseased cells.

[0240] The agents and compositions provided herein can be used alone or in combination with conventional treatment regimens such as surgery, radiation, chemotherapy, and / or bone marrow transplantation (autologous, syngeneic, allogeneic, or unrelated).

[0241] The terms "immunization" or "vaccination" describe the process of treating a subject with the goal of inducing an immune response for therapeutic or prophylactic reasons.

[0242] The term "in vivo" relates to a location inside a subject.

[0243] The terms "subject," "individual," "organism," or "patient" are used interchangeably and refer to vertebrates, preferably mammals. For example, in the context of the present invention, mammals are humans, non-human primates, domestic animals such as dogs, cats, sheep, cows, goats, pigs, horses, etc., experimental animals such as mice, rats, rabbits, guinea pigs, etc., and captive animals such as animals in zoos. As used herein, the term "animal" also includes humans. The term "subject" may also include patients, i.e., animals, preferably humans, suffering from a disease, preferably as described herein.

[0244] The term "autologous" is used to describe anything that originates from the same subject. For example, an "autologous transplant" refers to a transplant of tissue or organs that originates from the same subject. Such procedures are advantageous because they overcome immune barriers that would otherwise lead to rejection.

[0245] The term "heterologous" is used to describe something that is composed of multiple different elements. As an example, transplanting bone marrow from one individual into a different individual constitutes a heterologous transplant. Heterologous genes are genes that originate from a source other than the subject.

[0246] As part of a composition for immunization or vaccination, one or more agents described herein are preferably co-administered with one or more adjuvants to induce an immune response or increase an immune response. The term "adjuvant" refers to a compound that prolongs or enhances or promotes an immune response. The effects of the compositions of the present invention are preferably exerted without the addition of an adjuvant. However, the compositions of the present invention may contain any known adjuvant. Adjuvants include a heterogeneous group of compounds such as oil emulsions (e.g., Freund's adjuvant), inorganic compounds (e.g., alum), bacterial products (e.g., Bordetella pertussis toxin), liposomes, and immunostimulatory complexes. An example of an adjuvant is monophosphoryl lipid A (MPL SmithKline Beecham). Saponins such as QS21 (SmithKline Beecham), DQS21 (SmithKline Beecham; WO 96 / 33739), QS7, QS17, QS18 and QS-L1 (So et al., 1997, Mol. Cells 7:178-186), Freund's incomplete adjuvant, Freund's complete adjuvant, vitamin E, montanid, alum, CpG oligonucleotides (Krieg et al., 1995, Nature 374:546-549), and various water-in-oil emulsions prepared from biodegradable oils such as squalene and / or tocopherol.

[0247] Other substances that stimulate the patient's immune response can also be administered. It is possible, for example, to use cytokines in vaccination due to their regulatory properties on lymphocytes. Such cytokines include, for example, interleukin-12 (IL-12), GM-CSF, and IL-18, with interleukin-12 being shown to increase the protective effect of vaccines (cf. Science 268:1432-1434, 1995).

[0248] There are many compounds that enhance the immune response and therefore can be used in vaccination. These include co-stimulatory molecules such as B7-1 and B7-2 (CD80 and CD86, respectively) provided in protein or nucleic acid form.

[0249] According to the present invention, a body sample can be a tissue sample, including a body fluid and / or a cell sample. Such body samples can be obtained by conventional means, such as by biopsy, including punch biopsy, and by obtaining blood, bronchial aspirate, sputum, urine, feces, or other body fluids. According to the present invention, the term "sample" also includes processed samples such as fragments or isolates of biological samples, such as nucleic acid or cell isolates.

[0250] Reagents such as the vaccines and compositions described herein can be administered via any conventional route, including by injection or infusion. Administration can be, for example, orally, intravenously, intraperitoneally, intramuscularly, subcutaneously, or transdermally. In one embodiment, administration can be intranodal, such as by injection into a lymph node. Other forms of administration contemplate in vitro transfection of antigen-presenting cells, such as dendritic cells, with the nucleic acids described herein, followed by administration of the antigen-presenting cells.

[0251] The agents described herein are administered in an effective amount. An "effective amount" refers to an amount that, alone or in combination with other doses, achieves a desired response or desired effect. In the context of treatment of a specific disease or condition, the desired response preferably involves inhibition of the disease process. This includes slowing the progression of the disease, and in particular, blocking or reversing the progression of the disease. A desired response in the treatment of a disease or condition may also be delaying or preventing the onset of the disease or condition.

[0252] The effective amount of the agents described herein depends on the condition to be treated, the severity of the disease, the individual parameters of the patient including age, physical condition, size and weight, the duration of treatment, the type of concomitant treatment (if any), the specific route of administration, and similar factors. Therefore, the dosage of the agents described herein may depend on different such parameters. In cases where the initial dose does not induce an adequate response in the patient, a higher dose (or an effectively higher dose achieved by a different, more localized route of administration) may be used.

[0253] The pharmaceutical compositions described herein are preferably sterile and contain an effective amount of the therapeutically active substance to produce the desired response or desired effect.

[0254] Pharmaceutical compositions as described herein are generally administered in pharmaceutically compatible amounts and pharmaceutically compatible formulations. The term "pharmaceutically compatible" refers to a non-toxic substance that does not interact with the functions of the active ingredients in the pharmaceutical composition. Such formulations may generally contain salts, buffer substances, preservatives, carriers, supplementary immune enhancing substances such as adjuvants (e.g., CpG oligonucleotides, cytokines, chemokines, saponins, GM-CSF and / or RNA) and suitable other therapeutically active compounds. When used in medicine, the salt should be pharmaceutically compatible. However, non-pharmaceutically compatible salts can be used to prepare pharmaceutically compatible salts, and are included in the present invention. Such pharmacologically and pharmaceutically compatible salts include, in a non-limiting manner, those prepared from the following acids: hydrochloric acid, hydrobromic acid, sulfuric acid, phosphoric acid, maleic acid, acetic acid, salicylic acid, citric acid, formic acid, malonic acid, succinic acid, etc. Pharmaceutically compatible salts can also be prepared as alkaline metal or alkaline earth metal salts, such as sodium, potassium or calcium salts.

[0255] The pharmaceutical compositions described herein may include a pharmaceutically compatible carrier. The term "carrier" refers to an organic or inorganic component, natural or synthetic, with which the active ingredients are combined to facilitate application. According to the present invention, the term "pharmaceutically compatible carrier" includes one or more compatible solid or liquid fillers, diluents, or encapsulating substances suitable for administration to a patient. The components of the pharmaceutical compositions of the present invention generally do not interact with each other in a manner that would significantly impair the desired efficacy of the drug.

[0256] The pharmaceutical compositions described herein may contain suitable buffer substances such as acetate, citrate, borate, and phosphate.

[0257] The pharmaceutical compositions may also suitably contain suitable preservatives, such as benzalkonium chloride, chlorobutanol, parabens and thimerosal.

[0258] The pharmaceutical composition is usually provided in a uniform dosage form and can be prepared in a manner known per se. The pharmaceutical composition of the present invention can be in the form of, for example, a capsule, tablet, lozenge, solution, suspension, syrup, elixir or emulsion.

[0259] Compositions suitable for parenteral administration generally comprise a sterile aqueous or non-aqueous preparation of the active compound that is preferably isotonic with the blood of the recipient. Examples of compatible carriers and solvents are Ringer's solution and isotonic sodium chloride solution. In addition, sterile fixed oils are generally used as solution or suspension media.

[0260] The present invention is described in detail by the following figures and examples, which are for illustrative purposes only and not limiting. Due to the description and examples, other embodiments of the present invention are available to the skilled person.

[0261] This application also relates to the following embodiments:

[0262] Embodiment 1. A method for predicting immunogenic amino acid modifications, comprising the steps of:

[0263] a) determining a binding score of the modified peptide to one or more MHC molecules, and

[0264] b) determining the binding score of the non-modified peptide to one or more MHC molecules, and / or

[0265] c) determining the binding score of the modified peptide to one or more T cell receptors when present in an MHC-peptide complex.

[0266] Embodiment 2. The method of embodiment 1, wherein the modified peptide comprises a fragment of a modified protein, the fragment comprising the modification present in the protein.

[0267] Embodiment 3. The method of embodiment 1 or 2, wherein the non-modified peptide contains a germline amino acid at a position corresponding to the modification site in the modified peptide.

[0268] Embodiment 4. The method of any one of embodiments 1 to 3, wherein the non-modified peptide is identical to the modified peptide except for the modification.

[0269] Embodiment 5. The method of any one of embodiments 1 to 4, wherein the non-modified peptide and the modified peptide are 8 to 15 amino acids in length, preferably 8-12 amino acids in length.

[0270] Embodiment 6. The method of any one of embodiments 1 to 5, wherein the one or more MHC molecules comprise different MHC molecule types, in particular different MHC alleles.

[0271] Embodiment 7. The method of any one of embodiments 1 to 6, wherein the one or more MHC molecules are MHC class I molecules and / or MHC class II molecules.

[0272] Embodiment 8. The method of any one of embodiments 1 to 7, wherein the binding score to one or more MHC molecules is determined by a method comprising sequence comparison to an MHC binding motif database.

[0273] Embodiment 9. The method of any one of embodiments 1 to 8, wherein step a) comprises determining whether the score meets a predetermined threshold for binding to one or more MHC molecules.

[0274] Embodiment 10. The method of any one of embodiments 1 to 9, wherein step b) comprises determining whether the score meets a predetermined threshold for binding to one or more MHC molecules.

[0275] Embodiment 11. The method of any one of embodiments 1 to 10, wherein the threshold applied in step a) is different from the threshold applied in step b).

[0276] Embodiment 12. The method of any one of Embodiments 1 to 11, wherein the predetermined threshold value for binding to one or more MHC molecules reflects the likelihood of binding to one or more MHC molecules.

[0277] Embodiment 13. The method of any one of embodiments 1 to 12, wherein step c) comprises determining a chemical and physical similarity score between the non-modified amino acid and the modified amino acid.

[0278] Embodiment 14. The method of any one of embodiments 1 to 13, wherein step c) comprises determining whether the score meets a predetermined threshold of chemical and physical similarity between amino acids.

[0279] Embodiment 15. The method of any one of embodiments 1 to 14, wherein the chemical and physical similarity scores are determined based on the likelihood that amino acids are interchangeable in their natural state.

[0280] Embodiment 16. The method of any one of embodiments 1 to 15, wherein the higher the frequency of amino acid exchange in the natural state, the higher the amino acid similarity is considered to be.

[0281] Embodiment 17. The method of any one of Embodiments 1 to 16, wherein the chemical and physical similarities are determined using an evolution-based log-odds matrix.

[0282] Embodiment 18. The method of any one of embodiments 1 to 17, wherein the modification or modified peptide is predicted to be immunogenic if the binding score of the non-modified peptide to one or more MHC molecules meets a threshold value indicative of binding to one or more MHC molecules, and the binding score of the modified peptide to one or more MHC molecules meets a threshold value indicative of binding to one or more MHC molecules, provided that the chemical and physical similarity scores of the non-modified amino acid and the modified amino acid meet a threshold value indicative of chemical and physical dissimilarity.

[0283] Embodiment 19. The method of any one of embodiments 1 to 18, wherein the modification or modified peptide is predicted to be immunogenic if the non-modified peptide binds or has the potential to bind to one or more MHC molecules, and the modified peptide binds or has the potential to bind to one or more MHC molecules, with the proviso that the non-modified amino acid is chemically and physically dissimilar to the modified amino acid or has the potential to be chemically and physically dissimilar.

[0284] Embodiment 20. The method of embodiment 18 or 19, wherein the modification is not at an anchor site for binding to one or more MHC molecules.

[0285] Embodiment 21. The method of any one of embodiments 1 to 17, wherein the modification or modified peptide is predicted to be immunogenic if the binding score of the non-modified peptide to one or more MHC molecules meets a threshold value indicating no binding to one or more MHC molecules, and the binding score of the modified peptide to one or more MHC molecules meets a threshold value indicating binding to one or more MHC molecules.

[0286] Embodiment 22. The method of any one of embodiments 1 to 17 and 21, wherein the modification or modified peptide is predicted to be immunogenic if the non-modified peptide does not bind or has the potential to not bind to one or more MHC molecules, and the modified peptide binds or has the potential to bind to one or more MHC molecules.

[0287] Embodiment 23. The method of embodiment 21 or 22, wherein the modification is at an anchor site for binding to one or more MHC molecules.

[0288] Embodiment 24. The method of any one of Embodiments 1 to 23, comprising performing step a) on two or more differently modified peptides, wherein the two or more differently modified peptides comprise the same modification.

[0289] Embodiment 25. The method of embodiment 24, wherein the two or more differently modified peptides comprising the same modification comprise different fragments of a modified protein, said different fragments comprising the same modification present in said protein.

[0290] Embodiment 26. The method of embodiment 24 or 25, wherein the two or more differently modified peptides comprising the same modification comprise all possible MHC binding fragments of the modified protein comprising the same modification present in the protein.

[0291] Embodiment 27. The method of any one of embodiments 24 to 26, further comprising selecting the modified peptide from two or more differently modified peptides comprising the same modification having a likelihood or greatest likelihood of binding to one or more MHC molecules.

[0292] Embodiment 28. The method of any one of embodiments 24 to 27, wherein the two or more differently modified peptides comprising the same modification differ in length and / or site of modification.

[0293] Embodiment 29. The method of any one of embodiments 1 to 28, comprising performing step a), and optionally performing one or both of steps b) and c), on two or more differently modified peptides.

[0294] Embodiment 30. The method of embodiment 29, wherein the two or more differently modified peptides comprise the same modification and / or comprise different modifications.

[0295] Embodiment 31. The method of embodiment 30, wherein the different modifications are present in the same and / or different proteins.

[0296] Embodiment 32. The method of any one of embodiments 24 to 31, comprising comparing the scores of two or more of the differently modified peptides.

[0297] Embodiment 33. A method as described in embodiment 32, wherein the binding score of the modified peptide to one or more MHC molecules is weighted higher than the binding score of the modified peptide to one or more T cell receptors when present in an MHC-peptide complex, preferably higher than the chemical and physical similarity score between the non-modified amino acid and the modified amino acid, and the binding score of the modified peptide to one or more T cell receptors when present in an MHC-peptide complex, preferably the chemical and physical similarity score between the non-modified amino acid and the modified amino acid, is weighted higher than the binding score of the non-modified peptide to one or more MHC molecules.

[0298] Embodiment 34. The method of any one of embodiments 1 to 33, further comprising identifying non-synonymous mutations in one or more protein coding regions.

[0299] Embodiment 35. The method of any one of embodiments 1 to 34, wherein the modification is identified by partially or completely sequencing the genome or transcriptome of one or more cells, such as one or more cancer cells and optionally one or more non-cancerous cells, and identifying mutations in one or more protein-coding regions.

[0300] Embodiment 36. The method of embodiment 34 or 35, wherein the mutation is a somatic mutation.

[0301] Embodiment 37. The method of embodiments 34 to 36, wherein the mutation is a cancer mutation.

[0302] Embodiment 38. The method of any one of embodiments 1 to 37, for use in the preparation of a vaccine.

[0303] Embodiment 39. The method of embodiment 38, wherein the vaccine is derived from one or more modifications or one or more modified peptides predicted by the method to be immunogenic.

[0304] Embodiment 40. A method of providing a vaccine comprising the steps of:

[0305] One or more modifications or one or more modified peptides predicted to be immunogenic are identified by the method of any one of embodiments 1 to 37.

[0306] Embodiment 41. The method of embodiment 40, further comprising the steps of:

[0307] Vaccines comprising peptides or polypeptides containing the modifications predicted to be immunogenic or modified peptides, or nucleic acids encoding the peptides or polypeptides, are provided.

[0308] Embodiment 42. A vaccine produced according to the method of any one of Embodiments 38 to 41.

[0309] Attached photos

[0310] Figure 1 .A review of MHC binding prediction.

[0311] Figure 2 . as M of 50 optimized B16F10 mutations and 82 optimized CT26.WT mutations (total 132 mutations) mut Immunogenicity analysis of the function of the score, of which 30 were immunogenic. All vaccinations were performed with RNA. For B16F10, immunogenicity was analyzed by stimulating BMDC with RNA and measuring the immune response of splenocytes by ELISPOT and FACS. For CT26.WT, immunogenicity was analyzed by stimulating BMDC with RNA and peptide, respectively, and measuring the immune response of splenocytes by ELISPOT; if the peptide or RNA elicited an immune response, the mutation was considered immunogenic. The cumulative distribution of immunogenic mutations is M mut The graph shows the score below a given M mut The total number of mutations scored (red), the number of immunogenic mutations among these mutations (blue), and the percentage of immunogenic mutations in the total mutations (black). B The following ranges: ≤0.3, (0.3, 1], >1 for each M mut Histogram of the percentage of immunogenic mutations in the library. Errors shown are standard errors.

[0312] Figure 3 As M mut Immunogenicity analysis of B16F10 and CT26.WT as a function of M. Cumulative distribution of immunogenic mutations is shown in Figure 3. mut Scoring function. The following ranges for B16(B) and CT26(D): [0.1, 0.3], (0.3, 1], (1, ∞) for each M mut Histograms of the percentage of immunogenic mutations in the library. Panels A and B are based on analysis of 50 prioritized B16F10 mutations, 12 of which were immunogenic. Panels C and D are based on analysis of 82 prioritized B16F10 mutations, 30 of which were immunogenic. For more details, see Figure 2 Errors are standard errors.

[0313] Figure 4. Immunogenicity and control hypothesis model. A The class I immunogenicity shown assumes that both the WT and MUT epitopes are presented by cells and that the mutations alter the physicochemical properties of the amino acids sufficiently for the immune system to register the change and generate an immune response (shown as a lightning bolt). A Control H n The hypothesis is simply that H A The opposite of the hypothesis is that the mutation does not significantly change the physicochemical properties of the amino acid and is therefore less likely to be “detected” by the immune system and produce an immune response. B UH C ), the WT epitope is not presented, but the MUT epitope is presented. H B and H C . Note that for α * =α, immunogenic H BC1 Model (M mut <β) is a composite of all four groups: H BC1 =U[H A ,H B ,H C ,H n ].

[0314] Figure 5 .Hypothetical relationship between T score and immunogenicity. According to the Class I immunogenicity model, TCRs that bind strongly to wild-type epitopes are eliminated during T cell development. Existing TCRs should show only weak or no binding affinity to the wild-type epitope (A). Epitopes containing amino acid substitutions with high T scores have similar physicochemical properties to the wild-type amino acid and therefore may have little effect on binding to existing TCRs (B). Epitopes containing amino acid substitutions with T scores have a greater likelihood of increasing binding affinity to existing TCRs and therefore have a greater likelihood of becoming immunogenic (C). In this schematic, color coding is used to pair T cells with matching peptides. Orange / yellow mutations represent mutations with high T scores (similar to WT), while blue / purple mutations represent mutations with low T scores (significant physicochemical differences compared to WT).

[0315] Figure 6 As M mut Cumulative distribution of immunogenic mutations as a function of A satisfies the background control hypothesis H BC1 :{M mut The percentage of immunogenic mutations with ≤β} and the percentage of mutations satisfying the partial hypothesis H A’ :H BC1 ∩{T≤τ}, partial hypothesis H BC2 :HBC1 ∩{M mut ≤α} and the complete hypothesis H A :H BC1 ∩{M mut Comparison of the percentage of immunogenic mutations for {T≤τ}, α=1, τ=1. B satisfies the control hypothesis H BC1 :{M mut The percentage of immunogenic mutations with ≤β} satisfies the opposite hypothesis: H BC1 ∩{T>τ} and H BC1 ∩{M mut Comparison of the percentage of immunogenic mutations with a mutation count >α. The analyses in A and B are based on a combined B16F10 and CT26.WT dataset containing 132 mutations, 30 of which were immunogenic. Each data point in the graph is based on ≥4 mutations.

[0316] Figure 7 As M mut Cumulative distribution of immunogenic mutations as a function of the score. Satisfying the background control hypothesis H BC1 :{M mut The percentage of immunogenic mutations with ≤β} and the percentage of mutations satisfying the partial hypothesis H A’ :H BC1 ∩{T≤τ}, partial hypothesis H BC2 :H BC1 ∩{M mut ≤α} and the complete hypothesis H A :H BC1 ∩{M mut Comparison of the percentage of immunogenic mutations for B16 (A) and CT26 (C) with α = 1 and τ = 1. In B16 (B) and CT26 (D), the background control hypothesis H is satisfied. BC1 :{M mut The percentage of immunogenic mutations with ≤β} satisfies the opposite hypothesis: H BC1 ∩{T>τ} and H BC1 ∩{M mut Comparison of the percentage of immunogenic mutations of α>α}.

[0317] Panels A and B are based on analysis of 50 prioritized B16F10 mutations, 12 of which were immunogenic. Panels C and D are based on analysis of 82 prioritized B16F10 mutations, 30 of which were immunogenic. Each data point in the graph is based on ≥4 mutations.

[0318] Figure 8WT immunogenicity control. To test whether omitting the MUT+ / WT+ regimen had an impact on these results, we excluded from the dataset 9 MUT+ / WT+ mutations and 2 mutations for which WT measurements were not performed, leaving a total of 121 mutations (43 B16 and 78 CT26), of which 19 were MUT+ / WT- (5 B16 and 14 CT26). We again found the same trend as in the full dataset, i.e., as M mut Rating, H A The advantage of the hypothesis over the partial hypothesis and the opposite hypothesis and background control H BC1 Compared to the disadvantage of the highly nonlinear response of the function. A as M mut Cumulative distribution of immunogenicity as a function of score. B Each M mut Histogram of the percentage of immunogenic mutations in the library. C satisfies the background control hypothesis H BC1 The percentage of immunogenic mutations is consistent with the H A’ 、H BC2 and H A Comparison of the percentage of immunogenic mutations. D satisfies the background control hypothesis H BC1 Comparison with the percentage of immunogenic mutations that satisfy the opposite hypothesis. For further details see Figure 5 Description.

[0319] Figure 9 . The proportion of immunogenic mutations as a function of RPKM. Red: all 50 B16 mutations and 82 CT26 mutations (132 mutations in total) without filtering. Blue: H with α=1, β=0.5, τ=1 A Hypothesized mutations. B. Percentage of immunogenic mutations at different RPKM ranges without filtering. RPKM library is: 1 = (0, 1], 2 = (1, 5], 3 = (5, 50], 4 = (50, ∞). C. H with α = 1, β = 0.5, τ = 1 A The percentage of immunogenic mutations at different RPKM ranges for the hypothesized mutations. The RPKM library is: 1 = (0, 1], 2 = (1, ∞). Errors are SE.

[0320] Figure 10 Class II immunogenic epitopes of anchor and non-anchor site mutations. Anchor site motifs were analyzed using SYFPEITHI.

[0321] Figure 11 .Proposed model for immunogenic tumor-associated epitopes.

[0322] Figure 12 Example of a method for weighting mutation rankings. For each mutation, its ranking in the mutation ranking list can be further weighted by the patient's HLA type, the possible window length of the HLA type, and the probability of causing low Mmut The solution or the one that leads to H A and / or H B UH C The weighting factor is weighted by the number of combined solutions for the mutation sites within the classified epitope. Since all solutions for each mutation can potentially be proposed in parallel, this weighting factor can be an important contributor to mutation ranking.

[0323] Figure 13 .Targeting M mut and ΔM=M mut –M wt Example of a scatter plot of all epitope schemes for the mutation chr14_52837882 from CT26 Example

[0324] The techniques and methods used herein are described herein or in a manner known per se and as described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2 nd All methods, including kits and reagents, were performed according to the manufacturer's instructions unless otherwise specified.

[0325] Example 1: Establishment of a model for predicting the immunogenicity of T cell epitopes

[0326] Previously, we studied the immunogenicity of 50 somatic mutations identified in the B16F10 mouse melanoma cell line (JCCastle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012)). These 50 mutations were selected from a pool of 563 non-synonymous somatic mutations that maximized the expression of MHC class I (JCCastle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012)) (also see Example 2). When searching all possible MHC class I allele spaces, potential epitope lengths, and sequence windows (positioning mutation sites), we predicted the minimum epitope for each mutation, i.e., the epitope with the lowest MHC class I consistency score (Y.Kim et al., Nucleic Acids Research 40, W525 (2012)) (defined herein as M mut)(JCCastle et al.,Exploiting themutanome for tumor vaccination.Cancer Research 72,1081(2012)). The immunogenicity of these mutations was measured using RNA vaccination, and peptide vaccination was then used to confirm the previously discovered peptide reads (see Example 2)(JCCastle et al.,Exploiting the mutanome for tumor vaccination.Cancer Research 72,1081(2012)), showing that only 12 of the 50 mutations (24%) were immunogenic (Table 1), with the MUT+ / WT- sequence containing only 10% of all tested mutations.

[0327] Table 1. Number of immunogenic mutations following RNA vaccination of B16F10 and CT26.WT mouse strains

[0328]

[0329] *This table excludes two CT26 MUT+ mutations because their WT reactivity has not been measured. A total of 18 MUT+ mutations out of 82 CT26 mutations have been measured to date, yielding a 22% success rate.

[0330] Results from the B16F10 mouse test case showed that the initial selection had low M mut Non-synonymous mutations expressed with a score (≤3.9) yielded a lower success rate in predicting immunogenicity. Therefore, if personalized vaccines targeting tumor-specific neoantigens are to become effective therapies, a better understanding of the mechanisms driving immunogenicity is needed. In an attempt to discover additional variables that contribute to immunogenicity, we investigated the immunogenicity of non-synonymous somatic mutations expressed identified in the mouse colon cancer cell line CT26.WT. In total, the M-based mut 96 mutations were selected based on score (low vs. high), average RPKM (low vs. high), and cellular localization (intracellular vs. extracellular) and tested for immunogenicity using RNA vaccination with peptide and RNA readouts (see Example 2 for further details). Together with the B16F10 cell line, our dataset includes 132 epitopes whose immunogenicity was measured ex vivo on mouse splenocytes.

[0331] MHC identity score. In order to study the immunogenicity of M mut Based on the dependence of , we define the cumulative percentage of immunogenic mutations as M mut function, namely M mutThe percentage of mutations that scored less than a given threshold for immunogenicity (denoted as β). Analysis of the combined B16 and CT26 datasets across a total of 132 mutations showed that the immunogenicity success rate was significantly affected by the M mut The highly nonlinear dependence of Figure 2 A). Figure 2 A shows that immunogenic mutations are enriched in very low M mut Score (≤~0.2). mut ≤0.1, the percentage of immunogenic mutations peaked at ∼60% and decreased with M mut The increase in mut ≥2~25% or less. M mut The percentage of immunogenic mutations with a mutation ratio of ≤0.3 versus >0.3 was 44.4% versus 17.1%, a statistically significant difference (P value = 0.004, Fisher's exact test, one-tailed). mut The histogram of the percentage of immunogenic mutations ≤ 0.3, (0.3, 1] and > 1 shows the percentage of immunogenic mutations with M mut decrease with the increase of Figure 2 B). Figure 2 The lowest library in B (M mut ≤0.3) had a success rate of 44.4%, compared with 20.7% for the middle library and 20.7% for the highest library (M mut >1) was statistically significant (P values ​​= 0.05 and 0.004, respectively, Fisher's exact test, one-tailed), indicating that for M mut When the mutation groups of B16 and CT26 were analyzed separately, the success rate was also observed ( Figure 3 ).

[0332] To date, our criteria for selecting mutations have focused on presentation, and we have observed that limiting the MHC binding score to mutated epitopes allows for prediction of immunogenic epitopes with an accuracy of up to 60%. However, presentation is a necessary but not sufficient condition for inducing immunogenicity. By identifying additional criteria for TCR recognition, we may be able to further improve the accuracy of our predictions. We hypothesize two mutually exclusive mechanisms driving immunogenicity, which we refer to as type I and type II immunogenicity models.

[0333] Class I immunogenicity. In order for the TCR repertoire to recognize the mutated epitope and generate an immune response, we hypothesize that three conditions must be met (H A Figure 4): (i) the wild-type epitope is presented to the immune system at some point during the organism's development, and the matching TCR is eliminated by strong TCR / pMHC binding; (ii) the mutated epitope is presented; and (iii) the physicochemical properties of the mutated amino acid are sufficiently "different" from the wild-type amino acid (by some metric we define below) to enable the TCR repertoire to "detect" or "record" the substitution. Conditions (i) and (ii) ensure that the immune system is actually exposed to the change, i.e., the mutation. Condition (iii) requires that the mutation significantly alter the physicochemical properties of the wild-type amino acid so that the binding affinity of the mutated epitope to the existing (uncleaved) TCR is likely to increase, thereby initiating the signaling cascade that leads to an immune response ( Figure 5 ).

[0334] TCR recognition score. Class I immunogenicity models require a metric to estimate the physicochemical differences between two amino acids. It is known in molecular evolution that amino acids that are frequently interchanged may have chemical and physical similarities, while amino acids that are rarely interchanged may have different physicochemical properties. The log-probability matrix measures the likelihood of a given substitution occurring naturally compared to the likelihood of the substitution occurring by chance. The patterns observed in the log-probability matrix imposed by natural selection "reflect the similarity of amino acid residue functions in the weak interactions between amino acid residues and another amino acid residue in the three-dimensional conformation of the protein" (MO Dayhoff, RM Schwartz, BC Orcutt, A model for evolutionary change. MO Dayhoff, ed. Atlas of protein sequence and structure Vol. 5, 345 (1978)). Therefore, we use the evolution-based log-probability matrix, which we call "T score" to reflect TCR recognition, as an effective scoring matrix for cancer-associated amino acid substitutions. Substitutions with positive T scores (i.e., log-probability) are likely to occur naturally and therefore correspond to two amino acids with similar physicochemical properties. The Class I model predicts that substitutions with positive T scores are less likely to be immunogenic. Conversely, substitutions with negative T scores reflect that the substitution is likely to occur naturally and therefore correspond to two amino acids with significantly different physicochemical properties. According to our model, such substitutions are more likely to be immunogenic. We compared different methods for estimating the log-odds matrix and found that the results were very stable relative to the exact method selected. The maximum likelihood (ML) estimation method called WAG (S. Whelan, N. Goldman, Molecular biology and evolution 18, 691 (2001)) using a PAM (acceptable point mutation) distance of 250 appears to best separate predicted immunogenic from non-immunogenic mutations, so we used this matrix to represent the results (see Example 2 for further details).

[0335] Class II Immunogenicity. In the Class II model of immunogenicity, we hypothesize that a mutation may be immunogenic if the immune system has never seen the wild-type epitope before and is therefore stimulated by the mutated epitope. Therefore, in this model, in order for a mutation to be immunogenic, we hypothesize that two conditions must be met: (i) the wild-type epitope has never been presented to the immune system, and (ii) the mutated peptide is presented. These conditions can be met together if, for example, the mutation is located at an anchor site, thereby turning a "non-binding" epitope into a "binding epitope." Formally, Class II immunogenicity can be divided into two sub-hypotheses: high T score ( Figure 4 H in B) and low T score ( Figure 4 H in C However, since the wild-type epitope is assumed to be unpresented, the nature of the amino acid substitutions is not expected to affect TCR recognition, so we compared the class II immunogenicity with the consensus hypothesis: H B UH C are considered equivalent.

[0336] Test for Class I immunogenicity. Class I immunogenicity ( Figure 4 H in A ) Let's restate the assumption mathematically: we need the wild-type epitope to be presented (M wt ≤α), the mutated epitope is presented (M mut ≤β), and amino acid substitution is insignificant (T≤τ), where M wt is defined as the MHC identity score (same HLA allele and window length) of the mutant epitope with the wild-type amino acid replacing the mutant amino acid, and T represents the T score. Since all three conditions are required, we expect that the MHC identity score will be similar to that based on M alone. mut The classifier ( Figure 4 H in BC1 ) or compared with the partial hypothesis: H BC1 ∩{M wt ≤α} and H BC1 ∩{T≤τ} compared to H A Therefore, we calculated the percentage of immunogenic mutations (the number of true positives divided by the sum of true positives and false positives) as H BC1 、Partial hypothesis H A’ :H BC1 ∩{T≤τ} and partial hypothesis H BC2 :H BC1 ∩{M wt ≤α}. We found that conservative thresholds for τ in the range of ≈0.5 to 1 performed best (the range for the WAG250 matrix is ​​-5.1 ( Replace)-+5.4( We also found that α can be strictly conservative compared to β, setting α≈1. Figure 6 A does show that when considering the combined mutation groups of B16 and CT26, based on H BC2 and H A’ The classifier is better than the background control hypothesis H BC1 The obtained accuracy is higher. In addition, based on the complete hypothesis H A The classifier is better than the partial hypothesis H BC1 and H BC2Higher precision is obtained, thus showing an additive effect. When analyzing the B16 and CT26 datasets separately, the same conclusion is held ( Figure 7 ).

[0337] Since the conditions M wt ≤α and T≤τ are assumed to be necessary conditions for immunogenicity, it is expected that an analyzer based on the condition H BC1 ∩{T>τ} or the condition H BC1 ∩{M wt >α} (i.e., negating the second condition) performs worse than H BC1 . Indeed, we found that this is the case for B16 and CT26 when analyzed jointly ( Figure 6 B) or separately ( Figure 7 ). Therefore, we infer that the B16 and CT26 datasets support H A hypothesis either jointly or separately. Omitting mutations that also show reactivity for WT RNA does not affect these conclusions ( Figure 8 ).

[0338] H A hypothesis control. Although mutations with high T scores may still be immunogenic, hypotheses rich in such mutations should be statistically rich in non-immunogenic mutations. Thus, if we compare the H A hypothesis (H BC2 ∩{T≤τ}) with its opposite H BC2 ∩{T>τ} ( Figure 4 the H in n ), we should observe a statistically significant lack of immunogenic mutations. Table 2 indeed shows that for M mut ≤β = 0.5, M wt ≤α = 1, and T≤τ = 1, H A outperforms H n , with a success rate of 52.5% (n = 21) compared to 21.4% (n = 14; P = 0.068, one-tailed Fisher's exact test).

[0339] Table 2. Percentage of immunogenic mutations under different hypotheses based on the combined B16 and CT26 datasets containing 133 mutations.

[0340] ​​​​​​​​​​​The difference between the success rates becomes larger. For example, for β = 0.25, the difference between n The success rate of the group was 17% (n=6) compared with H A The group success rate was 67% (n=14) (P=0.066, one-tailed Fisher's exact test)—see Table 3.

[0342] Table 3. Ranked list of 133 measured B16F10 / CT26. Will satisfy the basic control hypothesis H BC1 (M mut ≤0.25) were divided into three disjoint hypothesis categories: immunogenic mutations H A Hypothesis (M wt ≤0.8, T≤0.5), H enriched in non-immunogenic mutations n / Opposite H A Hypothesis (M wt ≤0.8, T>0.5), and immunogenic mutant H B UH C Hypothesis (M wt >0.8). It is proposed to classify candidate H based on the relative importance of distinguishing variables. A and H B UH C For H A , the proposed order is: M mut (Decrease) → T score (Decrease) → M WT Reduce. For H B UH C , the proposed order is: M mut (Lower) → M WT (Increase). Errors are standard errors.

[0343]

[0344] Class I immunogenicity (H A ):67±14% success rate

[0345]

[0346] Class II immunogenicity (HBUHC): 0% success rate

[0347]

[0348] H n :17±15% success rate

[0349]

[0350] Examples of other weighting factors that may further improve the immunogenicity ranking are given in Example 3.

[0351] More generally, the condition satisfying H BC1 (M mut The mutation list of ≤β) is divided into three categories: H A 、H n and H B UH C (Table 3), where H A Rich in immunogenic mutations, H n enriched for non-immunogenic mutations. In the case of B16 and CT26, H B UH C All three candidates in the group were non-immunogenic, contrary to our expectations. wt Choose a more appropriate threshold α* so that α*>>α, then there will be no B UH C Test the predictions.

[0352] The highest accuracy of the immunogenicity classifier. According to Table 1, the average success rate of predicting immunogenicity in the combined dataset of B16 and CT26 is 22.7% (=30 / 132). mut Using the most stringent threshold (β = 0.1) for the scores, the accuracy of the immunogenicity classifier increased to 60% (= 6 / 10; H in Table 2 BC1 ). By BC1 With M wt The accuracy is increased to 66.7% (= 6 / 9) when the two criteria are combined. A The classifier resulted in an additive response, which increased the accuracy to 75% (= 6 / 8) (Table 2).

[0353] In the combined B16 / CT26 dataset, the highest-ranked H A The epitope was MUT33 of B16 (see Table 3). Further analysis showed that MUT33 indeed elicited an MHC class I restricted CD8+ response and exhibited ex vivo immunogenicity against the minimal predicted epitope (data not shown).

[0354] Effect of gene expression. The ratio of immunogenic mutations (the ratio of the number of immunogenic mutations with an RPKM value below a given threshold to the total number of immunogenic mutations) was plotted as a function of the RPKM values ​​of the B16 and CT26 candidates, showing that the ratio stagnates somewhat at very low RPKM values ​​( Figure 9 ). Regardless of whether H is applied AThe effect was observed in all cases. The percentage of immunogenic mutations in different RPKM libraries was plotted ( Figure 9 ), indicating that RPKM values ​​≤ 1 have a slightly lower success rate (with or without H A filtering hypothesis), although suggestive, it should be noted that these results are within the margin of error.

[0355] Study of published CD8+ epitopes. We next wanted to understand whether published T cell-restricted tumor antigens with single amino acid substitutions elicited a CD8+-restricted response that satisfied our immunogenicity model. Of the 17 published epitopes (P. Van der Bruggen, V. Stroobant, N. Vigneron, B. Van den Eynde. (Cancer Immun, http: / / www.cancerimmunity.org / peptide / , 2013)) (Table 4), 5 met the H A Standard (α=0.7, β=0.2, τ=0.5), four meet H C UH B Standard (α=2.2, β=0.4), and two satisfying H n Standard (α=0.6, β=0.3, τ=1.7).

[0356] Table 4. Published epitopes with single amino acid substitutions that generate CD8+ responses. See Example 2 for a list of references. B UH C Anchor site mutations in the group.

[0357]

[0358] *Based on the WAG250 log-probability matrix, color legend: T≤0.5,0.5<T≤1,T> 1

[0359] Therefore, H A With H C UH B The hypotheses together roughly account for 50% of the disclosed epitopes. Interestingly, due to anchor site mutations, H C UH B Condition (red box in Table 4) 3 of the 4 open epitopes have M wt Rating greater than 10 ( Figure 10 ). Due to H C UH B The hypothesis requires that the probability of any cell presenting the wild-type epitope during organismal development remains very small, thus expecting M wt The threshold should be kept high, i.e., α *>>α. Indeed, when α increased from 0.8 to >3, the B16 / CT26 false positives in Table 3 disappeared. C UH B M in the hypothesis wt A more appropriate threshold might be 3-10.

[0360] MZ7-MEL cell line. To test the ability of our immunogenicity model to predict immunogenic epitopes in a human tumor model setting, we studied the MZ7-MEL cell line established in 1988 from a spleen metastasis of a patient with malignant melanoma (V. Lennerz et al., Proceedings of the National Academy of Sciences of the United States of America 102, 16013 (2005)). Screening of a cDNA library from MZ7-MEL cells containing autologous tumor-reactive T cells revealed that at least 5 new antigens were able to generate a CD8+ response (V. Lennerz et al., Proceedings of the National Academy of Sciences of the United States of America 102, 16013 (2005)). This constitutes the largest collection of CD8+ new antigens derived from patients to date. Applying our immunogenicity model to these epitopes, we found that 3 new antigens were classified as H A epitope, and a neoantigen with an anchor site mutation was classified as H B UH C Epitopes (arrows in Table 4, and Figure 10 ). Therefore, 4 out of these 5 epitopes could be explained by our immunogenicity model.

[0361] To test our ability to re-predict these epitopes in the MZ7-MEL cell line, we sequenced the exome of the MZ7-MEL cell line (see Methods). A total of 743 expressed non-synonymous mutations were identified. All five mutations previously identified by Lennerz et al. (V. Lennerz et al., Proceedings of the National Academy of Sciences of the United States of America 102, 16013 (2005)) were found. We then calculated the T score, M score, and M score for each mutation. mut and M wtThe HLA alleles and epitopes that resulted in the minimum MHC concordance score for a given mutation were also reported. Mutations were classified into one of three groups using thresholds α = 0.8, β = 0.2, τ = 0.5 (plus the condition RPKM > 0.2): H A 、H B UH C and H n , and then ranked based on their potential to be immunogenic as described in Table 3. We found that out of 743 mutations, 32 mutations met the H A Standard (Table 5), 12 meet H B UH C Standards (Table 6) and 15 that meet H n standard.

[0362] Table 5.H A Using thresholds α = 0.8, β = 0.2, and τ = 0.5, 32 of the 743 non-synonymous mutations expressed in MZ7-MEL were classified as H mutations. A -Immunogenicity. Based on M mut The ranking is performed using the T score (descending) → T score (descending) ranking scheme. Neoantigens with immunogenicity determined by Lennerz et al. are highlighted in yellow. Furthermore, the RPKM must exceed 0.2.

[0363]

[0364] * Ranking based on Mmut and T score

[0365] ** T-score based on the WAG250 log-odds matrix

[0366] Table 6.H B UH C Class MZ7-MEL cell mutations. Using the threshold α * =0.8, β=0.2 and RPKM>2, 12 of the 743 expressed non-synonymous mutations in MZ7-MEL were classified as H B UH C -Immunogenicity. Based on M mut (Descending) → M wt Arrange by (ascending) sorting scheme.

[0367] Neoantigens identified as immunogenic by Lennerz et al. are highlighted in yellow.

[0368]

[0369] * Arrangement based on M mutand M wt

[0370] In the H A Among the 32 mutations of the class, using M mut →T score ranking scheme, three H identified by Lennerz et al. A Class II mutations (SIRT2, SNRPD1 and RBAF600) were ranked 2nd, 4th and 13th among the 18 arranged classes (see Table 3). B UH C Of the 12 mutations in this class, the previous Lennerz et al. mutation (SNRP116) ranked third. In addition, if a higher (more realistic) M wt If the threshold (e.g., α*~5) was set, the previous Lennerz et al. mutation was ranked first (along with one additional anchor site mutation—Table 7). Finally, four mutations were predicted to have the correct HLA allele, epitope length, and mutation site as reported by Lennerz et al.

[0371] Table 7.H B UH C Class MZ7-MEL cell mutations. Using the threshold α * =5, β =0.2 and RPKM>2, 2 mutations among the 743 expressed non-synonymous mutations of MZ7-MEL were classified as H B UH C -Immunogenicity. Based on M mut (Descending) → M wt Arrange by (ascending) sorting scheme.

[0372] Neoantigens identified as immunogenic by Lennerz et al. are highlighted in yellow.

[0373]

[0374] * Arrangement based on M mut and M wt

[0375] in conclusion

[0376] Analysis of the B16 and CT26 datasets supports the following model: immunogenicity is conferred if three conditions are met: the wild-type peptide is presented, the mutant peptide is presented, and the amino acid substitution has a sufficiently low log-odds score ( Figure 11A). This immunogenicity model, which we termed Class I immunogenicity, was further supported in the human melanoma cell line model MZ7-MEL. The MZ7-MEL model and published CD8+-restricted neoantigens support a second model, which we termed Class II immunogenicity, in which wild-type epitopes are not presented, but substitutions (e.g., at anchor sites) result in a significant increase in the MHC identity score (>5 to 10), generating novel, previously unseen epitopes ( Figure 11 B) This framework for defining immunogenicity consists of a three-variable classification scheme (M mut 、M wt Using this classification scheme, we were able to reduce the 743 mutations in MZ7-MEL to a list of 34 mutations, with 3 of the 5 Lennerz et al. epitopes ranked in the top 5.

[0377] Table 7 shows that class II immunogenic mutations are rare. Of the 743 mutations, only 2 were classified as class II immunogenic (using M wt In a mouse melanoma model, a small amount of H B UH C Class I mutations (Table 8). This observation emphasizes the importance of class I immunogenic mutations in personalized vaccines, which are expected to be the main type of mutations present in patient samples that can be used for vaccination. At the same time, the fact that one of the five epitopes found by Lennerz et al. is class II immunogenic may suggest that class II immunogenic mutations are more effective or are selected by the immune system in some way.

[0378] Table 8. Candidate H in different tumor models A and H B UH C Number of mutations. (α = 0.8, α* = 0.5, β = 0.2, τ = 0.5)

[0379]

[0380] Example 2: Materials and Methods

[0381] The materials and methods used in Example 1 are described below.

[0382] animal

[0383] C57BL / 6J and Balb / cJ mice were maintained at the University of Mainz (CRL) in accordance with federal and state policies on animal research.

[0384] Cells for melanoma and colorectal mouse tumor models

[0385] The B16F10 melanoma cell line (product: ATCC CRL-6475, lot number: 58078645) and the CT26.WT colon cancer cell line (product: ATCC CRL-2638, lot number: 58494154) were purchased from the American Type Culture Collection in 2010. Sequencing experiments were performed using early passage cells (passages 3 and 4). The cells were routinely tested for mycoplasma. The cells have not been requalified since receipt. The MZ7-MEL cell line (established in January 1988) and the autologous Epstein–Barr virus-transformed B cell line were obtained from Thomas et al. PhD (Department of Medicine, Hematology-Oncology, Johannes-Gutenberg University).

[0386] synthetic peptides

[0387] Peptides were purchased from Jerini Peptide Technologies (Berlin, Germany) or synthesized at the TRON peptide facility. The synthesized peptides were 27 amino acids in length, with either a mutated (MUT) or wild-type (WT) amino acid at position 14.

[0388] Immunization of mice

[0389] Age-matched female C57BL / 6 or Balb / c mice were intravenously injected with 20 μl of Lipofectamine in PBS. TM 20 μ g in vitro transcribed mRNA of RNAiMAX (Invitrogen), total injection volume is 200 μ l (3 mice per group). At the 0th, 3rd, 7th, 14th and 18th day immune mice. After the first injection 23 days, mice were killed and spleen cells were separated for immunological testing (see ELISPOT analysis). Utilize the sequence construction of 27 amino acids (aa) containing the mutation of site 14 to express one (single epitope) or two mutations (double epitope) DNA sequence, and clone it into pST1-2BgUTR-A120 main chain (S.Holtkamp et al., Blood 108,4009 (2006)). Previously described in vitro transcription and purification from this template (S.Kreiter et al., Cancer Immunology, Immunotherapy 56,1577 (2007)).

[0390] ELISpot assay

[0391] The enzyme-linked immunospot (ELISPOT) assay (S. Kreiter et al., Cancer Research 70, 9031 (2010)) and the generation of syngeneic bone marrow-derived dendritic cells (BMDC) as stimulators were previously described (L. MB et al., J. Immunol. Methods 223, 77 (1999)). For the B16F10 model, BMDC were peptide-pulsed (6 μg / ml) with peptides containing the indicated mutations, the corresponding wild-type peptides, or with a control peptide (VSV-NP). For the CT26 model, in addition to restimulation with peptides, BMDC were transfected with the corresponding in vitro transcribed mRNA, and the mRNA was also used for restimulation. For the analysis, 5×10 cells were plated in microtiter plates coated with anti-IFN-γ antibody (10 μg / mL, clone AN18; Mabtech). 4 BMDC with 5×10 5 After incubation at 37°C for 18 hours, the secretion of cytokines was detected using anti-IFN-γ antibody (clone R4-6A2; Mabtech). The number of spots was counted and the S5 VersaELISPOT Analyzer, ImmunoCapture™ Image Acquisition Software, and Version 5 Statistical analysis was performed using Student's t-test and Mann-Whitney test (nonparametric test). Responses were considered significant if p value < 0.05.

[0392] Intracellular cytokine analysis

[0393] Aliquots of splenocytes prepared for ELISPOT analysis were analyzed for cytokine production by intracellular flow cytometry. For this purpose, 2 × 10 6 Splenocytes were placed in 96-well plates in a medium (RPMI + 10% FCS) supplemented with the Golgi inhibitor Brefeldin A (10 μg / mL). 5Peptide-pulsed or RNA-transfected BMDCs were restimulated with cells from each animal for 5 h. After incubation, the cells were washed with PBS, resuspended in 50 μl PBS, and stained extracellularly with the following anti-mouse antibodies at 4 ° C for 20 minutes: anti-CD4 FITC, anti-CD8 APC-Cy7 (BD Pharmingen). After incubation, the cells were washed with PBS and then resuspended in 100 μL Cytofix / Cytoperm (BD Bioscience) solution at 4 ° C for 20 minutes to permeabilize the outer membrane. After permeabilization, the cells were washed with Perm / Wash-Buffer (BD Bioscience), resuspended in Perm / Wash-Buffer 50 μL / sample, and stained intracellularly with the following anti-mouse antibodies at 4 ° C for 30 minutes: anti-IFN-γ PE, anti-TNF-α PE-Cy7, anti-IL2 APC (BD Pharmingen). After washing with Perm / Wash-Buffer, cells were resuspended in PBS containing 1% paraformaldehyde for flow cytometric analysis. Samples were analyzed using a BD FACSCantoTMII cytometer and FlowJo (version 7.6.3).

[0394] Next-generation sequencing

[0395] Nucleic acid extraction: DNA and RNA were extracted from mixed cells and DNA was extracted from mouse tissues using the Qiagen DNeasy Blood and Tissue kit (for DNA) and the Qiagen RNeasy Micro kit (for RNA).

[0396] DNA exome sequencing: As previously described, exon capture of B16F10, C57BL / 6J, and CT26.WT, as well as DNA resequencing of Balb / cJ, was performed three times (JC Castle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012)). Exome capture for MZ7-MEL / EBV-B DNA resequencing was performed twice using a capture assay based on the Agilent XT Human all Exon 50Mb protocol to capture all protein coding regions. 3 μg of purified genomic DNA (gDNA) was fragmented into 150-200 bp using a Covaris S2 ultrasonic device. The fragments were end-repaired, 5' phosphorylated, and 3' adenylated according to the manufacturer's instructions. Agilent-indexed specific end-paired adapters were ligated to the gDNA fragments using a 10:1 adapter to gDNA molar ratio. Four cycles of pre-capture amplification were performed using Agilent InPE 1.0 with SureSelect Indexing Pre-Capture PCR primers and Herculase II polymerase. 500 ng of adaptor-ligated, PCR-enriched gDNA fragments were hybridized with Agilent exome capture bait at 65°C for 24 hours. The hybridized gDNA / RNA bait complexes were removed using streptavidin-coated magnetic beads, washed, and the RNA bait was sheared during elution in SureSelect Elution Buffer. Eluted gDNA fragments were amplified using SureSelect Indexing Post-Capture PCR with indexing PCR primers and Herculase II polymerase for 10 cycles. All cleanup was performed using 1.8 volumes of Agencourt AMPure XP magnetic beads. All quality controls were performed using the Invitrogen Qubit HS assay, and fragment size was determined using the Agilent 2100 Bioanalyzer HS DNA assay. Exome-enriched gDNA libraries were clustered on a cBot using the Truseq SR cluster kit v2.5 using 7 pM libraries and sequenced at 1 x 100 bps on an Illumina HiSeq2000 using the Truseq SBS kit.

[0397] RNA gene expression profiling (RNA-Seq): Two barcoded mRNA-seq cDNA libraries were prepared. mRNA was isolated from 5 μg of total RNA using Seramag Oligo(dT) magnetic beads (Thermo Scientific) (modified Illumina mRNA-seq protocol using NEB reagents) and fragmented using divalent cations and heat. The resulting fragments (160–220 bp) were converted to cDNA using random primers and SuperScript II (Invitrogen), followed by second-strand synthesis using DNA polymerase I and RNase H. The cDNA was end-repaired, 5' phosphorylated, and 3' adenylated according to the NEB RNA library kit instructions. Adapters specific for Illumina multiplex identifiers with single 3' T overhangs were ligated using T4 DNA ligase at a 10:1 adapter-to-cDNA insert molar ratio. The cDNA library was purified and size-selected to 300 bp (E-Gel 2% SizeSelect gel, Invitrogen). Enrichment was performed by PCR using Phusion DNA polymerase and Illumina-specific PCR primers, with the addition of an Illumina six-base index and flow cell-specific sequences. All cleanups prior to this step were performed using 1.8 volumes of Agencourt AMPure XP magnetic beads. All quality controls were performed using the Invitrogen Qubit HS assay, and fragment sizes were determined using the Agilent 2100 Bioanalyzer HS DNA assay. Barcoded RNA-Seq libraries were pooled and sequenced at 50 bp as described above.

[0398] NGS data analysis, gene expression: RNA sample output sequence reads were preprocessed according to the Illumina standard protocol, including filtering low-quality reads. Sequence reads were aligned to mm9 (AT Chinwalla et al., Nature 420, 520 (2002)) or hg18 (F. Collins, E. Lander, J. Rogers, R. Waterston, I. Conso, Nature 431, 931 (2004)) reference genome sequences using bowtie (version 0.12.5) (B. Langmead, C. Trapnell, M. Pop, SL Salzberg, Genome Biol 10, R25 (2009)). For genome alignments, two mismatches were allowed, and only the best alignment ("-v2-best") was reported; for transcript alignments, default parameters were used. Reads that did not match the genomic sequence were aligned to the UCSC database of all possible exon-exon junction sequences for known genes (F. Hsu et al., Bioinformatics 22, 1036 (2006)). Expression values ​​were determined by intersecting the read coordinates with those of the RefSeq transcript, counting overlapping exons and junction reads, and normalizing to RPKM expression units (number of reads mapping to each 1K base of the exon per 1 million mapped reads) (A. Mortazavi, BA Williams, K. McCue, L. Schaeffer, B. Wold, Nature methods 5, 621 (2008)).

[0399] NGS data analysis, exploration of somatic mutations: somatic mutations were determined as described previously (JC Castle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012)). Sequence reads were aligned with mm9 or hg18 reference genomes using bwa (default options, version 0.5.8c) (H. Li, R. Durbin, Bioinformatics 25, 1754 (2009)). Uncertain reads located at multiple sites of the genome were removed. Mutations were determined using a combination of two software programs: samtools (version 0.1.8) (H. Li, Bioinformatics 27, 1157 (2011)) and SomaticSniper (A. McKenna et al., Genome Research 20, 1297 (2010)). For B16F10 and C57BL / 6J, GATK (A. McKenna et al., Genome Research 20, 1297 (2010)) was also included. Potential somatic changes identified in all respective replicates were assigned a "false discovery rate" (FDR) confidence value (M. et al., PLoS computational biology 8, e1002714 (2012)) (CT26 and MZ7-MEL only).

[0400] Mutant selection and confirmation

[0401] The selection criteria for 50 B16F10 mutations used for immunogenicity testing have been described previously (JC Castle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012)). These mutation criteria include: (i) present in all three B16F10 replicates and absent in all three C57BL / 6 replicates; (ii) occurring in RefSeq transcripts; (iii) causing non-synonymous changes; (iv) occurring in B16F10-expressed genes (median RPKM in replicates>10, exon expression>0); and (v) for each mutation, M mut The score (see below) was required to be < 5. Of the remaining 59 mutations, quantile rankings of MHC class I score, MHC class II score, and transcript expression were generated, and the top 50 mutations (0.1 ≤ M mut≤3.9) were confirmed by PCR (for further details, see (JC Castle et al., Exploiting the mutanome for tumor vaccination. Cancer Research 72, 1081 (2012))). The selection criteria for the 96 CT26.WT mutations used for immunogenicity testing were further refined and included the following criteria: (i) present in all three CT26.WT replicates and absent in all three Balb / cJ replicates; (ii) FDR ≤ 0.05; (iii) occurring in UCSC known gene transcripts; (iv) causing non-synonymous changes; (v) not present in the dbSNP database; (vi) not located in genomic repeat regions. From the remaining 493 mutations, 8 12-member groups were identified based on three characteristics: M mut Mutations were selected based on a greedy algorithm based on score (lowest - [0.1, 1.9] vs. highest - [3.9-20.3]), protein compartment (extracellular vs. intracellular), and gene expression (median below 7.1 RPKM vs. above), and the threshold was adjusted accordingly. 94 of the 96 mutations generated were confirmed by PCR and then Sanger sequencing.

[0402] The selection criteria for MZ7-ML mutations for analysis included: (i) presence in two MZ7-MEL replicates and absence in two autologous EBV-B replicates, followed by steps (ii) to (vi) described above for CT26.WT. Applying steps (i)-(vi) reduced the initial list of ~8000 mutations to 743.

[0403] MHC binding prediction and M mut Rating calculation

[0404] Using IEDB to analyze resource consistency tools ( http: / / tools.immuneepitope.org / analyze / html / mhc_binding.html) to implement MHC binding prediction (Y. Kim et al., Nucleic Acids Research 40, W525 (2012)), which is based on a combination of the best performing prediction methods from benchmark studies (HHLin, S. Ray, S. Tongchusak, EL Reinherz, V. Brusic, BMC immunology 9, 8 (2008); B. Peters et al., PLoS computational biology 9, 8 (2008)), SMM (B. Peters, A. Sette, BMC bioinformatics 6, 132 (2005)), and a model that also combined some alleles (J. Sidney et al., Immunome Research 4, 2 (2008)). 2, e65 (2006)). The consensus method combines the prediction scores of all tools by generating a percentile ranking that reflects the combined prediction score of a given peptide against the peptide scores of 5 million random peptides from SWISSPROT.

[0405] For each mutation, we calculated the predicted MHC concordance score for all possible (i) sequence windows (location of the mutation), (ii) epitope determinations, and (iii) possible murine MHC class I alleles. The minimum score of all MHC concordance scores was defined as M mut score.

[0406] Calculate the log-odds matrix and T score

[0407] The log-odds matrix can be estimated from sequence alignments of large protein databases. Early log-odds matrices were based on pairwise comparisons of sequences (BLOSUM62 (S. Kreiter et al., Cancer Immunology, Immunotherapy 56, 1577 (2007))) and maximum parsimony (MP) estimation methods (e.g., PAM250 (MO Dayhoff, R.M. Schwartz, B.C. Orcutt, A model for evolutionary change. MO Dayhoff, ed. Atlas of proteins sequence and structure Vol. 5, 345 (1978)), JTT250 (SQLe, O. Gascuel, Molecular biology and evolution 25, 1307 (2008)) and Gonnet matrix (CC Dang, V. Lefort, V.S. Le, Q.S. Le, O. Gascuel, Bioinformatics 27, 2758 (2011))). Recently, methods based on maximum likelihood (ML) have been developed (e.g., VT160 (PG Higgs, TK Attwood, Bioinformatics and molecular evolution. (Wiley-Blackwell, 2009)), WAG (S. Whelan, N. Goldman, Molecular biology and evolution 18, 691 (2001)), and LG (V. Lennerz et al., Proceedings of the National Academy of Sciences of the United States of America 102, 16013 (2005)). Since ML is not limited to comparing only closely related sequences, as is the case with MP-based methods, this estimation method is expected to be the most accurate.

[0408] The calculation of the log-odds matrix has been described in detail elsewhere (CC Dang, V. Lefort, V S Le, Q S Le, O. Gascuel, Bioinformatics 27, 2758 (2011)). Briefly, the standard model of amino acid substitution assumes a 20×20 rate matrix Q ij represents a Markovian, time-continuous, time-reversible model, where q ij(i≠j) is the number of substitutions from amino acid i to j per unit time, and the diagonal elements are selected to satisfy Q can be decomposable so that for i≠j, Q ij =S ij π j , where S i,j is a symmetric commutative matrix, and π i is the probability of observing amino acid i (CC Dang, V. Lefort, V S Le, Q S Le, O. Gascuel, Bioinformatics 27, 2758 (2011)). Finally, Q is normalized so that Thus, the time unit t = 1.0 corresponds to 1.0 expected substitution per site or one "acceptable point mutation" per site, represented by a PAM distance of 100 (MO Dayhoff, RMS Schwartz, BC Orcutt, A model for evolutionary change. MO Dayhoff, ed. Atlas of protein sequence and structure Vol. 5, 345 (1978); SQLe, O. Gascuel, Molecular biology and evolution 25, 1307 (2008); CCDang, V. Lefort, VS Le, QS Le, O. Gascuel, Bioinformatics 27, 2758 (2011)). After time t, the probability that amino acid i is replaced by amino acid j is Pr(i→j|t) = P ij (t) is given by the 20×20 probability matrix P(t) = e tQ Given (using symbols to represent matrix exponentiation). The log-odds matrix for computing time t is given by the log-odds 20×20 matrix (MO Dayhoff, RMS Schwartz, BC Orcutt, A model for evolutionary change. MO Dayhoff, ed. Atlas of protein sequence and structure Vol. 5, 345 (1978, 1978)). Time reversible means π i P ij (t) = π j P ji (t), so T i,jis symmetrical (PG Higgs, TKAttwood, Bioinformatics and molecular evolution. (Wiley-Blackwell, 2009)).

[0409] Here we will The substituted T score is defined as T i,j , and depends on the evolutionary model and time t. We studied various models and PAM distances for T scores, including PAM, BLOSUM62, JTT, VT160, Gonnet, WAG, WAG* and LG (see references below). The figures in this report were generated using a T score based on the WAG model and a PAM distance of 250. This large PAM distance means that there is a high probability that the amino acid change exists (PG Higgs, TK Attwood, Bioinformatics and molecular evolution. (Wiley-Blackwell, 2009)), and can be used to detect distant relationships between sequences whose residues may be different but retain the physicochemical properties of amino acids (MO Dayhoff, RMS Schwartz, BC Orcutt, A model for evolutionary change. MO Dayhoff, ed. Atlas of proteins sequence and structure Vol. 5, 345 (1978); PG Higgs, TK Attwood, Bioinformatics and molecular evolution. (Wiley-Blackwell, 2009)).

[0410] Using the t-distribution test statistic, we compared the mean T scores of the WAG matrices for immunogenic and non-immunogenic epitopes in Table 3 using various PAM scores (1, 10, 25, 50, 100, 150, 200, and 250). As expected, analysis of the test statistic showed that the P value decreased monotonically with PAM distance, suggesting that a PAM distance of 250 was the optimal solution (data not shown). For all but the least accurate PAM matrix across all evolutionary models, the PAM distances assigned to the H A and H n Among all evolutionary models, the WAG250 model leads to H in Table 3 A Epitope and H n The maximum separation of epitopes was measured using the test statistic: [maximum T score (H A )-minimum T score (H n)] / σ(T score(H A ),T score (H n ) (data not shown). The same test statistic value is also largest for a PAM distance of 250 compared to smaller PAM distances.

[0411] Public CD8+ epitopes

[0412] From Cancer Immunity Journal (P. Van der Bruggen, V. Stroobant, N. Vigneron, B. Van den Eynde. (Cancer Immun, http: / / www.cancerimmunity.org / peptide / , 2013) ( http: / / cancerimmunity.org / peptide / mutations / ) CD8+ epitopes with single mutated amino acids were collected from published lists of tumor antigens resulting from mutations. HLA alleles were obtained from published lists or from the original papers (if the latter was more accurate). The references listed in Table 4 are as follows: (1) Lennerz et al. PNAS 102(44), pp.16013–16018(2005); (2) Karanikas et al. Cancer Res 61(9), pp.3718–3724(2001); (3) Sensi et al. Cancer Res 65(2), pp.632–640(2005); (4) Linard et al. J. Immunol 168(9), pp.4802–4808(2002); (5) Zorn et al. Eur. J. Immunol 29(2), pp.592–601(1999); (6) Graf et al. Blood 109(7), pp.2985–2988(2007); (7) Robbins et al. Cancer Res 65(2), pp.632–640(2005); (8) Linard et al. J. Immunol 168(9), pp.4802–4808(2002); (9) Zorn et al. Eur. J. Immunol 29(2), pp.592–601(1999); (10) Graf et al. Blood 109(7), pp.2985–2988(2007); (11) etal.J.Exp.Med 183(3),pp.1185–1192(1996);(8)Vigneron et al.Cancer Immun 2,pp.9(2002);(9)Echchakir et al.Cancer Res 61(10),pp.4078–4083(2001);(10)Hogan etal.Cancer Res 58(22), pp.5144–5150(1998); (11) Ito et al.Int.J.Cancer 120(12), pp.2618–2624(2007); (12) et al.Science 269(5228),pp.1281–1284(1995); (13)Gjertsen et al.Int.J.Cancer 72(5),pp.784–790(1997).

[0413] Example 3: Example of a strategy for weighting mutation scores to improve prioritization of immunogenic mutations

[0414] Once the RNA injected into the cell is translated and cleaved into short peptides, it can be presented by different HLA types in the cell. Therefore, it is obvious that the more HLA types are predicted to have low MHC identity (or similarity) scores, the more likely a given mutation is to be immunogenic, because it may be displayed on multiple HLA types in parallel. A and / or H B UH C weighting mutations by the number of HLA types present, or even by only those with low M mut The number of HLA types scored weights each mutation, potentially improving the immunogenicity ranking. In the most general scenario, when we inject 27-mer RNA or peptide into cells, we have the freedom to choose not only the HLA type, but also the length of the peptide and the mutation site within that peptide. Thus, we can look at all possible HLA types, all possible window lengths, and all possible mutation sites within the window, and calculate the number of mutations that are classified as HLA. A and / or H B UH C The number of solutions (for each given mutation) ( Figure 12 ). This may be an important weighting factor in prioritizing mutations and thus selecting the most effective epitopes for vaccination. Figure 13 Shown as M mut and ΔM=M mut –M wt Instances of the function scatter plot for all these scenarios. Sequence Listing <110> BIONTECH RNA PHARMACEUTICALS GMBH TRON-Medical Institute for Translational Oncology at the Johannes Gutenberg University Mainz (TRON-TRANSLATIONALE ONKOLOGIE AN DER UNIVERSITATSMEDIZIN DER JOHANNES GUTENBERG-UNIVERSITAT MAINZ GEMEINNUTZIGE GMBH) <120> Predicting the immunogenicity of T cell epitopes <130> 674-99 PCT <160> 94 <170> PatentIn version 3.5 <210> 1 <211> 15 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Linker sequence <220> <221> SITE <222> (1)..(3) <223> The partial sequence is repeated a times, wherein a is a number independently selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 <220> <221> SITE <222> (1)..(15) <223> a + b + c + d + e is not 0, and is preferably 2 or more, 3 or more, 4 or more, or 5 or more <220> <221> SITE <222> (4)..(6) <223> The partial sequence is repeated b times, wherein b is a number independently selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 <220> <221> SITE <222> (7)..(9) <223> The partial sequence is repeated c times, wherein c is a number independently selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 <220> <221> SITE <222> (10)..(12) <223> The partial sequence is repeated d times, wherein d is a number independently selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 <220> <221> SITE <222> (13)..(15) <223> The partial sequence is repeated e times, wherein e is a number independently selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 <400> 1 Gly Gly Ser Gly Ser Ser Ser Gly Gly Gly Ser Ser Gly Gly Ser Gly 1 5 10 15 <210> 2 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Linker sequence <400> 2 Gly Gly Ser Gly Gly Gly Gly Ser Gly 1 5 <210> 3 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 3 Ala Ala Val Ile Leu Arg Asp Ala Leu His Met 1 5 10 <210> 4 <211> 11 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 4 Ala Ala Val Ile Leu Arg Val Ala Leu His Met 1 5 10 <210> 5 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 5 Gly Gly Pro Gly Ser Glu Lys Ser Leu 1 5 <210> 6 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 6 Gly Gly Pro Gly Ser Gly Lys Ser Leu 1 5 <210> 7 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 7 Leu Ala Leu Pro Asn Asn Tyr Cys Asp Val 1 5 10 <210> 8 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 8 Leu Ala Leu Pro Asn Asn Tyr Cys Asp Phe 1 5 10 <210> 9 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 9 Ser His Leu Asn Asn Asp Val Trp Gln Ile 1 5 10 <210> 10 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 10 Ser His Leu Asn Asn Asp Phe Trp Gln Ile 1 5 10 <210> 11 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 11 Tyr Tyr Met Arg Asp Val Ile Ala Ile 1 5 <210> 12 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 12 Tyr Tyr Met Arg Asp Val Thr Ala Ile 1 5 <210> 13 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 13 Thr Tyr Leu Gln Pro Ala Gln Ala Gln Met 1 5 10 <210> 14 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 14 Ile Tyr Leu Gln Pro Ala Gln Ala Gln Met 1 5 10 <210> 15 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 15 Gln Ser Leu Gly Phe Thr Tyr Leu 1 5 <210> 16 <211> 8 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 16 Gln Arg Leu Gly Phe Thr Tyr Leu 1 5 <210> 17 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 17 Glu Tyr Trp Ala Ser Arg Ala Leu Asp Ser 1 5 10 <210> 18 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 18 Glu Tyr Trp Ala Ser Arg Ala Leu Gly Ser 1 5 10 <210> 19 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 19 Gly Tyr Leu Gln Phe Ala Tyr Glu Gly Cys 1 5 10 <210> 20 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 20 Gly Tyr Leu Gln Phe Ala Tyr Glu Gly Arg 1 5 10 <210> twenty one <211> 11 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> twenty one Val Thr Phe Gln Ala Phe Ile Asp Val Met Ser 1 5 10 <210> twenty two <211> 11 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> twenty two Val Thr Phe Gln Ala Phe Ile Asp Phe Met Ser 1 5 10 <210> twenty three <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> twenty three Pro Tyr Leu Thr Ala Leu Asp Asp Leu Leu 1 5 10 <210> twenty four <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> twenty four Pro Tyr Leu Thr Ala Leu Gly Asp Leu Leu 1 5 10 <210> 25 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 25 Ala Gly Gly Leu Phe Val Ala Asp Ala Ile 1 5 10 <210> 26 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 26 Ala Gly Gly Leu Phe Val Ala Asp Glu Ile 1 5 10 <210> 27 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 27 Val Gly Ile Asn Phe Leu Gln Ser Tyr Gln 1 5 10 <210> 28 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 28 Val Gly Ile Asn Ser Leu Gln Ser Tyr Gln 1 5 10 <210> 29 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 29 Thr Arg Pro Ala Arg Asp Gly Thr Phe 1 5 <210> 30 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 30 Thr Arg Pro Ala Gly Asp Gly Thr Phe 1 5 <210> 31 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 31 Glu Pro Gln Ile Ala Met Asp Asp Met 1 5 <210> 32 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 32 Glu Pro Gln Ile Asp Met Asp Asp Met 1 5 <210> 33 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 33 Ile Ala Met Gln Asn Thr Thr Gln Leu 1 5 <210> 34 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 34 Ile Ala Ile Gln Asn Thr Thr Gln Leu 1 5 <210> 35 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 35 Ala Ile Tyr His His Ala Ser Arg Ala Ile 1 5 10 <210> 36 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 36 Ala Ile Tyr Tyr His Ala Ser Arg Ala Ile 1 5 10 <210> 37 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 37 Ser Tyr Ile Ala Leu Val Asp Lys Asn Ile 1 5 10 <210> 38 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 38 Ser Tyr Leu Ala Leu Val Asp Lys Asn Ile 1 5 10 <210> 39 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 39 Val Ile Pro Ile Leu Glu Met Gln Phe 1 5 <210> 40 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 40 Val Ile Pro Ile Leu Glu Val Gln Phe 1 5 <210> 41 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 41 Gly Tyr Leu Gln Phe Ala Tyr Asp Gly Arg 1 5 10 <210> 42 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 42 Gly Tyr Leu Gln Phe Ala Tyr Glu Gly Arg 1 5 10 <210> 43 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 43 Val Tyr Leu Asn Leu Leu Leu Lys Phe Thr 1 5 10 <210> 44 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 44 Val Tyr Leu Asn Leu Phe Leu Lys Phe Thr 1 5 10 <210> 45 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 45 Lys Ile Phe Ser Glu Val Thr Leu Lys 1 5 <210> 46 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 46 Lys Ile Phe Ser Glu Val Thr Pro Lys 1 5 <210> 47 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 47 Ser His Glu Thr Val Ile Ile Glu Leu 1 5 <210> 48 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 48 Ser His Glu Thr Val Thr Ile Glu Leu 1 5 <210> 49 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 49 Phe Leu Asp Glu Phe Met Glu Gly Val 1 5 <210> 50 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 50 Phe Leu Asp Glu Phe Met Glu Ala Val 1 5 <210> 51 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 51 Arg Pro His Val Pro Glu Ser Ala Phe 1 5 <210> 52 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 52 Gly Pro His Val Pro Glu Ser Ala Phe 1 5 <210> 53 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 53 Leu Leu Leu Asp Asp Leu Leu Val Ser Ile 1 5 10 <210> 54 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 54 Leu Leu Leu Asp Asp Ser Leu Val Ser Ile 1 5 10 <210> 55 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 55 Ile Leu Asp Thr Ala Gly Arg Glu Glu Tyr 1 5 10 <210> 56 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 56 Ile Leu Asp Thr Ala Gly Gln Glu Glu Tyr 1 5 10 <210> 57 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 57 Glu Thr Val Ser Glu Gln Ser Asn Val 1 5 <210> 58 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 58 Glu Thr Val Ser Glu Glu Ser Asn Val 1 5 <210> 59 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 59 Lys Ile Leu Asp Ala Val Val Ala Gln Lys 1 5 10 <210> 60 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 60 Lys Ile Leu Asp Ala Val Val Ala Gln Glu 1 5 10 <210> 61 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 61 Lys Ile Asn Lys Asn Pro Lys Tyr Lys 1 5 <210> 62 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 62 Glu Ile Asn Lys Asn Pro Lys Tyr Lys 1 5 <210> 63 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 63 Tyr Val Asp Phe Arg Glu Tyr Glu Tyr Tyr 1 5 10 <210> 64 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 64 Tyr Val Asp Phe Arg Glu Tyr Glu Tyr Asp 1 5 10 <210> 65 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 65 Ser Tyr Leu Asp Ser Gly Ile His Phe 1 5 <210> 66 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 66 Ser Tyr Leu Asp Ser Gly Ile His Ser 1 5 <210> 67 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 67 Lys Glu Leu Glu Gly Ile Leu Leu Leu 1 5 <210> 68 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 68 Lys Glu Leu Glu Gly Ile Leu Leu Pro 1 5 <210> 69 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 69 Thr Leu Asp Trp Leu Leu Gln Thr Pro Lys 1 5 10 <210> 70 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 70 Thr Leu Gly Trp Leu Leu Gln Thr Pro Lys 1 5 10 <210> 71 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 71 Phe Ile Ala Ser Asn Gly Val Lys Leu Val 1 5 10 <210> 72 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 72 Phe Ile Ala Ser Lys Gly Val Lys Leu Val 1 5 10 <210> 73 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 73 Val Val Pro Cys Glu Pro Pro Glu Val 1 5 <210> 74 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 74 Val Val Pro Tyr Glu Pro Pro Glu Val 1 5 <210> 75 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 75 Ala Cys Asp Pro His Ser Gly His Phe Val 1 5 10 <210> 76 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 76 Ala Arg Asp Pro His Ser Gly His Phe Val 1 5 10 <210> 77 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 77 Val Val Val Gly Ala Val Gly Val Gly 1 5 <210> 78 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 78 Val Val Val Gly Ala Gly Gly Val Gly 1 5 <210> 79 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 79 Gly Leu Ala Gly Gly Leu Leu Ala Leu 1 5 <210> 80 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 80 Val Leu Phe Trp Gly Leu Leu Leu Leu 1 5 <210> 81 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 81 Gly Ile Leu Gly Trp Val Leu Tyr Leu 1 5 <210> 82 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 82 Val Leu Phe Trp Gly Leu Leu Phe Tyr 1 5 <210> 83 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 83 Gly Leu Ala Gly Gly Leu Leu Ala Tyr 1 5 <210> 84 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 84 Val Ile Phe Trp Gly Leu Leu Val Leu 1 5 <210> 85 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 85 Lys Ile Leu Asp Ala Val Val Ala Gln Glu 1 5 10 <210> 86 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 86 Lys Ile Leu Asp Ala Val Val Ala Gln Lys 1 5 10 <210> 87 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 87 Gln His Ser Ala Ala Pro Gly Pro Pro 1 5 <210> 88 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 88 Gln His Ser Ala Ala Pro Gly Pro Leu 1 5 <210> 89 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 89 Glu Ile Asn Lys Asn Pro Lys Tyr Lys 1 5 <210> 90 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 90 Lys Ile Asn Lys Asn Pro Lys Tyr Lys 1 5 <210> 91 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 91 Tyr Val Asp Phe Arg Glu Tyr Glu Tyr Asp 1 5 10 <210> 92 <211> 10 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 92 Tyr Val Asp Phe Arg Glu Tyr Glu Tyr Tyr 1 5 10 <210> 93 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 93 Ser Tyr Leu Asp Ser Gly Ile His Ser 1 5 <210> 94 <211> 9 <212> PRT <213> Artificial Sequence <220> <221> <222> <223> Epitope <400> 94 Ser Tyr Leu Asp Ser Gly Ile His Phe 1 5

Claims

1. A method for predicting immunogenic amino acid modifications, comprising: For each of a plurality of modified peptides, a binding score of the modified peptide to one or more MHC molecules is determined ( M mut ), wherein the modified peptide comprises one or more amino acid modifications at specified positions relative to a corresponding non-modified peptide, wherein the modified peptide is identical to the corresponding non-modified peptide except for the amino acid modifications, and wherein the modified peptide and the corresponding non-modified peptide are 8 to 15 amino acids in length; For each of a plurality of non-modified peptides corresponding to the modified peptide, a binding score ( M wt ); for each of the plurality of modified peptides, determining a binding score (T) for the modified peptide to one or more T cell receptors when present in an MHC-peptide complex, wherein the T score is based on chemical and physical similarity between one or more modified amino acids in the modified peptide and the amino acids at corresponding designated positions of the non-modified peptide, wherein the chemical and physical similarity is determined based on the likelihood that amino acids are interchangeable in nature; as well as Based on certain M mut 、 M wt The immunogenicity of the modified peptides was ranked by the T score, where M mut The weight is higher than T ,and T The weight is higher than M wt ; wherein the at least one candidate modified peptide is identified from the plurality of modified peptides if the at least one candidate modified peptide is more immunogenic than at least one other modified peptide in the plurality of modified peptides, wherein the modified peptide is predicted to be immunogenic if, (i) M mut indicating binding of the modified peptide to one or more MHC molecules, (ii) M wt indicates binding of the non-modified peptide to one or more MHC molecules, and (iii) T The scores indicate how chemically and physically different the modified and non-modified amino acids are.

2. The method of claim 1, wherein M mut and M wt The scores were each determined by a method involving sequence comparison to a database of MHC binding motifs.

3. The method of claim 1 or 2, wherein the binding score of the modified peptides in the plurality of modified peptides to one or more MHC molecules is ( M mut ) based on predicted MHC binding consensus scores calculated for: i) all possible positions of mutations within the modified peptide, (ii) all possible epitope lengths, and (iii) all possible murine MHC class I alleles.

4. The method of claim 1 or 2, wherein the binding score ( T ) calculated from the evolution-based log-odds matrix for T cell receptor recognition.

5. The method of claim 4, wherein the log-odds matrix is ​​calculated from a maximum likelihood based method for comparing peptide sequences.

6. The method of claim 1, 2, or 5, wherein the modified peptide comprises a fragment of a modified protein, the fragment comprising the modification present in the protein.

7. The method of claim 1, 2, or 5, wherein the non-modified peptide comprises germline amino acid residues at positions corresponding to the modifications in the modified peptide.

8. The method of claim 1, 2, or 5, wherein the one or more MHC molecules comprise different MHC molecule types.

9. The method of claim 8, wherein the one or more MHC molecules comprise different MHC alleles.

10. The method of claim 1, 2, 5 or 9, wherein the one or more MHC molecules are MHC class I molecules and / or MHC class II molecules.

11. The method of claim 1, 2, 5, or 9, further comprising determining a binding score of the modified peptide to one or more MHC molecules ( M mut ) meets a predetermined first threshold for binding to the one or more MHC molecules.

12. The method of claim 1, 2, 5, or 9, further comprising determining a binding score of the non-modified peptide to one or more MHC molecules. M wt ) meets a predetermined second threshold for binding to the one or more MHC molecules.

13. The method of claim 1, 2, 5 or 9, further comprising determining a binding score of the modified peptide to one or more T cell receptors when the modified peptide is present in an MHC-peptide complex. T ) satisfies a predetermined third threshold.

14. The method of claim 1, 2, 5 or 9, wherein the method is used to identify immunogenic peptide epitopes for use in preparing a vaccine.

15. The method of claim 14, wherein the immunogenic peptide epitope comprises one or more non-synonymous mutations corresponding to tumor-specific cancer mutations.

16. The method of claim 15, wherein the vaccine is a personalized cancer vaccine designed to target a patient's tumor-specific cancer mutation.

17. A method for designing a vaccine, comprising: Use of the method of any one of claims 1 to 13 to identify modified peptides predicted to be immunogenic.

18. The method of claim 17, wherein the vaccine comprises an immunogenic modified peptide.

19. The method of claim 18, further comprising providing a nucleic acid encoding the identified immunogenic modified peptide.

20. A method for providing a vaccine comprising the steps of: Identifying a modification or modified peptide predicted to be immunogenic by the method of any one of claims 1 to 13.

21. A method for preparing a vaccine, comprising: Producing a vaccine comprising: one or more modified peptides; one or more nucleic acids encoding said one or more modified peptides; or a plurality of cells expressing said one or more modified peptides, The one or more modified peptides are selected from a group of peptides ranked by immunogenicity based on the method according to any one of claims 1 to 13.

22. A computer-based analytical method for identifying and selecting modified peptides for use as epitopes for generating cancer vaccines, comprising: For each of a plurality of modified peptides, a binding score of the modified peptide to one or more MHC molecules is determined ( M mut ), wherein the modified peptide comprises one or more amino acid modifications at specified positions relative to a corresponding non-modified peptide, wherein the modified peptide is identical to the corresponding non-modified peptide except for the amino acid modifications, and wherein the modified peptide and the corresponding non-modified peptide are 8 to 15 amino acids in length; For each of a plurality of non-modified peptides corresponding to the modified peptide, a binding score ( M wt ); for each of the plurality of modified peptides, determining a binding score (T) for the modified peptide to one or more T cell receptors when present in an MHC-peptide complex, wherein the T score is based on chemical and physical similarity between one or more modified amino acids in the modified peptide and the amino acids at corresponding designated positions of the non-modified peptide, wherein the chemical and physical similarity is determined based on the likelihood that amino acids are interchangeable in nature; as well as Based on certain M mut 、 M wt The immunogenicity of the modified peptides was ranked by the T score, where M mut The weight is higher than T ,and T The weight is higher than M wt ; wherein the at least one candidate modified peptide is identified from the plurality of modified peptides if the at least one candidate modified peptide is more immunogenic than at least one other modified peptide in the plurality of modified peptides, wherein the modified peptide is predicted to be immunogenic if, (i) M mut indicating binding of the modified peptide to one or more MHC molecules, (ii) M wt indicates binding of the non-modified peptide to one or more MHC molecules, and (iii) T The score indicates that the modified amino acid and the non-modified amino acid are chemically and physically different; and One or more modified peptides are selected for use as epitopes for generating cancer vaccines based on ranking of immunogenicity.

Citation Information

Patent Citations

  • Vaccines containing a saponin and a sterol

    WO1996033739A1

  • Immunostimulation mediated by gene-modified dendritic cells

    WO1997024447A1

  • Individualized vaccines for cancer

    WO2012159754A2

Cited By

  • Predicting immunogenicity of T cell epitopes

    CN121130065A