Individualized vaccines for cancer

An individualized cancer vaccine targeting patient-specific mutations through neoepitopes stimulates a targeted immune response against both primary tumors and metastases, addressing the inefficacies of traditional cancer treatments.

JP2025109722APending Publication Date: 2025-07-25BIONTECH SE +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025071795
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2012-01-02
Filing Date
2025-04-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing cancer treatments are often ineffective due to the molecular diversity of tumors, with less than 25% of patients benefiting from approved therapies, and metastatic tumors often acquire additional gene mutations, making personalized immunotherapy challenging.

Method used

An individualized cancer vaccine is developed by identifying patient-specific cancer mutations through next-generation sequencing, targeting neoepitopes presented by MHC molecules, and administering RNA encoding polypeptides to stimulate a specific immune response against primary tumors and metastases.

Benefits of technology

The vaccine induces a targeted immune response against both primary tumors and metastases, leveraging mutations exclusive to the patient's cancer, reducing tumor immune escape and enhancing T cell activity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109722000017
    Figure 2025109722000017
  • Figure 2025109722000018
    Figure 2025109722000018
  • Figure 2025109722000019
    Figure 2025109722000019
Patent Text Reader

Abstract

To provide vaccines which are specific to a patient's tumor and are potentially useful for immunotherapy of a primary tumor and tumor metastasis.SOLUTION: In one aspect, the present invention relates to a method for producing individualized vaccines for cancer, comprising (a) identifying cancer specific somatic mutation in a tumor sample of a cancer patient to provide a patient's cancer mutation signature, and (b) providing a vaccine characterized by the cancer mutation signature obtained in step (a). In a further aspect, the invention relates to a vaccine that can be obtained by the method.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the provision of a vaccine that is specific for a patient's tumor and may be useful for immunotherapy of primary tumors and tumor metastases.

Background Art

[0002] Cancer is a major cause of death, accounting for one in four of all deaths. The treatment of cancer has traditionally been based on the law of averages, i.e., what works best for the majority of patients. However, due to the molecular diversity in cancer, it is often the case that less than 25% of treated individuals benefit from approved therapies. Personalized medicine based on a patient's individualized treatment is considered a possible solution to the low efficiency and high cost of innovation in drug development.

[0003] Antigen-specific immunotherapy aims to enhance or induce a specific immune response in a patient and has been successful in the control of cancer diseases. T cells play a central role in cell-mediated immunity in humans and animals. The recognition and binding of a specific antigen is mediated by the T cell receptor (TCR) expressed on the surface of T cells. The T cell receptor (TCR) of T cells can bind to major histocompatibility complex (MHC) molecules and interact with immunogenic peptides (epitopes) presented on the surface of target cells. Specific binding of the TCR initiates a signaling cascade within the T cell, leading to proliferation and differentiation into mature effector T cells.

[0004] The identification of increasing pathogens and tumor-associated antigens (TAAs) has led to a broad collection of suitable targets for immunotherapy. Cells presenting immunogenic peptides (epitopes) derived from these antigens can be specifically targeted by either active or passive immunization strategies. Active immunization can tend to induce and expand antigen-specific T cells that can specifically recognize and kill diseased cells in patients. A variety of antigen formats can be used for tumor vaccination, including whole cancer cells, proteins, peptides, or RNAs, DNAs, or viral vectors such as immunovectors, which can be administered either in vitro by pulsing DCs directly or in vivo after transfer into the patient.

[0005] Cancer can arise from the accumulation of genomic mutations and epigenetic changes, some of which may act as causative agents. In addition to tumor-associated antigens, human cancers on average carry 100 - 120 nonsynonymous mutations, many of which are targetable by vaccines. More than 95% of the mutations in tumors are private and patient-specific (Weide et al., 2008: J. Immunother. 31, 180 - 188). The number of somatic mutations that can give rise to tumor-specific T cell epitopes ranges from 30 to 400. It is predicted in silico that 40 - 60 HLA class I-restricted epitopes per patient are derived from tumor-specific somatic mutations (Azuma et al., 1993: Nature 366, 76 - 79). Furthermore, novel immunogenic HLA class II-restricted epitopes are also likely to result from tumor-associated mutations, but the number is still unknown.

[0006] It should be noted that some non-synonymous mutations are involved as a cause of neoplastic transformation, are important for the maintenance of the carcinogenic phenotype (driver mutations), and can present potential "Achilles' heels" of cancer cells. Since such non-synonymous mutations do not undergo central immune tolerance, they may be ideal candidates for the development of individual cancer vaccines. Mutations found in primary tumors may also be present during metastasis. However, several studies have demonstrated that metastatic tumors in patients acquire additional gene mutations during the evolution of individual tumors, and these are often clinically significant (Suzuki et al., 2007: Mol. Oncol. 1(2), pp. 172-180; Campbell et al., 2010: Nature 467(7319), pp. 1109-1113). Furthermore, the molecular characteristics of many metastases are also significantly different from those of primary tumors.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Patent Document 2

Patent Document 3

Non-Patent Documents

[0008]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Non-Patent Document 5

Non - Patent Document 6

Non - Patent Document 7

Non - Patent Document 8

Non - Patent Document 9

Non - Patent Document 10

Non - Patent Document 11

Non - Patent Document 12

Non - Patent Document 13

Non - Patent Document 14

Non - Patent Document 15

Non-Patent Document 25

Non-Patent Document 26

Non-Patent Document 27

Non-Patent Document 28

Non-Patent Document 29

Non-Patent Document 30

Non-Patent Document 31

Non-Patent Document 32

Non-Patent Document 33

Non-Patent Document 34

Non-Patent Document 35

Non-Patent Document 36

Non-Patent Document 37

Non-Patent Document 38

Non-Patent Document 39

Non-Patent Document 40

Non-Patent Document 41

Non-Patent Document 42

Non-Patent Document 43

Non-Patent Document 44

Non-Patent Document 45

Non-Patent Document 46

Non-Patent Document 47

Non-Patent Document 48

Non-Patent Document 49

Non-Patent Document 50

Non-Patent Document 51

Non-Patent Document 52

Non-Patent Document 53

Non-Patent Document 54

Non-Patent Document 55

Non-Patent Document 56

Non-Patent Document 57

Non-Patent Document 58

Non-Patent Document 59

Non-Patent Document 60

Non-Patent Document 61

Non-Patent Document 62

Non - Patent Document 63

Non - Patent Document 64

Non - Patent Document 65

Non - Patent Document 66

Non - Patent Document 67

Non - Patent Document 68

Non - Patent Document 69

Non - Patent Document 70

Non - Patent Document 71

Non - Patent Document 72

Non - Patent Document 73

Non - Patent Document 74

Non - Patent Document 75

Non - Patent Document 76

Non-Patent Document 77

Non-Patent Document 78

Non-Patent Document 79

Non-Patent Document 80

Non-Patent Document 81

Non-Patent Document 82

Non-Patent Document 83

Non-Patent Document 84

Non-Patent Document 85

Non-Patent Document 86

Non-Patent Document 87

Non-Patent Document 88

Non-Patent Document 89

Non-Patent Document 90

Non-Patent Document 91

Non-Patent Document 92

Non-Patent Document 93

Non-Patent Document 94

Non-Patent Document 95

Non-Patent Document 96

Summary of the Invention

Problems to be Solved by the Invention

[0009] The technical problem underlying the present invention is to provide an individualized cancer vaccine with high efficacy.

[0010] The present invention is based on the identification of patient-specific cancer mutations and the targeting of the individual cancer mutation “signatures” of the patient. Specifically, the present invention, which includes an approach to individualized immunotherapy based on the sequencing of the genome, preferably the exome, or the transcriptome, aims to immunotherapeutically target multiple individual mutations in cancer. Sequencing using next-generation sequencing (NGS) enables the rapid and cost-effective identification of patient-specific cancer mutations.

[0011] The identification of nonsynonymous point mutations that result in amino acid changes presented by the patient's major histocompatibility complex (MHC) molecules provides novel epitopes (neoepitopes) that are specific to the patient's cancer but not found in the patient's normal cells. By collecting a set of mutations from cancer cells such as circulating tumor cells (CTCs), it is possible to provide a vaccine that induces an immune response that can target the primary tumor even if it contains genetically distinct subpopulations and tumor metastases. In vaccination, such neoepitopes identified in accordance with the present application are provided to the patient in the form of a polypeptide containing the neoepitope, and after appropriate processing and presentation by MHC molecules, the neoepitopes are presented to the patient's immune system to stimulate appropriate T cells.

[0012] Preferably, such a polypeptide is provided to the patient by administering RNA encoding the polypeptide. Strategies for directly injecting in vitro transcribed RNA (IVT-RNA) into patients by various immunization routes have been successful in testing in various animal models. The protein translated and expressed in transfected cells after the RNA can be presented on MHC molecules on the cell surface after processing and induce an immune response.

[0013] The advantages of using RNA as a form of reversible gene therapy include transient expression and non-transforming characteristics. Since RNA does not need to enter the nucleus to be expressed and has no possibility of being integrated into the host genome, the risk of carcinogenesis is eliminated. The transfection rate achievable with RNA is relatively high. Furthermore, the amount of protein achieved corresponds to that in physiological expression.

[0014] The theoretical basis for the immunotherapeutic targeting of multiple individual mutations is that (i) these mutations are expressed exclusively, (ii) mutated epitopes can be expected to be ideal for T cell immunotherapy because the T cells that recognize them have not undergone thymic selection, (iii) tumor immune escape can be reduced by targeting, for example, "driver mutations" highly relevant to the tumor phenotype, and (iv) multi-epitope immune responses are more likely to bring about improved clinical benefits.

Means for Solving the Problems

[0015] The present invention relates to an efficient method for providing an individualized recombinant cancer vaccine that can induce an efficient and specific immune response in cancer patients and target primary tumors and tumor metastases. The cancer vaccine provided according to the present invention, when administered to a patient, provides a collection of epitopes presented by MHC that are specific to the patient's tumor and are appropriate for the stimulation, priming and / or expansion of T cells directed against cells expressing the antigen from which the epitopes presented by MHC are derived. Thus, the vaccine described herein can preferably induce or enhance a cellular response, preferably cytotoxic T cell activity, against a cancer disease characterized by the presentation of antigens expressed in one or more cancers by class I MHC. Since the vaccine provided according to the present invention targets cancer-specific mutations, this is specific to the patient's tumor.

[0016] In one aspect, the present invention is a method for providing an individualized cancer vaccine, comprising (a) Identifying cancer-specific somatic mutations in a tumor specimen of a cancer patient and providing a cancer mutation signature of the patient; (b) Providing a vaccine characterized by the cancer mutation signature obtained in step (a); The present invention relates to a method comprising these steps.

[0017] In one embodiment, the method of the present invention comprises: i) Providing a tumor specimen from a cancer patient and preferably a non-tumorigenous specimen derived from the cancer patient; ii) Identifying sequence differences between the genome, exome, and / or transcriptome of the tumor specimen and the genome, exome, and / or transcriptome of the non-tumorigenous specimen; iii) Designing a polypeptide comprising an epitope incorporating the sequence differences determined in step (ii); iv) Providing the polypeptide designed in step (iii) or a nucleic acid, preferably RNA, encoding said polypeptide; v) Providing a vaccine comprising the polypeptide or nucleic acid provided in step (iv). The present invention also relates to a method comprising these steps.

[0018] According to the present invention, the tumor specimen relates to any sample such as a body sample derived from a patient containing or expected to contain a tumor or cancer cells. The body sample can be any tissue sample such as blood, a tissue sample obtained from a primary tumor or tumor metastasis, or any other sample containing a tumor or cancer cells. Preferably, the body sample is blood, and the cancer-specific somatic mutations or sequence differences are determined in one or more circulating tumor cells (CTCs) contained in the blood. In another embodiment, the tumor specimen relates to a sample containing one or more isolated tumor or cancer cells such as circulating tumor cells (CTCs) or one or more isolated tumor or cancer cells such as circulating tumor cells (CTCs).

[0019] A non-tumorigenic sample relates to any sample, such as a bodily sample, from a patient or preferably another individual of the same species as the patient, preferably a healthy individual that does not contain or is expected not to contain tumors or cancer cells. The bodily sample can be any tissue sample, such as a sample from blood or non-tumorigenic tissue.

[0020] According to the present invention, the term "cancer mutation signature" can refer to all cancer mutations present in one or more cancer cells of a patient, or can refer to only a portion of the cancer mutations present in one or more cancer cells of a patient. Thus, the present invention can include the identification of all cancer-specific mutations present in one or more cancer cells of a patient, or can include the identification of only a portion of the cancer-specific mutations present in one or more cancer cells of a patient. Generally, the methods of the present invention provide the identification of several mutations that provide a sufficient number of neoepitopes to include in a vaccine. A "cancer mutation" relates to a sequence difference between the nucleic acid contained in a cancer cell and the nucleic acid contained in a normal cell.

[0021] Preferably, the mutations identified in the methods according to the present invention are non-synonymous mutations, preferably non-synonymous mutations of proteins expressed in tumors or cancer cells.

[0022] In one embodiment, the cancer-specific somatic mutations or sequence differences are determined in the genome of the tumor sample, preferably the entire genome. Thus, the methods of the present invention can include identifying the cancer mutation signature of the genome of one or more cancer cells, preferably the entire genome. In one embodiment, the step of identifying cancer-specific somatic mutations in a tumor sample of a cancer patient includes identifying a whole-genome cancer mutation profile.

[0023] In one embodiment, cancer-specific somatic mutations or sequence differences are determined in the exome of a tumor specimen, preferably the entire exome. An exome is a part of an organism's genome formed by exons, which are the coding portions of expressed genes. The exome provides the genetic blueprint used in the synthesis of proteins and other functional gene products. This is the most functionally relevant part of the genome and thus most likely to contribute to the phenotype of the organism. The exome of the human genome is estimated to constitute 1.5% of the entire genome (Ng, PC et al., PLoS Gen., 4(8): 1-15, 2008). Thus, the method of the present invention may include identifying the cancer mutation signature of the exome of one or more cancer cells, preferably the entire exome. In one embodiment, the step of identifying cancer-specific somatic mutations in a tumor specimen of a cancer patient includes identifying an entire exome cancer mutation profile.

[0024] In one embodiment, cancer-specific somatic mutations or sequence differences are determined in the transcriptome of a tumor specimen, preferably the entire transcriptome. A transcriptome is the set of all RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNAs, produced in one cell or cell population. In the context of the present invention, the transcriptome means the set of all RNA molecules produced in one cell, cell population, preferably a cancer cell population, or all cells of a given individual at a specific point in time. Thus, the method of the present invention may include identifying the cancer mutation signature of the transcriptome of one or more cancer cells, preferably the entire transcriptome. In one embodiment, the step of identifying cancer-specific somatic mutations in a tumor specimen of a cancer patient includes identifying an entire transcriptome cancer mutation profile.

[0025] In one embodiment, the step of identifying cancer-specific somatic mutations or identifying sequence differences includes single-cell sequencing of one or more, preferably 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more cancer cells. Thus, the method of the present invention may include identifying the cancer mutation signature of the one or more cancer cells. In one embodiment, the cancer cells are circulating tumor cells. Cancer cells such as circulating tumor cells can be isolated prior to single-cell sequencing.

[0026] In one embodiment, the step of identifying cancer-specific somatic mutations or identifying sequence differences includes using next-generation sequencing (NGS).

[0027] In one embodiment, the step of identifying cancer-specific somatic mutations or identifying sequence differences includes sequencing the genomic DNA and / or RNA of a tumor specimen.

[0028] To reveal cancer-specific somatic mutations or sequence differences, the sequence information obtained from the tumor specimen is preferably compared to a reference such as sequence information obtained from sequencing nucleic acids such as DNA or RNA of normal non-cancer cells such as germline cells obtained from either the patient or a different individual. In one embodiment, normal genomic germline DNA is obtained from peripheral blood mononuclear cells (PBMC).

[0029] The vaccine provided according to the method of the present invention, when administered to a patient, preferably provides a collection of epitopes presented by MHC that incorporate sequence changes based on the identified mutation or sequence difference, for example, 2 or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, preferably up to 60, up to 55, up to 50, up to 45, up to 40, up to 35 or up to 30 epitopes presented by MHC. Also, such epitopes presented by MHC that incorporate sequence changes based on the identified mutation or sequence difference are also referred to herein as "neoepitopes". Presentation of these epitopes by the patient's cells, particularly antigen-presenting cells, preferably results in targeting of the epitopes by T cells when bound to MHC, and thus the patient's tumor, preferably the primary tumor and tumor metastases, express the antigen from which the epitope presented by MHC is derived and present the same epitope on the surface of the tumor cells.

[0030] To provide a vaccine, the method of the invention may optionally comprise incorporating a sufficient number of neoepitopes (preferably in the form of coding nucleic acids) into the vaccine, or may further comprise the step of determining the utility of the identified mutations in the epitopes for cancer vaccination. Thus, the further step can comprise one or more of the following: (i) an assessment of whether the sequence change is located within an epitope presented by a known or predicted MHC, (ii) in vitro and / or in silico testing of whether the sequence change is located within an epitope presented by the MHC, for example testing whether the sequence change is part of the peptide sequence that is processed and / or presented as an epitope presented by the MHC, and (iii) an in vitro test of whether the putative mutated epitope, particularly when present in its native sequence configuration, for example when adjacent to the amino acid sequence adjacent to the epitope in a naturally occurring protein and when expressed in an antigen-presenting cell, can stimulate the patient's T cells with the desired specificity. Such adjacent sequences may each comprise 3 or more, 5 or more, 10 or more, 15 or more, 20 or more amino acids, preferably up to 50, 45, 40, 35 or 30 amino acids, and may be adjacent to the epitope sequence at the N-terminus and / or C-terminus.

[0031] The mutations or sequence differences determined according to the invention can be ranked for their utility as epitopes for cancer vaccination. Thus, in one aspect, the method of the invention comprises an analysis process, manual or computer-based, in which the identified mutations are analyzed and selected for their utility in each vaccine provided. In a preferred embodiment, the analysis process is a computer algorithm-based process. Preferably, the analysis process comprises one or more, preferably all, of the following steps: - identifying mutations that alter the expressed protein, for example by analyzing transcripts, - Identifying mutations that may be immunogenic, i.e., comparing the obtained data with available datasets of confirmed immunogenic epitopes, such as those contained in public immune epitope databases like the IMMUNE EPITOPE DATABASE AND ANALYSIS RESOURCE at http: / / www.immunoepitope.org, to identify mutations that may be immunogenic.

[0032] The step of identifying mutations that may be immunogenic may include determining and / or ranking epitopes according to their predicted MHC binding ability, preferably MHC class I binding ability.

[0033] In another embodiment of the invention, epitopes can be selected and / or ranked by using additional parameters such as protein impact, related gene expression, sequence uniqueness, predicted presentability, and association with oncogenes.

[0034] Also, multiple CTC analyses also enable the selection and prioritization of mutations. For example, mutations found in the majority of CTCs may be ranked higher than mutations found in the lower portion of CTCs.

[0035] The collection of mutation-based neoepitopes, identified according to the invention and provided by the vaccines of the invention, preferably exists in the form of a polypeptide comprising said neoepitopes (polyepitope polypeptide), or a nucleic acid encoding said polypeptide, in particular RNA. Furthermore, the neoepitope may exist in the form of a vaccine sequence in the polypeptide, i.e., in its native sequence configuration, adjacent to the amino acid sequence that is also adjacent to said epitope in, for example, a naturally occurring protein. Such adjacent sequences may each contain 5 or more, 10 or more, 15 or more, 20 or more amino acids, preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids, and may be adjacent to the epitope sequence at the N-terminus and / or C-terminus. Thus, the vaccine sequence may contain 20 or more, 25 or more, 30 or more, 35 or more, 40 or more amino acids, preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids. In one embodiment, the neoepitope and / or the vaccine sequence are arranged head-to-tail in the polypeptide.

[0036] In one embodiment, the neoepitope and / or vaccine sequences are separated by a linker, particularly a neutral linker. As used herein, the term "linker" refers to a peptide that is added between two peptide domains, such as an epitope or vaccine sequence, to connect the peptide domains. There are no particular restrictions on the linker sequence. However, the linker sequence preferably reduces steric hindrance between the two peptide domains, is well translated, and aids or permits processing of the epitope. Furthermore, the linker should have no or only few immunogenic sequence elements. The linker should preferably not create non-native neoepitopes, such as those resulting from junction stitching between neighboring neoepitopes that can give rise to unwanted immune responses. Thus, the polyepitope vaccine should preferably contain a linker sequence that can reduce the number of unwanted MHC-binding junction epitopes. Hoyt et al. (EMBO J. 25(8), 1720-1729, 2006) and Zhang et al. (J. Biol. Chem., 279(10), 8635-8641, 2004) have shown that glycine-rich sequences impair proteasomal processing, and thus the use of glycine-rich linker sequences serves to minimize the number of peptides contained in linkers that are processable by the proteasome. Furthermore, glycine has been observed to inhibit strong binding at MHC-binding groove positions (Abastado et al., J. Immunol. 151(7), 3569-3575, 1993). Schlessinger et al. (Proteins, 61(1), 115-126, 2005) discovered that the amino acids glycine and serine included in the amino acid sequence are more efficiently translated and processed by the proteasome, resulting in a more flexible protein that allows better access to the encoded neoepitope. The linker can contain 3 or more, 6 or more, 9 or more, 10 or more, 15 or more, 20 or more amino acids, preferably up to 50, up to 45, up to 40, up to 35 or up to 30 amino acids. Preferably, the linker is rich in glycine and / or serine amino acids.Preferably, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the amino acids of the linker are glycine and / or serine. In a preferred embodiment, the linker is substantially composed of the amino acids glycine and serine. In one embodiment, the linker has the amino acid sequence (GGS). a (GSS) b (GGG) c (SSG) d (GSG) e and wherein a, b, c, d, and e are independently numbers selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, and a + b + c + d + e is different from 0, preferably 2 or more, 3 or more, 4 or more, or 5 or more. In one embodiment, the linker includes sequences described herein including, for example, the linker sequence GGSGGGGSG described in the Examples.

[0037] In another embodiment of the invention, the collection of mutation-based neoepitopes identified according to the invention and provided by the vaccine of the invention is preferably in the form of a collection of polypeptides comprising said neoepitopes on various polypeptides (each of said polypeptides comprising one or more neoepitopes, which may also overlap), or a collection of nucleic acids, particularly RNA, encoding said polypeptides.

[0038] In a particularly preferred embodiment, the polyepitope polypeptide according to the invention is administered to a patient in the form of a nucleic acid capable of being expressed in the patient's cells, such as antigen-presenting cells, to produce the polypeptide, preferably in the form of RNA, such as in vitro transcribed or synthetic RNA. The invention also contemplates administering one or more multi-epitope polypeptides included by the term "polyepitope polypeptide" for the purposes of the present invention, preferably in the form of a nucleic acid capable of being expressed in the patient's cells, such as antigen-presenting cells, to produce one or more polypeptides, preferably in the form of RNA, such as in vitro transcribed or synthetic RNA. When administering multiple multi-epitope polypeptides, the neoepitopes provided by different multi-epitope polypeptides may be different or partially overlapping. Once present in the patient's cells, such as antigen-presenting cells, the polypeptide according to the invention is processed to produce the neoepitopes identified in accordance with the present invention. Administration of the vaccine provided in accordance with the present invention can provide epitopes presented by MHC class II that can induce a CD4+ helper T cell response against cells expressing the antigen from which the epitopes presented by MHC are derived. Alternatively or in addition, administration of the vaccine provided in accordance with the present invention can provide epitopes presented by MHC class I that can induce a CD8+ T cell response against cells expressing the antigen from which the epitopes presented by MHC are derived. Furthermore, administration of the vaccine provided in accordance with the present invention can provide one or more neoepitopes (including known neoepitopes and neoepitopes identified in accordance with the present invention), as well as one or more epitopes that do not contain cancer-specific somatic mutations but are expressed by cancer cells and preferably induce an immune response against cancer cells, preferably a cancer-specific immune response.In one embodiment, administration of the vaccine provided according to the present invention can induce a CD4+ helper T cell response against cells expressing an antigen that is an epitope presented by MHC class II and / or an antigen from which the epitope presented by MHC is derived, a neoepitope, and that does not contain cancer-specific somatic mutations, and can induce a CD8+ T cell response against cells expressing an antigen that is an epitope presented by MHC class I and / or an antigen from which the epitope presented by MHC is derived. In one embodiment, the epitope that does not contain cancer-specific somatic mutations is derived from a tumor antigen. In one embodiment, the neoepitope and epitope that do not contain cancer-specific somatic mutations have a synergistic effect in the treatment of cancer. Preferably, the vaccine provided according to the present invention is useful for polyepitope stimulation of cytotoxic and / or helper T cell responses.

[0039] In a further aspect, the present invention provides a vaccine obtainable by the method according to the present invention. Accordingly, the present invention relates to a vaccine comprising a recombinant polypeptide comprising a mutation-based neoepitope or a nucleic acid encoding said polypeptide, wherein said epitope is generated by cancer-specific somatic mutations in a tumor specimen of a cancer patient. Such a recombinant polypeptide may also include epitopes that do not contain the above-described cancer-specific somatic mutations.

[0040] Preferred embodiments of such vaccines are described above in the context of the method of the present invention.

[0041] The vaccine provided according to the present invention may contain a pharmaceutically acceptable carrier and optionally one or more adjuvants, stabilizers, etc. The vaccine can be in the form of a therapeutic vaccine or a prophylactic vaccine.

[0042] Another aspect relates to a method of inducing an immune response in a patient, comprising administering to the patient a vaccine provided according to the present invention.

[0043] Another aspect is a method of treating a cancer patient, comprising: (a) providing an individualized cancer vaccine by the method according to the present invention; and (b) administering said vaccine to the patient. The present invention also relates to a method comprising the above steps.

[0044] Another aspect is a method of treating a cancer patient, comprising administering to the patient a vaccine according to the present invention.

[0045] In a further aspect, the present invention provides a vaccine as described herein for use in the treatment methods described herein, particularly for the treatment or prevention of cancer.

[0046] The cancer treatment described herein can be combined with surgical resection and / or radiation and / or conventional chemotherapy.

[0047] Another aspect of the present invention is a method of determining a false discovery rate based on next-generation sequencing data, comprising: collecting a first sample of genetic material from an animal or a human; collecting a second sample of genetic material from the animal or the human; collecting a first sample of genetic material from tumor cells; collecting a second sample of genetic material from said tumor cells; determining a common coverage tumor comparison by counting all bases of a reference genome included in both the tumor and at least one of said first sample of genetic material from the animal or human and said second sample of genetic material from the animal or human; determining a same vs. same comparison by counting all bases of a reference genome included in both said first sample of genetic material from the animal or human and said second sample of genetic material from the animal or human; Forming normalization by dividing the common coverage tumor comparison by the common coverage versus identical comparison; Determining the false discovery rate by: 1) dividing the number of single nucleotide variations having a quality score higher than Q in the comparison of the first sample of genetic material from an animal or human with the second sample of genetic material from an animal or human by 2) the number of single nucleotide variations having a quality score higher than Q in the comparison of the first sample of genetic material from the tumor cells with the second sample of genetic material from the tumor cells, and 3) multiplying the result by the normalization; relates to a method comprising.

[0048] In one embodiment, the genetic material is DNA.

[0049] In one embodiment, Q is establishing a set of quality characteristics S = (s1,..., s n ) such that S is preferred over T = (t1,..., t i ) when s i > t n for all i = 1,..., n, as indicated by S > T; Defining an intermediate false discovery rate by: 1) dividing the number of single nucleotide variations having a quality score S > T in the comparison of the first DNA sample from an animal or human with the second DNA sample from an animal or human by 2) the number of single nucleotide variations having a quality score S > T in the comparison of the first DNA sample from the tumor cells with the second DNA sample from the tumor cells, and 3) multiplying the result by the normalization; Determining the value range of each of the m mutations, each having n quality characteristics; Sampling up to p values from the value range; Creating each possible combination of the sampled quality values, resulting in p n data points; Using the random sample of the data points as predictors for random forest training; Using the corresponding intermediate false discovery rate value as the response for the random forest training; The regression score of the random forest training determined thereby is Q.

[0050] In one embodiment, the second DNA sample from an animal or a human is allogeneic to the first DNA sample from the animal or the human. In one embodiment, the second DNA sample from an animal or a human is autologous to the first DNA sample from the animal or the human. In one embodiment, the second DNA sample from an animal or a human is xenogeneic to the first DNA sample from the animal or the human.

[0051] In one embodiment, the genetic material is RNA.

[0052] In one embodiment, Q is A step of establishing a set of quality characteristics S = (s1,..., s n ), where S is more preferable than T = (t1,..., t i ) when s i > t for all i = 1,..., n, indicated by S > T; n 1) Dividing the number of single nucleotide mutations having a quality score S > T in the comparison between the first RNA sample from an animal or a human and the second RNA sample from an animal or a human by the number of single nucleotide mutations having a quality score S > T in the comparison between the first RNA sample from the tumor cells and the second RNA sample from the tumor cells, and 2) multiplying the result by the normalization to define the intermediate false discovery rate; Determining the value range of each characteristic for m mutations each having n quality characteristics; Sampling up to p values from the value range; ​Create each possible combination of the sampled quality values, and p n resulting in n data points, and using a random sample of the data points as predictors for random forest training, and using the corresponding intermediate false discovery rate value as the response for the random forest training, and determined by, the resulting regression score of the random forest training is Q.

[0053] In one embodiment, the second RNA sample from the animal or human is allogeneic to the first RNA sample from the animal or human. In one embodiment, the second RNA sample from the animal or human is autologous to the first RNA sample from the animal or human. In one embodiment, the second RNA sample from the animal or human is heterologous to the first RNA sample from the animal or human.

[0054] In one embodiment, the false discovery rate is used to prepare a vaccine formulation. In one embodiment, the vaccine is deliverable intravenously. In one embodiment, the vaccine is deliverable transdermally. In one embodiment, the vaccine is deliverable intramuscularly. In one embodiment, the vaccine is deliverable subcutaneously. In one embodiment, the vaccine is customized for a particular patient.

[0055] In one embodiment, one of the first sample of genetic material from the animal or human and the second sample of genetic material from the animal or human is from the particular patient.

[0056] In one embodiment, the step of determining a common coverage tumor comparison by counting all bases of a reference genome included in both the tumor and at least one of the first sample of genetic material from the animal or human and the second sample of genetic material from the animal or human uses an automated system to count all bases.

[0057] In one embodiment, the step of determining the common coverage to identity comparison by counting all bases of a reference genome included in both the first sample of genetic material from an animal or a human and the second sample of genetic material from an animal or a human uses the automated system.

[0058] In one embodiment, the step of forming a normalization by dividing the common coverage tumor comparison by the common coverage to identity comparison uses the automated system.

[0059] In one embodiment, the step of determining the false discovery rate by: 1) dividing the number of single nucleotide variants having a quality score higher than Q in the comparison of the first sample of genetic material from an animal or a human and the second sample of genetic material from an animal or a human, 2) by the number of single nucleotide variants having a quality score higher than Q in the comparison of the first sample of genetic material from the tumor cells and the second sample of genetic material from the tumor cells, and 3) multiplying the result by the normalization uses the automated system.

[0060] Another aspect of the invention is a method of determining a putative receiver operating characteristic (ROC) curve, comprising: receiving a dataset of mutations, each mutation having an associated false discovery rate (FDR); for each mutation: determining a true positive rate (TPR) by subtracting the FDR from 1; determining a false positive rate (FPR) by setting the FPR equal to the FDR; forming a putative ROC by plotting, for each mutation, a point of the cumulative TPR and FPR values up to that mutation, divided by the sum of all TPR and FPR values. The method is related to.

[0061] Other features and advantages of the invention will be apparent from the following detailed description and claims.

Best Mode for Carrying Out the Invention

[0062] The present invention will be described in detail below. However, since methods, protocols, and reagents can vary, it should be understood that the present invention is not limited to the specific methods, protocols, and reagents described herein. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the present invention, which is limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0063] The elements of the present invention are described below. These elements are described using specific embodiments, but it should be understood that additional embodiments can be created by combining them in any manner and in any number. The various examples and preferred embodiments described should not be construed as limiting the present invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments that combine the explicitly described embodiments with any number of disclosed and / or preferred elements. Further, unless otherwise indicated by the context, any permutation and combination of all the described elements in this application should be considered to be disclosed by the description of this application. For example, in one preferred embodiment, the RNA contains a poly(A)-tail consisting of 120 nucleotides, and in another preferred embodiment, the RNA molecule contains a 5'-cap analog. In a preferred embodiment, the RNA contains a poly(A)-tail consisting of 120 nucleotides and a 5'-cap analog.

[0064] Preferably, the terms used herein are defined as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", edited by H.G.W. Leuenberger, B. Nagel, and H. Kolbl, (1995) Helvetica Chimica Acta, CH-4010, Basel, Switzerland.

[0065] Unless otherwise specified, conventional methods of biochemistry, cell biology, immunology, and recombinant DNA techniques described in the literature of the art (see, for example, Molecular Cloning: A Laboratory Manual, 2nd edition, edited by J. Sambrook et al., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989) are used in the practice of the present invention.

[0066] Unless the context requires otherwise, throughout this specification and the following claims, the word "comprise", and variations such as "comprises" and "comprising", are to be interpreted as implying the inclusion of the stated member, integer or step or group of members, integers or steps but not the exclusion of any other member, integer or step or group of members, integers or steps, provided that in some embodiments such other members, integers or steps or group of members, integers or steps may be excluded, that is, the subject is to be understood as consisting of the inclusion of the stated member, integer or step or group of members, integers or steps. The terms "a" and "an" and "the" and similar references used in connection with the description of the present invention (especially in connection with the claims) are to be construed to include both the singular and the plural unless otherwise specified herein or clearly contradicted by the context. The recitation of a range of values herein is merely intended to serve as a shorthand method of referring individually to each separate value within the range. Unless otherwise specified herein, each individual value is incorporated herein as if it were individually recited herein.

[0067] All of the methods described herein can be performed in any suitable order unless otherwise specified herein or clearly contradicted by the context. The use of any and all examples, or exemplary language (e.g., "such as") provided herein is intended merely to better illustrate the invention and does not pose a limitation on the scope of the invention as claimed elsewhere. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0068] Several documents are cited throughout the text of this specification. Each of the documents cited herein, whether supra or infra, including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc., is incorporated herein by reference in its entirety. Nothing herein should be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention.

[0069] The vaccines provided according to the present invention are recombinant vaccines.

[0070] The term "recombinant" in the context of the present invention means "created by genetic engineering". Preferably, a "recombinant entity" such as a recombinant polypeptide in the context of the present invention does not exist in nature and is preferably the result of a combination of entities such as amino acid or nucleic acid sequences that are not combined in nature. For example, a recombinant polypeptide in the context of the present invention may contain several amino acid sequences, such as neoepitopes or vaccine sequences derived from different proteins or different parts of the same protein fused together, for example by peptide bonds or suitable linkers.

[0071] As used herein, the term "naturally occurring" refers to an object that can be found in nature. For example, a peptide or nucleic acid that is present in an organism (including viruses), can be isolated from a natural source, and has not been intentionally modified in a laboratory by man, is naturally occurring.

[0072] According to the present invention, the term "vaccine" refers to a pharmaceutical preparation (pharmaceutical composition) or product that, upon administration, induces an immune response, in particular a cellular immune response, that recognizes and attacks pathogens or diseased cells, such as cancer cells. Vaccines can be used to prevent or treat diseases. The term "personalized cancer vaccine" refers to a specific cancer patient, meaning that the cancer vaccine is adapted to the needs or special circumstances of an individual cancer patient.

[0073] The term "immune response" refers to an integrated bodily response to an antigen, preferably a cellular immune response or a cellular and humoral immune response. The immune response can be protective / preventive and / or therapeutic.

[0074] "Inducing an immune response" can mean that there was no immune response to a specific antigen prior to induction, but it can also mean that there was a certain level of immune response to a specific antigen prior to induction and that the immune response was enhanced after induction. Thus, "inducing an immune response" includes "enhancing an immune response". Preferably, after inducing an immune response in a subject, the subject is protected from developing a disease such as a cancer disease, or the medical condition is alleviated by inducing the immune response. For example, an immune response against a tumor-expressed antigen can be induced in a patient suffering from a cancer disease or a subject at risk of developing a cancer disease. In this case, inducing an immune response can mean that the medical condition of the subject is alleviated, the subject does not develop metastases, or a subject at risk of developing a cancer disease does not develop a cancer disease.

[0075] "Cellular immune response", "cellular response", "cellular response to an antigen" or similar terms mean that they include a cellular response directed at cells, characterized by the presentation of an antigen using class I or class II MHC. The cellular response relates to cells called T cells or T-lymphocytes that act as either "helper" or "killer". Helper T cells (also called CD4 + T cells) play a central role by regulating the immune response, and killer cells (cytotoxic T cells, cytolytic T cells, also called CD8 + T cells or CTLs) kill diseased cells such as cancer cells and prevent the production of further diseased cells. In a preferred embodiment, the present invention includes stimulating an anti-tumor CTL response against tumor cells expressing one or more tumor-expressed antigens and preferably presenting such tumor-expressed antigens using class I MHC.

[0076] As used herein, the term "antigen" includes any substance that elicits an immune response. In particular, the term "antigen" relates to any substance that specifically reacts with an antibody or a T-lymphocyte (T cell), preferably a peptide or a protein. According to the present invention, the term "antigen" includes any molecule that contains at least one epitope. Preferably, the antigen is, in the context of the invention, a molecule that optionally, after processing, preferably induces an immune reaction that is specific for the antigen (including the cells expressing the antigen). According to the present invention, any suitable antigen that is a candidate for an immune reaction may be used, and the immune reaction is preferably a cellular immune reaction. In the context of embodiments of the present invention, the antigen is preferably presented by a cell, preferably an antigen-presenting cell, in the context of an MHC molecule, which includes diseased cells, particularly cancer cells, and which results in an immune reaction against the antigen. The antigen is preferably a product that corresponds to or is derived from a naturally occurring antigen. Such naturally occurring antigens include tumor antigens.

[0077] In a preferred embodiment, the antigen is a tumor antigen, i.e., a part of a tumor cell such as a protein or peptide expressed in a tumor cell, which can be derived from the cytoplasm, cell surface or cell nucleus, particularly those mainly present intracellularly or as surface antigens of tumor cells. For example, tumor antigens include carcinoembryonic antigen, α1-fetoprotein, isofetoprotein, and fetal sulfoglycoprotein, α2-H-iron protein and γ-fetoprotein. According to the present invention, the tumor antigen preferably includes any antigen that is expressed in a tumor or cancer and tumor or cancer cells and is optionally characteristic in terms of type and / or expression level. In one embodiment, the term "tumor antigen" or "tumor-associated antigen" relates to a protein that is specifically expressed in a limited number of tissues and / or organs or at a specific developmental stage under normal conditions. For example, the tumor antigen may be specifically expressed in gastric tissue, preferably gastric mucosa, in the genital organs, such as the testis, in trophoblast tissue, such as the placenta, or in germ line cells under normal conditions, and is expressed or abnormally expressed in one or more tumor or cancer tissues. In this context, "a limited number" preferably means three or less, more preferably two or less. In the context of the present invention, tumor antigens include, for example, differentiation antigens, preferably cell type-specific differentiation antigens, i.e., proteins that are specifically expressed at a specific differentiation stage in a specific cell type under normal conditions, cancer / testis antigens, i.e., proteins that are specifically expressed in the testis and optionally in the placenta under normal conditions, and germ line-specific antigens. Preferably, the tumor antigen or the abnormal expression of the tumor antigen identifies cancer cells. In the context of the present invention, the tumor antigen expressed by cancer cells in a subject, such as a patient suffering from a cancer disease, is preferably a self-protein in the subject. In a preferred embodiment, the tumor antigen, in the context of the present invention, is specifically expressed under normal conditions in a tissue or organ that is non-essential, i.e., a tissue or organ that does not cause the death of the subject when damaged by the immune system, or in a body organ or structure that is not accessible or only slightly accessible by the immune system.

[0078] According to the present invention, the terms "tumor antigen", "tumor-expressed antigen", "cancer antigen" and "antigen expressed in cancer" are equivalents and are used interchangeably herein.

[0079] The term "immunogenicity" relates to the relative effectiveness of an antigen in inducing an immune response.

[0080] The "antigenic peptide" according to the present invention preferably relates to a part or fragment of an antigen that can stimulate an immune response, preferably a cellular response, against the antigen or by the expression of the antigen, preferably against cells characterized by the presentation of the antigen, particularly diseased cells such as cancer cells. Preferably, the antigenic peptide can stimulate a cellular response against cells characterized by the presentation of the antigen using class I MHC, and preferably can stimulate antigen-responsive cytotoxic T-lymphocytes (CTL). Preferably, the antigenic peptide according to the present invention is a peptide presented by MHC class I and / or class II, or can be processed to produce a peptide presented by MHC class I and / or class II. Preferably, the antigenic peptide contains an amino acid sequence substantially corresponding to the amino acid sequence of a fragment of the antigen. Preferably, the fragment of the antigen is a peptide presented by MHC class I and / or class II. Preferably, the antigenic peptide according to the present invention contains an amino acid sequence substantially corresponding to the amino acid sequence of such a fragment and is processed to produce such a fragment, i.e., a peptide presented by MHC class I and / or class II derived from the antigen.

[0081] When the peptide is presented directly, i.e., without processing, particularly without cleavage, it has a length suitable for binding to MHC molecules, particularly class I MHC molecules, preferably a length of 7 to 20 amino acids, more preferably a length of 7 to 12 amino acids, more preferably a length of 8 to 11 amino acids, particularly a length of 9 or 10 amino acids.

[0082] When the peptide is part of a larger entity, such as a vaccine sequence or an additional sequence of a polypeptide, and is to be presented after processing, particularly after cleavage, the peptide produced by the processing has a length suitable for binding to MHC molecules, particularly class I MHC molecules, preferably a length of 7 to 20 amino acids, more preferably a length of 7 to 12 amino acids, more preferably a length of 8 to 11 amino acids, particularly a length of 9 or 10 amino acids. Preferably, the sequence of the peptide to be presented after processing is derived from the amino acid sequence of the antigen, i.e., its sequence substantially corresponds to, and preferably is identical to, a fragment of the antigen. Thus, in one embodiment, the antigen peptide or vaccine sequence according to the present invention comprises a sequence having a length of 7 to 20 amino acids, more preferably a length of 7 to 12 amino acids, more preferably a length of 8 to 11 amino acids, particularly a length of 9 or 10 amino acids, which substantially corresponds to, and preferably is identical to, a fragment of the antigen, and which constitutes the peptide presented after processing of the antigen peptide or vaccine sequence. According to the present invention, such peptides produced by processing contain the identified sequence variations.

[0083] According to the present invention, the antigen peptide or epitope may be present in the vaccine as part of a larger entity, such as a vaccine sequence and / or polypeptide comprising a plurality of antigen peptides or epitopes. The presented antigen peptide or epitope is produced after appropriate processing.

[0084] Peptides having an amino acid sequence substantially corresponding to the sequence of a peptide presented by class I MHC may differ by one or more residues that are not essential for TCR recognition of the peptide presented by class I MHC or for peptide binding to MHC. Such substantially corresponding peptides can also stimulate antigen-responsive CTLs and can be considered immunologically equivalent. Peptides having an amino acid sequence different from that of the presented peptide at residues that do not affect TCR recognition but improve the stability of binding to MHC may improve the immunogenicity of the antigen peptide and can be referred to herein as "optimized peptides". A rational approach can be used to design substantially corresponding peptides using existing knowledge regarding which of these residues are likely to affect binding to either MHC or TCR. The resulting functional peptides are contemplated as antigen peptides.

[0085] Antigen peptides should be recognizable by the T cell receptor when presented by MHC. Preferably, when recognized by the T cell receptor, antigen peptides can induce clonal expansion of T cells bearing a T cell receptor that specifically recognizes the antigen peptide in the presence of appropriate co-stimulatory signals. Preferably, antigen peptides can stimulate an immune response, preferably a cellular response, against the cells from which they are derived or characterized by the expression of the antigen, preferably characterized by the presentation of the antigen, particularly when presented in the context of MHC molecules. Preferably, antigen peptides can stimulate a cellular response against cells characterized by antigen presentation using class I MHC and preferably can stimulate antigen-responsive CTLs. Such cells are preferably target cells.

[0086] "Antigen processing" or "processing" refers to the degradation of a polypeptide or antigen into processing products of fragments of said polypeptide or antigen (e.g., degradation of a polypeptide into peptides), and the association (e.g., by binding) of one or more of these fragments with MHC molecules for presentation to specific T cells by a cell, preferably an antigen-presenting cell.

[0087] "Antigen-presenting cell" (APC) refers to a cell that presents peptide fragments of protein antigens associated with MHC molecules on its cell surface. Some APCs can activate antigen-specific T cells.

[0088] Professional antigen-presenting cells are very efficient at internalizing antigens either by phagocytosis or receptor-mediated endocytosis and then displaying fragments of the antigen bound to class II MHC molecules on their membranes. T cells recognize and interact with the antigen-class II MHC molecule complex on the membrane of the antigen-presenting cell. Subsequently, additional co-stimulatory signals are generated by the antigen-presenting cell, leading to the activation of the T cell. The expression of co-stimulatory molecules is a defining feature of professional antigen-presenting cells.

[0089] The main types of professional antigen-presenting cells are dendritic cells, macrophages, B-cells, and certain activated epithelial cells, which have the broadest range of antigen presentation and are probably the most important antigen-presenting cells.

[0090] Dendritic cells (DCs) are a population of leukocytes that present antigens captured in peripheral tissues to T cells by both the MHC class II and I antigen presentation pathways. It is well known that dendritic cells are powerful inducers of the immune response and that the activation of these cells is an important step in the induction of anti-tumor immunity.

[0091] Dendritic cells can be conveniently classified into "immature" and "mature" cells, which can be used as a simple way to distinguish between two well-characterized phenotypes. However, this nomenclature should not be interpreted as excluding all possible intermediate stages of differentiation.

[0092] Immature dendritic cells are characterized as antigen-presenting cells with a high capacity for antigen uptake and processing, which correlates with high expression of Fcγ receptors and mannose receptors. The mature phenotype typically has lower expression of these markers but is characterized by high expression of cell surface molecules that govern T cell activation, such as class I and class II MHC, adhesion molecules (e.g., CD54 and CD11), and costimulatory molecules (e.g., CD40, CD80, CD86, and 4-1BB).

[0093] Maturation of dendritic cells refers to the state of dendritic cell activation in which such antigen-presenting dendritic cells bring about the primary stimulation of T cells, while presentation by immature dendritic cells results in tolerance. Maturation of dendritic cells is mainly induced by biomolecules with microbial features detected by innate receptors (such as bacterial DNA, viral RNA, endotoxin, etc.), pro-inflammatory cytokines (TNF, IL-1, IFN), ligation of CD40 on the dendritic cell surface by CD40L, and substances released from cells undergoing stress-induced cell death. Dendritic cells can be induced by culturing bone marrow cells with cytokines such as granulocyte-macrophage colony-stimulating factor (GM-CSF) and tumor necrosis factor alpha in vitro.

[0094] Non-professional antigen-presenting cells do not constitutively express MHC class II proteins required for interaction with naive T cells. These are expressed only upon stimulation of non-professional antigen-presenting cells by specific cytokines such as IFNγ.

[0095] "Antigen-presenting cells" can be loaded with peptides presented by MHC class I by transducing cells with a nucleic acid encoding a peptide or a polypeptide containing the peptide to be presented, preferably RNA, such as a nucleic acid encoding an antigen.

[0096] In some embodiments, a pharmaceutical composition of the invention comprising a gene delivery vehicle targeting dendritic cells or other antigen-presenting cells can be administered to a patient to effect in vivo transduction. For example, in vivo transduction of dendritic cells can generally be performed using any method known in the art, such as those described in WO 97 / 24447 or the gene gun technique described by Mahvi et al., Immunology and cell Biology 75:456-460, 1997.

[0097] According to the present invention, the term "antigen-presenting cell" also includes target cells.

[0098] "Target cells" are meant to mean cells that are the target of an immune response, such as a cellular immune response. Target cells include cells that present an antigen or antigen epitope, i.e., a peptide fragment derived from an antigen, and include all undesirable cells such as cancer cells. In a preferred embodiment, target cells are cells that express the antigen described herein and preferably present said antigen using class I MHC.

[0099] The term "epitope" refers to an antigenic determinant in a molecule such as an antigen, i.e., a part or a fragment of a molecule that is recognized by the immune system, for example, recognized by T cells, when presented particularly in the context of association with MHC molecules. Epitopes of proteins such as tumor antigens preferably comprise a continuous or discontinuous portion of said protein, and preferably have a length of 5 to 100, preferably 5 to 50, more preferably 8 to 30, and most preferably 10 to 25 amino acids. For example, an epitope can preferably have a length of 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. It is particularly preferred that the epitope in the context of the present invention is a T cell epitope.

[0100] According to the present invention, an epitope may bind to an MHC molecule such as an MHC molecule on the surface of a cell, and thus can be an "MHC-binding peptide" or an "antigenic peptide". The term "MHC-binding peptide" relates to a peptide that binds to MHC class I and / or MHC class II molecules. In the case of a class I MHC / peptide complex, the binding peptide typically has a length of 8 to 10 amino acids, although longer or shorter peptides can be effective. In the case of a class II MHC / peptide complex, the binding peptide typically has a length of 10 to 25 amino acids, particularly 13 to 18 amino acids, while longer or shorter peptides can be effective.

[0101] The terms "epitope", "antigenic peptide", "antigenic epitope", "immunogenic peptide" and "MHC-binding peptide" are used interchangeably herein, and preferably relate to an incomplete presentation of an antigen that can preferably induce an immune response against the antigen, preferably an antigen-expressing or antigen-containing, preferably presenting cell.

[0102] Preferably, this term relates to the immunogenic portion of an antigen. Preferably, this is a portion of the antigen that is recognized (i.e., specifically bound) by a T cell receptor when presented, particularly in the context of association with MHC molecules. Preferred such immunogenic portions bind to MHC class I or class II molecules. An immunogenic portion as used herein is said to "bind" to an MHC class I or class II molecule if such binding can be detected using any assay known in the art.

[0103] As used herein, the term "neoepitope" refers to an epitope that is not present in a reference such as normal non-cancerous cells or germline cells, but is found in cancer cells. This includes, in particular, situations where a corresponding epitope is found in normal non-cancerous cells or germline cells, but a neoepitope is generated by a change in the sequence of the epitope due to one or more mutations in the cancer cells.

[0104] The term "portion" refers to a fraction. With respect to a specific structure such as an amino acid sequence or a protein, the term, its "portion", can designate a continuous or discontinuous fraction of said structure. Preferably, a portion of an amino acid sequence comprises at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, preferably at least 40%, preferably at least 50%, more preferably at least 60%, more preferably at least 70%, even more preferably at least 80%, most preferably at least 90% of the amino acids of said amino acid sequence. Preferably, when a portion is a discontinuous fraction, said discontinuous fraction is composed of 2, 3, 4, 5, 6, 7, 8, or more parts of the structure, each part being a continuous element of the structure. For example, a discontinuous fraction of an amino acid sequence may be composed of 2, 3, 4, 5, 6, 7, 8, or more, preferably 4 or less parts of said amino acid sequence, each part preferably comprising at least 5 consecutive amino acids, at least 10 consecutive amino acids, preferably at least 20 consecutive amino acids, preferably at least 30 consecutive amino acids of the amino acid sequence.

[0105] The terms "part" and "fragment" are used interchangeably herein and refer to a continuous element. For example, a part of a structure such as an amino acid sequence or a protein refers to a continuous element of said structure. A part, portion or fragment of a structure preferably comprises one or more functional properties of said structure. For example, a part, portion or fragment of an epitope, peptide or protein is preferably immunologically equivalent to the epitope, peptide or protein from which it is derived. In the context of the present invention, a "part" of a structure such as an amino acid sequence preferably comprises at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98%, at least 99% of the whole structure or amino acid sequence and preferably consists of it.

[0106] The term "immunoreactive cell" in the context of the present invention relates to cells that exert effector functions during an immune response. An "immunoreactive cell" is preferably a cell characterized by an antigen or antigen presentation, or a cell that can bind to an antigen peptide derived from an antigen and mediate an immune response. For example, such cells secrete cytokines and / or chemokines, secrete antibodies, recognize cancer cells, and optionally eliminate such cells. For example, immunoreactive cells include T cells (cytotoxic T cells, helper T cells, tumor-infiltrating T cells), B cells, natural killer cells, neutrophils, macrophages, and dendritic cells. Preferably, in the context of the present invention, the "immunoreactive cell" is a T cell, preferably CD4 + and / or CD8 + T cell.

[0107] Preferably, the "immunoreactive cell" recognizes an antigen or an antigen peptide derived from an antigen with a certain degree of specificity, particularly when presented on the surface of antigen-presenting cells or diseased cells such as cancer cells in the context of MHC molecules. Preferably, the recognition enables the cell that recognizes the antigen or the antigen peptide derived from the antigen to become responsive or reactive. When the cell is a helper T cell (CD4 + T cell) that possesses a receptor for recognizing an antigen or an antigen peptide derived from an antigen in the context of MHC class II molecules, such responsiveness or reactivity is the release of cytokines and / or CD8 +It may involve the activation of lymphocytes (CTL) and / or B cells. When the cell is a CTL, such responsiveness or reactivity may involve the elimination of cells presented in the context of MHC class I molecules, i.e., cells characterized by antigen presentation using class I MHC, such as by apoptosis or perforin-mediated cytolysis. According to the present invention, CTL responsiveness may include sustained calcium flux, cell division, production of cytokines such as IFN-γ and TNF-α, upregulation of activation markers such as CD44 and CD69, and specific cytolytic killing of antigen-expressing target cells. Also, CTL responsiveness can be determined using an artificial reporter that accurately indicates CTL responsiveness. CTLs that recognize an antigen or an antigen-derived antigen peptide and are responsive or reactive are also referred to herein as "antigen-responsive CTLs". When the cell is a B cell, such responsiveness may include the release of immunoglobulins.

[0108] The terms "T cell" and "T lymphocyte" are used interchangeably herein and include cytotoxic T cells (CTL, CD8+ T cells) including T helper cells (CD4+ T cells).

[0109] T cells belong to a group of white blood cells known as lymphocytes and play a central role in cell-mediated immunity. They can be distinguished from other lymphocyte populations, such as B cells and natural killer cells, by the presence on their cell surface of a special receptor called the T cell receptor (TCR). The thymus is the major organ that governs the maturation of T cells. Several different subsets of T cells have been discovered, each with clearly distinct functions.

[0110] T helper cells assist other white blood cells in immune processes, including, among other functions, the maturation of B cells into plasma cells and the activation of cytotoxic T cells and macrophages. These cells are also known as CD4+ T cells because they express the CD4 protein on their surface. Helper T cells are activated when peptide antigens are presented by MHC class II molecules expressed on the surface of antigen-presenting cells (APCs). After activation, they rapidly divide and secrete small proteins called cytokines that regulate or assist the active immune response.

[0111] Cytotoxic T cells destroy virus-infected cells and tumor cells and are also associated with transplant rejection. These cells are also known as CD8+ T cells because they express the CD8 glycoprotein on their surface. These cells recognize their targets by binding to antigens associated with MHC class I, which is present on the surface of almost all cells in the body.

[0112] In the majority of T cells, the T cell receptor (TCR) exists as a complex of several proteins. The actual T cell receptor is composed of two separate peptide chains produced from the independent T cell receptor alpha and beta (TCRα and TCRβ) genes and is called the α- and β-TCR chains. Gamma delta T cells (γδ T cells) represent a small subset of T cells that possess a distinctly different T cell receptor (TCR) on their surface. However, in γδ T cells, the TCR consists of one γ-chain and one δ-chain. This group of T cells is much rarer than αβ T cells (2% of all T cells).

[0113] The first signal in T cell activation is provided by the binding of the T cell receptor to short peptides presented by the major histocompatibility complex (MHC) on another cell. This ensures that only T cells with a TCR specific for that peptide are activated. The partner cell is usually a professional antigen-presenting cell (APC), typically a dendritic cell in the case of a naive response, although B cells and macrophages can be important APCs. Peptides presented to CD8+ T cells by MHC class I molecules are typically 8 - 10 amino acids in length. Since the ends of the binding cleft of MHC class II molecules are open, peptides presented to CD4+ T cells by MHC class II molecules are typically longer.

[0114] According to the present invention, the T cell receptor has a significant affinity for a pre-determined target in a standard assay and can bind to the pre-determined target when binding to the pre-determined target. "Affinity" or "binding affinity" is often measured by the equilibrium dissociation constant (K D ). The T cell receptor does not have a significant affinity for a target in a standard assay and cannot (substantially) bind to the target when not significantly binding to the target.

[0115] The T cell receptor can preferably specifically bind to a pre-determined target. The T cell receptor can bind to a pre-determined target while not being able to (substantially) bind to other targets, i.e., not having a significant affinity for other targets in a standard assay and not significantly binding to other targets, and is specific for the pre-determined target.

[0116] Cytotoxic T lymphocytes can be produced in vivo by causing an antigen or antigen peptide to be taken up in vivo by an antigen-presenting cell. The antigen or antigen peptide can be represented as a protein, as DNA (e.g., within a vector), or as RNA. The antigen can be processed to yield a peptide partner of an MHC molecule, while its fragments can be presented without further processing, particularly in the latter situation when they can bind to MHC molecules. Generally, administration to a patient by intradermal injection is possible. However, intranodal injection into lymph nodes can also be performed (Maloy et al. (2001), Proc Natl Acad Sci USA 98: 3299 - 303). The resulting cells present the complex of interest and are recognized by autologous cytotoxic T lymphocytes, which then proliferate.

[0117] Specific activation of CD4+ or CD8+ T cells can be detected in various ways. Methods for detecting specific T cell activation include detecting the proliferation of T cells, the production of cytokines (e.g., lymphokines), or the development of cytolytic activity. In CD4+ T cells, a preferred method for detecting specific T cell activation is detection of T cell proliferation. In CD8+ T cells, a preferred method for detecting specific T cell activation is detection of the development of cytolytic activity.

[0118] The terms "major histocompatibility complex" and the abbreviation "MHC" include MHC class I and MHC class II molecules and relate to a complex of genes present in all vertebrates. MHC proteins or molecules are important in signal transduction between lymphocytes and antigen-presenting cells or diseased cells in the immune response, and MHC proteins or molecules bind peptides and present them for recognition by T cell receptors. The proteins encoded by MHC are expressed on the surface of cells and display both self-antigens (peptide fragments from the cell itself) and non-self antigens (e.g., fragments of invading microorganisms) to T cells.

[0119] The MHC region is divided into three subgroups: class I, class II, and class III. MHC class I proteins contain an α-chain and β2-microglobulin (not part of the MHC encoded by chromosome 15). They present antigen fragments to cytotoxic T cells. On most immune system cells, specifically antigen-presenting cells, MHC class II proteins contain α- and β-chains and present antigen fragments to T-helper cells. The MHC class III region encodes other immune components such as complement components and some that encode cytokines.

[0120] In humans, the genes in the MHC region that encode antigen-presenting proteins on the cell surface are called human leukocyte antigen (HLA) genes. However, the abbreviation MHC is often used to refer to HLA gene products. The HLA genes include nine so-called classical MHC genes, namely HLA-A, HLA-B, HLA-C, HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRA, and HLA-DRB1.

[0121] In a preferred embodiment of all aspects of the present invention, the MHC molecule is an HLA molecule.

[0122] The term "cell characterized by antigen presentation" or "cell presenting an antigen" or similar expressions means, in the context of MHC molecules, particularly MHC class I molecules, a cell that presents an antigen or a fragment derived from said antigen, for example a diseased cell presenting an antigen or a fragment thereof by antigen processing, such as a cancer cell, or an antigen-presenting cell. Similarly, the term "disease characterized by antigen presentation" refers to a disease that includes cells characterized by antigen presentation, particularly using class I MHC. Antigen presentation by a cell can be achieved by transfecting the cell with a nucleic acid such as RNA encoding the antigen.

[0123] The term "fragment of the antigen being presented" or similar expressions means, for example, a fragment that can be presented by MHC class I or class II, preferably MHC class I, when added directly to an antigen-presenting cell. In one embodiment, the fragment is a fragment that is naturally presented by a cell expressing the antigen.

[0124] The term "immunologically equivalent" means that immunologically equivalent molecules, such as immunologically equivalent amino acid sequences, exhibit the same or essentially the same immunological properties and / or exert the same or essentially the same immunological effects with respect to, for example, the type of immunological effect such as induction of humoral and / or cellular immune responses, the strength and / or duration of the induced immune response, or the specificity of the induced immune response. In the context of the present invention, the term "immunologically equivalent" is preferably used with respect to the immunological effects or properties of the peptides used for immunization. For example, an amino acid sequence is immunologically equivalent to a reference amino acid sequence when it induces an immune response having a specificity that reacts with the reference amino acid sequence when the amino acid sequence is exposed to the immune system of a subject.

[0125] The term "immune effector function" in the context of the present invention includes any function mediated by components of the immune system that results in, for example, inhibition of tumor growth and / or inhibition of tumorigenesis, including killing of tumor cells, or inhibition of intravasation and metastasis of tumors. Preferably, the immune effector function in the context of the present invention is an effector function mediated by T cells. Such functions, in the case of helper T cells (CD4 + T cells), include recognition of an antigen or an antigen peptide derived from the antigen by the T cell receptor in the context of MHC class II molecules, release of cytokines and / or CD8 +Activation of lymphocytes (CTL) and / or B cells, in the case of CTL, recognition of an antigen or an antigen-derived antigen peptide by a T cell receptor in the context of an MHC class I molecule, for example elimination of cells presented in the context of an MHC class I molecule, i.e., cells characterized by antigen presentation using class I MHC, by apoptosis or perforin-mediated cytolysis, production of cytokines such as IFN-γ and TNF-α, and specific cytolytic killing of antigen-expressing target cells.

[0126] The term "genome" relates to the total amount of genetic information in the chromosomes of an organism or cell. The term "exome" refers to the coding regions of the genome. The term "transcriptome" relates to the set of all RNA molecules.

[0127] According to the present invention, "nucleic acid" is preferably deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), more preferably RNA, and most preferably in vitro transcribed RNA (IVT RNA) or synthetic RNA. According to the present invention, nucleic acids include genomic DNA, cDNA, mRNA, recombinantly produced and chemically synthesized molecules. According to the present invention, nucleic acids can exist as single-stranded or double-stranded linear or covalently closed circular molecules. According to the present invention, nucleic acids can be isolated. According to the present invention, the term "isolated nucleic acid" means that the nucleic acid has been (i) amplified in vitro, for example by polymerase chain reaction (PCR), (ii) recombinantly produced by cloning, (iii) purified by separation, for example by cleavage and gel electrophoresis, or (iv) synthesized, for example by chemical synthesis. Nucleic acid can be used in the form of RNA, which can be prepared, in particular, by in vitro transcription from a DNA template for introduction into cells, i.e., transfection. RNA can be further modified by sequence stabilization, capping, and polyadenylation prior to administration.

[0128] The term "genetic material" refers to an isolated nucleic acid of either DNA or RNA, a section of a double helix, a section of a chromosome, or the entire genome of an organism or cell, particularly its exome or transcriptome.

[0129] The term "mutation" refers to a change or difference in a nucleic acid sequence (nucleotide substitution, addition, or deletion) compared to a reference. A "somatic mutation" can occur in any of the body's cells other than germ cells (sperm and eggs) and thus is not passed on to offspring. These changes can (but do not necessarily) cause cancer or other diseases. Preferably, the mutation is a non-synonymous mutation. The term "non-synonymous mutation" refers to a mutation that results in an amino acid change such as an amino acid substitution in the translation product, preferably a nucleotide substitution.

[0130] According to the present invention, the term "mutation" includes point mutations, indels, fusions, chromothripsis, and RNA editing.

[0131] According to the present invention, the term "indel" describes a special class of mutations defined as mutations that result in co-existing insertions and deletions and a net increase or decrease in nucleotides. In the coding regions of the genome, these result in frameshift mutations unless the length of the indel is a multiple of three. Indels can be contrasted with point mutations. Where indels insert and delete nucleotides from a sequence, a point mutation is a form of substitution that replaces one of the nucleotides.

[0132] Fusion can result in hybrid genes formed from two previously separate genes. This can occur as a result of translocation, interstitial deletion, or chromosomal inversion. Often, fusion genes are oncogenes. Oncogenic fusion genes can give rise to gene products with new or different functions from the two fusion partners. Alternatively, a proto-oncogene is fused to a strong promoter and thus begins to function through upregulation caused by the strong promoter of the upstream fusion partner. Also, oncogenic fusion transcripts can be caused by trans-splicing or read-through events.

[0133] According to the present invention, the term "chromothripsis" refers to a genetic phenomenon in which a specific region of the genome is shattered by a single disruptive event and then stitched back together.

[0134] According to the present invention, the term "RNA editing" or "editing RNA" refers to a molecular process in which the information content in an RNA molecule is altered by chemical changes in its base composition. RNA editing includes nucleoside modifications such as cytidine (C) to uridine (U) and adenosine (A) to inosine (I), deamination, as well as non-template nucleotide addition and insertion. RNA editing of mRNA effectively changes the amino acid sequence of the encoded protein so that it differs from that predicted by the genomic DNA sequence.

[0135] The term "cancer mutation signature" refers to a set of mutations present in cancer cells when compared to non-cancerous reference cells.

[0136] According to the present invention, a "reference" can be used to correlate and compare the results obtained by the methods of the present invention from tumor specimens. Typically, a "reference" can be obtained based on one or more normal specimens, particularly specimens not affected by cancer disease, obtained from either the patient or one or more different individuals, preferably healthy individuals, especially individuals of the same species. A "reference" can be determined empirically by testing a sufficiently large number of normal specimens.

[0137] Any suitable sequencing method can be used in accordance with the present invention, with next-generation sequencing (NGS) technology being preferred. In order to increase the speed of the sequencing step of the method, in the future, third-generation sequencing methods may replace NGS technology. For the purpose of clarity, the term "next-generation sequencing" or "NGS" in the context of the present invention refers to all new high-throughput sequencing technologies, as opposed to the "conventional" sequencing method known as Sanger chemistry, which reads nucleic acid templates randomly and in parallel along the entire genome by cutting the entire genome into small pieces. Such NGS technologies (also known as massively parallel sequencing technologies) can deliver nucleic acid sequence information of the entire genome, exome, transcriptome (all transcribed sequences of the genome), or methylome (all methylated sequences of the genome) in a very short period of time, for example, within 1 to 2 weeks, preferably within 1 to 7 days, or most preferably within less than 24 hours, and in principle enable single-cell sequencing techniques. A plurality of NGS platforms, which are commercially available or mentioned in the literature, for example, Zhang et al., 2011: The impact of next-generation sequencing on genomics. J. Genet Genomics 38(3), pp. 95-109 or Voelkerding et al., 2009: Next generation sequencing: From basic research to diagnostics. Clinical chemistry 55, pp. 641-658, can be used in the context of the present invention. Non-limiting examples of such NGS technologies / platforms are as follows. 1) For example, first, the sequencing by synthesis technique known as pyrosequencing implemented in the GS-FLX 454 Genome Sequencer (trademark) of Roche affiliate 454 Life Sciences (Branford, Connecticut), as described by Ronaghi et al., 1998: A sequencing method based on real-time pyrophosphate. Science 281(5375), pp. 363-365. This technique uses emulsion PCR, in which single-stranded DNA-binding beads are encapsulated by vigorously vortexing an aqueous micelle containing PCR reactants surrounded by oil for emulsion PCR amplification. During the pyrosequencing step, the light released from phosphate molecules during nucleotide incorporation as polymerase synthesizes the DNA strand is recorded. 2) Sequencing by synthesis developed by Solexa (now part of Illumina Inc., San Diego, California), which is based on reversible dye-terminators and implemented, for example, in the Illumina / Solexa Genome Analyzer (trademark) and the Illumina HiSeq 2000 Genome Analyzer (trademark). In this technique, all four nucleotides are added simultaneously, along with DNA polymerase, to oligonucleotide-primed cluster fragments in the flow cell channels. The cluster strands with all four fluorescently labeled nucleotides are extended for sequencing by bridge amplification. 3) For example, sequencing by ligation implemented on the SOLid (trademark) platform of Applied Biosystems (now Life Technologies Corporation, Carlsbad, California). In this technique, a pool of all possible oligonucleotides of a fixed length is labeled according to the sequenced position. The oligonucleotides are annealed and ligated, and the preferential ligation of the matching sequences by DNA ligase results in a signal that gives the information of the nucleotide at that position. Before sequencing, the DNA is amplified by emulsion PCR. The resulting beads, each containing only copies of the same DNA molecule, are placed on a slide glass. As a second example, the Polonator (trademark) G.007 platform of Dover Systems (Salem, New Hampshire) also uses emulsion PCR based on randomly arrayed beads to amplify DNA fragments for parallel sequencing, thereby using sequencing by ligation. 4) For example, single molecule sequencing technologies implemented on the PacBio RS system of Pacific Biosciences (Menlo Park, California) or the HeliScope (trademark) platform of Helicos Biosciences (Cambridge, Massachusetts). A distinct feature of this technology is its ability to sequence a single DNA or RNA molecule without amplification, defined as single molecule real-time (SMRT) DNA sequencing. For example, HeliScope uses a highly sensitive fluorescence detection system to directly detect each nucleotide as it is being synthesized. Similar techniques based on fluorescence resonance energy transfer (FRET) have been developed by Visigen Biotechnology (Houston, Texas). Other fluorescence-based single molecule techniques are from U.S. Genomics (GeneEngine (trademark)) and Genovoxx (AnyGene (trademark)). 5) Nanotechnologies for single molecule sequencing that use various nanostructures placed on a chip, for example, to monitor the movement of polymerase molecules on a single strand during replication. Non-limiting examples of techniques based on nanotechnology are the GridON™ platform of Oxford Nanopore Technologies (Oxford, UK), the hybridization-assisted nanopore sequencing (HANS™) platform developed by Nabsys (Providence, Rhode Island), and a ligase-based DNA sequencing platform with a trademark using DNA nanoball (DNB) technology called combinatorial probe anchor ligation (cPAL™). 6) Electron microscopy-based techniques for single molecule sequencing, such as those developed by LightSpeed Genomics (Sunnyvale, California) and Halcyon Molecular (Redwood City, California). 7) Ion semiconductor sequencing based on the detection of hydrogen ions released during DNA polymerization. For example, Ion Torrent Systems (San Francisco, California) uses a high-density array of microscale measurement wells to perform this biochemical process in a massively parallel fashion. Each well has a different DNA template. There is an ion-sensitive layer under the well and an ion sensor with a trademark under that.

[0138] Preferably, the DNA and RNA preparations serve as starting materials for NGS. Such nucleic acids can be readily obtained from samples such as biological materials, for example, from fresh, flash-frozen or formalin-fixed paraffin-embedded tumor tissue (FFPE), or from freshly isolated cells or from CTCs present in a patient's peripheral blood. Normal non-mutated genomic DNA or RNA can be extracted from normal somatic tissues, but in the context of the present invention, germline cells are preferred. Germline DNA or RNA is extracted from peripheral blood mononuclear cells (PBMCs) in patients suffering from non-hematological malignancies. Nucleic acids extracted from FFPE tissue or freshly isolated single cells are highly fragmented, but these are suitable for NGS applications.

[0139] Several targeted NGS methods for exome sequencing are described in the literature (for reviews, see, for example, Teer and Mullikin 2010: Human Mol Genet 19(2), pp. R145-51), and all of these can be used in combination with the present invention. Many of these methods (described, for example, as genome capture, genome partitioning, genome enrichment, etc.) use hybridization techniques and include array-based (for example, Hodges et al., 2007: Nat. Genet. 39, pp. 1522-1527) and solution-based (for example, Choi et al., 2009: Proc. Natl. Acad. Sci USA 106, pp. 19096-19101) hybridization approaches. Also, commercially available kits for the preparation of DNA samples and subsequent capture of exomes are available; for example, Illumina Inc. (San Diego, California) offers the TruSeq™ DNA Sample Preparation Kit and the TruSeq™ Exome Enrichment Kit.

[0140] For example, when comparing the sequence of a tumor sample with the sequence of a reference sample such as the sequence of a germline sample, in order to reduce the number of false positive findings in the detection of cancer-specific somatic mutations or sequence differences, it is preferable to determine the sequences during replication of one or both of these sample types. Accordingly, it is preferable to determine the sequence of a reference sample such as the sequence of a germline sample two, three, or more times. Alternatively or in addition thereto, the sequence of the tumor sample is determined two, three, or more times. Also, the sequence of a reference sample such as the sequence of a germline sample and / or the sequence of the tumor sample can also be determined multiple times by determining the sequence in genomic DNA at least once and determining the sequence in RNA of the reference sample and / or the tumor sample at least once. For example, by determining the mutations between replicates of a reference sample such as a germline sample, the false discovery rate (FDR) of expected somatic mutations can be estimated as a statistical quantity. Technical replicates of one sample should yield the same result, and all mutations detected during this "comparison to the identical" are false positives. In particular, technical replicates of the reference sample can be used as a reference for estimating the number of false positives in order to determine the false discovery rate of somatic mutation detection in the tumor sample relative to the reference sample. Further, various quality-related metrics (e.g., coverage or SNP quality) can be combined into a single quality score using machine learning techniques. For a given somatic mutation, all other mutations having a quality score above it can be counted, thereby enabling the ranking of all mutations in the dataset.

[0141] According to the present invention, a high-throughput whole-genome single-cell genotyping method can be applied.

[0142] In one embodiment of high-throughput whole-genome single-cell genotyping, a Fluidigm platform can be used. Such a method can include the following steps: 1. Sampling tumor tissue / cells and healthy tissue from a given patient. 2. Extract genetic material from cancerous and healthy cells and then sequence its exome (DNA) using standard next-generation sequencing (NGS) protocols. The NGS coverage is such that it can detect heterozygous alleles with a frequency of at least 5%. Also, extract the transcriptome (RNA) from cancer cells, convert it to cDNA, and sequence it to determine which genes are expressed by the cancer cells. 3. Identify non-synonymous expressed single nucleotide variants (SNVs) as described herein. Filter out sites that are SNPs in healthy tissue. 4. Select N = 96 mutations from (3) spanning various frequencies. Design a SNP genotyping assay based on fluorescence detection and synthesize for these mutations (examples of such assays include TaqMan-based SNP assays by Life Technologies or SNP type assays by Fluidigm). The assay includes specific target amplification (STA) primers for amplifying the unit replication sequences containing the given SNV, which are standards in TaqMan and SNP type assays. 5. Isolate individual cells from tumor and healthy tissue either by laser microdissection (LMD) or disaggregation into single cell suspensions and then sort as previously described (Dalerba P. et al. (2011) Nature Biotechnology 29:1120 - 1127). Cells can be selected without pre-selection (i.e., unbiasedly) or cancer cells can be enriched. Enrichment methods include specific staining, sorting by cell size, histological examination during LMD, etc. 6. Isolate individual cells in PCR tubes containing master mix and STA primers, and amplify the unit replication sequences containing SNVs. Alternatively, amplify the genome of a single cell by whole genome amplification (WGA) as previously described (Frumkin D. et al. (2008) Cancer Research 68:5924). Cell lysis is achieved either by a heating step at 95°C or with a dedicated lysis buffer. 7. Dilute the samples amplified with STA and load them onto a Fluidigm genotyping array. 8. Use samples from healthy tissue as positive controls to determine homozygous allele clusters (without mutations). Since NGS data indicate that homozygous mutations are very rare, typically only two clusters, namely XX and XY, are predicted, where X = healthy. 9. The number of arrays that can be run is not limited, and in practice, it is possible to assay up to about 1000 single cells (about 10 arrays). When performed on 384-well plates, sample preparation can be shortened to several days. 10. Then, determine the SNVs of each cell.

[0143] In another embodiment of high-throughput whole-genome single-cell genotyping, an NGS platform can be used. Such a method may include the following steps: 1. Steps 1 - 6 above are the same except that N (the number of SNVs assayed) can be much larger than 96. In the case of WGA, several cycles of STA are performed thereafter. The STA primers contain two universal tag sequences on each primer. 2. After STA, PCR amplify the barcode primers to the unit replication sequences. The barcode primers contain unique barcode sequences and the above universal tag sequences. Thus, each cell contains a unique barcode. 3. Mix the unit replication sequences from all cells and sequence them by NGS. A practical limit on the number of cells that can be multiplexed is the number of plates that can be prepared. Since samples can be prepared in 384-well plates, the practical limit is approximately 5000 cells. 4. Detect SNVs (or other structural aberrations) of individual cells based on the sequence data.

[0144] For antigen prioritization, tumor phylogenetic reconstruction based on single-cell genotyping ("phylogenetic antigen prioritization") can be used according to the present invention. In addition to antigen prioritization based on criteria such as expression, type of mutation (nonsynonymous vs. others), MHC binding characteristics, additional dimensions of prioritization designed to address intra-tumor and inter-tumor heterogeneity as well as biopsy bias can be used, for example, as described below.

[0145] 1. Identification of the most abundant antigens Based on the above single-cell assays associated with high-throughput whole-genome single-cell genotyping methods, the frequency of each SNV can be accurately estimated, and the most abundant SNVs present can be selected to provide an individualized cancer vaccine (IVAC).

[0146] 2. Identification of primary basal antigens based on rooted tree analysis NGS data from tumors suggest that homozygous mutations (hits in both alleles) are rare events. Therefore, there is no need for haplotyping, and a phylogenetic tree of tumor somatic mutations can be created from a single-cell SNV dataset. Use the germline sequence to root the tree. Use an algorithm to duplicate the sequences of the nodes near the root of the tree to duplicate the ancestral sequences. These sequences contain the earliest mutations (defined herein as primary basal mutations / antigens) predicted to be present in the primary tumor. Since the probability of two mutations occurring in the same allele at the same position in the genome is low, the mutations in the ancestral sequences are predicted to be fixed in the tumor.

[0147] Prioritizing a primary basal antigen is not equivalent to prioritizing the most frequent mutations during biopsy (although the primary basal mutations are expected to be one of the most frequent during biopsy). The reasons are as follows. Assuming that two SNVs appear to be present in all cells from which the biopsy is derived (and thus have the same frequency, i.e., 100%), but one mutation is basal and the other is not, the basal mutation should be selected for IVAC. This is because the basal mutation is likely to be present in all regions of the tumor, while the latter mutation may be a more recent mutation that was accidentally fixed in the region from which the biopsy was taken. Furthermore, basal antigens are likely to be present in metastatic tumors derived from the primary tumor. Therefore, by prioritizing basal antigens for IVAC, the possibility that IVAC can eradicate not only part of the tumor but also the entire tumor can be greatly increased.

[0148] If secondary tumors are present and these are also sampled, the evolutionary tree of all tumors can be inferred. This may improve the robustness of the tree and enable the detection of the basal mutations of all tumors.

[0149] 3. Identification of antigens spanning the tumor maximally Another approach to obtain antigens that target all tumor sites maximally is to take several biopsies from the tumor. One strategy would be to select antigens identified as present in all biopsies by NGS analysis. To improve the probability of identifying basal mutations, phylogenetic analysis based on single-cell mutations from all biopsies can be performed.

[0150] In the case of metastases, biopsies from all tumors can be obtained and the mutations common to all tumors identified by NGS can be selected.

[0151] 4. Prioritization of antigens that inhibit metastasis using CTCs Metastatic tumors are thought to originate from a single cell. Thus, by combining genotyping of individual cells isolated from various tumors of a given patient with genotyping of the patient's circulating tumor cells (CTCs), the history of cancer evolution can be reconstructed. The prediction is to observe metastatic tumors that evolve through a clade of CTCs derived from the primary tumor from the original tumor.

[0152] The following (an unbiased method for identifying, counting, and gene-probing CTCs) describes an extension of the above high-throughput whole-genome single-cell genotyping method for unbiased isolation and genomic analysis of CTCs. Using the above analysis, phylogenetic trees of the primary tumor, CTCs, and secondary tumors (if present) arising from metastases can then be reconstructed. Based on this tree, mutations (passengers or drivers) that occurred at or immediately after the time when CTCs first detached from the primary tumor can be identified. The prediction is that the genome of CTCs arising from the primary tumor is evolutionarily more similar to the primary tumor genome than the secondary tumor genome. Furthermore, it is predicted that the genome of CTCs arising from the primary tumor contains unique mutations that are likely to be fixed in the secondary tumor or, if the secondary tumor is to be formed in the future. These unique mutations can be prioritized for IVAC that targets (or prevents) metastasis.

[0153] The advantage of prioritizing CTC mutations over primary basal mutations is that antigens derived from CTCs can mobilize T cells to specifically target metastases, thus serving as an independent arm (using different antigens) from T cells that target the primary tumor. Furthermore, since the probability of tumor escape should correspond to the number of cancer cells carrying a given antigen, the likelihood of immune escape from antigens derived from CTCs is expected to be lower when secondary tumors are few (or absent).

[0154] 5. Identification of antigens coexisting on the same cell ("cocktail" IVAC) Tumors are thought to evolve to suppress mutations caused by immune system and treatment selection pressures. Cancer vaccines that coexist on the same cells and target multiple antigens frequently present in tumors are more likely to circumvent tumor escape mechanisms and thus reduce the likelihood of relapse. Such "cocktail vaccines" would be analogous to antiretroviral combination therapies in HIV+ patients. Coexisting mutations can be identified by phylogenetic analysis or examination of SNV alignments of all cells.

[0155] Furthermore, according to the present invention, an unbiased method can be used to identify, count, and gene probe CTCs. Such a method may include the following steps: 1. Obtain a tumor biopsy and determine the somatic mutation profile. 2. Option 1: Select N≥96 mutations for further investigation based on a previously established prioritization scheme. Option 2: Perform a single-cell assay (see the high-throughput whole-genome single-cell genotyping method described above), followed by phylogenetic analysis, to select N≥96 primary basal mutations and optionally more recent mutations to maximize diversity. The former mutations are useful for identifying CTCs (see below), and the latter are useful for generating phylogenetic analysis (see the section "Identification of antigens coexisting on the same cell ("cocktail" IVAC)"). 3. Obtain whole blood from a cancer patient 4. Lyse red blood cells 5. Enrich CTCs by removing white blood cells by depleting CD45+ cells (e.g., by sorting, magnetic beads conjugated to anti-CD45 antibody, etc.). 6. Remove free DNA by DNAase digestion. The origin of the free DNA may be DNA present in the blood or DNA from dead cells. 7. Select the remaining cells into a PCR tube, perform STA (based on the selected mutations), and screen on Fluidigm (using the high-throughput whole-genome single-cell genotyping method described above). CTCs should generally be positive for multiple SNVs. 8. Subsequently, the cells identified as cancerous (= CTCs) can be further phylogenetically analyzed based on the screened SNV panel (see the section "Identification of Antigens Coexisting on the Same Cell" ("Cocktail" IVAC)).

[0156] It is also possible to combine this method with previously established methods for isolating CTCs. For example, EpCAM+ cells, or cells positive for cytokeratin can be selected (Rao CG. et al. (2005) International journal of oncology 27:49, Allard WJ. et al. (2004) Clinical Cancer Research 10:6897 - 6904). Subsequently, these putative CTCs can be confirmed / profiled on Fluidigm / NGS to induce their mutations.

[0157] CTCs can be counted using this method. This method depends on the mutation profile of cancer somatic cell mutations specific to the patient rather than one specific marker that may or may not be expressed by cancer cells, so this is an unbiased method for detecting and counting CTCs.

[0158] According to the present invention, a technique (phylogenetic filtering) including tumor phylogenetic reconstruction based on single-cell genotyping for enriching driver mutations can be used.

[0159] In one embodiment of this technique, a pan-tumor phylogenetic analysis is performed to recover driver mutations.

[0160] For example, driver mutations from n = 1 tumor can be detected.

[0161] In the above section "Identification of Primary Basal Antigens Based on Rooted Tree Analysis", methods for recovering ancestral sequences and / or methods for identifying cells having sequences close to the root of the tree are described. By definition, since these are sequences close to the root of the tree, the number of mutations in these sequences is expected to be significantly less than the number of mutations in the bulk cancer sample. Therefore, by selecting sequences close to the root of the tree, many passenger mutations are expected to be "phylogenetically filtered" out. This procedure has the potential to greatly enrich for driver mutations. Subsequently, the driver mutations can be used to identify / select patient treatments or can be used as a lead for new therapies.

[0162] In another example, driver mutations from n > 1 tumors of a given type can be detected.

[0163] By reconstructing primary basal mutations from many tumors of a particular type, the likelihood of detecting driver mutations can be greatly enhanced. Since basal sequences close to the root of the tree filter out many passenger mutations, the signal-to-noise ratio in the detection of driver mutations is predicted to increase significantly. Therefore, this method has the potential to (1) detect driver mutations with lower frequencies and (2) detect frequent driver mutations from fewer samples.

[0164] In another embodiment of a technique ("phylogenetic filtering") that includes tumor phylogenetic reconstruction based on single-cell genotyping to enrich for driver mutations, a phylogenetic analysis is performed to recover metastases that cause driver mutations.

[0165] In the above section "Prioritization of antigens that inhibit metastasis using CTCs", a method for detecting CTC-related mutations is described. This method can also be used to enrich for driver mutations that cause metastasis. For example, by mapping the combined phylogeny of primer tumors, secondary tumors, and CTCs, CTCs derived from the primary tumor should connect between the primary-secondary tumor clades. Such phylogenetic analysis can help identify mutations specific to this transition between the primer and secondary tumors. A fraction of these mutations can be driver mutations. Furthermore, by comparing unique CTC mutations (i.e., n>1 tumors) from different examples of the same cancer, unique driver mutations that cause metastasis can be further enriched.

[0166] According to the present invention, phylogenetic analysis can be used to identify primary versus secondary tumors.

[0167] In the case of metastasis, if all tumors are sampled, a rooted tree can be used to predict the chronological order in which the tumors appeared, i.e., which tumor is the primary tumor (the node closest to the root of the tree) and which tumors are the most recent. This can be useful when it is difficult to determine which tumor is primary.

[0168] In the context of the present invention, the term "RNA" relates to a molecule containing at least one ribonucleotide residue and preferably consisting entirely or substantially of ribonucleotide residues. "Ribonucleotide" relates to a nucleotide having a hydroxyl group at the 2'-position of the β-D-ribofuranosyl group. The term "RNA" includes double-stranded RNA, single-stranded RNA, isolated RNA such as partially or completely purified RNA, essentially pure RNA, synthetic RNA, and recombinantly produced RNA such as modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution and / or alteration of one or more nucleotides. Such alterations can include, for example, the addition of non-nucleotide substances to the ends or internally of the RNA, such as at one or more nucleotides of the RNA. Also, the nucleotides in an RNA molecule can include non-standard nucleotides such as non-naturally occurring nucleotides or chemically synthesized nucleotides or deoxynucleotides. These altered RNAs can be referred to as analogs or analogs of naturally occurring RNA.

[0169] According to the present invention, the term "RNA" includes and preferably relates to "mRNA". The term "mRNA" means "messenger RNA" and relates to a "transcript" produced by using a DNA template and encoding a peptide or polypeptide. Typically, mRNA includes a 5'-UTR, a protein-coding region, and a 3'-UTR. mRNA has a limited half-life in cells and in vitro. In the context of the present invention, mRNA can be produced by in vitro transcription from a DNA template. In vitro transcription methods are known to those skilled in the art. For example, various in vitro transcription kits are commercially available.

[0170] According to the present invention, the stability and translation efficiency of RNA can be modified as needed. For example, one or more modifications having an RNA stabilizing effect and / or increasing the translation efficiency can stabilize the RNA and increase its translation. Such modifications are described, for example, in PCT / EP2006 / 009448, which is incorporated herein by reference. In order to increase the expression of RNA used according to the present invention, within the coding region, i.e., the sequence encoding the expressed peptide or protein, preferably without changing the sequence of the expressed peptide or protein, the GC content can be increased by modification to increase the stability of mRNA, codon optimization can be performed, and thus the translation in cells can be enhanced.

[0171] The term "modification" in the context of RNA used in the present invention includes any modification of the RNA that is not naturally present in said RNA.

[0172] In one embodiment of the present invention, the RNA used according to the present invention does not have a non-cap 5'-triphosphate. Removal of such non-cap 5'-triphosphate can be achieved by treating the RNA with phosphatase.

[0173] The RNA according to the present invention can have modified ribonucleotides in order to increase its stability and / or reduce its cytotoxicity. For example, in one embodiment, in the RNA used according to the present invention, cytidine is partially or completely, preferably completely, replaced with 5-methylcytidine. Alternatively or in addition, in one embodiment, in the RNA used according to the present invention, uridine is partially or completely, preferably completely, replaced with pseudouridine.

[0174] In one embodiment, the term "capping" relates to providing a 5'-cap or 5'-cap analog to RNA. The term "5'-cap" refers to the cap structure found on the 5' end of an mRNA molecule, and generally consists of a guanosine nucleotide linked to the mRNA via an unusual 5' to 5' triphosphate bond. In one embodiment, this guanosine is methylated at the 7 position. The term "conventional 5'-cap" refers to a naturally occurring RNA 5'-cap, preferably a 7-methylguanosine cap (m 7 G). In the context of the present invention, the term "5'-cap" includes 5'-cap analogs that are similar to the RNA cap structure and are preferably modified to have the ability to stabilize RNA and / or enhance RNA translation when attached thereto in vivo and / or intracellularly.

[0175] Preferably, the 5' end of the RNA comprises a Cap structure having the following general formula.

[0176]

Chemical formula

[0177] Wherein R1 and R2 are independently hydroxy or methoxy, and W - , X - and Y - are independently oxygen, sulfur, selenium, or BH3. In a preferred embodiment, R1 and R2 are hydroxy, and W - , X - and Y - are oxygen. In a further preferred embodiment, one of R1 and R2, preferably R1, is hydroxy and the other is methoxy, and W - , X - and Y - are oxygen. In a further preferred embodiment, R1 and R2 are hydroxy, and one of W - , X - and Y - , preferably X -is sulfur, selenium, or BH3, preferably sulfur, and the others are oxygen. In a further preferred embodiment, one of R1 and R2, preferably R2, is hydroxy and the other is methoxy, and W - , X - and Y - of which one, preferably X - is sulfur, selenium, or BH3, preferably sulfur, and the others are oxygen.

[0178] In the above formula, the right nucleotide is linked to the RNA strand via its 3' group.

[0179] W - , X - and Y - at least one of which is sulfur, i.e., a Cap structure having phosphorothioate moieties exists in various diastereomeric forms, all of which are included herein. Further, the present invention includes all tautomers and stereoisomers of the above formula.

[0180] For example, a Cap structure having the above structure where R1 is methoxy, R2 is hydroxy, X - is sulfur, and W - and Y - are oxygen exists in two diastereomeric forms (Rp and Sp). These can be separated by reverse-phase HPLC and are named D1 and D2 according to their elution order from a reverse-phase HPLC column. According to the present invention, the D1 isomer of m2 7,2'-O GppspG is particularly preferred.

[0181] Providing a 5'-cap or 5'-cap analog to RNA may be achieved by in vitro transcription of a DNA template in the presence of said 5'-cap or 5'-cap analog, where said 5'-cap is co-transcriptionally incorporated into the RNA strand being made, or, the RNA may be made, for example, by in vitro transcription, and the 5'-cap may be attached to the RNA post-transcriptionally using a capping enzyme, such as the capping enzyme of vaccinia virus.

[0182] The RNA may contain further modifications. For example, further modifications of the RNA used in the present invention include elongation or cleavage of the naturally occurring poly(A) tail, or introduction of a UTR not associated with the coding region of the RNA, for example, replacement or insertion of an existing 3'-UTR with one or more, preferably 2 copies of a 3'-UTR derived from a globin gene, such as alpha2-globin, alpha1-globin, beta-globin, preferably beta-globin, more preferably human beta-globin, which can be a change in the 5'- or 3'-untranslated region (UTR).

[0183] RNA having an unmasked poly-A sequence is translated more efficiently than RNA having a masked poly-A sequence. The term "poly(A) tail" or "poly-A sequence" typically refers to a sequence of adenylyl (A) residues located at the 3'-end of an RNA molecule, and "unmasked poly-A sequence" means that the poly-A sequence at the 3'-end of the RNA molecule ends with an A of the poly-A sequence and no nucleotides other than A located downstream, i.e., at the 3'-end of the poly-A sequence, follow. Furthermore, a poly-A sequence about 120 base pairs in length results in optimal transcriptional stability and translation efficiency of the RNA.

[0184] Therefore, to increase the stability and / or expression of the RNA used according to the present invention, it can be modified to be present with a poly-A sequence preferably having a length of 10 to 500, more preferably 30 to 300, even more preferably 65 to 200, particularly 100 to 150 adenosine residues. In a particularly preferred embodiment, the poly-A sequence has a length of about 120 adenosine residues. To further increase the stability and / or expression of the RNA used according to the present invention, the poly-A sequence can be left unmasked.

[0185] Furthermore, the incorporation of a 3'-untranslated region (UTR) into the 3'-untranslated region of an RNA molecule can result in enhanced translation efficiency. By incorporating two or more such 3'-untranslated regions, a synergistic effect can be achieved. The 3'-untranslated regions can be self or heterologous to the RNA into which they are introduced. In certain embodiments, the 3'-untranslated regions are derived from the human β-globin gene.

[0186] The combinations of modifications described above, namely the incorporation of a poly-A sequence, the demasking of the poly-A sequence, and the incorporation of one or more 3'-untranslated regions, have a synergistic effect on increasing the stability and translation efficiency of the RNA.

[0187] The term "stability" of RNA relates to the "half-life" of the RNA. The "half-life" relates to the period required to eliminate half of the activity, amount, or number of the molecule. In the context of the present invention, the half-life of the RNA is an indicator of the stability of the RNA. The half-life of the RNA can affect the "duration of expression" of the RNA. An RNA having a long half-life can be predicted to be expressed for a long period of time.

[0188] Of course, if it is desirable to decrease the stability and / or translation efficiency of the RNA according to the present invention, it is possible to modify the RNA so as to interfere with the functions of the above-described elements that increase the stability and / or translation efficiency of the RNA.

[0189] The term "expression" is used in its most general sense according to the present invention and includes, for example, the production of RNA and / or peptide or polypeptide by transcription and / or translation. With respect to RNA, the terms "expression" or "translation" relate in particular to the production of peptide or polypeptide. This also includes partial expression of the nucleic acid. Furthermore, expression can be transient or stable.

[0190] According to the present invention, the term "expression" also includes "abnormal expression" or "abnormal overexpression". According to the present invention, "abnormal expression" or "abnormal overexpression" means that the expression is altered, preferably increased, compared to the state of a subject not suffering from a disease associated with abnormal expression or abnormal overexpression of a reference, for example, a specific protein, such as a tumor antigen. An increase in expression means an increase of at least 10%, particularly at least 20%, at least 50% or at least 100%, or more. In one embodiment, the expression is found only in diseased tissue and is suppressed in healthy tissue.

[0191] The term "specifically expressed" means that a protein is expressed essentially only in a specific tissue or organ. For example, a tumor antigen specifically expressed in the gastric mucosa means that the protein is mainly expressed in the gastric mucosa and is not expressed or is not significantly expressed in other tissues or other tissue or organ types. Thus, a protein that is exclusively expressed in cells of the gastric mucosa and is expressed to a significantly lesser extent in any other tissue, such as the testis, is specifically expressed in cells of the gastric mucosa. In some embodiments, the tumor antigen may also be specifically expressed in a plurality of tissue types or organs, for example, two or three tissue types or organs, but preferably three or fewer different tissue or organ types, under normal conditions. In this case, the tumor antigen is specifically expressed in these organs. For example, if a tumor antigen is preferably expressed to approximately the same extent in the lung and the stomach under normal conditions, the tumor antigen is specifically expressed in the lung and the stomach.

[0192] In the context of the present invention, the term "transcription" relates to the process by which the genetic code of a DNA sequence is transcribed into RNA. Subsequently, the RNA can be translated into protein. According to the present invention, the term "transcription" includes "in vitro transcription", and the term "in vitro transcription" relates to the process by which RNA, particularly mRNA, is synthesized in vitro in a cell-free system, preferably using an appropriate cell extract. Preferably, a cloning vector is applied to the production of the transcript. These cloning vectors are generally named transcription vectors and are encompassed by the term "vector" according to the present invention. According to the present invention, the RNA used in the present invention is preferably in vitro transcribed RNA (IVT-RNA) and can be obtained by in vitro transcription of an appropriate DNA template. The promoter for controlling transcription can be any promoter of any RNA polymerase. Specific examples of RNA polymerases are T7, T3, and SP6 RNA polymerases. Preferably, in vitro transcription is controlled by a T7 or SP6 promoter according to the present invention. The DNA template for in vitro transcription can be obtained by cloning a nucleic acid, particularly cDNA, and introducing it into an appropriate vector for in vitro transcription. cDNA can be obtained by reverse transcription of RNA.

[0193] The term "translation" according to the present invention relates to the process in the ribosome of a cell by which a strand of messenger RNA directs the assembly of a sequence of amino acids to produce a peptide or polypeptide.

[0194] An expression control sequence or regulatory sequence that may be functionally linked to a nucleic acid according to the present invention can be homologous or heterologous with respect to the nucleic acid. A coding sequence and a regulatory sequence are "functionally" linked together when they are covalently bonded together such that the transcription or translation of the coding sequence is under the control or influence of the regulatory sequence. When using the functional linkage of a regulatory sequence and a coding sequence to translate the coding sequence into a functional protein, induction of the regulatory sequence results in transcription of the coding sequence without causing a shift in the reading frame of the coding sequence or preventing the coding sequence from being translated into the desired protein or peptide.

[0195] According to the present invention, the terms "expression control sequence" or "regulatory sequence" include promoters, ribosome binding sequences and other control elements that control the transcription of nucleic acids or the translation of the resulting RNA. In certain embodiments of the present invention, the regulatory sequences can be controlled. The exact structure of the regulatory sequences may vary depending on the species or cell type, but generally includes 5'-untranscribed as well as 5'- and 3'-untranslated sequences involved in the initiation of transcription or translation, such as the TATA box, capping sequence, CAAT sequence, etc. In particular, the 5'-untranscribed regulatory sequences include a promoter region containing the promoter sequence for transcriptional control of the functionally linked gene. The regulatory sequences can also include enhancer sequences or upstream activation sequences.

[0196] Preferably, according to the present invention, the RNA to be expressed in the cell is introduced into the cell. In one embodiment of the method according to the present invention, the RNA introduced into the cell is obtained by in vitro transcription of an appropriate DNA template.

[0197] According to the present invention, terms such as "RNA capable of being expressed" and "RNA encoding" are used interchangeably herein, and with respect to a particular peptide or polypeptide, it means that the RNA can be expressed to produce the peptide or polypeptide when present in an appropriate environment, preferably within a cell. Preferably, the RNA according to the present invention can interact with the translation machinery of the cell to provide a peptide or polypeptide that it is capable of expressing.

[0198] Terms such as "transfer," "introduce," or "transfect" are used interchangeably herein and relate to introducing nucleic acids, particularly exogenous or heterologous nucleic acids, particularly RNA, into cells. According to the present invention, cells can form organs, tissues, and / or parts of an organism. According to the present invention, administration of the nucleic acid can be achieved either as naked nucleic acid or in combination with an administration reagent. Preferably, the administration of the nucleic acid is in the form of naked nucleic acid. Preferably, the RNA is administered in combination with a stabilizing substance such as an RNase inhibitor. The present invention also contemplates repeatedly introducing the nucleic acid into cells to enable long-term sustained expression.

[0199] Cells can be transfected using any carrier that can associate with RNA, for example, by forming a complex with RNA or forming vesicles that encapsulate or encapsulate RNA, resulting in increased RNA stability compared to naked RNA. Carriers useful according to the present invention include, for example, cationic lipids, liposomes, particularly cationic liposomes, and lipid-containing carriers such as micelles, as well as nanoparticles. Cationic lipids can form complexes with negatively charged nucleic acids. Any cationic lipid can be used according to the present invention.

[0200] Preferably, introduction of RNA encoding a peptide or polypeptide into cells, particularly cells present in vivo, results in expression of the peptide or polypeptide in the cells. In certain embodiments, it is preferred to target the nucleic acid to specific cells. In such embodiments, the carrier (e.g., retrovirus or liposome) applied to administer the nucleic acid to the cells displays a targeting molecule. For example, a molecule such as an antibody specific for a surface membrane protein on the target cell or a ligand of a receptor on the target cell can be incorporated into or conjugated to the nucleic acid carrier. When the nucleic acid is administered by liposome, a protein that binds to a surface membrane protein associated with endocytosis can be incorporated into the liposome formulation to enable targeting and / or uptake. Such proteins include capsid proteins or fragments thereof specific for a particular cell type, an antibody against an intracellularly translocated protein, a protein targeting an intracellular location, and the like.

[0201] According to the present invention, the term "peptide" refers to a substance comprising two or more, preferably three or more, preferably four or more, preferably six or more, preferably eight or more, preferably ten or more, preferably thirteen or more, preferably sixteen or more, preferably twenty-one or more, and preferably up to 8, 10, 20, 30, 40 or 50, particularly 100 amino acids, covalently linked by peptide bonds. The terms "polypeptide" or "protein" refer to large peptides, preferably peptides having more than 100 amino acid residues, but generally, the terms "peptide", "polypeptide" and "protein" are synonymous and are used interchangeably herein.

[0202] According to the present invention, the term "sequence variation" with respect to a peptide or protein relates to amino acid insertion mutants, amino acid addition mutants, amino acid deletion mutants and amino acid substitution mutants, preferably amino acid substitution mutants. According to the present invention, all of these sequence variations have the potential to create new epitopes.

[0203] An amino acid insertion mutant contains the insertion of one or two or more amino acids in a specific amino acid sequence.

[0204] An amino acid addition mutant contains the amino and / or carboxy-terminal fusion of one or more amino acids, such as 1, 2, 3, 4, or 5 amino acids, or more.

[0205] An amino acid deletion mutant is characterized by the removal of one or more amino acids from the sequence, for example, the removal of 1, 2, 3, 4, or 5 amino acids or more.

[0206] An amino acid substitution mutant is characterized in that at least one residue in the sequence is removed and another residue is inserted in its place.

[0207] According to the present invention, the term "derived from" means that a particular entity, particularly a particular sequence, is present in the object from which it is derived, particularly an organism or molecule. In the case of an amino acid sequence, particularly a specific sequence region, "derived from" particularly means that the relevant amino acid sequence is derived from the amino acid sequence in which it is present.

[0208] The term "cell" or "host cell" preferably refers to an untreated cell, i.e., a cell having an untreated membrane that has not released its normal intracellular components such as enzymes, organelles, or genetic material. The untreated cell is preferably a living cell, i.e., a living cell capable of performing any normal metabolic function. Preferably, according to the present invention, the term relates to any cell that can be transformed or transfected with exogenous nucleic acid. According to the present invention, the term "cell" includes prokaryotic cells (e.g., Escherichia coli (E. coli)) or eukaryotic cells (e.g., dendritic cells, B cells, CHO cells, COS cells, K562 cells, HEK293 cells, HELA cells, yeast cells, and insect cells). The exogenous nucleic acid can be found intracellularly (i) freely dispersed by itself, (ii) incorporated into a recombinant vector, or (iii) incorporated into the host cell genome or mitochondrial DNA. Mammalian cells such as cells from humans, mice, hamsters, pigs, goats, and primates are particularly preferred. The cells can be derived from a number of tissue types and include primary cells and cell lines. Specific examples include keratinocytes, peripheral blood leukocytes, bone marrow stem cells, and embryonic stem cells. In a further embodiment, the cells are antigen-presenting cells, particularly dendritic cells, monocytes, or macrophages.

[0209] A cell containing a nucleic acid molecule preferably expresses a peptide or polypeptide encoded by the nucleic acid.

[0210] The term "clonal expansion" refers to the process by which a particular entity increases. In the context of the present invention, this term is preferably used in relation to an immunological response in which lymphocytes are stimulated by an antigen, proliferate, and the specific lymphocytes that recognize the antigen are amplified. Preferably, clonal expansion results in the differentiation of lymphocytes.

[0211] Terms such as "reduce" or "inhibit" preferably relate to the ability to cause a decrease in the overall level of at least 5%, at least 10%, at least 20%, more preferably at least 50%, and most preferably at least 75%. The term "inhibit" or similar phrases include complete or substantially complete inhibition, i.e., a decrease to zero or substantially zero.

[0212] Terms such as "increase", "enhance", "promote", or "extend" preferably relate to an increase, enhancement, promotion, or extension of at least about 10%, preferably at least 20%, preferably at least 30%, preferably at least 40%, preferably at least 50%, preferably at least 80%, preferably at least 100%, preferably at least 200%, particularly at least 300%. These terms can also relate to an increase, enhancement, promotion, or extension from zero or an immeasurable or undetectable level to a level higher than zero or a measurable or detectable level.

[0213] Using the agents, compositions, and methods described herein, a subject suffering from a disease, such as a disease characterized by the presence of diseased cells that express an antigen and present antigenic peptides, can be treated. Particularly preferred diseases are cancer diseases. Also, immunization or vaccination can be performed using the agents, compositions, and methods described herein to prevent the diseases described herein.

[0214] According to the present invention, the term "disease" refers to any pathological condition, including cancer diseases, particularly the forms of cancer diseases described herein.

[0215] The term "normal" refers to a healthy state or a healthy subject or tissue, i.e., a state in a non-diseased state, and "healthy" preferably means non-cancerous.

[0216] According to the present invention, the "disease comprising cells expressing an antigen" means that the expression of the antigen is detected in the cells of the affected tissue or organ. The expression in the cells of the affected tissue or organ may be increased compared to the state of healthy tissue or organ. The increase means at least 10%, particularly at least 20%, at least 50%, at least 100%, at least 200%, at least 500%, at least 1000%, at least 10000% or even more increase. In one embodiment, the expression is found only in the affected tissue and the expression in healthy tissue is suppressed. According to the present invention, diseases comprising or related to cells expressing an antigen include cancer diseases.

[0217] Cancer (the medical term is malignant neoplasm) is a class of diseases in which a group of cells shows uncontrolled growth (division beyond normal limits), invasion (invasion and destruction of adjacent tissues), and in some cases metastasis (spread to other locations in the body via lymph or blood). Due to these three malignant characteristics of cancer, they are distinguished from benign tumors that are self-limiting and do not infiltrate or metastasize. Most cancers form tumors, but some, such as leukemia, do not.

[0218] Malignant tumor is essentially synonymous with cancer. Malignant disease, malignant neoplasm, and malignant tumor are essentially synonymous with cancer.

[0219] According to the present invention, the term "tumor" or "tumor disease" preferably refers to abnormal growth of cells (referred to as neoplastic cells, tumor-forming cells or tumor cells) that form a swelling or lesion. "Tumor cells" mean abnormal cells that grow by rapid uncontrolled cell proliferation and continue to grow after the stimulus that initiated the new growth has ended. Tumors show partial or complete lack of structural organization and functional cooperation with normal tissues and usually form distinct tissue masses that can be either benign, pre-malignant or malignant.

[0220] A benign tumor is a tumor that lacks all three of the malignant characteristics of cancer. Thus, by definition, a benign tumor does not grow in an unlimited aggressive manner, does not invade surrounding tissues, and does not spread (metastasize) to non-adjacent tissues.

[0221] A neoplasm is an abnormal mass of tissue as a result of neoplasia. Neoplasia (Greek for new growth) is the abnormal proliferation of cells. The growth of cells exceeds that of the normal tissue around it and is not coordinated. The growth persists in the same excessive manner even after the stimulus has ended. This usually causes a mass or tumor. A neoplasm can be benign, pre-malignant, or malignant.

[0222] As used in the present invention, "tumor growth" or "tumor growth" refers to the tendency of a tumor to increase in size and / or the tendency of tumor cells to proliferate.

[0223] For the purposes of the present invention, the terms "cancer" and "cancer disease" are used interchangeably with the terms "tumor" and "tumor disease".

[0224] Cancers are classified by the type of cells that resemble tumors and thus the type of tissue from which the tumor is presumed to originate. These are histological and positional respectively.

[0225] As used in the present invention, the term "cancer" includes leukemia, seminoma, melanoma, teratoma, lymphoma, neuroblastoma, glioma, rectal cancer, endometrial cancer, kidney cancer, adrenal cancer, thyroid cancer, blood cancer, skin cancer, brain cancer, cervical cancer, intestinal cancer, liver cancer, colon cancer, gastric cancer, intestinal cancer, head and neck cancer, gastrointestinal cancer, lymph node cancer, esophageal cancer, colorectal cancer, pancreatic cancer, ear, nose and throat (ENT) cancer, breast cancer, prostate cancer, uterine cancer, ovarian cancer and lung cancer and their metastases. Examples thereof are lung tumors, breast tumors, prostate tumors, colon tumors, renal cell tumors, cervical tumors, or metastases of the above-mentioned cancer types or tumors. Further, according to the present invention, the term cancer also includes cancer metastases and cancer relapses.

[0226] The main types of lung cancer are small cell lung carcinoma (SCLC) and non-small cell lung carcinoma (NSCLC). There are mainly three subtypes of non-small cell lung carcinoma, namely squamous cell lung carcinoma, adenocarcinoma, and large cell lung carcinoma. Adenocarcinoma accounts for about 10% of lung cancer. This cancer is usually found in the periphery of the lung, in contrast to both small cell lung cancer and squamous cell lung cancer, which tend to be more centrally located.

[0227] Skin cancer is a malignant growth on the skin. The most common types of skin cancer are basal cell carcinoma, squamous cell carcinoma, and melanoma. Malignant melanoma is a serious type of skin cancer. It is caused by the uncontrolled growth of pigment cells called melanocytes.

[0228] According to the present invention, a "carcinoma" is a malignant tumor derived from epithelial cells. This group represents the most common cancers, including the common forms of breast, prostate, lung, and colon cancers.

[0229] "Bronchioloalveolar carcinoma" is a lung carcinoma thought to be derived from the epithelium of the terminal bronchioles, where the neoplastic tissue extends along the alveolar walls and grows in small masses within the alveoli. Mucin can be demonstrated in a part of the cells including exfoliated cells and in the substances in the alveoli.

[0230] "Adenocarcinoma" is cancer that originates from glandular tissue. This tissue is also part of a larger tissue classification known as epithelial tissue. Epithelial tissue includes the skin, glands, and various other tissues that line the cavities and organs of the body. Epithelium is embryologically derived from the ectoderm, endoderm, and mesoderm. To be classified as an adenocarcinoma, cells do not necessarily have to be part of a gland as long as they have secretory properties. This type of carcinoma can occur in some higher mammals, including humans. Well-differentiated adenocarcinomas tend to resemble the glandular tissue from which they originate, while poorly differentiated ones may not. By staining cells from a biopsy, a pathologist can determine whether a tumor is an adenocarcinoma or some other type of cancer. Due to the ubiquitous nature of glands in the body, adenocarcinomas can occur in many tissues of the body. Each gland may not secrete the same substance, but as long as the cells have an exocrine function, they are considered glands, and thus their malignant form is named adenocarcinoma. Malignant adenocarcinomas tend to metastasize if given enough time to invade other tissues. Ovarian adenocarcinoma is the most common type of ovarian carcinoma. This includes serous and mucinous adenocarcinomas, clear cell adenocarcinomas, and endometrioid adenocarcinomas.

[0231] Renal cell carcinoma, also known as renal cell adenocarcinoma, is a type of kidney cancer that originates from the lining of the proximal convoluted tubule, a very small tube in the kidney that filters blood and removes waste. Renal cell carcinoma is by far the most common type of kidney cancer in adults and is the most lethal of all genitourinary tumors. The distinct subtypes of renal cell carcinoma are clear cell renal cell carcinoma and papillary renal cell carcinoma. Clear cell renal cell carcinoma is the most common form of renal cell carcinoma. When viewed under a microscope, the cells that make up clear cell renal cell carcinoma appear very pale or transparent in color. Papillary renal cell carcinoma is the second most common subtype. These cancers form small finger-like projections (called papillae) in some, if not most, of the tumor.

[0232] Lymphomas and leukemias are malignant tumors that originate from hematopoietic (blood-forming) cells.

[0233] A blastoma or blastocytoma is a tumor (usually malignant) that resembles immature or embryonic tissue. Many of these tumors are most common in children.

[0234] "Metastasis" means the spread of cancer cells from their original site to another part of the body. The formation of metastases is a very complex process that depends on the detachment of malignant cells from the primary tumor, invasion of the extracellular matrix, penetration of the endothelial basement membrane to enter body cavities and blood vessels, and then infiltration of the target organ after being transported by the blood. Finally, the growth of new tumors at the target site, i.e., secondary or metastatic tumors, depends on angiogenesis. Tumor metastasis often occurs even after removal of the primary tumor because tumor cells or components may remain and develop metastatic potential. In one embodiment, the term "metastasis" according to the present invention relates to "distant metastasis" related to metastases away from the primary tumor and the associated lymph node system.

[0235] The cells of secondary or metastatic tumors resemble those of the original tumor. This means that, for example, when ovarian cancer metastasizes to the liver, the secondary tumor is composed of abnormal ovarian cells rather than abnormal liver cells. In this case, the tumor in the liver is called metastatic ovarian cancer rather than liver cancer.

[0236] In ovarian cancer, metastasis can occur in the following ways: by direct contact or extension, it can infiltrate adjacent tissues or organs located near or around the ovaries, such as the fallopian tubes, uterus, bladder, rectum, etc.; by seeding or exfoliation into the peritoneal cavity, which is the most common way of spreading ovarian cancer, where cancer cells detach from the surface of the ovarian mass and "fall" onto other structures in the abdomen, such as the liver, stomach, colon, or diaphragm; by detaching from the ovarian mass, infiltrating the lymphatic vessels and then moving to other regions or distant organs of the body, such as the lungs or liver; by detaching from the ovarian mass, infiltrating the blood system and moving to other regions or distant organs of the body.

[0237] According to the present invention, metastatic ovarian cancer includes cancers of the fallopian tubes, cancers of abdominal organs such as cancers of the intestine, uterus, bladder, rectum, liver, stomach, colon, diaphragm, lung, cancers of the lining (peritoneum) of the abdomen or pelvis, and cancers of the brain. Similarly, metastatic lung cancer refers to cancer that has spread from the lung to distal and / or several sites of the body and includes cancers of the liver, adrenal gland, bone, and brain.

[0238] The term "circulating tumor cell" or "CTC" refers to cells that detach from a primary tumor or tumor metastasis and circulate in the bloodstream. CTCs can constitute the seeds for the growth of further tumors (metastases) in various subsequent tissues. Circulating tumor cells are found at a frequency of approximately 1 to 10 CTCs per milliliter of whole blood in patients suffering from metastatic disease. Research methods for isolating CTCs have been developed. Several research methods for isolating CTCs in the art are described, for example, techniques that utilize the fact that epithelial cells generally express the cell adhesion protein EpCAM, which is not present in normal blood cells. Immunomagnetic bead-based capture involves treating a blood sample with an antibody against EpCAM conjugated to magnetic particles, followed by separating the tagged cells in a magnetic field. Subsequently, rare CTCs are discriminated from contaminating white blood cells by staining the isolated cells with antibodies against another epithelial marker, cytokeratin, and a common white blood cell marker, CD45. This robust and semi-automated approach identifies CTCs with an average yield of approximately 1 CTC / mL and a purity of 0.1% (Allard et al., 2004: Clin Cancer Res 10, pp. 6897-6904). A second method for isolating CTCs involves using a microfluidics-based CTC capture device that includes flowing whole blood through a chamber embedded with 80,000 microposts functionalized by coating with an antibody against EpCAM. Subsequently, CTCs are stained with a secondary antibody against either cytokeratin or a tissue-specific marker such as PSA in prostate cancer or HER2 in breast cancer, and visualized by automatically scanning the microposts along three-dimensional coordinates in multiple planes. The CTC-chip can identify cytokerating-positive circulating tumor cells in patients with a median yield of 50 cells / ml and a purity range of 1-80% (Nagrath et al., 2007: Nature 450, pp. 1235-1239). Another possibility for isolating CTCs is to use the Veridex, LLC (Raritan, NJ) CellSearch™ Circulating Tumor Cell (CTC) Test, which captures, identifies, and counts CTCs in a blood tube.The CellSearch (trademark) system is a method approved by the US Food and Drug Administration (FDA) for counting CTCs in whole blood and is based on a combination of immunomagnetic labeling and automated digital microscopy. Other methods for isolating CTCs are described in the literature, and all of them can be used in combination with the present invention.

[0239] Relapse or recurrence occurs when a person is affected again by a condition that previously affected them. For example, if a patient has a tumor disease, undergoes treatment for the disease successfully, and develops the disease again, the newly developed disease can be regarded as a relapse or recurrence. However, according to the present invention, a relapse or recurrence of a tumor disease may occur at the site of the original tumor, but it is not necessarily so. Thus, for example, if a patient has an ovarian tumor and the treatment received is successful, a relapse or recurrence can be the occurrence of an ovarian tumor or the occurrence of a tumor at a site different from the ovary. Also, a relapse or recurrence of a tumor includes situations where the tumor occurs at a site different from the site of the original tumor and situations where the tumor occurs at the site of the original tumor. Preferably, the original tumor from which the patient received treatment is a primary tumor, and the tumor at a site different from the site of the original tumor is a secondary or metastatic tumor.

[0240] "To treat" means administering to a subject a compound or composition described herein for preventing or eliminating a disease, including reducing the size or number of tumors in the subject, arresting or delaying the disease in the subject, inhibiting or delaying the onset of a new disease in the subject, reducing the frequency or severity of symptoms and / or recurrence in a subject currently or previously afflicted with the disease, and / or prolonging, i.e., increasing, the lifespan of the subject. In particular, the term "treatment of a disease" includes curing the onset or symptoms of a disease, shortening the duration, alleviating, preventing, delaying or inhibiting progression or exacerbation, or preventing or delaying.

[0241] "At risk" means a subject, i.e., a patient, identified as having a higher than normal likelihood of developing a disease, particularly cancer, compared to the general population. Further, a subject who has had or currently has a disease, particularly cancer, is at increased risk of developing the disease because such a subject can continue to develop the disease. Also, a subject who currently has or has had cancer is also at increased risk of cancer metastasis.

[0242] The term "immunotherapy" relates to treatments involving the activation of a specific immune response. In the context of the present invention, terms such as "protect", "prevent", "preventive", "preventatively" or "protective" relate to preventing or treating or both the occurrence and / or proliferation of a disease in a subject, particularly to minimizing the likelihood that a subject will develop the disease or delaying the onset of the disease. For example, a person at risk of the above-described tumor is a candidate for a treatment that prevents the tumor.

[0243] Prophylactic administration of immunotherapy, such as prophylactic administration of a composition of the present invention, preferably protects the recipient from the onset of the disease. Therapeutic administration of immunotherapy, such as therapeutic administration of a composition of the present invention, can result in inhibition of disease progression / growth. This preferably includes decelerating the progression / growth of the disease, particularly disrupting the progression of the disease, which results in the elimination of the disease.

[0244] Immunotherapy can be performed using any of a variety of techniques in which the drugs provided herein function to remove diseased cells from a patient. Such removal can occur as a result of enhancing or inducing an immune response specific to an antigen or cells expressing an antigen in the patient.

[0245] In certain embodiments, the immunotherapy can be active immunotherapy, and the treatment relies on in vivo stimulation of the host's innate immune system using administration of immune response modifiers (such as polypeptides and nucleic acids provided herein) to react to diseased cells.

[0246] The drugs and compositions provided herein can be used alone or in combination with conventional treatment regimens such as surgery, irradiation, chemotherapy, and / or bone marrow transplantation (autologous, syngeneic, allogeneic or unrelated).

[0247] The terms "immunization" or "vaccination" describe the process of treating a subject for therapeutic or prophylactic reasons to induce an immune response.

[0248] The term "in vivo" relates to the situation within a subject.

[0249] The terms "subject", "individual", "organism" or "patient" are used interchangeably and relate to vertebrates, preferably mammals. For example, mammals in the context of the present invention include humans, non-human primates, domestic animals such as dogs, cats, sheep, cattle, goats, pigs, horses, etc., laboratory animals such as mice, rats, rabbits, guinea pigs, etc., and captive animals such as zoo animals. Also, the term "animal" as used herein includes humans. Also, the term "subject" can include a patient suffering from a disease, preferably a disease described herein, i.e., an animal, preferably a human.

[0250] The term "autologous" is used to describe all things derived from the same subject. For example, "autotransplantation" refers to the transplantation of tissue or organs derived from the same subject. Such procedures are advantageous for overcoming immunological barriers that would otherwise result in rejection.

[0251] The term "heterologous" is used to describe something consisting of multiple different elements. As an example, transplanting the bone marrow of one individual into another individual constitutes heterologous transplantation. A heterologous gene is a gene derived from a source other than the subject.

[0252] As part of a composition for immunization or vaccination, preferably one or more of the agents described herein are administered together with one or more adjuvants for inducing or increasing an immune response. The term "adjuvant" relates to a compound that prolongs or enhances or accelerates an immune response. The compositions of the present invention preferably exert their effects without adding an adjuvant. Nevertheless, the compositions of the present application may contain any known adjuvant. Adjuvants include a heterogeneous group of compounds such as oil emulsions (e.g., Freund's adjuvant), inorganic compounds (such as alum), bacterial products (such as Bordetella pertussis toxin), liposomes, and immunostimulating complexes. Examples of adjuvants are saponins such as monophosphoryl-lipid-A (MPL SmithKline Beecham), QS21 (SmithKline Beecham), DQS21 (SmithKline Beecham, WO96 / 33739), QS7, QS17, QS18, and QS-L1 (So et al., 1997, Mol. Cells 7:178-186), incomplete Freund's adjuvant, complete Freund's adjuvant, vitamin E, montanid, alum, CpG oligonucleotides (Krieg et al., 1995, Nature 374:546-549), and various water-in-oil emulsions prepared from biodegradable oils such as squalene and / or tocopherol.

[0253] Other substances that stimulate the patient's immune response can also be administered. For example, due to its regulatory properties on lymphocytes, it is possible to use cytokines during vaccination. Such cytokines include, for example, interleukin-12 (IL-12) (see Science 268: pages 1432-1434, 1995), GM-CSF, and IL-18, which have been shown to increase the protective effect of vaccines.

[0254] There are several compounds that enhance the immune response and can thus be used in vaccination. The compounds include costimulatory molecules provided in the form of proteins such as B7-1 and B7-2 (CD80 and CD86, respectively) or nucleic acids.

[0255] According to the present invention, a "tumor specimen" is a sample such as a body sample containing tumor or cancer cells such as circulating tumor cells (CTCs), particularly a tissue sample and / or a cell sample including body fluids. According to the present invention, a "non-tumorigenic specimen" is a sample such as a body sample that does not contain tumor or cancer cells such as circulating tumor cells (CTCs), particularly a tissue sample and / or a cell sample including body fluids. Such body samples can be obtained in conventional manners such as by tissue biopsy including punch biopsy and by collecting blood, bronchial aspirate, sputum, urine, feces, or other body fluids. According to the present invention, the term "sample" also includes fractions or isolates of biological samples, such as processed samples such as nucleic acid or cell isolates.

[0256] The medicaments, vaccines, and compositions having therapeutic activity described herein can be administered via any conventional route, including those by injection or infusion. Administration can be carried out, for example, orally, intravenously, intraperitoneally, intramuscularly, subcutaneously, or transdermally. In one embodiment, administration is carried out intranodally, such as by injection into a lymph node. Other administration forms envision ex vivo transfection of antigen-presenting cells such as dendritic cells using the nucleic acids described herein, followed by administration of the antigen-presenting cells.

[0257] The agents described in this specification are administered in an effective amount. "Effective amount" means an amount that, alone or together with additional dosages, achieves a desired response or desired effect. In the case of treating a particular disease or a particular condition, the desired response preferably relates to inhibiting the course of the disease. This includes delaying the progression of the disease, particularly interrupting or reversing the progression of the disease. Also, the desired response in the treatment of a disease or condition may be delaying or preventing the onset of said disease or said condition.

[0258] An effective amount of the agent described in this specification depends on parameters specific to the patient, including the condition being treated, the severity of the disease, age, physiological state, size and weight, the duration of the treatment, the type of concomitant treatment (if any), the particular route of administration and similar factors. Accordingly, the dosage of the agent described in this specification may depend on various such parameters. If the response in a patient is insufficient at the initial dosage, higher dosages (or effectively higher dosages achieved by a different, more local route of administration) may be used.

[0259] The pharmaceutical composition of the present invention is preferably sterile and contains a substance having a therapeutically active amount to produce a desired response or desired effect.

[0260] The pharmaceutical compositions of the present invention are generally administered in a pharmaceutically compatible amount in a pharmaceutically compatible preparation. The term "pharmaceutically compatible" refers to non-toxic substances that do not interact with the active components of the pharmaceutical composition. Such preparations usually contain salts, buffering substances, preservatives, carriers, adjuvants, such as immunopotentiating substances such as CpG oligonucleotides, cytokines, chemokines, saponins, GM-CSF and / or RNA, and, where appropriate, other compounds having therapeutic activity. When used in medicine, the salts should be pharmaceutically compatible. However, pharmaceutically incompatible salts may be used in the preparation of pharmaceutically compatible salts and are included in the present invention. Such pharmacologically and pharmaceutically compatible salts include, in a non-limiting manner, those prepared from acids such as hydrochloric acid, hydrobromic acid, sulfuric acid, nitric acid, phosphoric acid, maleic acid, acetic acid, salicylic acid, citric acid, formic acid, malonic acid, succinic acid, etc. Pharmaceutically compatible salts can also be prepared as alkali metal salts or alkaline earth metal salts such as sodium salts, potassium salts or calcium salts.

[0261] The pharmaceutical compositions of the present invention may contain a pharmaceutically compatible carrier. The term "carrier" refers to an organic or inorganic component of natural or synthetic nature that combines the active components to facilitate application. According to the present invention, the term "pharmaceutically compatible carrier" includes one or more compatible solid or liquid fillers, diluents or encapsulating substances suitable for administration to a patient. The components of the pharmaceutical compositions of the present invention are usually such that no interactions occur that substantially impair the desired pharmaceutical effectiveness.

[0262] The pharmaceutical compositions of the present invention may contain suitable buffering substances such as acetic acid in salts, citric acid in salts, boric acid in salts and phosphoric acid in salts.

[0263] Also, the pharmaceutical composition may, where appropriate, also contain suitable preservatives such as benzalkonium chloride, chlorobutanol, parabens and thimerosal.

[0264] Pharmaceutical compositions are usually provided in a uniform dosage form and can be prepared in a manner known per se. The pharmaceutical compositions of the present invention can be, for example, in the form of capsules, tablets, lozenges, solutions, suspensions, syrups, elixirs, or in the form of emulsions.

[0265] Compositions suitable for parenteral administration usually include sterile aqueous or non-aqueous preparations of the active compound, which are preferably isotonic with the recipient's blood. Examples of compatible carriers and solvents are Ringer's solution and isotonic sodium chloride solution. Furthermore, usually sterile, non-volatile oils are used as a medium for solutions or suspensions.

[0266] The present invention will be described in detail by the following figures and examples, which are used for illustrative purposes only and are not intended to be limiting. Thanks to the description and examples, further embodiments included in the present invention are also available to those skilled in the art.

Brief Description of the Drawings

[0267]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21

[0268] (Example) The techniques and methods used in this specification are carried out in the manner described in this specification or known per se and, for example, as described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd Edition, (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York. All methods, including the use of kits and reagents, are carried out according to the manufacturer's information unless specifically indicated otherwise.

[0269] (Example 1) Detection and prioritization of mutations The inventors first demonstrate the sequence profiling of tumor and normal samples to identify somatic mutations in an unbiased manner. The inventors demonstrate for the first time not only that this can be done with bulk tumor samples, but also that mutations from individual circulating tumor cells can be identified. Next, the inventors prioritize the mutations for inclusion in a poly-neo-epitope vaccine based on the predicted immunogenicity of the mutations and demonstrate that the identified mutations are indeed immunogenic.

[0270] Detection of mutations Theoretical basis for using CTC: Detection of circulating tumor cells (CTC) from the peripheral blood of cancer patients is an established independent prognostic marker of the clinical course of tumors (Pantel et al., Trends Mol Med 2010;16(9):398-406). For years, the clinical significance of CTC has been the subject of intensive scientific and clinical research in oncology. Detection of CTC in the blood of patients with metastatic breast, prostate, and colorectal cancers has prognostic relevance and has been shown to provide additional information to conventional imaging techniques and other prognostic tumor biomarkers. Serial blood samples taken from patients before treatment with therapeutic agents (systemic or targeted), during the initial stages of that treatment, and after that treatment provide information on treatment response / failure. Molecular analysis of drug-resistant CTC can provide further insights into resistance mechanisms (e.g., mutations in specific signaling pathways or loss of target expression) in individual patients. Further possibilities from CTC profiling and genetic characterization are the identification of novel cancer targets for the development of new targeted therapies. This new diagnostic strategy is called "liquid tumor biopsy". This profiling can be done quickly and repeatedly and requires only the patient's blood and no surgery, so it will provide a "real-time" view of the tumor state.

[0271] Mutations from tumor cells: The inventors demonstrated that mutations can be identified using B16 melanoma cells, exome capture for extracting protein-coding regions, next-generation sequencing using the inventors' HiSeq 2000, and then bioinformatics analysis using the inventors' "iCAM" software pipeline (Figure 1). 2448 non-synonymous mutations were identified and 50 were selected for confirmation. All 50 somatic mutations could be confirmed.

[0272] The following is an example of the somatic mutation protein impact found in B16 melanoma cells. Kif18b, NM_197959, exon 3 Mutation (+15 aa) SPSKPSFQEFVDWEN VSPELNSTDQPFLPS Wild type (+15 aa) SPSKPSFQEFVDWE K VSPELNSTDQPFLPS

[0273] Mutations from individual circulating tumor cells (CTCs): Next, the inventors were able to identify somatic mutations specific to the tumor from NGS profiling of RNA from a single CTC. Labeled B16 melanoma cells were intravenously injected into the tail of a mouse, the mouse was sacrificed, blood was collected from the heart, the cells were sorted to recover labeled circulating B16 cells (CTCs), RNA was extracted, SMART-based cDNA synthesis and non-specific amplification were performed, and then NGS RNA-Seq assay and partial sequence data analysis were performed (below).

[0274] The inventors profiled 8 individual CTCs and identified somatic mutations. Furthermore, in 8 out of 8 cells, previously identified somatic mutations were identified. In multiple cases, the data showed heterogeneity at the individual cell level. For example, in the gene Snx15 at position 144078227 on chromosome 2 (assembly mm9), 2 cells showed the reference nucleotide (C), while 2 cells showed the mutant nucleotide (T).

[0275] This demonstrates that somatic mutations can be identified by profiling individual CTCs, which is the fundamental path to "real-time" iVAC ( individualized vaccine) where the patient is repeatedly profiled and the results reflect the current state of the patient rather than the state at a previous time point. Furthermore, this demonstrates that it is possible to identify heterogeneous somatic mutations present in subpopulations of tumor cells and to evaluate the frequency of mutations such as to identify major and rare mutations.

[0276] Method Samples: For the profiling experiments, the samples included 5-10 mm tail samples from C57BL / 6 mice ("Black6") and highly aggressive B16F10 murine melanoma cells ("B16") originally derived from Black6 mice.

[0277] Circulating tumor cells (CTCs) were generated using fluorescently labeled B16 melanoma cells. B16 cells were resuspended in PBS and an equal volume of freshly prepared CFSE solution (5 μM in PBS) was added to the cells. The samples were gently mixed by vortexing and then incubated at room temperature for 10 minutes. To stop the labeling reaction, an equal volume of PBS containing 20% FSC was added to the samples and gently mixed by vortexing. After incubation at room temperature for 20 minutes, the cells were washed twice with PBS. Finally, the cells were resuspended in PBS and injected intravenously (i.v.) into the mice. Three minutes later, the mice were sacrificed and blood was collected.

[0278] Red blood cells were lysed from the blood samples by adding 1.5 ml of freshly prepared PharmLyse Solution (Beckton Dickinson) per 100 μl of blood. After one wash step, 7-AAD was added to the samples and incubated at room temperature for 5 minutes. Two wash steps were then performed following the incubation and the samples were resuspended in 500 μl of PBS.

[0279] Circulating B16 cells labeled with CFSE were sorted using an Aria I cell sorter (BD). Single cells were sorted onto 96-well V-bottom plates plated with 50 μl / well of RLT buffer (Quiagen). After sorting, the plates were stored at -80 °C until nucleic acid extraction and sample preparation were initiated.

[0280] Nucleic acid extraction and sample preparation: Nucleic acids from B16 cells (DNA and RNA) and Black6 tail tissue (DNA) were extracted using the Qiagen DNeasy Blood and Tissue kit (for DNA) and the Qiagen RNeasy Micro kit (for RNA).

[0281] For each individually sorted CTC, RNA was extracted and SMART-based cDNA synthesis and non-specific amplification were performed. RNA from the sorted CTC cells was extracted using the RNeasy Micro Kit (Qiagen, Hilden, Germany) according to the supplier's instructions. The modified BD SMART protocol was used for cDNA synthesis. Mint Reverse Transcriptase (Evrogen, Moscow, Russia) was combined with TS-Short (Eurogentec S.A., Seraing, Belgium) that introduced an oligo(dT)-T-primer long and an oligo(riboG) sequence to enable the generation of an extended template and template switching by the terminal transferase activity of the reverse transcriptase [Chenchik, A., Y. et al., 1998. Generation and use of high quality cDNA from small amounts of total RNA by SMART PCR., Gene Cloning and Analysis by RT-PCR., Edited by P. L. J. Siebert, BioTechniques Books, Natick, Massachusetts. pp. 305-319]. The first-strand cDNA synthesized according to the manufacturer's instructions was subjected to 35 amplification cycles in the presence of 200 μM dNTPs using 5 U of PfuUltra Hotstart High-Fidelity DNA Polymerase (Stratagene, La Jolla, California) and 0.48 μM of the primer TS-PCR primer (cycling conditions: 2 minutes at 95°C, 30 seconds at 94°C, 30 seconds at 65°C, 1 minute at 72°C, and a final extension of 6 minutes at 72°C). The success of CTC gene amplification was monitored by actin and GAPDH controlled with specific primers.

[0282] Next-generation sequencing, DNA sequencing: In this case, an Agilent Sure-Select solution-based capture assay designed to capture all mouse protein-coding regions was used to perform exome capture for DNA re-sequencing [Gnirke A et al., Solution hybrid selection with ultra-long oligonucleotides for massively parallel targeted sequencing. Nat Biotechnol 2009, 27:182-189].

[0283] Briefly, 3 μg of purified genomic DNA was fragmented to 150-200 bp using a Covaris S2 sonicator. The gDNA fragments were end-repaired using T4 DNA polymerase and Klenow DNA polymerase and 5'-phosphorylated using T4 polynucleotide kinase. The blunt-ended gDNA fragments were 3'-adenylated using Klenow fragment (3' to 5' exo-minus). T4 DNA ligase was used to ligate the 3' single T-overhang Illumina paired-end adapter to the gDNA fragments using a 10:1 adapter:genomic DNA insert molar ratio. The adapter-ligated gDNA fragments were concentrated prior to capture, and flow cell-specific sequences were added using 4 PCR cycles with Illumina PE PCR primers 1.0 and 2.0 and Herculase II polymerase (Agilent).

[0284] The PCR-enriched gDNA fragments ligated with 500 ng of adapter were hybridized with Agilent's SureSelect biotinylated mouse whole exome RNA library bait at 65 °C for 24 hours. The hybridized gDNA / RNA bait complex was removed using streptavidin-coated magnetic beads. The gDNA / RNA bait complex was washed, and the RNA bait was cleaved during elution in SureSelect elution buffer, leaving the PCR-enriched gDNA fragments ligated to the captured adapter. The gDNA fragments were PCR amplified after capture using Herculase II DNA polymerase (Agilent) and SureSelect GA PCR primers for 10 cycles.

[0285] All cleanups were performed using 1.8× volume of AMPure XP magnetic beads (Agencourt). All quality controls were performed using Invitrogen's Qubit HS assay, and the fragment sizes were determined using Agilent's 2100 Bioanalyzer HS DNA assay.

[0286] The exome-enriched gDNA library was clustered at 7 pM on the cBot using the Truseq SR Cluster Kit v2.5 and sequenced on the Illumina HiSeq2000 at 50 bp using the Truseq SBS Kit-HS 50 bp.

[0287] Next-generation sequencing, RNA sequencing (RNA-Seq): Barcoded mRNA-seq cDNA libraries were prepared from 5 μg of total RNA using a modified version of the Illumina mRNA-seq protocol. mRNA was isolated using Seramag oligo(dT) magnetic beads (Thermo Scientific). The isolated mRNA was fragmented using divalent cations and heat, resulting in fragments in the range of 160 - 220 bp. The fragmented mRNA was converted to cDNA using random primers and SuperScriptII (Invitrogen), and then the second strand was synthesized using DNA polymerase I and RNaseH. The cDNA was end-repaired using T4 DNA polymerase and Klenow DNA polymerase and 5'-phosphorylated using T4 polynucleotide kinase. The blunt-ended cDNA fragments were 3'-adenylated (3' to 5' exonuclease) using Klenow fragment. The 3' single T-overhang Illumina multiplex-specific adapters were ligated using T4 DNA ligase at a molar ratio of adapter:cDNA insert of 10:1.

[0288] The cDNA libraries were purified and size-selected at 200 - 220 bp using a 2% SizeSelect gel of E-Gel (Invitrogen). Concentration, addition of Illumina hexamer indexes and flow cell-specific sequences were performed by PCR using Phusion DNA polymerase (Finnzymes). All cleanups were performed using 1.8× volume of Agencourt AMPure XP magnetic beads. All quality controls were performed using Invitrogen's Qubit HS assay, and the fragment sizes were determined using Agilent's 2100 Bioanalyzer HS DNA assay.

[0289] Barcoded RNA-Seq libraries were clustered at 7 pM using the Truseq SR Cluster Kit v2.5 on the cBot and sequenced at 50 bp using the Truseq SBS Kit - HS 50 bp on the Illumina HiSeq2000.

[0290] CTC: For RNA-Seq profiling of CTCs, a modified version of this protocol using 500 - 700 ng of SMART-amplified cDNA was used, paired-end adapters were ligated, and PCR enrichment was performed using Illumina PE PCR primers 1.0 and 2.0.

[0291] NGS data analysis, gene expression: To determine the expression values, the output sequence reads from RNA samples from Illumina HiSeq 2000 were preprocessed according to the Illumina standard protocol. This included filtering of low-quality reads and demultiplexing. In RNA-Seq transcriptome analysis, the parameter "-v 2 -best" was used for genome alignment and the default parameters were used for transcript alignment, and bowtie (version 0.12.5) [Langmead B. et al., Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10:R25] was used to align the sequence reads to the reference genome sequence [Mouse Genome Sequencing Consortium. Initial sequencing and comparative analysis of the mouse genome. Nature, 420, 520-562 (2002)]. The alignment coordinates were compared with the exon coordinates of RefSeq transcripts [Pruitt KD. et al., NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins. Nucleic Acids Res. January 1, 2005;33(Database issue):D501-4], and for each transcript, the number of overlapping alignments was recorded. Sequence reads that could not be aligned to the genome sequence were aligned to a database of all possible exon-exon junction sequences of RefSeq transcripts.The number of reads aligned to the splice junctions was summed with the respective transcript counts obtained in the previous step, and for each transcript, it was normalized against RPKM (reads mapped / exon model in kilobases / millions of mapped reads [Mortazavi, A. et al. (2008). Mapping and quantifying mammalian transcriptomes by rna-seq. Nat Methods, 5(7):621 - 628]). Values for gene expression and exon expression were each calculated based on the normalized number of reads overlapping with the respective gene or exon.

[0292] Mutation discovery, bulk tumors: Using bwa (version 0.5.8c) [Li H. and Durbin R. (2009) Fast and accurate short read alignment with Burrows - Wheeler Transform. Bioinformatics, 25:1754 - 60], with default options, 50 - nt single - end reads from Illumina HiSeq 2000 were aligned to the reference mouse genome assembly mm9. Ambiguous reads - reads mapped to multiple positions in the genome - were removed, the remaining alignments were filtered, indexed, converted to binary compressed format (BAM), and a shell script was used to convert the read quality scores from Illumina standard phred + 64 to standard Sanger quality scores.

[0293] For each sequencing lane, mutations were identified using three software programs including samtools (version 0.1.8) [Li H., Improving SNP discovery by base alignment quality. Bioinformatics. April 15, 2011;27(8):1157 - 1158. Epub February 13, 2011], GATK (version 1.0.4418) [McKenna A. et al., The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. September 2010;20(9):1297 - 1303. Epub July 19, 2010], and SomaticSniper (http: / / genome.wustl.edu / software / somaticsniper). For samtools, the author-recommended options and filtering criteria were used, including the first round of filtering and a maximum coverage of 200. In the second round of filtering with samtools, the minimum indel quality score was 50, and the minimum quality for point mutations was 30. For GATK mutation calling, the best practice guidelines designed by the authors as presented in the GATK user guide were followed (http: / / www.broadinstitute.org / gsa / wiki / index.php / The_Genome_Analysis_Toolkit). The variant score recalibration step was omitted and replaced by hard filtering options. For SomaticSniper mutation calling, the default options were used, and only predicted mutations with a "somatic score" of 30 or more were further considered.

[0294] Somatic Mutation Discovery, CTC: Similar to the iCAM process for bulk tumors, 50-nt single-end reads from Illumina HiSeq 2000 were aligned to the reference mouse genome assembly mm9 using bwa (version 0.5.8c) with default options. Since the NGS reads of CTC were derived from an RNA-Seq assay, bowtie (described above) was used to also align the reads to the transcriptome sequences including exon-exon junctions. Using all the alignments, the nucleotide sequences from the reads were compared to both the reference genome and the B16 mutations derived from the bulk tumor. The identified mutations were evaluated using both Perl scripts and manual methods, using the software programs samtools and IGV (Integrated Genome Viewer) to image the results.

[0295] The output of "Somatic Mutation Discovery" is the identification of somatic mutations in tumor cells from NGS data to a list of mutations. In the B16 samples, 2448 somatic mutations were identified using exome resequencing.

[0296] Prioritization of Mutations Next, the potential of a pipeline for prioritizing mutations for vaccine inclusion is demonstrated. This method, called the "Individual Cancer Mutation Detection Pipeline" (iCAM), identifies and prioritizes somatic mutations through a series of steps incorporating multiple state-of-the-art algorithms and bioinformatics methods. The output of this process is a list of somatic mutations prioritized based on potential immunogenicity.

[0297] Identification of somatic mutations: For both B16 and Black6 samples (mutation discovery, as above), mutations are identified using three different algorithms. The first iCAM step is to create a high-confidence list of somatic mutations by combining the output lists from each algorithm. GATK and samtools report variants in one sample compared to the reference genome. To select high-confidence mutations with few false positives for a given sample (i.e., tumor or normal), mutations identified in all replicates are selected. Then, variants that are present in the tumor sample but not in the normal sample are selected. SomaticSniper automatically reports potential somatic mutations from pairs of tumor and normal data. Results are further filtered through the intersection of results obtained from replicates. To remove as many false positive calls as possible, the lists of mutations derived from using all three algorithms and all replicates were intersected. The final step for each somatic mutation is to assign a confidence value (p-value) for each mutation based on coverage depth, SNP quality, consensus quality, and mapping quality.

[0298] Mutation impact: The impact of filtering, consensus, and somatic mutations is determined by scripts within the iCaM mutation pipeline. First, mutations present in genomic regions that are not unique within the genome, such as those occurring in some protein paralogs and pseudogenes, are excluded from the analysis, and similarly, sequence reads that align to multiple positions are removed. Second, it is determined whether the mutation is present in the transcript. Third, it is determined whether the mutation is present in the protein-coding region. Fourth, the transcribed sequences with and without the mutation are translated to determine whether there is a change in the amino acid sequence.

[0299] Expression of mutations: The iCAM pipeline selects somatic mutations found in genes and exons expressed in tumor cells. The expression level is determined by NGS RNA-Seq of tumor cells (described above). The number of overlapping reads in genes and exons indicates the expression level. These counts are normalized against RPKM (reads per kilobase of exon model per million mapped reads, [Mortazavi A. et al., Mapping and quantifying mammalian transcriptomes by RNA-Seq. Nat Methods. July 2008;5(7):621-628. Epub May 30, 2008]), and those expressed above 10 RPKM are selected.

[0300] MHC binding: To determine the likelihood that epitopes containing mutant peptides bind to MHC molecules, the iCAM pipeline runs a modified version of MHC prediction software from the Immune Epitope Database (http: / / www.iedb.org / ). Local installations include modifications to optimize the data flow through the algorithm. For B16 and Black6 data, predictions were performed using all available black6 MHC class I alleles and all epitopes for each peptide length. Considering all MHC alleles and all potential epitopes overlapping with the mutations, mutations within the range of epitopes ranked within the 95th percentile of the prediction score distribution of the IEDB training data (http: / / mhcbindingpredictions.immuneepitope.org / dataset.html) are selected.

[0301] Mutation selection criteria: Somatic mutations were selected based on the following criteria: a) having a unique sequence content, b) identified by all three programs, c) high mutation confidence, d) non-synonymous protein changes, e) high transcript expression, f) and favorable MHC class I binding predictions.

[0302] The output of this process is a list of somatic mutations prioritized based on potential immunogenicity. There are 2,448 somatic mutations in B16 melanoma cells. Of these mutations, 1,247 are found in gene transcripts. Of these, 734 cause non-synonymous protein changes. Of these, 149 are in genes expressed in tumor cells. Of these, 102 of these are expressed and predicted to have non-synonymous mutations presented on MHC molecules. These 102 potential immunogenic mutations are then sent for mutation confirmation (below).

[0303] Mutation Confirmation Somatic mutations from DNA exome re-sequencing were confirmed by either of two methods: re-sequencing of the mutated region and RNA-Seq analysis.

[0304] For mutation confirmation by re-sequencing, the genomic region containing the mutation was amplified by standard PCR from both 50 ng of tumor DNA and normal control DNA. The size of the amplification products ranged from 150 to 400 nucleotides. The specificity of the reaction was controlled by loading the PCR products on a Qiaxel device (Qiagen). The PCR products were purified using the minElute PCR purification kit (Qiagen). Specific PCR products were sequenced using the standard Sanger sequencing method (Eurofins), followed by electropherogram analysis.

[0305] Moreover, the confirmation of mutations was also achieved by examining tumor RNA. Tumor gene and exon expression values were generated from RNA-Seq (RNA NGS) that yields nucleotide sequences mapped to transcripts and counted. The inventors examined the sequence data itself to identify mutations in tumor samples [Berger MF. et al., Integrative analysis of the melanoma transcriptome. Genome Res. April 2010;20(4):413-27. Epub February 23, 2010], resulting in an independent confirmation of the identified somatic mutations derived from DNA.

[0306]

Table 1A

[0307]

Table 1B

[0308] (Example 2) The IVAC selection algorithm enables the detection of immunogenic mutations To investigate whether specific T cell responses can be induced against the confirmed mutations from B16F10 melanoma cells, naive C57BL / 6 mice (n = 5 mice / peptide) were immunized subcutaneously twice (d0, d7) with 100 μg of peptide containing either the mutant or wild-type amino acid sequence (+50 μg of poly I:C, as an adjuvant) (see Table 2). All peptides had a length of 27 amino acids, and the mutant / wild-type amino acids were located centrally. On day 12, the mice were sacrificed and spleen cells were collected. As a readout, 5×10 5 spleen cells / well were used as effectors, and 5×10 4IFNγ ELISpot was performed using individual bone marrow dendritic cells (2 μg / ml) as target cells. Effector spleen cells were tested against mutant peptides, wild-type peptides, and control peptides (vesicular stomatitis virus nucleoprotein, VSV-NP).

[0309] In 44 sequences tested, 6 of them were observed to induce T cell immunity only directed against the mutant sequences and not against those directed against the wild-type peptides (Figure 3).

[0310] The data demonstrate that tumor-specific T cell immunity can be induced after utilization as a peptide vaccine in naïve mice using the identified and prioritized mutations.

[0311]

Table 2

[0312] (Example 3) The identified mutations can provide therapeutic anti-tumor immunity To verify whether the identified mutations have the potential to confer anti-tumor immunity after vaccination of naïve mice, the inventors investigated this question using the peptide of mutation 30, which has been shown to induce a T cell reactivity selective for the mutation. On day 0, B16F10 cells (7.5×10 4 cells) were subcutaneously inoculated. On days -4, +2, and +9, the mice were vaccinated with peptide 30 (see Table 1, 100 μg of peptide + 50 μg of poly I:C, subcutaneous). The control group received poly I:C only (50 μg, subcutaneous). Tumor growth was monitored every other day. On day +16, it was observed that only 1 out of 5 mice in the peptide vaccine group developed a tumor, while 4 out of 5 mice in the control group showed tumor growth.

[0313] The data demonstrate that a peptide sequence incorporating B16F10-specific mutations can confer anti-tumor immunity capable of efficiently destroying tumor cells (see Figure 4). Since B16F10 is a very aggressive tumor cell line, the finding that the method applied to identify and prioritize mutations ultimately led to the selection of mutations that are already potent as vaccines in themselves is an important proof of concept for the entire process.

[0314] (Example 4) Data supporting polyepitope antigen presentation Validated mutations from the patient's protein-coding regions constitute a pool from which candidates can be selected for the assembly of poly-neoepitope vaccine templates to be used as precursors for the GMP production of RNA vaccines. Suitable vector cassettes for the vaccine backbone have already been described (Holtkamp, S. et al., Blood, 108:4009-4017, 2006, Kreiter, S. et al., Cancer Immunol. Immunother., 56:1577-1587, 2007, Kreiter, S. et al., J. Immunol., 180:309-318, 2008). Preferred vector cassettes have their coding and untranslated regions (UTRs) modified to ensure maximum translation of the encoded protein over a long period (Holtkamp, S. et al., Blood, 108:4009-4017, 2006, Kuhn, A. N. et al., Gene Ther., 17:961-971, 2010). Furthermore, the vector backbone contains an antigen routing module for the co-expansion of cytotoxic and helper T cells (Kreiter, S. et al., Cancer Immunol. Immunother., 56:1577-1587, 2007, Kreiter, S. et al., J. Immunol., 180:309-318, 2008, Kreiter, S. et al., Cancer Research, 70(22), 9031-9040, 2010) (Figure 5). Importantly, the inventors have demonstrated that such RNA vaccines can be used to simultaneously present multiple MHC class I and class II epitopes.

[0315] The IVAC poly-neoepitope RNA vaccine sequences are constructed from stretches of up to 30 amino acids that contain mutations in the center. These sequences are linked head-to-tail via short linkers to form poly-neoepitope vaccines that encode up to 30 or more selected mutations and their flanking regions. These patient-specific, individually tailored inserts are codon-optimized and cloned into the RNA backbone described above. Quality control of such constructs includes in vitro transcription and expression in cells for verification of functional transcription and translation. Analysis of translation is performed using an antibody against the C-terminal targeting domain.

[0316] (Example 5) Scientific proof of concept of RNA poly-neoepitope constructs The concept of RNA poly-neoepitopes is based on long in vitro transcribed mRNAs consisting of sequentially arranged sequences encoding mutant peptides linked by linker sequences (see Figure 6). The coding sequences are selected from non-synonymous mutations and are always constructed from codons of mutant amino acids flanked by regions of 30 - 75 base pairs from the original sequence configuration. The linker sequences preferentially encode amino acids that are not processed by the cell's antigen processing machinery. The in vitro transcription construct is based on the pST1-A120 vector containing the T7 promoter, tandem beta-globin 3'UTR sequence and a 120 bp poly(A) tail, which has been shown to increase RNA stability and translation efficiency and thus enhance the T cell stimulating capacity of the encoded antigen (Holtkamp S. et al., Blood 2006; PMID:16940422). Furthermore, an MHC class I signal peptide fragment containing a stop codon adjacent to the poly-linker sequence for epitope cloning as well as the transmembrane and cytosolic domains (MHC class I transport signal or MITD) were inserted (Kreiter S. et al., J. Immunol., 180:309 - 318, 2008). The latter has been shown to increase antigen presentation and thus enhance the expansion of antigen-specific CD8+ and CD4+ T cells and improve effector function.

[0317] For the initial proof of concept, a bivalent epitope vector, i.e., one encoding a single polypeptide containing two mutated epitopes, was used. A codon-optimized sequence encoding (i) a mutated epitope of 20-50 amino acids, (ii) a glycine / serine-rich linker, (iii) a second mutated epitope of 20-50 amino acids, and (iv) a further glycine / serine-rich linker adjacent to appropriate recognition sites for restriction endonucleases for cloning into the construct based on pST1 described above was designed and synthesized by a commercial supplier (Geneart, Regensburg, Germany). After sequence confirmation, these were cloned into the pST1-based vector backbone to obtain the construct shown in FIG. 6.

[0318] The plasmid based on pST1-A120 described above was linearized using a class II restriction endonuclease. The linearized plasmid DNA was purified by phenol-chloroform extraction and ethanol precipitation. The linearized vector DNA was quantified by spectrophotometer and subjected to in vitro transcription essentially as described by Pokrovskaya and Gurevich (1994, Anal. Biochem. 220:420-423). A cap analog was added to the transcription reaction to obtain RNA with a correspondingly modified 5'-cap structure. During the reaction, GTP was present at 1.5 mM and the cap analog was present at 6.0 mM. All other NTPs were present at 7.5 mM. At the end of the transcription reaction, the linearized vector DNA was digested with 0.1 U / μl of TURBO DNase (Ambion, Austin, Tex., USA) for 15 minutes at 37°C. RNA was purified from these reactions using the MEGAclear kit (Ambion, Austin, Tex., USA) according to the manufacturer's protocol. The concentration and quality of the RNA were evaluated by spectrophotometry and analysis on a 2100 Bioanalyzer (Agilent, Santa Clara, Calif., USA).

[0319] To demonstrate that the mutant amino acids are incorporated and that the sequences adjacent to the linker sequence at the 5'- and 3'-positions can be processed, presented, and recognized by antigen-specific T cells, the inventors used T cells from peptide-vaccinated mice as effector cells. In an IFNγ ELISpot, it was tested whether the T cells induced by the above-described peptide vaccination could recognize either target cells (bone marrow dendritic cells, BMDC) pulsed with the peptide (2 μg / ml for 2 hours at 37 °C and 5% CO2) or transfected with RNA by electroporation (20 μg, generated as described above). As illustrated in Figure 7, for mutations 12 and 30 (see Table 2), it was observed that the RNA construct could generate epitopes recognized by mutant-specific T cells.

[0320] Using the provided data, it was demonstrated that poly-neoepitopes encoded in RNAs containing glycine / serine-rich linkers can be translated and processed within antigen-presenting cells, resulting in the presentation of the correct epitopes recognized by antigen-specific T cells.

[0321] (Example 6) Design of Poly-Neoepitope Vaccines - Relevance of Linkers The poly-neoepitope RNA construct contains a backbone construct in which a peptide encoding multiple somatic mutations linked to a linker peptide sequence is disposed therein. In addition to codon optimization by the backbone and increased RNA stability and translation efficiency, one embodiment of the RNA poly-neoepitope vaccine contains a linker designed to increase MHC class I and II presentation of the antigen peptide and decrease the presentation of harmful epitopes.

[0322] Linker: The linker sequence is designed to link multiple mutant-containing peptides. While enabling the generation and presentation of mutant epitopes, the linker should prevent the generation of harmful epitopes such as those created by junction stitching between adjacent peptides or between the linker sequence and endogenous peptides. These "junction" epitopes can not only compete with the intended epitopes to be presented on the cell surface and reduce the effectiveness of the vaccine, but also give rise to unwanted autoimmune reactions. Therefore, the inventors designed the linker sequence to a) avoid the generation of "junction" peptides that bind to MHC molecules, b) avoid proteasome processing that generates "junction" peptides, and c) be efficiently translated and processed by the proteasome.

[0323] To avoid the generation of "junction" peptides that bind to MHC molecules, the inventors compared various linker sequences. Glycine, for example, inhibits strong binding at MHC-binding groove positions [Abastado JP. et al., J Immunol. October 1, 1993;151(7):3569-3575]. The inventors examined multiple linker sequences and multiple linker lengths and calculated the number of "junction" peptides that bind to MHC molecules. Software tools from the Immune Epitope Database (IEDB, http: / / www.immuneepitope.org / ) were used to calculate the likelihood that a given peptide sequence contains ligands that bind to MHC class I molecules.

[0324] In the B16 model, 102 expressed nonsynonymous somatic mutations predicted to be presented on MHC class I molecules were identified. Using 50 confirmed mutations, the inventors computationally designed various vaccine constructs including either not using a linker or using various linker sequences, and calculated the number of harmful "junction" peptides using the IEDB algorithm (Figure 8).

[0325] Table 5 (Table 5) shows the results for several different linkers, different linker lengths, and with and without the use of linkers. The range of the number of MHC-binding junction peptides was 2 - 91 for epitope predictions of 9 and 10 amino acids (top and middle). The size of the linker affects the number of junction peptides (bottom). In this sequence, the fewest 9-amino acid epitopes are predicted with the 7-amino acid linker sequence GGSGGGG.

[0326] Also, Linker 1 and Linker 2 (see below) used in the experimentally tested RNA poly-neoepitope vaccine constructs also fortuitously had a low number of predicted junction neoepitopes. This also applies to the 9-mer and 10-mer predictions.

[0327] This demonstrates that the linker sequence is very important for the generation of poor MHC-binding epitopes. Furthermore, the length of the linker sequence affects the number of poor MHC-binding epitopes. The inventors have found that G-rich sequences interfere with the generation of MHC-binding ligands.

[0328]

Table 3

[0329]

Table 4

[0330]

Table 5

[0331] To avoid proteasome processing that can generate the "junction" peptide, we explored the use of various amino acids in the linker. Glycine-rich sequences impair proteasome processing [Hoyt MA et al. (2006). EMBO J 25(8):1720 - 1729, Zhang M. and Coffino P. (2004) J Biol Chem 279(10):8635 - 8641]. Thus, glycine-rich linker sequences act to minimize the number of linker-containing peptides that can be processed by the proteasome.

[0332] The linker should enable the mutant-containing peptide to be efficiently translated and processed by the proteasome. The amino acids glycine and serine are flexible [Schlessinger A and Rost B., Proteins. Oct 1, 2005;61(1):115 - 126], and including them in the linker results in a more flexible protein. We incorporated glycine and serine into the linker to increase protein flexibility, which should allow for more efficient translation and processing by the proteasome, which in turn should allow for better access to the encoded antigenic peptide.

[0333] Thus, the linker should be glycine-rich to prevent the generation of MHC-binding poor epitopes, should impede the proteasome's ability to process the linker peptide (which can be achieved by including glycine), and should be flexible to increase access to the mutant-containing peptide (which can be achieved by a combination of glycine and serine amino acids). Thus, in one embodiment of the vaccine construct of the present invention, the sequences GGSGGGGSGG and GGSGGGSGGS are preferably included as linker sequences.

[0334] (Example 7) RNA Poly - neoepitope vaccine The RNA poly-neoepitope vaccine construct is based on the pST1-A120 vector containing a T7 promoter, a tandem beta-globin 3'UTR sequence and a 120 bp poly(A) tail, which has been shown to increase the stability and translation efficiency of the RNA and thus enhance the T cell-stimulating ability of the encoded antigen ((Holtkamp S. et al., Blood 2006; PMID: 16940422). Furthermore, an MHC class I signal peptide fragment containing a stop codon adjacent to a poly-linker sequence for cloning epitopes, as well as a transmembrane and cytosolic domain (MHC class I transport signal or MITD) were inserted (Kreiter S. et al., J. Immunol., 180: 309-318, 2008). The latter has been shown to increase antigen presentation and thus enhance the expansion of antigen-specific CD8+ and CD4+ T cells and improve effector function.

[0335] To provide RNA poly-neoepitope constructs for 50 identified and validated mutations of B16F10, three RNA constructs were generated. The constructs consist of codon-optimized sequences encoding (i) a 25-amino acid mutated epitope, (ii) a glycine / serine-rich linker, (iii) a repeat of the mutated epitope sequence, followed by a glycine / serine-rich linker. Appropriate recognition sites for restriction endonucleases for cloning into the pST1-based construct described above flank the mutated epitope-containing sequences and the linker strands. The vaccine constructs were designed and synthesized by GENEART. After sequence confirmation, these were cloned into the pST1-based vector backbone to obtain RNA poly-neoepitope vaccine constructs.

[0336] Description of clinical procedures Clinical applications cover the following steps. · Eligible patients must consent to DNA analysis by next-generation sequencing. · Obtain tumor specimens and peripheral blood cells from routine diagnostic procedures (formalin-fixed paraffin-embedded tissues) and use them for mutation analysis as described. · Confirm the discovered mutations. · Design vaccines based on prioritization. For RNA vaccines, prepare the master plasmid template by gene synthesis and cloning. · Use the plasmid for clinical-grade RNA production, quality control, and release of RNA vaccines. · Send the vaccine drug products to each clinical trial center for clinical application. · RNA vaccines can be used as naked vaccines in formulation buffers or encapsulated in nanoparticles or liposomes for direct injection, for example, into lymph nodes, subcutaneously, intravenously, or intramuscularly. Alternatively, RNA vaccines can be used, for example, for in vitro transfection of adoptively transferred dendritic cells.

[0337] The entire clinical process takes less than 6 weeks. The "lag period" between patient informed consent and drug availability, including allowing the continuation of the standard treatment regimen until the investigational drug product becomes available, is carefully addressed in the clinical trial protocol.

[0338] (Example 8) Identification of Tumor Mutations and Their Utilization in Tumor Vaccination We applied NGS exome resequencing to discover mutations in the B16F10 murine melanoma cell line and identified 962 non-synonymous point mutations, which were in 563 expressed genes. Potential driver mutations were present in classical tumor suppressor genes (Pten, Trp53, Tp63, Pml), as well as genes involved in proto-oncogenic signaling pathways that control cell proliferation (e.g., Mdm1, Pdgfra), cell adhesion and migration (e.g., Fdz7, Fat1), or apoptosis (Casp9). Furthermore, B16F10 harbors mutations in Aim1 and Trrap, which have been previously described as being frequently altered in human melanoma.

[0339] The immunogenicity and specificity of 50 validated mutations were assayed using C57BL / 6 mice immunized with long peptides encoding the mutated epitopes. One-third (16 / 50) of them were shown to be immunogenic. Of these, 60% induced an immune response preferentially directed to the mutated sequences compared to the wild-type sequences.

[0340] The inventors tested this hypothesis in a tumor transplantation model. Immunization with the peptide gave in vivo tumor control in protective and therapeutic settings, whereby mutated epitopes containing single amino acid substitutions were identified as effective vaccines.

[0341] Animals C57BL / 6 mice (Jackson Laboratories) were housed at the University of Mines in accordance with federal and state policies in animal research.

[0342] Cells The B16F10 melanoma cell line was purchased from the American Type Culture Collection in 2010 (product: ATCC CRL-6475, lot number: 58078645). Early passages (passages 3 and 4) of the cells were used for tumor experiments. The cells were routinely tested for Mycoplasma. The cells have not been re-authenticated since receipt.

[0343] Next-generation sequencing Nucleic acid extraction and sample preparation: DNA and RNA from bulk B16F10 cells and DNA from C57BL / 6 tail tissue were extracted in triplicate using the Qiagen DNeasy Blood and Tissue Kit (for DNA) and the Qiagen RNeasy Micro Kit (for RNA).

[0344] DNA Exome Sequencing: Exome capture for DNA re-sequencing was performed in triplicate using an Agilent Sure-Select mouse solution-based capture assay designed to capture all mouse protein-coding regions (Gnirke A et al., Nat Biotechnol 2009;27:182-189). 3 μg of purified genomic DNA (gDNA) was fragmented to 150-200 bp using a Covaris S2 sonicator. The fragments were end-repaired, 5'-phosphorylated, and 3'-adenylated according to the manufacturer's instructions. Illumina paired-end adapters were ligated to the gDNA fragments using a 10:1 adapter:gDNA molar ratio. Illumina PE PCR primers 1.0 and 2.0 were used in 4 PCR cycles to add enrichment pre-capture and flow cell-specific sequences. 500 ng of adapter-ligated, PCR-enriched gDNA fragments were hybridized with Agilent's SureSelect biotinylated mouse whole exome RNA library bait at 65 °C for 24 hours. The hybridized gDNA / RNA bait complexes were removed using streptavidin-coated magnetic beads, washed, and the RNA bait was cleaved during elution in SureSelect elution buffer. These eluted gDNA fragments were PCR amplified for 10 cycles after capture. The exome-enriched gDNA library was clustered at 7 pM using the Truseq SR Cluster Kit v2.5 on a cBot and sequenced on an Illumina HiSeq2000 using the Truseq SBS Kit-HS 50 bp for 50 bp.

[0345] RNA Gene Expression “Transcriptome” Profiling (RNA-Seq): Barcoded mRNA-seq cDNA libraries were prepared in triplicate from 5 μg of total RNA (modified Illumina mRNA-seq protocol). mRNA was isolated using Seramag oligo(dT) magnetic beads (Thermo Scientific) and fragmented using divalent cations and heat. The resulting fragments (160 - 220 bp) were converted to cDNA using random primers and SuperScriptII (Invitrogen), and then the second strand was synthesized using DNA polymerase I and RNaseH. cDNA was end-repaired, 5'-phosphorylated, and 3'-adenylated according to the manufacturer's instructions. The 3'-single T-overhang Illumina multiplex-specific adapter was ligated using T4 DNA ligase with a 10:1 adapter:cDNA insert molar ratio. The cDNA library was purified and size-selected at 200 - 220 bp (E-Gel 2% SizeSelect gel, Invitrogen). Concentration, addition of Illumina hexamer indexes, and flow cell-specific sequences were performed by PCR using Phusion DNA polymerase (Finnzymes). All cleanups up to this step were performed using 1.8× volume of Agencourt AMPure XP magnetic beads. All quality control was performed using Invitrogen's Qubit HS assay, and fragment sizes were determined using Agilent's 2100 Bioanalyzer HS DNA assay. Barcoded RNA-Seq libraries were clustered and sequenced as described above.

[0346] NGS data analysis, gene expression: The output sequences read from the RNA samples were preprocessed according to the Illumina standard protocol, including filtering of low-quality reads. The sequence reads were aligned to the mm9 reference genome sequence (Waterston RH et al., Nature 2002;420:520-562) using bowtie (version 0.12.5) (Langmead B et al., Genome Biol 2009;10:R25). In genome alignment, two mismatches were allowed and only the best alignment ("-v2 - best") was recorded. For transcriptome alignment, the default parameters were used. Reads that could not be aligned to the genomic sequence were aligned to a database of all possible exon-exon junction sequences of RefSeq transcripts (Pruitt KD et al., Nucleic Acids Res 2007;35:D61-D65). Expression values were determined by intersecting the read coordinates with those of the RefSeq transcripts, counting the overlapping exon and junction reads, and normalizing to RPKM expression units (mapped reads / kilobase of exon model / million mapped reads) (Mortazavi A et al., Nat Methods 2008;5:621-628).

[0347] NGS Data Analysis, Somatic Mutation Discovery: Somatic mutations were identified as described in Example 9. Single-end reads of 50 nucleotides (nt) were aligned to the mm9 reference mouse genome using bwa (default options, version 0.5.8c) (Li H and Durbin R, Bioinformatics 2009;25:1754-1760). Ambiguous reads mapped to multiple positions in the genome were removed. Mutations were identified using three software programs: samtools (version 0.1.8) (Li H, Bioinformatics 2011;27:1157-1158), GATK (version 1.0.4418) (McKenna A et al., Genome Res 2010;20:1297-1303), and SomaticSniper (http: / / genome.wustl.edu / software / somaticsniper) (Ding L et al., Hum Mol Genet 2010;19:R188-R196). A "false discovery rate" (FDR) confidence value was assigned to potential mutations identified in all B16F10 triplicates (see Example 9).

[0348] Selection, Validation, and Function of Mutations Selection: Mutations had to meet the following criteria to be selected: (i) present in all B16F10 and absent in all C57BL / 6 triplicates, (ii) FDR ≤ 0.05, (iii) homogeneous in C57BL / 6, (iv) present in RefSeq transcripts, and (v) cause non-synonymous changes that are scored as reliable mutations. Selection for validation and immunogenicity testing required that the mutation be in a gene that was expressed (median RPKM > 10 across replicates).

[0349] Verification: Mutations derived from DNA were classified as verified if confirmed by either Sanger sequencing or B16F10 RNA-Seq reads. All selected mutants were amplified using adjacent primers from 50 ng of DNA from B16F10 cells and C57BL / 6 tail tissue, the products were visualized (QIAxcel system, Qiagen), and purified (QIAquick PCR Purification Kit, Qiagen). The unit replicated sequences of the expected size were excised from the gel, purified (QIAquick Gel Extraction Kit, Qiagen), and subjected to Sanger sequencing using the forward primer used for PCR amplification (Eurofins MWG Operon, Ebersberg, Germany).

[0350] Functional impact: The impact of the selected mutations was evaluated using the programs SIFT (Kumar P et al., Nat Protoc 2009;4:1073 - 1081) and POLYPHEN-2 (Adzhubei IA et al., Nat Methods 2010;7:248 - 249), which predict the functional significance of amino acids in protein function based on the location of protein domains and interspecies sequence conservation. Gene functions were inferred using the Ingenuity IPA tool.

[0351] Synthetic peptides and adjuvants Ovalbumin class I (OVA 258~265 )), class II (OVA class II 330~338 ), influenza nucleoprotein (Inf-NP 366~374 ), vesicular stomatitis virus nucleoprotein (VSV-NP 52~59 ), and tyrosinase-related protein 2 (Trp2 180~188All peptides, including ), were purchased from Jerini Peptide Technologies (Berlin, Germany). The synthetic peptides were 27 amino acids in length and had either the mutant (MUT) or wild-type (WT) amino acid at position 14. Polyinosinic acid:polycytidylic acid (poly(I:C), InvivoGen) was used as a subcutaneous adjuvant. Inf-NP 366~374 MHC pentamers specific for the peptides were purchased from ProImmune Ltd.

[0352] Immunization of Mice Female age-matched C57BL / 6 mice were injected subcutaneously in the outer flanks with 100 μg of peptide formulated in PBS (total volume 200 μl) and 50 μg of poly(I:C) (5 mice per group). All groups were immunized on days 0 and 7 with peptides encoding two different mutations, one peptide per flank. Mice were sacrificed 12 days after the first injection, and splenocytes were isolated for immunological assays.

[0353] Alternatively, female C57BL / 6 mice of the same age were intravenously injected with 20 μg of in vitro transcribed RNA formulated with 20 μl of Lipofectamine(™) RNAiMAX (Invitrogen) in PBS at a total injection volume of 200 μl (3 mice per group). All groups were immunized on days 0, 3, 7, 14, and 18. Mice were sacrificed 23 days after the first injection, and splenocytes were isolated for immunological assays. DNA sequences representing 1 (monoepitope), 2 (biepitope), or 16 mutations (polyepitope) were generated using 50 amino acids (aa) with the mutation at position 25 (biepitope), or 27 amino acids with the mutation at position 14 (mono- and polyepitopes), separated by a 9 amino acid glycine / serine linker, and cloned into the pST1-2BgUTR-A120 backbone (Holtkamp et al., Blood 2006;108:4009-17). In vitro transcription and purification from this template have been described previously (Kreiter et al., Cancer Immunol Immunother 2007;56:1577-87).

[0354] Enzyme-linked immunospot assay Enzyme-linked immunosorbent spot (ELISPOT) assay (Kreiter S et al., Cancer Res 2010;70:9031-40) and generation of syngeneic bone marrow-derived dendritic cells (BMDC) as stimulators have been described previously (Lutz MB et al., J Immunol Methods 1999;223:77-92). BMDC were either peptide pulsed (2 μg / ml) or transfected with in vitro transcribed (IVT) RNA encoding the indicated mutation or control RNA (eGFP-RNA). Sequences representing two mutations, each containing 50 amino acids and having a mutation at position 25 and separated by a 9-amino acid glycine / serine linker, were cloned into the pST1-2BgUTR-A120 backbone (Holtkamp S et al., Blood 2006;108:4009-17). In vitro transcription and purification from this template have been described previously (Kreiter S et al., Cancer Immunol Immunother 2007;56:1577-87). For the assay, 5×10 4 BMDC engineered with peptide or RNA were co-incubated in microtiter plates coated with anti-IFN-γ antibody (10 μg / mL, clone AN18, Mabtech) together with 5×10 5 freshly isolated splenocytes. After 18 h at 37 °C, cytokine secretion was detected with anti-IFN-γ antibody (clone R4-6A2, Mabtech). Spot numbers were counted and analyzed with an ImmunoSpot® S5 Versa ELISPOT Analyzer, ImmunoCapture™ Image Acquisition software, and ImmunoSpot® analysis software version 5. Statistical analysis was performed by Student's t-test and the Mann-Whitney test (non-parametric test). A response was considered significant if the test gave a p-value < 0.05 and the mean spot number was > 30 spots / 5×10 5 effector cells. Reactivity was assessed by the mean spot number (-: < 30, +: > 30, ++: > 50, +++ > 200 spots / well).

[0355] Intracellular cytokine assay 6 Aliquots of splenocytes prepared for the ELISPOT assay were subjected to analysis of cytokine production by intracellular flow cytometry. For this purpose, 2 × 10 5 cells / sample were plated in 96-well plates in culture medium (RPMI + 10% FCS) supplemented with the Golgi inhibitor brefeldin A (10 μg / mL). Cells from each animal were restimulated for 5 h at 37 °C with 2 × 10

[0356] peptide-pulsed BMDC. After incubation, the cells were washed with PBS, resuspended in 50 μL of PBS, and extracellularly stained for 20 min at 4 °C with the following anti-mouse antibodies: anti-CD4 FITC, anti-CD8 APC-Cy7 (BD Pharmingen). After incubation, the cells were washed with PBS and subsequently resuspended in 100 μL of Cytofix / Cytoperm (BD Bioscience) solution for 20 min at 4 °C for permeabilization of the outer membrane. After permeabilization, the cells were washed with Perm / Wash buffer (BD Bioscience), resuspended in Perm / Wash buffer at 50 μL / sample, and intracellularly stained for 30 min at 4 °C with the following anti-mouse antibodies: anti-IFN-γ PE, anti-TNF-α PE-Cy7, anti-IL2 APC (BD Pharmingen). After washing with Perm / Wash buffer, the cells were resuspended in PBS containing 1% paraformaldehyde for flow cytometry analysis. Samples were analyzed using a BD FACSCanto™ II cell meter and FlowJo (version 7.6.3).B16 melanoma tumor model In the tumor vaccination experiment, 7.5 × 10 4Individual B16F10 melanoma cells were subcutaneously inoculated into the flank of C57BL / 6 mice. In the preventive setting, immunization with the mutation-specific peptide was performed 4 days before tumor inoculation and on days 2 and 9 after tumor inoculation. In the therapeutic experiment, the peptide vaccine was administered on days 3 and 10 after tumor injection. The tumor size was measured every 3 days, and the mice were sacrificed when the tumor diameter reached 15 mm.

[0357] Alternatively, in the tumor vaccination experiment, 1 × 10 5 Individual B16F10 melanoma cells were subcutaneously inoculated into the flank of age-matched female C57BL / 6 mice. Peptide vaccination was performed on days 3, 10, and 17 after tumor inoculation using 100 μg of peptide and 50 μg of poly(I:C) formulated in PBS (total volume 200 μl) injected subcutaneously into the outer flank. RNA immunization was performed using 20 μg of in vitro transcribed mutant-encoded RNA formulated with 20 μl of Lipofectamine™ RNAiMAX (Invitrogen) in PBS with a total injection volume of 200 μl. As a control, one group of animals was injected with RNAiMAX (Invitrogen) in PBS. The animals were immunized on days 3, 6, 10, 17, and 21 after tumor inoculation. The tumor size was measured using calipers every 3 days, and the mice were sacrificed when the tumor diameter reached 15 mm.

[0358] Identification of Nonsynonymous Mutations in B16F10 Mouse Melanoma The aim of the present inventors was to identify somatic point mutations with potential immunogenicity in B16F10 mouse melanoma by NGS, test them for in vivo immunogenicity by peptide vaccination of mice, and measure the induced T cell responses by ELISPOT assay (Figure 9A). The exomes of the C57BL / 6 wild-type background genome and that of B16F10 cells were sequenced using extraction and capture in triplicate, respectively. For each sample, over 100 million single-end 50 nt reads were generated. Of these, 80% aligned uniquely to the mouse mm9 genome and 49% aligned on target, demonstrating the success of target enrichment and resulting in coverage of over 20-fold for 70% of the target nucleotides in each of the triplicate samples. RNA-Seq of B16F10 cells profiled in triplicate also yielded a median of 30 million single-end 50 nt reads, 80% of which aligned to the mouse transcriptome.

[0359] DNA reads (exome capture) from B16F10 and C57BL / 6 were analyzed to identify somatic mutations. Copy number variation analysis (Sathirapongsasuti JF et al., Bioinformatics 2011;27:2648-2654) demonstrated deletions in B16F10 including DNA amplification and homozygous deletion of the tumor suppressor Cdkn2a (cyclin-dependent kinase inhibitor 2A, p16Ink4A). Focusing on point mutations to identify possible immunogenic mutations, the present inventors identified 3,570 somatic point mutations with FDR ≤ 0.05 (Figure 9B). The most frequent mutation class was C>T / G>A transversions, which typically arise from ultraviolet light (Pfeifer GP et al., Mutat Res 2005;571:19-31). Of these somatic mutations, 1,392 were present in transcripts and 126 mutations were in untranslated regions. Of the 1,266 mutations in the coding region, 962 caused non-synonymous protein changes, 563 of which were present in expressed genes (Figure 9B).

[0360] Assignment and verification of identified mutations to carrier genes Notably, many of the mutant genes (962 genes containing non-synonymous point mutations) have previously been associated with the cancer phenotype. Mutations were found in established tumor suppressor genes, including Pten, Trp53 (also known as p53), and Tp63. In Trp53, the most established tumor suppressor (Zilfou JT et al., Cold Spring Harb Perspect Biol 2009;1:a001883), a mutation from asparagine to aspartic acid at protein position 127 (p.N127D) is localized within the DNA binding domain and is predicted by SIFT to alter function. Pten contains two mutations (p.A39V, p.T131P), both of which are predicted to have a deleterious effect on protein function. The p.T131P mutation is in proximity to a mutation (p.R130M) that has been shown to abolish phosphatase activity (Dey N et al., Cancer Res 2008;68:1862-1871). Furthermore, mutations have been found in genes related to the DNA repair pathway, such as Brca2 (breast cancer 2, early onset), Atm (ataxia telangiectasia mutated), Ddb1 (damage-specific DNA binding protein 1), and Rad9b (RAD9 homolog B). Additionally, mutations are present in other tumor-related genes, including Aim1 (tumor suppressor "absent in melanoma 1"), Flt1 (oncogene Vegr1, fms-related tyrosine kinase 1), Pml (tumor suppressor "promyelocytic leukemia"), Fat1 ("FAT tumor suppressor homolog 1"), Mdm1 (TP53-binding nuclear protein), Mta3 (metastasis-associated 1 family, member 3), and Alk (anaplastic lymphoma receptor tyrosine kinase). The inventors discovered a mutation to p.S144F in Pdgfra (platelet-derived growth factor receptor, alpha polypeptide), a cell membrane-bound receptor tyrosine kinase of the MAPK / ERK pathway that was previously identified in tumors (Verhaak RG et al., Cancer Cell 2010;17:98-110). A mutation is present at p.L222V in Casp9 (caspase 9, apoptosis-related cysteine peptidase).CASP9 proteolytically cleaves poly(ADP-ribose) polymerase (PARP), regulates apoptosis, and has been associated with several cancers (Hajra KM et al., Apoptosis 2004;9:691-704). The mutations discovered by the inventors may potentially affect PARP and apoptosis signaling. Most interestingly, no mutations were found in Braf, c-Kit, Kras, or Nras. However, mutations were identified in Rassf7 (RAS-related protein) (p.S90R), Ksr1 (kinase suppressor of ras1) (p.L301V), and Atm (PI3K pathway) (p.K91T), all of which are predicted to have a significant impact on protein function. Trrap (transformation / transcription domain-associated protein) has been identified as a new potential melanoma target in human melanoma specimens so far this year (Wei X et al., Nat Genet 2011;43:442-6). In B16F10, the Trrap mutation occurs at p.K2783R and is predicted to disrupt the overlapping phosphatidylinositol kinase (PIK)-related kinase FAT domain.

[0361] From the 962 non-synonymous mutations identified using NGS, 50 mutations including 41 with FDR < 0.05 were selected for PCR-based validation and immunogenicity testing. The selection criteria were the position in the expressed gene (RPKM > 10) and predicted immunogenicity. Notably, all 50 mutations could be confirmed (Table 6, Figure 9B).

[0362]

Table 6A

[0363]

Table 6B

[0364]

Table 6C

[0365] Figure 9C shows the positions of the B16F10 chromosome, gene density, gene expression, mutations, and filtered mutations (inner ring).

[0366] In vivo test of immunogenicity test using long peptides representing mutations To provide antigens for immunogenicity testing of these mutations, the inventors used long peptides that have many advantages over other peptides for immunization (Melief CJ and van der Burg SH, Nat Rev Cancer 2008;8:351-60). Long peptides can induce antigen-specific CD8+ and CD4+ T cells (Zwaveling S et al., Cancer Res 2002;62:6187-93, Bijker MS et al., J Immunol 2007;179:5033-40). Furthermore, long peptides require processing to be presented on MHC molecules. Such uptake is most efficiently performed by optimal dendritic cells to prime a strong T cell response. In contrast, fitting peptides do not require trimming and are exogenously loaded onto all cells expressing MHC molecules, including inactivated B and T cells, leading to the induction of tolerance and fratricide (Toes RE et al., J Immunol 1996;156:3911-8, Su MW et al., J Immunol 1993;151:658-67). For each of the 50 validated mutations, peptides 27 amino acids in length with the mutated or wild-type amino acid located centrally were designed. Thus, any potential MHC class I and class II epitopes 8-14 amino acids in length carrying the mutation can be processed from this precursor peptide. As an adjuvant for peptide vaccination, poly(I:C), which is known to promote cross-presentation and increase the efficacy of the vaccine, was used (Datta SK et al., J Immunol 2003;170:4102-10, Schulz O et al., Nature 2005;433:887-92). The 50 mutations were tested in vivo in mice for T cell induction. Impressively, 16 of the 50 mutated coding peptides were found to induce an immune response in immunized mice. The induced T cells showed different response patterns (Table 7).

[0367] [Table 7]

[0368] Eleven peptides induced an immune response that preferentially recognized mutated epitopes. This is illustrated in mice immunized with mutations 30 (MUT30, Kif18b) and 36 (MUT36, Plod2) (Figure 10A). The ELISPOT assay revealed a strong mutation-specific immune response without cross-reactivity to wild-type peptides or unrelated control peptides (VSV-NP). For five peptides, including mutations 05 (MUT05, Eef2) and 25 (MUT25, Plod2) (Figure 10A), an immune response with comparable recognition of both mutant and wild-type peptides was obtained. As exemplified by mutations 01 (MUT01, Fzd7), 02 (MUT02, Xpot), and 07 (MUT07, Trp53), the majority of mutant peptides were unable to induce a significant T cell response. The immune responses induced by some of the mutations found were well within the range of immunogenicity generated by immunizing mice with the described MHC-class I epitope from murine melanoma tumor antigen tyrosinase-related protein 2 (Trp2180-188, Figure 10A) as a positive control (500 spots / 5×10 5individual cells) (Bloom MB et al., Exp Med 1997;185:453-459, Schreurs MW et al., Cancer Res 2000;60:6995-7001). For selected peptides that induce a strong mutation-specific T cell response, immune recognition was confirmed by independent methods. Instead of long peptides, in vitro transcribed RNA (IVT RNA) encoding mutant peptide fragments MUT17, MUT30, and MUT44 was used for immunological readout. BMDCs transfected with RNA encoding the mutation or unrelated RNA served as antigen-presenting cells (APCs) in the ELISPOT assay, while splenocytes from immunized mice served as the effector cell population. BMDCs transfected with mRNA encoding MUT17, MUT30, and MUT44 were specifically and strongly recognized by splenocytes from mice immunized with the respective long peptides (Figure 10B). Significantly lower reactivity was recorded against BMDCs transfected with control RNA, which is likely due to non-specific activation of BMDCs by single-stranded RNA (Student's t-test, MUT17: p = 0.0024, MUT30: p = 0.0122, MUT44: p = 0.0075). These data confirm that the induced mutation-specific T cells recognize epitopes that are processed endogenously. Two mutations that induce preferential recognition of the mutated epitope are in the genes Actn4 and Kif18b. The somatic mutation in ACTN4 (actinin, alpha 4) is at p.F835V in the calcium-binding "EF-hand" protein domain. Both SIFT and PolyPhen predict a significant effect of this mutation on protein function, although this gene is not an established cancer gene. However, mutation-specific T cells against ACTN4 have recently been associated with positive patient outcomes (Echchakir H et al., Cancer Res 2001;61:4078-4083).KIF18B (kinesin family member 18B) is a kinesin with microtubule motility activity and ATP and nucleotide binding that is involved in the regulation of cell division (Lee YM et al., Gene 2010;466:16-25) (Figure 10C). The DNA sequence at the position encoding p.K739 is homologous in reference C57BL / 6, while a heterozygous somatic mutation has been revealed in the DNA reading of B16F10. Both nucleotides were detected in the B16F10 RNA-Seq reads and verified by Sanger sequencing. KIF18B has not previously been associated with the cancer phenotype. The mutation p.K739N is not localized to a known functional or conserved protein domain (Figure 10C, bottom), and thus is likely to be a passenger mutation rather than a driver. These examples suggest a lack of correlation between the ability to induce a mutation-recognizing immune response and functional or immunological relevance.

[0369] In Vivo Evaluation of the Antitumor Activity of Vaccine Candidates To evaluate whether the immune responses induced in vivo translate into antitumor effects in mice bearing tumors, the inventors selected MUT30 (mutated in Kif18b) and MUT44 as examples. These mutations have been shown to preferentially induce a strong immune response against the mutant peptides and to be endogenously processed (Figure 10A, B). The therapeutic potential of vaccination with mutant peptides was investigated by immunizing mice with either MUT30 or MUT44 and an adjuvant 3 and 10 days after transplantation of 7.5×10 5 individual B16F10s. Tumor growth was inhibited by vaccination with either peptide compared to the control group (Figure 11A). Since B16F10 is a very aggressively growing tumor, protective immune responses were also tested. Mice were immunized with the MUT30 peptide and 4 days later 7.5×10 5Individual B16F10 cells were subcutaneously inoculated and boosted with MUT30 2 and 9 days after tumor challenge. Complete tumor protection and 40% survival were observed in mice treated with MUT30, while all mice in the control treatment group died within 44 days (left in Figure 11B). In mice that developed tumors despite immunization with MUT30, tumor growth was slower, resulting in a median survival period extension of 6 days compared to the control group (right in Figure 11B). These data imply that vaccination against a single mutation can already confer an antitumor effect.

[0370] Immunization with RNA encoding the mutation Various RNA vaccines were generated using 50 validated mutations from the B16F10 melanoma cell line. DNA sequences representing 1 (monoepitope), 2 (biepitope), or 16 different mutations (polyepitope) were created using 50 amino acids (aa) for the mutations at position 25 (biepitope), or 27 amino acids for the mutations at position 14 (mono and polyepitopes), separated by a 9 - amino - acid glycine / serine linker. These constructs were cloned into the pST1 - 2BgUTR - A120 backbone for in vitro transcription of mRNA (Holtkamp et al., Blood 2006;108:4009 - 4017).

[0371] To test the in vivo ability to induce T - cell responses against various RNA - vaccines, groups of 3 C57BL / 6 mice were immunized by formulating the RNA with RNAiMAX lipofectamine and subsequently injecting it intravenously. After 5 immunizations, the mice were sacrificed and splenocytes were analyzed for mutation - specific T - cell responses using intracellular cytokine staining and IFN - γ ELISPOT assays after restimulation with peptides encoding the corresponding mutations or control peptides (VSV - NP).

[0372] Figure 12 shows an example of each vaccine design. In the upper panel, mice were immunized with CD4 specific for MUT30 +Vaccinated with a monoepitope-RNA encoding MUT30 (mutation in Kif18b) that induces T cells (see exemplary FACS plots). In the middle panel, the graphs and FACS plots show CD4 + T cell induction specific for MUT08 (mutation in Ddx23) after immunization with a bivalent epitope encoding MUT33 and MUT08. In the bottom panel, mice were immunized with a polyepitope encoding 16 different mutations including MUT08, MUT33 and MUT27 (see Table 8). The graphs and FACS plots illustrate that T cells reactive to MUT27 are of the CD8 phenotype.

[0373]

Table 8

[0374] The data shown in Figure 13 were generated using the same polyepitope. The graphs show ELISPOT data after restimulation of splenocytes with control (VSV-NP), MUT08, MUT27 and MUT33 peptides, demonstrating that the polyepitope vaccine can induce specific T cell responses against several different mutations.

[0375] Collectively, the data indicate the potential to induce mutation-specific T cells using RNA-encoded mono-, bi- and polyepitopes. Furthermore, the data show induction of CD4 + and CD8 + T cells and induction of several different specificities from one construct.

[0376] Immunization with model epitopes To further characterize the polyepitope RNA-vaccine design, a DNA sequence containing five different known model epitopes including one MHC class II epitope (ovalbumin class I (SIINFEKL), class II (OVA class II), influenza nucleoprotein (Inf-NP), vesicular stomatitis virus nucleoprotein (VSV-NP) and tyrosinase-related protein 2 (Trp2)) was generated. The epitopes were separated by glycine / serine linkers of the same nine amino acids used in the mutant polyepitope. This construct was cloned into the pST1-2BgUTR-A120 backbone for in vitro transcription of mRNA.

[0377] Using in vitro transcribed RNA, five C57BL / 6 mice were vaccinated by intranodal immunization (four immunizations with 20 μg of RNA in the inguinal lymph nodes). Five days after the last immunization, blood samples and splenocytes were collected from the mice for analysis. Figure 14A shows an IFN-γ ELISPOT analysis of splenocytes restimulated with the indicated peptides. All three MHC-class I epitopes (SIINFEKL, Trp2 and VSV-NP) were clearly seen to induce very large numbers of antigen-specific CD8 + T cells. Also, the MHC-class II epitope OVA class II induced a strong CD4 + T cell response. The fourth MHC class I epitope was analyzed by staining of CD8 + T cells specific for Inf-NP using fluorescently labeled pentameric MHC-peptide complexes (pentamers) (Figure 14B).

[0378] These data demonstrate that the polyepitope design using glycine / serine linkers to separate different immunogenic MHC-class I and -class II epitopes can induce specific T cells against all encoded epitopes regardless of their immunodominance.

[0379] Anti-tumor response after treatment with mutant polyepitope RNA vaccine Using the same polyepitope analyzed for immunogenicity in FIG. 13, the antitumor activity against B16F10 tumor cells of RNA encoding the mutations was investigated. Specifically, a group of C57BL / 6 mice (n = 10) was subcutaneously inoculated in the flank with 1 × 10 5 B16F10 melanoma cells. On days 3, 6, 10, 17, and 21, the mice were immunized with polyepitope RNA using a liposomal transfection reagent. The control group was injected with liposomes alone.

[0380] FIG. 21 shows the survival curves of the groups, revealing a significantly improved median survival period of 27 days, with 1 out of 10 mice surviving without tumors, compared to a median survival period of 18.5 days in the control group.

[0381] Antitumor response after treatment with a combination of mutant and normal peptides The antitumor activity of the validated mutations was evaluated by a therapeutic in vivo tumor experiment using MUT30 as a peptide vaccine. Specifically, a group of C57BL / 6 mice (n = 8) was subcutaneously inoculated in the flank with 1 × 10 5 B16F10 melanoma cells. On days 3, 10, and 17, the mice were immunized with MUT30, tyrosinase-related protein 2 (Trp2 180~188 ) or a combination of both peptides using poly I:C as an adjuvant. Trp2 is a known CD8 + epitope expressed by B16F10 melanoma cells.

[0382] FIG. 15A shows the mean tumor growth of the groups. The known CD8 + T cell epitope and CD4 +In the group immunized with the combination of T cells and MUT30, it is clearly seen that tumor growth is almost completely inhibited until day 28. The known Trp2 epitope is not sufficient to provide a good anti-tumor effect alone in this setting, yet both single-treatment groups (MUT30 and Trp2) provide inhibition of tumor growth compared to the untreated group from the early stage of the experiment until day 25. These data are strengthened by the survival curves shown in Figure 15B. Clearly, the median survival time is increased in mice injected with a single peptide, and 1 / 8 of the mice in the Trp2-vaccinated group are alive. Furthermore, the group treated with both peptides shows an even better median survival time, and 2 / 8 of the mice survive.

[0383] Overall, both epitopes act in a synergistic manner to provide a potent anti-tumor effect.

[0384] (Example 9) Framework for reliability-based somatic mutation detection and application to B16-F10 melanoma cells NGS is unbiased in that it enables high-throughput discovery of mutations across the genome or within target regions such as exons encoding proteins.

[0385] However, while revolutionary, NGS platforms still tend to have errors, leading to false mutation calls. Furthermore, the quality of the results depends on experimental design parameters and analysis methods. Mutation calls typically include scores designed to distinguish true mutations from errors, but the usefulness and interpretation of these scores are not fully understood with respect to experimental optimization. This is especially true when comparing tissue states such as comparing tumors and normal for somatic mutations. As a result, researchers are forced to rely on personal experience to determine experimental parameters and discretionary filtering thresholds for selecting mutations.

[0386] This study aims to a) establish a framework for comparing parameters and a method for identifying somatic mutations, and b) assign confidence values to the identified mutations. The inventors sequence triplicate samples from C57BL / 6 mice and the B16-F10 melanoma cell line. The false discovery rate of somatic mutations detected using these data is formulated, which is then used as an indicator to evaluate existing mutation discovery software and laboratory protocols.

[0387] Various experimental and algorithmic factors contribute to the false positive rate of mutations found by NGS [Nothnagel. M. et al., Hum. Genet. February 23, 2011 [Epub before print]]. Error causes include PCR artifacts, biases in first stimulation [Hansen, K.D. et al., Nucleic. Acids. Res. 38, e131 (2010), Taub, M.A. et al., Genome Med. 2, 87 (2010)] and targeted enrichment [Bainbridge, M.N. et al., Genome Biol. 11, R62 (2010)], sequence effects [Nakamura, K. et al., Acids Res. (2011), first published online on May 16, 2011, doi:10.1093 / nar / gkr344], base calling that causes sequence errors [Kircher, M. et al., Genome Biol. 10, R83 (2009). Epub August 14, 2009], and read alignment [Lassmann, T. et al., Bioinformatics 27, 130 - 131 (2011)], which further cause coverage fluctuations affecting downstream analysis, such as variant calling around indels [Li, H., Bioinformatics 27, 1157 - 1158 (2011)] and sequencing errors.

[0388] No general statistical model has been described that accounts for the effects of the various error sources on somatic mutation calling, and only individual aspects have been covered without removing all biases. Recent computer methods for measuring the expected amount of false-positive mutation calls include the use of the translocation / base substitution ratio of a set of mutations [Zhang, Z., Gerstein, M., Nucleic Acids Res 31, 5338-5348 (2003), DePristo, M.A. et al., Nature Genetics 43, 491-498 (2011)], machine learning [DePristo, M.A. et al., Nature Genetics 43, 491-498 (2011)], and inheritance errors when working with family genomes [Ewen, K.R. et al., Am. J. Hum. Genet. 67, 727-736 (2000)] or pooled samples [Druley, T.E. et al., Nature Methods 6, 263-265 (2009), Bansal, V., Bioinformatics 26, 318-324 (2010)]. For optimization purposes, Druley et al. [Druley, T.E. et al., Nature Methods 6, 263-265 (2009)] relied on short plasmid sequence fragments, which may not be representative of the samples. In a set of single nucleotide variants (SNVs) and selected experiments, comparison with SNVs identified by other techniques is possible [Van Tassell, CP. et al., Nature Methods 5, 247-252 (2008)], but it is difficult to evaluate for novel somatic mutations.

[0389] Using exome sequencing projects as an example, the inventors propose a calculation of the false discovery rate (FDR) based on NGS data alone. This method not only is applicable to the selection and prioritization of diagnostic and therapeutic targets, but also aids in the development of algorithms and methods by making it possible to define confidence-driven recommendations for similar experiments.

[0390] To discover mutations, DNA from tail tissues of three C57BL / 6 (black6) mice (littermates) and DNA from B16-F10 (B16) melanoma cells were individually enriched (Agilent Sure Select Whole Mouse Exome) in triplicate for exons encoding proteins, resulting in six samples. RNA was extracted from B16 cells in triplicate. Single-end 50nt (1×50nt) and paired-end 100nt (2×100nt) reads were generated on an Illumina HiSeq 2000. Each sample was loaded onto a separate lane, resulting in an average of 104 million reads per lane. DNA reads were aligned to the mouse reference genome using bwa [Li, H., Durbin, R., Bioinformatics 25, 1754 - 1760 (2009)], and RNA reads were aligned to bowtie [Langmead, B. et al., Genome Biol. 10, R25 (2009)]. An average coverage of 38-fold was achieved in the 1×50nt library at 97% of the target region, and an average coverage of 165-fold was obtained in the 2×100nt experiment at 98% of the target region.

[0391] Somatic mutations were independently identified using the software packages SAMtools [Li, H. et al., Bioinformatics 25, 2078 - 2079 (2009)], GATK [DePristo, M.A. et al., Nature Genetics 43, 491 - 498 (2011)] and SomaticSNiPer [Ding, L. et al., Hum. Mol. Genet (2010), first published online on September 15, 2010] (Figure 16) by comparing single nucleotide mutations found in the B16 samples to the corresponding loci in the black6 samples (B16 cells are originally derived from black6 mice). Potential mutations were filtered respectively according to the recommendations by the authors of each software (SAMtools and GATK), or by selecting an appropriate lower cut-off value for the somatic score of SomaticSNiPer.

[0392] To create the false discovery rate (FDR) of somatic mutation discovery, the inventors first intersected the mutation sites and obtained 1,355 high-quality somatic mutations as the consensus of all three programs (Figure 17). However, the differences observed in the results of the applied software tools were significant. To avoid false conclusions, the inventors developed a method to assign FDR to each mutation using replicates. Technical replicates of the samples should yield the same results, and all detected mutations in this "comparison to identicals" are false positives. Therefore, to determine the false discovery rate of somatic mutation detection in tumor samples compared to normal samples ("tumor comparison"), technical replicates of the normal samples can be used as a reference to estimate the number of false positives.

[0393] Figure 18A shows examples of mutations found in the black6 / B16 data, including somatic mutations (left), non-somatic mutations relative to the reference (center), and potential false positives (right). Each somatic mutation can be associated with a quality score Q. The number of false positives in the tumor comparison indicates some false positives in the comparison to identicals. Therefore, for a given mutation with quality score Q detected in the tumor comparison, the false discovery rate is estimated by calculating the ratio of the number of mutations with score Q or better in the comparison to identicals to the total number of mutations with score Q or better found in the tumor comparison.

[0394] Most mutation detection frameworks calculate multiple quality scores, thus presenting challenges in the definition of Q. Here, a random forest classifier [Breiman, L., Statist. Sci. 16, 199 - 231 (2001)] is applied to combine multiple scores into a single quality score Q. See the methods section for details on the calculation of quality scores and FDR.

[0395] Since a potential bias in the comparison method is uneven coverage, the false discovery rate of coverage is normalized.

[0396]

Number

[0397] Calculate the common coverage by counting each of all the bases of the reference genome included in both tumor and normal samples or both "matched" samples respectively.

[0398] For each mutation discovery method, generate a receiver operating characteristic (ROC) curve by estimating the number of false positives and positives at each false discovery rate (FDR) (see method), calculate the area under the curve (AUC), thereby enabling comparison of the mutation discovery strategies (Figure 18B).

[0399] Furthermore, the selection of reference data can affect the calculation of FDR. Using the available black6 / B16 data, it is possible to create 18 replicates (combinations of black6 vs black6 and black6 vs b16). When comparing the resulting FDR distributions of somatic mutation sets, the results are consistent (Figure 18B).

[0400] Using this definition of false discovery rate, the inventors established a general framework for evaluating the effects of various experimental and algorithm parameters on the resulting sets of somatic mutations. Next, the inventors applied this framework to study the effects of software tools, coverage, paired-end sequencing, and the number of technical replicates on the identification of somatic mutations.

[0401] First, the selection of software tools has a clear impact on the identified somatic mutations (Figure 19A). Among the tested data, SAMtools yields the highest enrichment of true positives in a set of somatic mutations ranked by FDR. However, note that all tools provide many parameters and quality scores for individual mutations. Here, the default settings specified by the algorithm developers were used. It is anticipated that the parameters can be optimized, and it is emphasized that the FDR framework defined herein is designed to perform and evaluate such optimizations.

[0402] In the described B16 sequencing experiment, each sample was sequenced in an individual flow cell lane, achieving an average target region base coverage of 38-fold for each sample. However, this coverage may not be necessary to obtain equally good sets of somatic mutations, and in some cases, costs can be reduced. Also, the impact of coverage (caverage) depth on whole-genome SNV detection has recently been discussed [Ajay, S.S. et al., Genome Res. 21, 1498 - 1505 (2011)]. To study the effect of coverage on exon capture data, the number of aligned sequence reads of all 1×50nt libraries was downsampled to generate coverages of approximately 5, 10, and 20-fold, respectively, and then the mutation calling algorithm was applied again. As expected, higher coverage results in better (i.e., fewer false positives) sets of somatic mutations, but the improvement from 20-fold coverage to the maximum is marginal (Figure 19B).

[0403] It is clear that various experimental settings are simulated and ranked using available data and frameworks. Comparing duplicates and triplicates, triplicates do not offer advantages compared to duplicates (Figure 19C), but duplicates offer clear improvements compared to studies that do not use any replication at all. Regarding the proportion of somatic mutations in a given cohort, enrichment is seen from 24.2% in runs without replication to 71.2% in duplicates and 85.8% in triplicates at a 5% FDR. Despite this enrichment, as shown by the lower ROC AUC and shift of the curve to the left, the use of triplicate crosses removes more mutations with low FDR than those with high FDR (Figure 19C). Specificity is slightly increased at the expense of lower sensitivity.

[0404] Additional 2×100nt libraries sequenced were used to simulate 1×100, two 2×50, and two 1×50nt libraries, respectively, by a second read and / or in silico removal of the 3' and 5' ends of the reads, resulting in a total of five simulated libraries. These libraries were compared using the calculated FDR of the predicted mutations (Figure 19D). Despite much higher average coverage (higher than 77 vs 38), somatic mutations found using the 2×50 5' and 1×100nt libraries had lower ROC AUC and thus a poorer FDR distribution than the 1×50nt library. This phenomenon results from the accumulation of high-FDR mutations in low-coverage regions, because the sets of low-FDR mutations found are very similar. As a result, the optimal sequencing length should be small such that the sequenced bases are concentrated around the capture probe sequence (although this may lose information on the somatic state of mutations not included in the region), or approximately the fragment length (in this case, a 2×100nt = 200nt total length with a ~250nt fragment), with coverage gaps effectively filled. This is also supported by the higher ROC AUC of the 2×50nt 3' library (simulated using only the 3' end of the 2×100nt library) compared to the ROC AUC of the 2×50nt 5' library (simulated using only the 5' end of the 2×100nt library), despite the lower base quality at the 3' read end.

[0405] These observations make it possible to define the best practice procedures for discovering somatic mutations. Across all evaluated parameters, 20-fold coverage and the use of technical duplicates in both samples achieve results that are close to optimal while considering cost in these relatively homogeneous samples. A 1×50 nt library that yields approximately 100 million reads is considered the most practical option for achieving this coverage. This continues to be faithful to all pairs of possible datasets. The inventors retrospectively applied these parameter settings and calculated the FDR of 50 selected mutations from the intersection of all three methods, as shown in Figure 17, without using additional filtering of the raw variant calls. All mutations were confirmed by a combination of Sanger re-sequencing and reads of the B16 RNA-Seq sequences. 44 of these mutations would have been found using a 5% FDR cutoff (Figure 20). As a negative control, the loci of 44 predicted mutations with a high FDR (>50%) were re-sequenced and each sequence in the RNA-Seq data was examined. While 37 of these mutations were not verified, it was found that the remaining 7 loci of potential mutations were not only not covered by RNA-Seq reads but also not obtained in the sequencing reaction.

[0406] Shows the application of the framework to four specific questions, but is not limited to these parameters in any way and can be applied to study the effects of all experimental or algorithmic parameters, such as the effects of alignment software, options for mutation measurement criteria, or options for the source of exome selection.

[0407] All experiments were performed on a set of B16 melanoma cells, but this method is not limited to these data. The only requirement is that a "matched" reference data set be available, which means that at least a single technical replicate of non-tumor samples should be done for each new protocol. Since it has been shown to be robust within certain limits regarding the options for technical replicates, replicates are not necessarily required in every single experiment. However, this method does require that various quality measures be comparable between the reference training set and the remaining data sets.

[0408] Within the scope of this contribution, the inventors developed a statistical framework for the detection of somatic mutations driven by the false discovery rate. This framework is applicable not only to diagnostic or therapeutic target selection, but also enables general comparison of the steps of the experimental and computer protocols of the generated quasi ground truth data. Here, this idea was applied to make decisions on software tools, coverage, duplication, and paired-end sequencing protocols.

[0409] Method Library capture and sequencing Next-generation sequencing, DNA sequencing: Exome capture for DNA re-sequencing was performed using an Agilent Sure-Select solution-based capture assay designed to capture all known mouse exons in this example [Gnirke, A. et al., Nat. Biotechnol. 27, 182-189 (2009)]. 3 μg of purified genomic DNA was fragmented to 150-200 nt using a Covaris S2 sonicator. The gDNA fragments were end-repaired using T4 DNA polymerase and Klenow DNA polymerase, and 5'-phosphorylated using T4 polynucleotide kinase. The blunt-ended gDNA fragments were 3'-adenylated using Klenow fragment (3'-to-5' exo-minus). T4 DNA ligase was used to ligate the 3'-single T-overhang Illumina paired-end adapter to the gDNA fragments using a 10:1 adapter:genomic DNA insert molar ratio. The adapter-ligated gDNA fragments were concentrated prior to capture, and flow cell-specific sequences were added using 4 PCR cycles with Illumina PE PCR primers 1.0 and 2.0 and Herculase II polymerase (Agilent).

[0410] 500 ng of the PCR-enriched gDNA fragments with adapter ligation were hybridized with the Agilent SureSelect biotinylated mouse whole exome RNA library bait at 65 °C for 24 hours. The hybridized gDNA / RNA bait complex was removed using magnetic beads coated with streptavidin. The gDNA / RNA bait complex was washed, and the RNA bait was cleaved during elution in SureSelect elution buffer, leaving the captured PCR-enriched gDNA fragments ligated to the adapter. The gDNA fragments were PCR amplified after capture using Herculase II DNA polymerase (Agilent) and SureSelect GA PCR primers for 10 cycles.

[0411] Cleanup was performed using 1.8× volume of AMPure XP magnetic beads (Agencourt). Quality control was performed using Invitrogen's Qubit HS assay, and fragment size was determined using Agilent's 2100 Bioanalyzer HS DNA assay.

[0412] Exome-enriched gDNA libraries were clustered at 7 pM using the Truseq SR Cluster Kit v2.5 on the cBot and sequenced using the Truseq SBS Kit on the Illumina HiSeq2000.

[0413] Exome data analysis Sequenced reads were aligned to the reference mouse genome assembly mm9 [Mouse Genome Sequencing Consortium, Nature 420, 520 - 562 (2002)] using bwa (version 0.5.8c) [Li, H., Durbin, R., Bioinformatics 25, 1754 - 1760 (2009)] with default options. Ambiguous reads, reads that mapped to multiple positions in the genome provided by the -bwa output, were removed. The remaining alignments were sorted, indexed, converted to binary compressed format (BAM), and the read quality scores were converted from the Illumina standard phred+64 to the standard Sanger quality scores using a shell script.

[0414] For each sequencing lane, mutations were identified using three software programs: SAMtools pileup (version 0.1.8) [Li, H. et al., Bioinformatics 25, 2078 - 2079 (2009)], GATK (version 1.0.4418) [DePristo, M.A. et al., Nature Genetics 43, 491 - 498 (2011)], and SomaticSniper [Ding, L. et al., Hum. Mol. Genet (2010), first published online on September 15, 2010]. For SAMtools, the recommended options and filtering criteria by the authors were used, including the first round of filtering with a maximum coverage of 200 (http: / / sourceforge.net / apps / mediawiki / SAMtools / index.php?title=SAM_FAQ, accessed in September 2011). In the second round of filtering with SAMtools, the minimum indel quality score was 50 and the minimum quality for point mutations was 30. For GATK mutation calling, the best practice guidelines designed by the authors as presented in the GATK user guide were followed (http: / / www.broadinstitute.org / gsa / wiki / index.php / The_Genome_Analysis_Toolkit, accessed in October 2010). For each sample, local realignment around indel sites was then followed by base quality recalibration. The UnifiedGenotyper module was applied to the alignment data file that resulted. If necessary, the known polymorphisms in dbSNP [Sherry, S.T. et al., Nucleic Acids Res. 29, 308 - 311 (2009)] (version 128 for mm9) were provided at each step. The variant score recalibration step was omitted and replaced by hard filtering options. For SomaticSniper mutation calling, the default options were used and only predicted mutations with a "somatic score" of 30 or more were further considered.Furthermore, for each potentially mutated locus, non-zero coverage in normal tissue was required, and all mutations located in repetitive sequences defined by the RepeatMasker track of the UCSC Genome Browser for the mouse genome assembly mm9 were removed [Fujita, P.A. et al., Nucleic Acids Res. 39, pp. 876-882 (2011)].

[0415] RNA-Seq Barcoded mRNA-seq cDNA libraries were prepared from 5 μg of total RNA using a modified version of the Illumina mRNA-seq protocol. mRNA was isolated using Seramag Oligo(dT) magnetic beads (Thermo Scientific). The isolated mRNA was fragmented using divalent cations and heat, resulting in fragments in the range of 160-200 bp. The fragmented mRNA was converted to cDNA using random primers and SuperScriptII (Invitrogen), and then the second strand was synthesized using DNA polymerase I and RNaseH. The cDNA was end-repaired using T4 DNA polymerase and Klenow DNA polymerase, and 5'-phosphorylated using T4 polynucleotide kinase. The blunt-ended cDNA fragments were 3'-adenylated (3' to 5' exonucleased) using Klenow fragment. A 3'-single T-overhang Illumina multiplex-specific adapter was ligated onto the cDNA fragments using T4 DNA ligase. The cDNA libraries were purified and size-selected at 300 bp using a 2% SizeSelect E-Gel (Invitrogen). Enrichment, addition of Illumina hexamer indexes and flow cell-specific sequences were performed by PCR using Phusion DNA polymerase (Finnzymes). All cleanups were performed using 1.8× volume of Agencourt AMPure XP magnetic beads.

[0416] Barcode RNA-seq libraries were clustered at 7 pM on the cBot using the Truseq SR Cluster Kit v2.5 and sequenced on the Illumina HiSeq2000 using the Truseq SBS Kit.

[0417] Raw output data from the HiSeq were processed according to the Illumina standard protocol, including removal of low-quality reads and demultiplexing. Subsequently, sequence reads were aligned to the reference genome sequence [Mouse Genome Sequencing Consortium, Nature 420, 520 - 562 (2002)] using bowtie [Langmead, B. et al., Genome Biol. 10, R25 (2009)]. Alignment coordinates were compared to the exon coordinates of RefSeq transcripts [Pruitt, K.D. et al., Nucleic Acids Res. 33, 501 - 504 (2005)], and for each transcript, the number of overlapping alignments was recorded. Sequence reads that did not align to the genome sequence were aligned to a database of all possible exon-exon junction sequences of RefSeq transcripts [Pruitt, K.D. et al., Nucleic Acids Res. 33, 501 - 504 (2005)]. Alignment coordinates were compared to RefSeq exon and junction coordinates, reads were counted, and for each transcript, they were normalized to RPKM (reads per kilobase of transcript per million mapped reads [Mortazavi, A. et al., Nat. Methods 5, 621 - 628 (2008)]).

[0418] Verification of SNV SNVs were selected for Sanger re-sequencing and validation by RNA. SNVs predicted by all three programs, being non-synonymous and found in transcripts having a minimum of 10 RPKM, were identified. Of these, 50 with the highest SNP quality scores provided by the programs were selected. As negative controls, 44 SNVs having an FDR of over 50%, present only in one cell line sample and predicted by only one mutation calling program, were selected. The selected variants were verified by using DNA, PCR amplification of regions using 50 ng of DNA followed by Sanger sequencing (Eurofins MWG Operon, Ebersberg, Germany). The reactions were successful at 50 and 32 loci for the positive and negative controls, respectively. Verification was also performed by examination of tumor RNA-Seq reads.

[0419] Calculation of FDR and machine learning Calculation of random forest quality scores: Commonly used algorithms for mutation calling (DePristo, M.A. et al., Nature Genetics 43, pp. 491 - 498 (2011), Li, H. et al., Bioinformatics 25, pp. 2078 - 2079 (2009), Ding, L. et al., Hum. Mol. Genet (2010), first published online on September 15, 2010) output multiple scores, all of which potentially affect the quality of mutation calling. These include, but are not limited to, the quality of the targeted base assigned by the instrument, the quality alignment at this position, the number of reads targeting this position, or the score of the difference between the two genomes compared at this position. Calculation of the false discovery rate requires ordering of the mutations, but this is not directly achievable for all mutations as there may be conflicting information from the various quality scores.

[0420] The inventors use the following strategy to achieve a complete ordering. In the first step, a very strict definition of superiority is applied by assuming that, and only if, one mutation has better quality than another in all categories. Thus, a set of quality characteristics S = (s1, …, s n ) is preferred to T = (t1, …, t n ), and is denoted by S > T if and only if s i > t i for all i = 1, …, n. Define the intermediate FDR (IFDR) as follows.

[0421]

Number

[0422] However, in many closely related cases no comparison is feasible, and thus no benefit can be drawn from the vast amount of available data, so the IFDR is considered only as an intermediate step. Therefore, the good generalization properties of random forest regression [Breiman, L., Statist. Sci. 16, pp. 199 - 231 (2001)] are exploited, and the random forest is trained as implemented in R (R Development Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria, 2010, Liaw, A., Wiener, M., R News 2, pp. 18 - 22 (2002)).

[0423] For m input mutations, each having n quality characteristics, the value range of each characteristic was determined, and values up to p at uniform intervals were sampled from within this range. If the set of values of the quality characteristics was smaller than p, this set was used instead of the sampled set. Subsequently, all possible combinations of the sampled or selected quality values were created, resulting in a maximum of p data points in the n-dimensional quality space. A random sample of 1% of these points and the corresponding IFDR values were used as predictors and responses, respectively, for random forest training. n The resulting regression score is the generalized quality score Q, which can be regarded as a locally weighted combination of the individual quality scores. This enables the direct comparison of single values of any two mutations and the calculation of the actual false discovery rate.

[0424] To train the random forest model used to generate the results of this study, after calculating the IFDR of the samples for all somatic mutations of the samples, a random 1% subset was selected. This ensures that the entire available quality space is mapped to FDR values. The quality characteristics "SNP quality", "coverage depth", "consensus quality", and "RMS mapping quality" (SAMtools, p = 20); "SNP quality", "coverage depth", "variant confidence / non-filtered depth", and "RMS mapping quality" (GATK, p = 20); or "SNP quality", "coverage depth", "consensus quality", "RMS mapping quality", and "somatic score" (SomaticSNiPer, p = 12) were used respectively. Different p-values ensure comparable sizes of sets.

[0425]

Number

[0426]

[0427] ​Calculation of common coverage: The number of possible mutation calls can introduce a large bias into the definition of the false discovery rate. The number of called mutations can be comparable and can serve as the basis for calculating the false discovery rate only when the number of possible positions where mutations can occur is the same for tumor comparison and autologous comparison. To correct for this potential bias, the common coverage ratio is used. As the common coverage, the number of bases having at least 1 coverage in both samples is defined and used for mutation calling. The common coverage is calculated separately for tumor comparison and autologous comparison.

[0428] Estimation of ROC The receiver operating characteristic (ROC) curve and the corresponding area under the curve (AUC) are useful for organizing classifiers and visualizing their performance [Fawcett, T., Pattern Recogn. Lett. 27, pp. 861-874 (2006)]. This concept is extended to evaluate the performance of experiments and computer procedures. However, plotting the ROC graph requires knowledge of all true and false positive (TP and FP) examples in the dataset, which is usually not given in high-throughput data (such as NGS data) and is difficult to establish. Therefore, the inventors estimate the ratio of each TP and FP using the calculated FDR, plot the ROC graph, and calculate the AUC. The central concept is that the FDR of a single mutation in the dataset gives the ratio of how much this mutation contributes to the total of TP / FP mutations, respectively. Also, in a list of random assignments to TP and FP, the resulting ROC AUC is equal to 0.5 in this method, indicating a completely random prediction.

[0429] Two conditions, namely

[0430]

Number

[0431] and FPR + TPR = 1 [2] Starting with, the FPR and TPR are the necessary false positive and true positive ratios for a given mutation, respectively, defining the corresponding points in the ROC space. [1] and [2] can be rearranged as follows. TPR = 1 - FPR [3] and FPR = FDR [4]

[0432] To obtain the estimated ROC curve, mutations in the dataset are screened by FDR, and for each mutation, the point is plotted where the cumulative TPR and FPR values up to this mutation are divided by the sum of all TPR and TPR values, respectively. The AUC is calculated by summing the areas of all consecutive trapezoids between the curve and the x-axis.

Claims

**Claim 1** A method of providing an individualized cancer vaccine, comprising: (a) identifying cancer-specific somatic mutations in a tumor specimen of a cancer patient to provide a cancer mutation signature of the patient; and (b) providing a vaccine characterized by the cancer mutation signature obtained in step (a). A method comprising the above steps. **Claim 2** The method according to claim 1, wherein the step of identifying cancer-specific somatic mutations comprises identifying a cancer mutation signature of the exome of one or more cancer cells. **Claim 3** The method according to claim 1 or 2, wherein the step of identifying cancer-specific somatic mutations comprises single cell sequencing of one or more cancer cells. **Claim 4** The method according to claim 3, wherein the cancer cells are circulating tumor cells. **Claim 5** The method according to any one of claims 1 to 4, wherein the step of identifying cancer-specific somatic mutations comprises using next generation sequencing (NGS). **Claim 6** The method according to any one of claims 1 to 5, wherein the step of identifying cancer-specific somatic mutations comprises sequencing genomic DNA and / or RNA of a tumor specimen. **Claim 7** The method according to claim 6, wherein the step of identifying cancer-specific somatic mutations is repeated at least in duplicate. **Claim 8** The method according to any one of claims 1 to 7, further comprising the step of determining the utility of the identified mutations in an epitope for cancer vaccination. **Claim 9** The method according to any one of claims 1 to 8, wherein the vaccine characterized by the patient's mutation signature comprises a polypeptide comprising a neoepitope based on a mutation, or a nucleic acid encoding said polypeptide. **Claim 10** The method according to claim 9, wherein the polypeptide comprises neoepitopes based on up to 30 mutations. **Claim 11** The method according to claim 9 or 10, wherein the polypeptide further comprises an epitope that does not contain cancer-specific somatic mutations expressed by cancer cells. **Claim 12** The method according to any one of claims 9 to 11, wherein the epitope has a natural sequence configuration such that it forms a vaccine sequence. **Claim 13** The method according to claim 12, wherein the vaccine sequence is about 30 amino acids in length. **Claim 14** The method according to any one of claims 9 to 13, wherein the neoepitopes, epitopes and / or vaccine sequences are arranged head-to-tail.

15. The method according to any one of claims 9 to 14, wherein the neoepitope, epitope and / or vaccine sequence are separated by a linker.

16. The method according to any one of claims 1 to 15, wherein the vaccine is an RNA vaccine.

17. The method according to any one of claims 1 to 15, wherein the vaccine is a prophylactic vaccine and / or a therapeutic vaccine.

18. A vaccine obtainable by the method according to any one of claims 1 to 17.

19. A vaccine comprising a recombinant polypeptide comprising a neoepitope based on a mutation, or a nucleic acid encoding said polypeptide, wherein said neoepitope is generated by a cancer-specific somatic mutation in a tumor specimen of a cancer patient.

20. The vaccine according to claim 19, wherein the polypeptide further comprises an epitope that does not contain a cancer-specific somatic mutation expressed by cancer cells.

21. A method of treating a cancer patient, comprising: (a) providing an individualized cancer vaccine by the method according to any one of claims 1 to 17; and (b) administering said vaccine to a patient. The method comprising.

22. A method of treating a cancer patient, comprising administering to the patient a vaccine according to any one of claims 18 to 20.

23. A method of determining a false discovery rate based on next-generation sequencing data, comprising: collecting a first sample of genetic material from an animal or a human; collecting a second sample of genetic material from an animal or a human; collecting a first sample of genetic material from tumor cells; collecting a second sample of genetic material from said tumor cells; determining a common coverage tumor comparison by counting all bases of a reference genome included in both the tumor and at least one of said first sample of genetic material from an animal or a human and said second sample of genetic material from an animal or a human; determining a common coverage versus identity comparison by counting all bases of a reference genome included in both said first sample of genetic material from an animal or a human and said second sample of genetic material from an animal or a human; forming a normalization by dividing said common coverage tumor comparison by said common coverage versus identity comparison. 1) determining the number of single nucleotide mutations having a quality score higher than Q in the comparison of the first sample of genetic material from an animal or a human and the second sample of genetic material from an animal or a human; 2) dividing by the number of single nucleotide mutations having a quality score higher than Q in the comparison of the first sample of genetic material from the tumor cells and the second sample of genetic material from the tumor cells; and 3) multiplying the result by the normalization to determine the false discovery rate A method comprising the steps of: **Claim 24** The method according to claim 23, wherein the genetic material is DNA. **Claim 25** wherein Q is A set of quality characteristics S = (s 1 ,…, s n ), for all i=1,...,n, i >t i If S is T=(t 1 ,…, t n ), denoted by S>T; 1) determining the number of single nucleotide mutations having a quality score S>T in the comparison of the first DNA sample from an animal or a human and the second DNA sample from an animal or a human; 2) dividing by the number of single nucleotide mutations having a quality score S>T in the comparison of the first DNA sample from the tumor cells and the second DNA sample from the tumor cells; and 3) multiplying the result by the normalization to define the intermediate false discovery rate; determining the value range of each of m mutations having n quality characteristics respectively; sampling up to p values from the value range; Create each possible combination of the sampled quality values, resulting in p n data points, and using the random sample of the data points as predictors for random forest training; using the corresponding intermediate false discovery rate value as the response for the random forest training and determined thereby, and the resulting regression score of the random forest training is Q. The method according to claim 23. **Claim 26** The method according to claim 24, wherein the second DNA sample from an animal or a human is allogeneic to the first DNA sample from an animal or a human. **Claim 27** The method according to claim 24, wherein the second DNA sample from an animal or a human is autologous to the first DNA sample from an animal or a human. **Claim 28** The method according to claim 24, wherein the second DNA sample from an animal or a human is heterologous to the first DNA sample from an animal or a human. **Claim 29** The method according to claim 23, wherein the genetic material is RNA. **Claim 30** wherein Q is A step of establishing a set of quality characteristics S = (s 1 , …, s n ), wherein for all i = 1, …, n, if s i > t i , then S is more preferable than T = (t 1 , …, t n ), indicated by S > T, and 1) determining the number of single nucleotide mutations having a quality score S>T in a comparison of said first RNA sample from an animal or human and said second RNA sample from an animal or human; 2) dividing by the number of single nucleotide mutations having a quality score S>T in a comparison of said first RNA sample from said tumor cells and said second RNA sample from said tumor cells; 3) multiplying the result by said normalization to define an intermediate false discovery rate; determining, for each of m mutations having n quality characteristics each, the value range of each characteristic; sampling up to p values from said value ranges; Create each possible combination of the sampled quality values, resulting in p n number of data points, and using a random sample of said data points as predictors for random forest training; using the corresponding intermediate false discovery rate value as the response for said random forest training; The method according to claim 29, wherein the regression score of said random forest training resulting therefrom is Q. **Claim 31** The method according to claim 30, wherein said second RNA sample from an animal or human is allogeneic to said first RNA sample from an animal or human. **Claim 32** The method according to claim 30, wherein said second RNA sample from an animal or human is autologous to said first RNA sample from an animal or human. **Claim 33** The method according to claim 30, wherein said second RNA sample from an animal or human is heterologous to said first RNA sample from an animal or human. **Claim 34** The method according to claim 23, wherein a vaccine formulation is produced using said false discovery rate. **Claim 35** The method according to claim 34, wherein said vaccine is deliverable intravenously. **Claim 36** The method according to claim 34, wherein said vaccine is deliverable transdermally. **Claim 37** The method according to claim 34, wherein said vaccine is deliverable intramuscularly. **Claim 38** The method according to claim 34, wherein said vaccine is deliverable subcutaneously. **Claim 39** The method according to claim 34, wherein said vaccine is customized for a specific patient. **Claim 40** The method according to claim 39, wherein one of said first sample of genetic material from an animal or human and said second sample of genetic material from an animal or human is from said specific patient. **Claim 41** The method of claim 23, wherein the step of determining a common coverage tumor comparison by counting all bases of a reference genome included in both the tumor and at least one of the first sample of genetic material from an animal or human and the second sample of genetic material from an animal or human uses an automated system to count all bases.

42. The method of claim 41, wherein the step of determining a common coverage versus identity comparison by counting all bases of a reference genome included in both the first sample of genetic material from an animal or human and the second sample of genetic material from an animal or human uses the automated system.

43. The method of claim 41, wherein the step of forming a normalization by dividing the common coverage tumor comparison by the common coverage versus identity comparison uses the automated system.

44. The method of claim 41, wherein the step of determining a false discovery rate comprises: 1) dividing the number of single nucleotide variants having a quality score higher than Q in a comparison of the first sample of genetic material from an animal or human and the second sample of genetic material from an animal or human by the number of single nucleotide variants having a quality score higher than Q in a comparison of the first sample of genetic material from the tumor cells and the second sample of genetic material from the tumor cells, and 3) multiplying the result by the normalization.

45. A method for determining a putative receiver operating characteristic (ROC) curve, comprising: receiving a dataset of mutations, each mutation having an associated false discovery rate (FDR); for each mutation: determining a true positive rate (TPR) by subtracting the FDR from 1; determining a false positive rate (FPR) by setting the FPR equal to the FDR; and forming a putative ROC by plotting, for each mutation, the point of the cumulative TPR and FPR values up to that mutation, divided by the sum of all TPR and FPR values. The method includes the above steps.

Citation Information

Patent Citations

  • CH-4010、

  • PCT/EP2006/009448

  • US983299303

  • Vaccines containing a saponin and a sterol

    WO1996033739A1

  • Immunostimulation mediated by gene-modified dendritic cells

    WO1997024447A1