Selection of neoepitopes as disease-specific targets for enhanced efficacy therapies

JP7900449B2Active Publication Date: 2026-08-04BIONTECH SE
View PDF 13 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BIONTECH SE
Filing Date
2024-08-05
Publication Date
2026-08-04

AI Technical Summary

Benefits of technology

【0063】 本発明の他の特長及び利点は、以下の詳細な説明及び特許請求の範囲から明らかであろう。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007900449000001
    Figure 0007900449000001
  • Figure 0007900449000002
    Figure 0007900449000002
  • Figure 0007900449000003
    Figure 0007900449000003
Patent Text Reader

Abstract

To provide: methods for determining whether neoepitopes that are only expressed in or on diseased cells are suitable disease-specific targets, such that the diseased cell is less likely to be able to escape immune surveillance; and use of the neoepitopes in providing an immune response against diseased cells expressing the neoepitopes.SOLUTION: The invention relates to a method for determining the suitability of a neoepitope resulting from a disease-specific mutation at an allele in a gene (mutated allele) as a disease-specific target, the method comprising determining, in a diseased cell or population of diseased cells, the copy number of the mutated allele encoding the neoepitope.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for determining the suitability of disease-specific neoepitopes as disease-specific targets, and to the use of such identified suitable neoepitopes in immunotherapy specifically targeted to affected tissues of a patient, such as tumor tissue, that express one or more of the identified suitable neoepitopes. [Background technology]

[0002] Cancer is the leading cause of mortality, accounting for one in four all-cause deaths. Cancer treatment has traditionally followed the law of averages, which best applies to the majority of patients. However, due to molecular heterogeneity in cancer, often less than 25% of treated individuals benefit from approved therapies. Personalized medicine, based on individually tailored treatments for each patient, is considered a potential solution to the low efficacy and high costs associated with drug discovery innovation.

[0003] Personalized cancer immunotherapy is emerging as a promising breakthrough in cancer treatment, with the potential to transform the standard of care for millions of cancer patients diagnosed worldwide each year. The integrative aspect of personalized cancer immunotherapy is that it allows the immune system to target genetic abnormalities (mutations) specific to a patient's cancer. Such disease-specific mutations can encode neoepitopes, which are disease-specific targets. The most prevalent genetic abnormalities in cancer genomes that can be used as disease-specific targets for personalized immunotherapy are non-synonymous single nucleotide polymorphisms (SNVs). Therefore, accurate and thorough identification of a patient's SNVs in the coding regions of their genome is a critical step in the process of producing personalized cancer immunotherapy.

[0004] However, as described herein, knowing the identity of disease-specific mutations is only part of the picture. Rather, complete genetic profiling of mutations requires knowledge of the exact number of copies of the gene containing the mutation in affected cells, e.g., tumor cells (including both wild-type and mutated alleles), the number of copies of the mutated allele in tumor cells (referred herein to as the mutational conjugation status), and the degree of subclonality of the mutation in samples of affected cells, such as tumor samples. In fact, copy number diversity occurring in affected cells is an important component of genetic diversity in affected cells across most disease manifestations. Furthermore, the degree of copy number diversity, the identity of the genes affected by copy number diversity, and the exact genetic composition of copy number diversity are unique to each individual and can vary widely between individuals. For general information, see Shlien and Malkin, 2009, Genome Med. 1:62; Yang et al., 2013, Cell 153:919-929. Precise knowledge of these genetic characteristics can be crucial in selecting mutations that, when targeted, have the potential to provide immunity against tumor evasion and thus confer overall tumor control. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] European Patent No. 2 198 292(B1) [Patent Document 2] European Patent No. 2 002 016(B1) [Patent Document 3] European Patent Application Publication No. 2 835 752(A) [Patent Document 4] International Patent Application Publication Number WO 2014 / 014497 [Patent Document 5] WO 2014 / 138153 [Patent Document 6] WO 97 / 24447 [Patent Document 7] PCT / EP2006 / 009448

Patent document 8

Non-licensed literature

[0006]

Non-licensed literature 1

Non-licensed Document 2

Non-licensed Document 4

Non-licensed Document 5

Non-licensed Document 6

Non-licensed Document 7

Non-licensed Document 8

[0007] To maximize the effectiveness of personalized cancer immunotherapy and confer sustained tumor control to the majority of treated patients, therapies must circumvent the tumor's ability to evade immune surveillance in some way, for example, by silencing the expression of mutated targets or by deleting genes. If the mutation is not expressed, for example, deleted from the genome, immunotherapy cannot target the mutation, and if this problem is not addressed, immunotherapy will risk recurrence. The selection of suitable neoepitopes that enhance tumor control will benefit any personalized immunotherapy approach that targets neoepitopes, regardless of how it is implemented. Therefore, in this field, there is a need for a method for selecting neoepitopes resulting from disease-specific mutations that lead to enhanced tumor control. [Means for solving the problem]

[0008] This invention provides a method to overcome the limitations of the current state of the art by providing a method for determining the suitability of neoepitopes resulting from disease-specific mutations in genes as disease-specific targets, which would result in enhanced tumor control in cancer, where the affected tissue cannot easily evade immune surveillance. Once a suitable neoepitope is identified, such a suitable epitope can be used as a disease-specific target to induce a specific immune response in patients with the disease. For example, the disease may be cancer and potentially a primary tumor, and tumor metastases expressing a suitable neoepitope can be targeted for more effective treatment.

[0009] The present invention relates to a method for determining the suitability of a neoepitope resulting from a disease-specific mutation in an allele (mutated allele) in a gene as a disease-specific target, the method comprising the step of determining the copy number of the mutated allele encoding the neoepitope in an affected cell or a population of affected cells. As used herein, the copy number may also be referred to as the conjugation state, for example, if the copy number of the mutated allele is 4, the mutated allele has a conjugation state of 4. As used herein, an allele is a site in the genome having a specific nucleotide identity, which may be the same in both the maternal and paternal copies of the genome (homozygous genotype), or the identity may differ in the maternal and paternal copies of the genome (heterozygous genotype). A mutated allele is an allele that, due to a disease-specific mutation, has a different identity from the site in the corresponding normal genome, for example, in unaffected cells of the same individual, preferably in a genome derived from unaffected cells of the same tissue type as the affected cells (matched genome). A neoepitope suitable as a disease-specific target (suitable neoepitope) is, as used herein, a neoepitope that is unlikely to be downregulated or silenced (e.g., by deletion) by the affected tissue when targeted by the immune system, such that the affected tissue is less likely to be able to avoid a response, preferably an immunological response, generated against the neoepitope by, for example, vaccination against the neoepitope or administration of immune cells capable of targeting (binding to) the neoepitope. In some embodiments, the copy number of a mutated allele may be the same as the copy number of the gene containing the mutated allele, and therefore the present invention also relates to a method for determining the suitability of a neoepitope resulting from a disease-specific mutation in a gene as a disease-specific target, the method comprising the step of determining the copy number of the gene having the disease-specific mutation in an affected cell or a population of affected cells.

[0010] In one embodiment, a high copy number of a mutated allele or gene with disease-specific mutations indicates suitability as a disease-specific target for the neoepitope, such that the higher the copy number of the mutated allele or gene with disease-specific mutations, the higher the suitability of the neoepitope as a disease-specific target. In one embodiment, if the copy number of a mutated allele or gene with disease-specific mutations in affected cells is greater than 2, this indicates suitability as a disease-specific target for the neoepitope. In one embodiment, if the copy number of a gene with disease-specific mutations is greater than 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or greater than 100, this indicates suitability as a disease-specific target for the neoepitope.

[0011] If not all copies of a gene having at least one mutated allele have the mutation, it is preferable that many copies of the gene have the mutated allele, i.e., a higher proportion of the gene copies, rather than a lower proportion, have the mutated allele (a higher fractional zygosity, rather than a lower one). Therefore, in certain embodiments, the mutated allele is found in a high proportion of copies of the gene having at least one mutated allele (zygosity), and the zygosity is the ratio of the number of copies of the mutated allele (zygosity of the mutated allele) to the total number of copies of the nucleotide site to which the mutated allele is mapped, in particular, to the reference genome or the corresponding wild-type genome or matched genome, i.e., the wild-type genome from the same individual (zygosity of the mutated allele). The higher the zygosity of the copies of the mutated allele, the better the neoepitope is as a disease-specific target. Preferably, the junction rate can exceed 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 0.95, and most preferably, the junction rate is 1, i.e., all copies of the gene in the affected cell have the mutated allele. When the junction rate is 1, there are no wild-type copies of the gene so that the affected cell cannot revert to the expression of the corresponding wild-type epitope. As used herein, a ratio of 1 means that the hypothesis that the genetic arrangement of the mutated allele / gene, e.g., copy number, and junction status are the same cannot be disproven by the data, i.e., it is statistically consistent.

[0012] It is known that diseased tissues such as tumors can have heterogeneous genetic makeup and gene expression. Therefore, not all diseased cells in diseased tissue have the same number of copies of a gene and / or a gene with at least one mutated allele (total number of gene copies) and / or a copy of the gene with the mutated allele. This is even more true in the case of disease-specific mutations themselves. Therefore, it is desirable (a high clonal proportion rather than a low one) if, for example, the number of copies of the mutated allele, and / or the junction rate, and / or the total number of copies of the nucleotide site to which the mutated allele is mapped is the same or similar in a high percentage of diseased cells in the diseased tissue, rather than in a low percentage of diseased cells. For example, the higher the percentage of diseased cells that have the same or similar number of copies of the mutated allele, and / or the junction rate, and / or the total number of copies of the nucleotide site to which the mutated allele is mapped, the more suitable the neoepitope is as a disease-specific target. For example, the clonal percentage may be at least 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, or at least 0.9. In a preferred embodiment, all affected cells in the affected tissue have the same or similar copy number, i.e., the clonal percentage is 1, i.e., statistically consistent. As used herein, the same or similar copy number includes the same copy number, or copy numbers within 30%, 25%, 20%, 15%, 10%, 5%, 4%, 3%, 2% or less of the copy number or absolute copy number, for example, with or without error correction.

[0013] Preferably, the clonal proportion of a mutation can be obtained by the proportion of affected cells having the same or similar genetic configuration of the mutation, where the genetic configuration of the mutation includes the total number of copies of the nucleotide site to which the mutation is mapped and the number of copies of the mutated allele. A feature is said to be fixed in the population of affected cells if it is present in all affected cells to such an extent that it cannot be statistically disproven by the available data. Preferably, a clonal proportion of 1 means that the genetic configuration of the mutation is fixed in the population of affected cells. Preferably, the genetic configuration of the mutation is fixed in the population of affected cells if the mutation is fixed in the population of affected cells and the CNV affecting the site encoding the mutation is fixed in the population of affected cells. Preferably, the genetic configuration of the mutation is fixed if the total number of copies of the nucleotide site to which the mutation is mapped is 2 and the mutation present in the equilibrium region of the affected (tumor) genome is determined to be fixed in the population of affected cells.

[0014] Genes in which disease-specific mutations are found can potentially exist in any gene in the genome. Preferred types of genes in which mutations resulting in suitable neoepitopes are those whose expression leads to the transformation of cells into an oncological phenotype, or whose absence of expression leads to cancer cells that lose that oncological phenotype; that is, genes whose expression contributes to tumor progression. Such genes are known as driver genes. Examples of driver genes for many types of tumors are well known. For example, a list of 291 highly reliable cancer driver genes acting on 3,205 tumors derived from 12 different cancer types is disclosed in Tamborero et al., 2013, Comprehensive identification of mutational cancer driver genes across 12 tumor types, Scientific Reports 3:2650. Further driver genes were identified using the methods disclosed in Youn et al., 2011, Identifying cancer driver genes in tumor genome sequencing studies, Bioinformatics 27(2):175-181; Sakoparnig et al., 2015, Identification of constrained cancer driver genes based on mutation timing, PLoS Comput. Biol. 11(1):e1004027; and Forbes et al., 2008, Current protocols in human genetics, pp. 10-11. Disease-specific mutations in driver genes may or may not contribute to the oncological phenotype. Preferably, all copies of the driver gene found in affected cells have disease-specific mutations. Also preferably, all cells in affected tissue are affected cells in which all copies of the driver gene have disease-specific mutations.

[0015] Another preferred type of gene is an essential gene. In one embodiment, an essential gene is a gene that, when silenced or its expression is reduced (e.g., by deletion), results in at least impaired growth or reduced fitness in cells, preferably diseased cells. Such genes are referred to herein as essential genes. In one embodiment, an essential gene is a gene that shows at least 10% reduced growth or reduced fitness in diseased cells when the gene is silenced or its expression is reduced compared to cells in which the gene is neither silenced nor reduced in expression. In one embodiment, the reduced growth or reduced fitness is at least 20%, 30%, 40%, 50%, 60%, 75%, 80%, 90%, or at least 95%, and most preferably, the silencing or reduction in expression of the essential gene results in lethality of the diseased cells. Preferably, the entire copy of the essential gene found in the diseased cells has a disease-specific mutation.

[0016] Essential genes are well known in the art, and for example, lists of essential genes in humans (e.g., in human cell lines or inferred from other organisms) are disclosed in Liao et al., 2008, Proc. Nat. Acad. Sci. USA 105: pp. 6987-6992 and Georgi et al., 2013, PLoS Genetics 9 (5): e1003484, along with those for mice (Liao et al., 2007, Trends Genet. 23: pp. 378-381), fruit flies (Spradling et al., 1999, Genetics 153: pp. 135-177), nematodes (C. elegans) (Kamath et al., 2003, Nature 421: pp. 231-237), and zebrafish (Amsterdam et al., 2004, Proc. Natl. Acad. Corresponding orthologues in other eukaryotes have also been disclosed, such as Arabidopsis thaliana (Tzafrir et al., 2004, Plant Physiol. 135:1206-1220), and yeast (Kim et al., 2010, Nat. Biotechnol. 28:617-623). A list of essential genes derived from human cancer cell lines is disclosed in Wang et al., 2015, Science 350:1096-1101, and a list of essential genes can be found in the essential gene database, DEG5.0 (Zhang et al., 2009, Nucleic Acids Res. 37:D455-D458).

[0017] Furthermore, a list of essential genes whose deletion / silencing significantly reduces the fitness of a cell line cohort can be empirically generated from multiple healthy tissues and / or cancer cell lines, which may originate from a donor or patient. Gene deletion / silencing can be performed experimentally using various molecular biology techniques, such as CRISPR technology and RNA interference, and cell viability or fitness is determined by the presence or absence of expression of the putative essential genes. The list of essential genes can also be determined experimentally from cells or cell lines, or it can be obtained by a bioinformatics approach. Cells or cell lines may be affected cells or cell lines (tumor cells or cell lines) or unaffected (healthy / normal) cells or cell lines, and may be obtained from a donor or a patient with the disease. Preferably, unaffected cells or cell lines originate from the same histological type as affected cells, and more preferably from the same patient. In embodiments where the disease is cancer, cells or cell lines may be obtained from a primary tumor or any metastases, if present. Furthermore, the list of essential genes may be essentially the same as the minimum set of genes expressed in a wide variety of tissues within the body. For example, essential genes are genes expressed in a wide variety of different tissues and are expressed at RPKM (minimum reads per kilobase of transcript per million mapped reads) thresholds greater than 0, preferably 0.1, 0.5, 1, 2, 3, 4, 5, 10, 20, and 25. Such a list of essential genes can be obtained by analyzing RNA expression data (e.g., RNA-seq) obtained from a panel of cell samples from at least 5, 6, 7, 8, 9, 10, 15, 20, and 25 different tissues. Furthermore, if patient-derived tumor cell lines are available, the growth rate of each modified cell line can be measured by deleting one copy at a time of genes containing mutations encoding neoepitopes.Such measurements / analyses can be performed by high-throughput methods known in the art, which allow for the screening of at least one gene, preferably many genes, at once to evaluate their effect on the fitness of affected or unaffected cells. Such methods also allow for the detection of synthetic pathogenic or lethal combinations of genes, as described later. Briefly, deletions of one or more candidate genes can be tested using a library of cell lines, each lacking one gene, so that the effect of gene deletions on cells can be determined.

[0018] The present invention further relates to a method for determining the suitability of a neoepitope resulting from a disease-specific mutation in a gene as a disease-specific target, comprising the step of determining the copy number of a gene in an affected cell or a population of affected cells, that is, determining the copy number of a gene in which at least one copy has a disease-specific mutation. When a gene has a high copy number in affected cells such as tumor cells, for example, due to localized amplification, the probability that the gene can become a driver gene is very high. Therefore, a high copy number of a gene indicates the suitability of the neoepitope as a disease-specific target, and the higher the copy number of the gene, the higher the suitability of the neoepitope as a disease-specific target. For example, a high copy number can be a copy number greater than 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90 or greater than 100. A high copy number may also be at least 50% greater than the copy number of the gene in the corresponding unaffected cells. High copy number may refer to a case where the copy number of a gene having at least one copy of a disease-specific mutation is at least 2×, 3×, 5×, 10×, 15×, 20×, 25×, 30×, 40×, 50×, 60×, 70×, 80×, 90×, or at least 100× greater than the copy number of the gene in the corresponding unaffected cell. Due to copy number diversity that can also exist in normal genomes, the copy number of a gene in a normal genome is not necessarily 2. Furthermore, localized amplification is known to be observed more frequently in certain diseases than in others, such as in glioblastoma, where the epidermal growth factor receptor gene is frequently locally amplified; therefore, this embodiment is well suited for use in these diseases.

[0019] Furthermore, it is preferable that the gene copy number is found to be the same or similar in a high percentage of affected cells rather than in a low percentage of affected cells, so that the higher the percentage of affected cells having the same or similar copy number, the greater the suitability of the neoepitope as a disease-specific target. In a preferred embodiment, all affected cells in the affected tissue have the same or similar copy number of the gene having at least one copy of the disease-specific mutation, i.e., the clonal proportion is 1. As used herein, the same or similar copy number includes the same copy number, or copy numbers within 30%, 25%, 20%, 15%, 10%, 5%, 4%, 3%, and 2% or less of the copy number or absolute copy number, for example, with or without error correction.

[0020] Furthermore, a high-copy-number gene in which at least one copy has a disease-specific mutation resulting in a neoepitope is preferably a driver gene, such as a driver gene known in the art, whose expression results in the transformation of cells into an oncological phenotype, or whose lack of expression results in cancer cells that lose that oncological phenotype. Alternatively, it may be an essential gene, for example, a gene that, when silenced or its expression is reduced, results in at least impaired growth or reduced fitness of affected cells.

[0021] The present invention also relates to a method for determining the suitability of a neoepitope resulting from a disease-specific mutation in a gene as a disease-specific target, the method comprising the step of determining whether the gene having the disease-specific mutation is an essential gene in affected cells or a population of affected cells. In one embodiment, an essential gene is a gene that, when silenced or its expression is reduced (e.g., by deletion of the gene), results in at least impaired growth or reduced fitness of affected cells. In this embodiment, a neoepitope shows favorable suitability as a disease-specific target if the gene is an essential gene and all copies of the essential gene have the disease-specific mutation (conjugation rate of 1). In one embodiment, an essential gene is a gene expressed in a wide variety of different tissues and expressed at an RPKM (minimum number of reads per kilobase of transcript per million mapped reads) threshold greater than 0, preferably greater than 0.1, 0.5, 1, 2, 3, 4, 5, 10, 20, or 25. Preferably, all copies of the essential gene contain the mutation. Furthermore, it is preferable that a high proportion of affected cells contain copies of essential genes with disease-specific mutations (a high clonal proportion, not a low one), so that the higher the proportion of affected cells, the more suitable the neoepitope is as a disease-specific target. In a more preferred embodiment, all affected cells in the affected tissue have copies of essential genes with disease-specific mutations, i.e., the clonal proportion is 1.

[0022] It is known that when certain genes are individually silenced or their expression is individually reduced, they may have only a very slight effect, if any, on the fitness or growth capacity of affected cells. However, it has been observed that when two such genes are both silenced or their expression is reduced, it can result in a much stronger growth impairment that leads to death. Such genetic combinations are called synthetic lethal or synthetic pathogenic / impairmental. For a discussion of synthetic lethal and synthetic pathogenic genes, and methods for identifying such genes, see Nijman, 2011, Synthetic lethality: General principles, utility and detection using genetic screens in human cells, FEBS Lett. 585: pp. 1-6. Since both genes are required for cell survival, it is unlikely that cells will silence or reduce the expression of both genes. Therefore, suitable combinations of neoepitopes as disease-specific targets may result from disease-specific mutations in at least two genes, both of which are synthetically lethal or synthetically pathological. Considering this, the present invention further relates to a method for determining the suitability of a combination of at least two neoepitopes resulting from disease-specific mutations in at least two genes as a combination of disease-specific targets, the method comprising the step of determining whether the combination of at least two genes, each having a disease-specific mutation, is a synthetically lethal or synthetically pathological gene. If the combination of at least two genes results in a synthetically lethal or synthetically pathological phenotype, this indicates that the resulting neoepitope is a suitable combination of disease-specific targets. In preferred embodiments, synthetic pathology results in at least a better effect on cell growth / fitness than would be expected from the additive effect of deletion / decreased expression of each gene individually. Since the number of possible combinations increases with the number of neoepitopes, this approach is preferred when a large number of suitable neoepitopes exist.For example, 10 mutations correspond to 45 possible combinations, 100 mutations correspond to 4950 combinations, and 1000 mutations correspond to approximately 500,000 combinations. In certain embodiments, at least two genes each have a higher conjugation rate rather than a lower one, preferably a conjugation rate of 1, and / or each have a higher clonal proportion rather than a lower one, preferably a clonal proportion of 1, both of which are found in affected cells and affected cells of affected tissue. As used herein, each neoepitope found in a suitable neoepitope combination is considered to be a neoepitope suitable for the purposes of the present invention.

[0023] As referenced herein, the copy number of a gene in either an infected or uninfected cell may be a relative copy number, but is preferably an absolute copy number, and more preferably a polyploid copy number, such as an absolute copy number normalized to the polyploidy of the genome of the infected cell, i.e., the copy number of the genome. Even more preferably, the relative, absolute, and normalized copy numbers are error-corrected.

[0024] For example, when estimating the absolute copy number or conjugation status of a mutated allele or gene, the estimation may be inaccurate, and it is desirable to correct the absolute copy number to account for sources of error. In embodiments in which next-generation sequencing is used to obtain sequence information from genomes and exomes, sources of error may include: bias in the estimated purity of a sample of affected tissue such as a tumor sample; bias in the estimation parameters required to derive purity and / or absolute copy number; probabilistic error due to finite coverage of the sample being sequenced; limited detection capability due to low purity, low clonal proportion, etc.

[0025] In some embodiments, where the absolute copy number is determined using a balanced heterozygous segment containing a heterozygous SNP, such as that disclosed in the international PCT patent application titled "Tumor Modeling Based on Primary Balanced Heterozygous Segments," filed on the same date as this specification, which incorporates its entire disclosure herein by reference, errors in the absolute copy number of the segment may be transmitted to other estimated parameters, such as the absolute copy number of the mutated allele, e.g., the SNV encoding the neoepitope, the conjugation status of the mutated allele, the clonal proportion, etc. Since such downstream estimated parameters have clinical significance for determining the suitability of the neoepitope as a disease-specific target for the patient described herein, it is desirable to correct for errors in the absolute copy number to obtain the most accurate value of the absolute copy number. Furthermore, mutations encoding neoepitopes can be prioritized for their suitability for inclusion in vaccines administered to patients using the criteria described herein, and especially when the gene containing the mutation is an essential gene, a proportion 1 zygote state is qualitatively superior to a proportion zygote state of less than 1, so it is beneficial to have the most accurate estimates of the absolute copy number and any parameters derived therefrom.

[0026] In a preferred embodiment, the mutated allele, such as the absolute copy number of an SNV and / or the conjugation status of the SNV, can be error-corrected. In one embodiment, the absolute copy number of the SNV is first error-corrected, and then the conjugation status of the SNV is corrected to reflect the error-corrected absolute copy number of the SNV. Once the absolute copy number of the SNV and / or the conjugation status of the SNV have been error-corrected, the estimate of the clonal proportion can also be corrected to reflect the error-corrected absolute copy number, including a determination of whether the clonal proportion is statistically consistent with a value of 1.

[0027] The absolute copy number of an SNV is preferably obtained from the absolute copy number of the segment to which the SNV is mapped. The absolute copy number of all segments in the affected genome, e.g., in a tumor genome, can be error-corrected, and this includes the absolute copy number of the SNV. Once the absolute copy number of all segments in the genome has been error-corrected, the error-corrected ploidy can be calculated based on the error-corrected absolute copy number of the segment in the affected genome, e.g., in a tumor genome.

[0028] In specific embodiments, if the absolute copy number of an SNV is error-corrected so that the new absolute copy number differs from the original absolute copy number, this can be interpreted as an indication that the estimated absolute copy number of the SNV is unreliable.

[0029] An example of error correction for the absolute copy number of a segment is parity error correction, which includes correction for odd absolute copy numbers of a segment relative to even absolute copy numbers when the segment lies in an equilibrium region. An equilibrium region is a region of the affected genome, such as a tumor genome, where the maternal and paternal alleles within this region underwent equivalent (equilibrium) amplification, or neither the maternal nor paternal alleles underwent any amplification at all.

[0030] The decision to error-correct the odd absolute copy number of a segment for the nearest higher even absolute copy number or the nearest lower even absolute copy number may depend on a comparison with disease and normal read data mapped to the segment, as well as with a predicted boundary defining the absolute copy number of the segment. If normal read data is related to the sequenced normal sample, then disease read data is related to the sequenced affected sample. In particular, if the disease is cancer, tumor read data is related to the sequenced tumor sample.

[0031] For example, in the first type of parity error correction, CN mutHowever, if the mutation is the absolute copy number in the affected genome of the segment to which it is mapped, then CN mut The absolute number of copies of a segment that is predicted to have a value of r is r > ρ th If so, CN mut It can be corrected to +1, so r < ρ th If so, CN mut It can be corrected to -1, where r is the ratio of disease segment read data counts to normal (the ratio of the number of disease, e.g., tumor read data mapped to a segment to the number of normal read data mapped to a segment), and ρ th This is the predicted decision boundary, and its value also depends on the purity of the diseased tissue sample, such as a tumor sample.

[0032] The allele-specific copy number of a segment is the number of copies of either the maternal or paternal allele of the segment in the affected genome. If a segment contains heterozygous SNPs (heterozygous segments), the allele-specific copy number of the segment can be determined using the heterozygous SNPs. Heterozygous segments can be assigned to preferred nodes, which can be defined as a specific combination of the absolute copy number of the heterozygous segment and the allele-specific copy number of the heterozygous segment. Even nodes are a subset of nodes in which the absolute copy number of the segment is even. If a heterozygous segment contains more than one heterozygous SNP, a group of two or more heterozygous SNPs may be represented by a single member of the group, or the allele frequencies of all members of the group may be averaged, or the median may be adopted, as long as the allele frequency of each heterozygous SNP is calculated in a consistent manner for any allele that has more or fewer copies in the affected genome.

[0033] The first type of parity error correction for heterozygous segments may involve finding the even node most likely to correspond to the heterozygous segment based on a maximum likelihood framework that considers the measured disease (e.g., tumor) read data and normal read data mapped to the segment.

[0034] In the second type of parity error correction, the nearest upstream and downstream segments that do not require parity error correction are preferably identified within 10 Mb, 5 Mb, and 1 Mb of the segment containing the SNV. If the absolute copy numbers of both the nearest upstream and downstream segments are the same, the absolute copy number of the segment containing the SNV is changed to the absolute copy number of the nearest segment. Generally, the second type of parity error correction is preferred over the first type of parity error correction unless it cannot be performed due to the inability to identify a suitable adjacent segment, in which case the first type of parity error correction can be provided.

[0035] In lieu of or in addition to parity error correction, other forms of error correction for the absolute copy number of a segment may be introduced. For example, a method may be used in which the absolute copy number of the segment immediately adjacent to the gene containing the mutation in the affected genome is considered, preferably when the change in absolute copy number is 3, 2 or less, preferably 1, and preferably when the majority of adjacent segments (50%, 60%, 70%, 80%, 90%, 100%) have an absolute copy number equivalent to the mode, and the absolute copy number of the segment containing the SNV is changed to the mode of absolute copy number of the adjacent segments.

[0036] Different error correction schemes for absolute copy numbers can be combined. In a preferred embodiment, a first parity error correction is applied as a first layer of error correction, on which additional error correction methods can be applied.

[0037] As used herein, a segment may be a predetermined genomic region, for example, a predetermined one based on a reference genome. A segment may span, for example, genes defined in the reference genome to which the read data are aligned. A segment may be a gene fragment, an exon, a union of exons, or a union of exons related to a given gene. A segment may be another set of a predetermined region in the reference genome (with or without introns), or another set of a predetermined region in the reference genome based on a normal genome. In specific embodiments, a segment may be a region of the reference genome having a given constant copy number and / or a given allele-specific copy number in the affected, e.g., tumor genome, or a gene fragment having a given constant copy number and / or allele-specific copy number in the affected, e.g., tumor genome. A segment may be defined as having or excluding introns.

[0038] Numerous copies of a segment in a given genome (e.g., a normal genome or a tumor genome) can be defined as the total frequency of occurrence of the segment's nucleotide sequence in the genome, ignoring diversity resulting from SNPs and / or SNVs and / or other cancer-related changes, including but not limited to mutations, insertions, deletions and / or other cancer-related genetic variants. Preferably, different copies of a segment in a given genome have the same or approximately the same length.

[0039] The number of copies of a segment in a genome can mean the number of physical copies of the segment in the cell containing the genome. The absolute copy number of a segment in a normal genome can be defined as the number of physical copies of a given segment in healthy cells. The absolute copy number of a segment in an affected genome, e.g., a tumor genome, can be defined as the number of physical copies of a given segment in affected cells, e.g., tumor cells. The number of copies of a segment in a genome can be referred to as the absolute copy number of the segment in the genome. The term "copy number" can mean absolute copy number.

[0040] In specific embodiments, if only a portion of a segment is amplified or deleted in the genome, such partial copies of the segment may or may not be counted as copies of the segment. In preferred embodiments, copies of a segment that represent less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of the segment length may be ignored.

[0041] The reference genome is used to map the read data and to provide a coordinate system for normal genomes and diseased genomes, such as tumor genomes. The coordinate system may include providing chromosome numbers, nucleotide positions on the chromosomes, and orientation of the read data, with chromosomal positions indicated by lines.

[0042] The reference genome may be based on the genome of one or more members of the same species as the subject providing the affected tissue sample, or it may be based on the normal genome of the subject.

[0043] Tumor samples and other samples of affected tissue may contain contamination from normal genomes, particularly from the normal genome of the same patient from whom the sample was taken, and / or, in cases of intratumor heterogeneity, more than one tumor genome. Purity, tumor sample purity, tumor purity, and sample purity are all interpreted as equivalent terms and preferably mean the proportion of tumor cells present in the tumor sample. Normal contamination preferably means the proportion of normal cells present in the tumor sample and can be obtained by a purity of 1 minus 1.

[0044] Normalization to cell ploidy controls the presence of gene copies due to genome duplication events. In one embodiment, the absolute copy number can be normalized to genome ploidy, where ploidy is the average of the absolute copy numbers of all segments in a given genome in a given cell, weighted by the length of each segment. In another embodiment, the absolute copy number can be normalized to chromosome ploidy containing the mutated gene of interest (including the mutation), where ploidy is the average of the absolute copy numbers of all segments on a given chromosome in a given cell, weighted by the length of each segment on the chromosome. In yet another embodiment, the absolute copy number can be normalized to adjacent regions of the chromosome containing the mutated gene of interest, where ploidy is the average of each segment in a given region in a given cell, weighted by the length of each segment in the region. The adjacent region may be within a predetermined distance of the gene containing the disease-specific mutation, for example, within 100 megabases (Mb), 75 Mb, 50 Mb, 25 Mb, 10 Mb, 5 Mb, 4 Mb, 3 Mb, 2 Mb, or 1 Mb of the gene containing the disease-specific mutation. The copy number of the segment can be routinely calculated both experimentally and computationally by methods known in the art. For example, European Patent No. 2 198 292(B1) and European Patent No. 2 002 016(B1) disclose methods for determining the relative copy number and copy number frequency of nucleic acid sequences, respectively. Furthermore, European Patent Application Publication No. 2 835 752(A) and International Patent Application Publication Nos. WO 2014 / 014497 and WO 2014 / 138153 also disclose methods for determining copy number diversity. See also Machado et al., 2013, Copy Number Variation of Fc Gamma Receptor Genes in HIV-Infected and HIV-Tuberculosis Co-Infected Individuals in Sub-Saharan Africa, PLoS, 8(11):e78165. Other methods include the use of FACS, FISH or other fluorescence-based methods, spectral karyotyping (SKY), and digital PCR.The segment may be a gene.

[0045] Disease-specific mutations are preferably any mutations that result in the expression of a neoepitope on the surface of affected cells. In particular, the mutation may be an indel or a gene fusion event, or it may be a single nucleotide variant (point mutation). Preferably, disease-specific mutations are non-synonymous mutations, preferably non-synonymous mutations of proteins expressed in tumor or cancer cells. Any method known in the art can be used to determine disease-specific mutations, and in particular, a method using next-generation sequencing data to determine any changes between the genome / exome of affected cells and the genome / exome of corresponding unaffected wild-type cells is preferred. For example, Carter et al., 2012, Absolute quantification of somatic DNA alterations in human cancer, Nature Biotechnology 30:413-421; Cibulskis et al., 2013, Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples, Nature Biotechnology 31:213-219; and Li and Li, 2014, A general framework for analyzing tumor subclonality using SNP array and DNA sequencing data, Genome Biology 15:473-495, disclose not only methods for identifying disease-specific mutations but also methods for determining gene copy number, conjugation rate, and conjugation status, as well as proportional subclonality and conjugation rate.Another method for determining copy number, such as absolute copy number, involves the use of genomic segments, each segment containing at least one heterozygous single nucleotide polymorphism (SNP), which is balanced (equal numbers of each version of the heterozygous SNP) and shares a common number of copies (primary copy number), which is preferably the most frequently observed absolute copy number of all balanced segments of the genome, which is disclosed in the international PCT patent application titled "Tumor Modeling Based on Primary Balanced Heterozygous Segments," filed on the same date as this specification, and which is incorporated herein by reference in its entirety. Furthermore, in addition to determining absolute copy number, this application can also determine the conjugation status, conjugation rate, and subclonality of genes containing mutated alleles or alleles with at least one mutated copy. Moreover, the methodology also performs error correction for absolute copy number, which improves the accuracy of absolute copy number and conjugation status, as well as the parameters derived therefrom, such as subclonality, ploidy, etc.

[0046] Generally, the total number of copies of the nucleotide site to which the mutated allele is mapped can represent the absolute copy number of the mutation, which in particular can represent the absolute copy number of the SNV if the mutation is an SNV. Generally, the absolute copy number of a mutation can preferably be obtained from the absolute copy number of the segment to which the mutation is mapped (the absolute copy number is that of the affected genome, e.g., in the tumor genome).

[0047] In general, the copy number of a mutated allele encoding a neoepitope can mean the absolute copy number of the mutated allele of the mutation, which can mean the absolute copy number of the alternative allele of the SNV (the conjugation state of the SNV), and in particular, if the mutation is an SNV, the alternative allele of the SNV is the mutated allele.

[0048] Preferably, if the mutation is not an SNV, the absolute copy number of the mutated allele of the mutation can be estimated in a manner similar to that applied to SNVs.

[0049] Generally, the copy number of a gene can refer to the absolute copy number of a segment, and a segment can be a gene or contain a gene. The copy number of a gene can refer to the absolute copy number of a gene.

[0050] Preferably, the disease may be any disease in which an immune response to affected cells / tissues, such as virus-infected cells, is desired. Preferably, the disease is cancer.

[0051] The method of the present invention may include further steps to determine the usefulness / validity of the suitable neoepitope identified by the method of the present invention as a disease-specific target suitable for use in methods to induce an immune response to a suitable neoepitope, such as the inclusion of a suitable neoepitope in a cancer vaccine. Thus, further steps may include one or more of the following: determining the antigenicity and / or immunogenicity of the suitable neoepitope; evaluating whether the suitable neoepitope is expressed on the surface of affected cells; the ability of a peptide containing the suitable neoepitope to be presented as an MHC-presenting epitope; determining the effectiveness of expression of the suitable neoepitope from an encoding nucleic acid; and determining whether the assumed suitable neoepitope can stimulate T cells, such as patient T cells, with desired specificity, particularly when it exists in the context of its natural sequence, for example, when it is similarly adjacent to the neoepitope by an amino acid sequence in a naturally occurring protein, and when expressed in antigen-presenting cells.

[0052] Once a neoepitope is determined to be suitable for use as a target, taking into account its antigenicity / immunogenicity, expressive capacity, and ability to be presented as an MHC-presenting epitope, the identified suitable neoepitopes can be ranked, or prioritized, in terms of their potential to be downregulated or deleted from affected cells, with the affected tissue being less likely to evade targeting of the neoepitope. For example, one prioritization might start with the "best" neoepitope, which is a neoepitope encoded by an essential gene, where every copy of the essential gene has a mutation encoding the neoepitope; next, a pair of synthetic pathogenic or lethal genes, where every copy of each gene has a mutation encoding the neoepitope; next, a neoepitope encoded by a known driver gene with a very high absolute copy number, where every copy of this gene has a mutation; next, a neoepitope encoded by a gene not known to be a driver gene with a very high absolute copy number and high zygote status; and so on.

[0053] In some embodiments, a neoepitope encoded by an essential gene having a junction rate of 1 is preferred over other neoepitopes not encoded by an essential gene. In some embodiments, among two essential genes encoding a neoepitope with a junction rate of 1, the neoepitope encoded by the gene with a higher absolute copy number is preferred. In some embodiments, among two essential genes encoding a neoepitope, the neoepitope encoded by the gene whose deletion results in lower fitness is preferred. In some embodiments, among genes where all genes encode a neoepitope with a junction rate of 1, the neoepitope encoded by the gene with a higher absolute copy number is preferred. In some embodiments, when the conjugation rate is less than 1, neoepitopes encoded by genes with a high conjugation rate are preferred over genes with a high conjugation rate, and when the conjugation rates are the same or similar, genes with a high absolute copy number are preferred over genes with a high conjugation rate (10 copies of mutated alleles / 20 total copies of nucleotide site are better than 3 / 4 for a higher conjugation rate; 10 / 100 are better than 10 / 20 because the former may be a driver gene; 9 / 100 are better than 10 / 20 because although the conjugation rates are similar, the former may be a driver gene). In some embodiments, disease-specific mutations in neoepitopes encoded by driver genes that cause cell transformation to an oncological phenotype are preferred over mutations that do not play a role in cell transformation to an oncological phenotype. Furthermore, it is preferable that the neoepitopes have a higher clonal proportion rather than a lower one.

[0054] Further embodiments of the present invention relate to the use of a method for determining the suitability of neoepitopes as disease-specific targets for the manufacture of pharmaceuticals, such as vaccines, for example, personalized cancer vaccines. The vaccine may be derived from one or more suitable neoepitopes, or a combination of suitable neoepitopes identified by the method of the present invention. In preferred embodiments, the vaccine may include a peptide or polypeptide comprising one or more suitable neoepitopes, or a combination of suitable neoepitopes identified by the method of the present invention, or a nucleic acid encoding the peptide or polypeptide.

[0055] In particular, a recombinant vaccine can be provided that, when administered to a patient, preferably yields a collection of MHC-presented epitopes, where at least one of the collection is a suitable neoepitope, or at least two of the collection are suitable combinations of neoepitopes identified by the method of the present invention, such as two or more, five or more, ten or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, preferably up to sixty, five or more, five or more, four or more, four or more, four or more, or three or more MHC-presented epitopes. The presentation of these epitopes by the patient's cells, particularly antigen-presenting cells, preferably yields T cells that target the patient's tumor, preferably the primary tumor and tumor metastases, which express the epitope when bound to MHC, and thus the antigen from which the MHC-presented epitope originates, and present the epitope on the surface of tumor cells.

[0056] The method of the present invention is also useful in producing recombinant immune cells that express an antigen receptor targeted to one neoepitope in a suitable neoepitope or a combination of suitable neoepitopes. Preferably, the immune cells are T cells and the antigen receptor is a T cell receptor.

[0057] The present invention also relates to recombinant immune cells produced by a method for producing recombinant immune cells targeted to a suitable neoepitope or one epitope in a combination of suitable neoepitopes, the method comprising the step of transfecting immune cells with a recombinant antigen receptor targeted to a suitable neoepitope or one epitope in a combination of suitable epitopes, which is identified by the present invention's method for determining the suitability of a neoepitope as a disease-specific target.

[0058] The present invention also provides a method for targeting a population of cells or tissues expressing one or more neoepitopes. For example, an antibody against one or more neoepitopes can be used to target cells or tissues expressing one or more neoepitopes identified by the method described herein. In one embodiment, the present invention provides a method for eliciting an immune response against a target population of cells or tissue expressing one or more neoepitopes in a mammal, comprising the steps of (a) administering one or more immune cells expressing one or more antigen receptors targeted to one or more neoepitopes to a mammal; (b) administering a nucleic acid encoding one or more neoepitopes; or (c) administering a peptide or polypeptide containing one or more neoepitopes, wherein the neoepitopes are identified according to the method of the present invention for determining the suitability of the neoepitopes as disease-specific targets. In one embodiment, a method for eliciting an immune response against a target cell population or target tissue expressing one or more neoepitopes in a mammal includes the steps of: (i) determining the copy number (disease-specific mutation) of a mutated allele in a gene encoding a neoepitope in an affected cell or affected cell population; and (ii) (a) administering immune cells expressing an antigen receptor targeted to a neoepitope resulting from a disease-specific mutation; (b) administering a nucleic acid encoding a neoepitope resulting from a disease-specific mutation; or (c) administering a peptide or polypeptide containing a neoepitope resulting from a disease-specific mutation.

[0059] In one embodiment, a method for eliciting an immune response against a target cell population or target tissue expressing one or more neoepitopes in a mammal includes the steps of: (i) determining the copy number of a gene in an affected cell or affected cell population, wherein at least one copy of the gene has a disease-specific mutation resulting in a neoepitope; and (ii) (a) administering immune cells expressing an antigen receptor targeted to a neoepitope resulting from a disease-specific mutation; (b) administering a nucleic acid encoding a neoepitope resulting from a disease-specific mutation; or (c) administering a peptide or polypeptide containing a neoepitope resulting from a disease-specific mutation. In one embodiment, a method for inducing an immune response against a target cell population or target tissue expressing one or more neoepitopes in a mammal includes the steps of: (i) determining whether a gene having a disease-specific mutation resulting in a neoepitope is an essential gene in the affected cells or population of affected cells; and (ii) (a) administering immune cells expressing an antigen receptor targeted to a neoepitope resulting from a disease-specific mutation; (b) administering a nucleic acid encoding a neoepitope resulting from a disease-specific mutation; or (c) administering a peptide or polypeptide containing a neoepitope resulting from a disease-specific mutation. Preferably, all copies of the essential gene have a disease-specific mutation, i.e., the conjugation rate is 1.

[0060] In one embodiment, a method for inducing an immune response against a target cell population or target tissue expressing one or more neoepitopes in a mammal includes: (i) determining whether a combination of at least two genes, each having a disease-specific mutation resulting in a neoepitope, is a synthetic lethal or synthetic pathogenic gene in an affected cell or cell population; and (ii) (a) administering one or more immune cells expressing one or more antigen receptors targeted to one or more neoepitopes resulting from disease-specific mutations in at least two genes; (b) administering a nucleic acid encoding one or more neoepitopes resulting from disease-specific mutations in at least two genes; or (c) administering a peptide or polypeptide containing one or more neoepitopes resulting from disease-specific mutations in at least two genes. Preferably, all neoepitopes resulting from disease-specific mutations in at least two genes are targeted by the administered immune cells, encoded by the administered nucleic acid, or contained within the administered peptide or polypeptide.

[0061] Furthermore, an immune response can be induced in mammals having a disease, disorder, or condition associated with the expression of a neoepitope resulting from a disease-specific mutation, so that the disease, disorder, or condition can be treated or prevented. Preferably, the disease, disorder, or condition is cancer.

[0062] Preferably, the immune cells are T cells, the antigen receptor is a T cell receptor, and the immune response is a T cell-mediated immune response. More preferably, the immune response is an antitumor immune response, and the target cell population or target tissue expressing one or more suitable neoepitopes is tumor cells or tumor tissue.

[0063] Other features and advantages of the present invention will become apparent from the following detailed description and claims. [Modes for carrying out the invention]

[0064] The present invention is described in detail below, but please understand that the methods, protocols, and reagents may vary, and therefore the present invention is not limited to the specific methods, protocols, and reagents described herein. Furthermore, please understand that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the present invention, which is limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art.

[0065] The elements of the present invention are described below. These elements are described using specific embodiments, but it should be understood that additional embodiments can be created by combining them in any manner and in any number. The various preferred embodiments described should not be construed as limiting the present invention to only the embodiments expressedly described. This description should be understood as supporting and encompassing embodiments that combine the expressly described embodiments with any number of disclosed and / or preferred elements. Furthermore, unless otherwise indicated by the content, any permutations and combinations of all described elements in this application should be considered disclosed by the description of this application. For example, in a preferred embodiment, the neoepitope has a high conjugation state rather than a high conjugation rate, and in a preferred embodiment, the neoepitope is the result of a mutation in an essential gene, in a preferred embodiment, a suitable neoepitope has a high conjugation rate and is the result of a mutation in an essential gene, and in a more preferred embodiment, the conjugation rate is equal to 1.

[0066] Preferably, terms used herein are defined as those found in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)," edited by HGW Leuenberger, B. Nagel, and H. Kolbl, (1995) Helvetica Chimica Acta, CH-4010, Switzerland, Basel.

[0067] Unless otherwise specified, the implementation of this invention will utilize conventional methods of biochemistry, cell biology, immunology, and recombinant DNA techniques described in the literature of the art (see, for example, Molecular Cloning: A Laboratory Manual, 2nd edition, edited by J. Sambrook et al., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989).

[0068] Unless otherwise required by context, throughout this specification and the subsequent claims, the word “comprise,” and variations such as “comprises” and “comprising,” imply the inclusion of the described member, integer, or step or group of members, integers, or steps, but not any other member, integer, or step or group of members, integers, or steps, except in some embodiments, such other members, integers, or steps or groups of members, integers, or steps may be excluded, i.e., the subject matter should be understood to consist of the inclusion of the described member, integer, or step or group of members, integers, or steps. The terms “a,” “an,” and “the,” and similar references used in connection with the description of the invention (particularly in connection with the claims), should be interpreted as covering both singular and plural forms unless otherwise specified herein or unless it is clearly inconsistent by context. Enumerations of value ranges in this specification are intended solely as a way of abbreviating each separate value within the range individually. Unless otherwise specified herein, each individual value is incorporated herein as if it were individually listed herein.

[0069] All methods described herein may be performed in any suitable order, unless otherwise specified herein or if there is a clear contradiction in the content. For example, determining whether a neoepitope is a suitable disease-specific target by determining the copy number of the coding gene may be determined before, after, or simultaneously with determining whether the gene is a driver gene or an essential gene, or before, after, or simultaneously with determining whether the neoepitope is expressed on the surface of cells or induces a favorable immune response suitable for use in a vaccine.

[0070] The use of any and all examples or illustrative words (e.g., “such as”) provided herein is intended solely to better illustrate the invention and does not imply any limitation to the scope of the invention as otherwise stated in the claims. No words herein should be construed as indicating any non-claim elements essential to the practice of the invention.

[0071] Several documents are referenced throughout this specification. Each of the documents referenced herein (including all patents, patent applications, scientific publications, manufacturer specifications, instructions, etc.) is incorporated herein by reference in its entirety, whether above or below. Nothing in this specification should be construed as an acknowledgment that the present invention has no prior rights to such disclosures based on prior art.

[0072] The present invention envisions treatments for diseases, including immunotherapy and radiotherapy, in certain cancers by targeting neoepitopes ("suitable neoepitopes") that are expressed only inside or on the surface of affected cells and are expressed from genes that are unlikely to be silenced by affected cells, thus reducing the likelihood that affected cells can evade immune surveillance by targeted neoepitopes. Immunotherapy may be delivered by active and / or passive immunotherapy methods. For example, in one embodiment, cells can be targeted and killed by using an antibody or other molecule conjugated with a toxic agent that can specifically target neoepitopes and kill cells expressing neoepitopes, in accordance with the present invention.

[0073] The present invention particularly focuses on the identification of such suitable neoepitopes as disease-specific targets in immunotherapy. Once a suitable neoepitope is identified, it can be used in a vaccine to induce an immune response against the neoepitope by inducing and / or activating suitable effector cells, such as T cells that recognize the identified suitable neoepitope, particularly when presented in the context of MHC, via a suitable antigen receptor such as a T cell receptor or an artificial T cell receptor, which results in the death of affected cells expressing the suitable neoepitope. Alternatively, or even better, immune cells that recognize the identified suitable neoepitope can be administered via a suitable antigen receptor, which will also result in the death of cells expressing the suitable neoepitope.

[0074] An immunotherapy approach according to the present invention includes ii) immunization with a peptide or polypeptide containing a suitable neoepitope, iii) nucleic acids encoding a peptide or polypeptide containing a suitable neoepitope, iv) recombinant cells encoding a peptide or polypeptide containing a suitable neoepitope, iv) recombinant viruses encoding a peptide or polypeptide containing a neoepitope, and v) antigen-presenting cells pulsed with a peptide or polypeptide containing a neoepitope, or transfected with nucleic acids encoding a peptide or polypeptide. Another immunotherapy approach according to the present invention includes vi) a T cell receptor that recognizes a neoepitope, and vii) the transfer of effector cells (such as T cells) encoding a receptor that recognizes a neoepitope, in particular, when presented in the context of MHC.

[0075] In the context of this invention, the term "disease-specific mutation" refers to a somatic mutation present in the nucleic acid of affected cells but not in the nucleic acid of corresponding normal, unaffected cells. Since the disease can be cancer, the term "tumor-specific mutation" or "cancer-specific mutation" refers to a somatic mutation present in the nucleic acid of tumor or cancer cells but not in the nucleic acid of corresponding normal, i.e., non-tumor or non-cancerous cells. The terms "tumor-specific mutation" and "tumor mutation," and the terms "cancer-specific mutation" and "cancer mutation" are used interchangeably herein.

[0076] As used herein, a single nucleotide polymorphism (SNP) is a site in the normal genome in which at least one of two alleles (maternal or paternal) is different from that in the normal genome or has a different identity, for example, with respect to a reference genome.

[0077] As used herein, a heterozygous single nucleotide polymorphism (heterozygous SNP) is defined as a site in a normal genome where two alleles (maternal and paternal alleles) have different identities.

[0078] As used herein, the term “conjugation rate” refers to the proportion of the number of copies of a gene that have a disease-specific mutation, taking into account the total number of copies of the gene, regardless of whether the gene has a mutation or not. For example, if there are a total of 20 copies of a gene and 10 of the copies have a disease-specific mutation, the conjugation rate is 0.5. If all copies of the gene have a disease-specific mutation, the conjugation rate is 1. The conjugation rate may be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or at least 0.95. In the context of the present invention, a higher conjugation rate is preferred over a lower one. In some embodiments, the conjugation rate of a mutated allele encoding an epitope, preferably a neoepitope, is, for example, the ratio of the number of copies of the mutated allele to the total number of copies of the nucleotide site to which the mutated allele is mapped in the reference genome.

[0079] As used herein, the term “clonal proportion” refers to the proportion of affected cells that contain the same disease-specific mutation in the same gene, as well as their genetic characteristics such as copy number and junctionality, considering the total number of affected cells, regardless of whether the affected cells have the same mutation in the same gene. The term can also be applied to tumor tissue, in that the clonal proportion is the proportion of affected cells in tumor tissue that contain the same disease-specific mutation in the same gene, as well as their genetic characteristics such as copy number, considering the total number of cells in tumor tissue. For example, in a sample obtained from a tumor in which only half of the total number of tumor cells have the same mutation in the same gene, the clonal proportion is 0.5. The clonal proportion may be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or at least 0.95. In the context of the present invention, a higher clonal proportion is preferred over a lower one. A clonal proportion of 1 is most preferred, where all affected cells have the same disease-specific mutation in the same gene. The terms "clonal proportion," "proportional clonality," and "proportional subclonality" are used interchangeably in this specification.

[0080] The term "localized amplification" refers to the amplification or copy number increase of a portion of the genome, for example, amplification of one or more genes located together on the same chromosome, which, if no deletion event occurs for the same portion of the genome, results in a copy number greater than 2, preferably more than 5, 10, 15, 20, 25, 50, 75, or 100, with respect to that portion of the genome. Therefore, for the purposes of this invention, a gene that is locally amplified in an affected cell is a gene that may have an increased copy number compared to 2 wild-type copies or, in the case of such genes on the X and Y chromosomes in males, 1 wild-type copy. Localized amplification is distinct from whole-genome duplication and / or amplification events.

[0081] The term "immune response" refers to an integrated physical response to an antigen, preferably a cellular immune response or a cellular and humoral immune response. Immune responses can be protective, preventive / prophylactic, and / or therapeutic.

[0082] "Induction of an immune response" can mean that no immune response to a particular antigen existed before induction, but it can also mean that a certain level of immune response to a particular antigen existed before induction, and that the immune response was enhanced after induction. Therefore, "inducing an immune response" also includes "enhancing an immune response." Preferably, after inducing an immune response in a subject, the subject is protected from developing a disease such as cancer, or the disease is relieved by inducing an immune response. For example, an immune response to a tumor-expressing antigen can be induced in a patient with cancer or a subject at risk of developing cancer. In this case, inducing an immune response may mean that the subject's disease is relieved, the subject does not develop metastases, or the subject at risk of developing cancer does not develop cancer.

[0083] The terms "cellular immune response", "cellular response", "cellular response to an antigen", or similar terms are meant to include a cellular response to cells characterized by the presentation of an antigen having class I or class II MHC. The cellular response relates to cells called T cells or T lymphocytes that act as either "helper" or "killer". Helper T cells (also called CD4 + T cells) play a central role by regulating the immune response, and killer cells (also called cytotoxic T cells, cytolytic T cells, CD8 + T cells or CTLs) kill diseased cells such as cancer cells and prevent the production of further diseased cells. Preferably, the anti-tumor CTL response is stimulated against tumor cells that express one or more tumor-expressed antigens, preferably tumor-expressed antigens having class I MHC.

[0084] The "antigen" according to the present invention covers any substance that is a target of and / or induces an immune response, such as a specific reaction with an antibody or a T lymphocyte (T cell), preferably a peptide or a protein. Preferably, the antigen contains at least one epitope, such as a T cell epitope. Preferably, this antigen is a molecule that optionally induces an immune reaction that is specific for an antigen (including cells expressing the antigen) after processing, preferably in the context of an MHC molecule, which results in an immune response against the antigen (including cells expressing the antigen), preferably cells, preferably presented by antigen-presenting cells, including diseased cells, particularly cancer cells.

[0085] Preferably, in the context of the present invention, the antigen is a molecule that, after processing, optionally induces an immune response preferably specific to the antigen. According to the present invention, any suitable antigen that is a candidate for the immune response can be used, and the immune response may be both humoral and cellular. In the context of the present invention, the antigen is preferably presented by cells, preferably antigen-presenting cells, in the context of MHC molecules, which results in an immune response to the antigen. The antigen is preferably a product corresponding to or derived from a naturally occurring antigen. Such naturally occurring antigens may include or be derived from allergens, viruses, bacteria, fungi, parasites, and other infectious pathogens and pathogens, or the antigen may be a tumor antigen. According to the present invention, the antigen may correspond to a product of natural origin, for example, a viral protein or a portion thereof. In a preferred embodiment, the antigen is a surface polypeptide, i.e., a polypeptide that is naturally displayed on the surface of cells, pathogens, bacteria, viruses, fungi, parasites, allergens, or tumors. Antigens can trigger an immune response against cells, pathogens, bacteria, viruses, fungi, parasites, allergens, or tumors.

[0086] The terms “disease-associated antigen” or “disease-specific antigen” are used in their broadest sense to refer to any antigen that is associated with or specific to a disease. Such antigens are molecules containing epitopes that will stimulate the host’s immune system to produce a cellular antigen-specific immune response and / or humoral antibody response to the disease. Therefore, disease-associated antigens can be used for therapeutic purposes. Disease-associated antigens are preferably associated with microbial infections, typically microbial antigens, or cancer, typically tumors.

[0087] The term "pathogen" refers to pathogenic biological material that can cause disease in living organisms, preferably vertebrates. Pathogens include viruses, along with microorganisms such as bacteria, single-celled eukaryotes (protists), and fungi.

[0088] In the context of the present invention, the terms “tumor antigen” or “tumor-associated antigen” refer to proteins that are specifically expressed in a limited number of tissues and / or organs or at specific developmental stages under normal conditions. For example, tumor antigens may be specifically expressed in gastric tissue, preferably gastric mucosa, genital organs, e.g., testes, trophoblast tissue, e.g., placenta or germline cells, under normal conditions, and may be expressed or abnormally expressed in one or more tumor or cancerous tissues. In this context, “limited number” means preferably three or fewer, more preferably two or fewer. In relation to the present invention, tumor antigens include, for example, differentiation antigens, preferably cell-type-specific differentiation antigens, i.e., proteins specifically expressed in specific cell types at specific differentiation stages under normal conditions, cancer / testicular antigens, i.e., proteins specifically expressed in the testes and possibly the placenta under normal conditions, and germline-specific antigens. In the context of the present invention, tumor antigens are preferably associated with the cell surface of cancer cells and preferably not expressed or expressed very rarely in normal tissues. Preferably, the expression of tumor antigens or abnormal tumor antigens identifies cancer cells. In relation to the present invention, the tumor antigen expressed by cancer cells in a subject, for example, a patient suffering from cancer, is preferably an autoprotein in the subject. In a preferred embodiment, the tumor antigen is specifically expressed in tissues or organs that are not essential under normal conditions, i.e., tissues or organs that would not result in the death of the subject if damaged by the immune system, or in organs or structures of the body that are inaccessible or only partially accessible by the immune system, in relation to the present invention. Preferably, the amino acid sequence of the tumor antigen is identical between the tumor antigen expressed in normal tissue and the tumor antigen expressed in cancerous tissue.

[0089] According to the present invention, the terms “tumor antigen,” “tumor-expressed antigen,” “cancer antigen,” and “antigen expressed in cancer” are equivalent and are used interchangeably herein.

[0090] The terms “epitope,” “antigen peptide,” “antigen epitope,” “immunogenic peptide,” and “MHC-binding peptide” are used interchangeably herein and refer to an antigenic determinant in a molecule such as an antigen, i.e., a portion or fragment of an immunoactive compound that is recognized by the immune system, for example, particularly by T cells when presented in the context of an MHC molecule. A protein epitope preferably comprises a continuous or discontinuous portion of the protein and is preferably 5 to 100, preferably 5 to 50, more preferably 8 to 30, and most preferably 10 to 25 amino acids long. For example, an epitope may preferably be 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids long. According to the present invention, an epitope may be an “MHC-binding peptide” or “antigen peptide” because it can bind to an MHC molecule, such as an MHC molecule on the surface of a cell. The term “Major Histocompatibility Complex” and the abbreviation “MHC” refer to a complex of genes present in all vertebrates, comprising MHC class I and MHC class II molecules. MHC proteins or molecules are important for signaling between lymphocytes and antigen-presenting or infected cells in immune responses, and MHC proteins or molecules bind to peptides and present them for recognition by T cell receptors. Proteins encoded by MHC are expressed on the surface of cells and display both self-antigens (peptide fragments derived from the cell itself) and non-self-antigens (e.g., fragments of invading microorganisms) to T cells. A preferred immunogenic moiety of such a moiety binds to MHC class I or class II molecules. As used herein, an immunogenic moiety is said to “bind” to an MHC class I or class II molecule if the binding is detectable using any assay known in the art. The term “MHC-binding peptide” refers to a peptide that binds to MHC class I and / or MHC class II molecules. In the case of class I MHC / peptide complexes, the bound peptide is typically 8-10 amino acids long, but longer or shorter peptides may also be effective.In the case of class II MHC / peptide complexes, the bound peptide is typically 10–25 amino acids long, especially 13–18 amino acids long, although longer and shorter peptides may also be effective.

[0091] As used herein, the term “neoepitope” refers to an epitope that is not present in reference to normal non-cancerous or germline cells, etc., but is found in affected cells, such as cancer cells. This includes situations in which a corresponding epitope is found in normal non-cancerous or germline cells, but one or more mutations in the cancer cell alter the sequence of the epitope, resulting in the creation of a neoepitope. Furthermore, the neoepitope may be specific not only to affected cells but also to patients with the disease. Since the neoepitopes and suitable neoepitopes identified by the methods of the present invention are subsets of epitopes, the disclosures herein regarding epitopes as immunological targets in general apply equally to neoepitopes and suitable neoepitopes.

[0092] In one particularly preferred embodiment of the present invention, the epitope or neoepitope is a T cell epitope. As used herein, the term “T cell epitope” means a peptide that binds to an MHC molecule in a configuration recognized by a T cell receptor. Typically, T cell epitopes are presented on the surface of antigen-presenting cells.

[0093] As used herein, the term “prediction of immunogenic amino acid modification” means a prediction of whether a peptide containing such amino acid modification will be immunogenic in vaccination and, therefore, useful as an epitope, particularly a T-cell epitope.

[0094] According to the present invention, T cell epitopes can be present in a vaccine as part of a larger entity, such as a vaccine sequence and / or polypeptide containing two or more T cell epitopes. The presented peptide or T cell epitope is produced after appropriate processing.

[0095] T cell epitopes can be modified at one or more residues that are not essential for TCR recognition or MHC binding. Such modified T cell epitopes can be considered immunologically homogeneous.

[0096] Preferably, when a T cell epitope is presented by MHC and recognized by a T cell receptor, it can induce clonal growth of T cells possessing a T cell receptor that specifically recognizes the peptide / MHC complex, in the presence of appropriate co-stimulatory signals.

[0097] Preferably, the T cell epitope comprises an amino acid sequence substantially corresponding to the amino acid sequence of the antigen fragment. Preferably, the antigen fragment is an MHC class I and / or class II presenting peptide.

[0098] The T cell epitope according to the present invention preferably relates to a portion or fragment of an antigen that can stimulate an immune response, preferably a cellular response, against cells characterized by the expression of an antigen, preferably by the presentation of an antigen, such as an antigen or disease cells, particularly cancer cells. Preferably, the T cell epitope can stimulate a cellular response against cells characterized by the presentation of an antigen having class I MHC, and preferably can stimulate antigen-responsive cytotoxic T lymphocytes (CTLs).

[0099] In some embodiments, the antigen is an autoantigen, in particular a tumor antigen. Tumor antigens and their determination are known to those skilled in the art.

[0100] The term “immunogenicity” preferably relates to the relative effectiveness of inducing an immune response associated with a therapeutic treatment, such as treatment for cancer. As used herein, the term “immunogenic (is)” relates to the characteristic of having immunogenicity. For example, the term “immunogenic modification,” when used in the context of a peptide, polypeptide, or protein, relates to the effectiveness of the peptide, polypeptide, or protein in inducing an immune response resulting from and / or directed toward such modification. Preferably, an unmodified peptide, polypeptide, or protein does not induce an immune response, induces a different immune response, or induces a different level, preferably a lower level, of immune response.

[0101] According to the present invention, the terms “immunogenicity” or “being immunogenic” preferably relate to a biologically relevant immune response, particularly relative effectiveness in inducing an immune response useful for vaccination. Thus, in a preferred embodiment, an amino acid modification or modified peptide is immunogenic if it induces an immune response to a targeted modification in a subject, and this immune response may be beneficial for therapeutic or preventive purposes.

[0102] "Antigen processing" or "processing" means the degradation of a polypeptide or antigen into processing products of the polypeptide or antigen fragments (e.g., degradation of a polypeptide into a peptide), and the association (e.g., by binding) of one or more of these fragments with MHC molecules by a cell, preferably an antigen-presenting cell, for presentation to specific T cells.

[0103] Antigen-presenting cells (APCs) are cells that present peptide fragments of protein antigens associated with MHC molecules on their cell surface. Some APCs can activate antigen-specific T cells.

[0104] Professional antigen-presenting cells are highly efficient at internalizing antigens either through phagocytosis or receptor-mediated endocytosis, and then displaying antigen fragments bound to class II MHC molecules on their membranes. T cells recognize and interact with the antigen-class II MHC molecule complex on the membrane of the antigen-presenting cell. Subsequently, further costimulatory signals are generated by the antigen-presenting cell, leading to T cell activation. The expression of costimulatory molecules is a defining characteristic of professional antigen-presenting cells.

[0105] The main types of professional antigen-presenting cells are dendritic cells, macrophages, B-cells, and certain active epithelial cells, which have the broadest range of antigen presentation and are perhaps the most important antigen-presenting cells.

[0106] Dendritic cells (DCs) are a population of leukocytes that present antigens captured in peripheral tissues to T cells via both MHC class II and I antigen presentation pathways. It is well known that dendritic cells are potent inducers of the immune response, and their activation is a crucial step in inducing antitumor immunity. Dendritic cells are conveniently classified into “immature” and “mature” cells, which can be used as a simple way to distinguish between two well-characterized phenotypes. However, this nomenclature should not be interpreted as excluding all possible intermediate stages of differentiation. Immature dendritic cells are characterized as antigen-presenting cells with a high capacity for antigen uptake and processing, which correlates with high expression of Fcγ receptors and mannose receptors. The mature phenotype is typically characterized by lower expression of these markers but higher expression of cell surface molecules that govern T cell activation, such as class I and class II MHCs, adhesion molecules (e.g., CD54 and CD11), and costimulatory molecules (e.g., CD40, CD80, CD86, and 4-1BB).

[0107] Dendritic cell maturation refers to a state of dendritic cell activation in which such antigen-presenting dendritic cells provide initial stimulation to T cells, while presentation by immature dendritic cells leads to tolerance. Dendritic cell maturation is primarily triggered by biomolecules with microbial characteristics detected by innate receptors (bacterial DNA, viral RNA, endotoxins, etc.), pro-inflammatory cytokines (TNF, IL-1, IFN), ligation of CD40 on the dendritic cell surface by CD40L, and substances released from cells experiencing stress-induced cell death. Dendritic cells can be induced in vitro by culturing bone marrow cells with cytokines such as granulocyte-macrophage colony-stimulating factor (GM-CSF) and tumor necrosis factor alpha.

[0108] Nonprofessional antigen-presenting cells do not constitutively express MHC class II proteins necessary for interaction with naive T cells. These proteins are expressed only when nonprofessional antigen-presenting cells are stimulated by specific cytokines such as IFNγ.

[0109] "Antigen-presenting cells" can be loaded with MHC class I-presenting peptides by transduction of a nucleic acid, preferably RNA, that encodes a peptide or polypeptide containing the peptide to be presented, for example, a nucleic acid encoding an antigen.

[0110] In some embodiments, a pharmaceutical composition or vaccine of the present invention, comprising a gene delivery medium targeting dendritic or other antigen-presenting cells, can be administered to a patient to induce in vivo transfection. For example, in vivo transfection of dendritic cells can be carried out using any method known in the art, such as the gene gun method described in WO 97 / 24447 or by Mahvi et al., Immunology and Cell Biology 75: pp. 456-460, 1997.

[0111] The term "antigen-presenting cell" also includes target cells.

[0112] "Target cells" means cells that are the target of an immune response, such as a cellular immune response. Target cells include cells that present an antigen or antigen epitope, i.e., a peptide fragment derived from an antigen, and include all undesirable cells, such as cancer cells. In a preferred embodiment, target cells are cells that express the antigen described herein and preferably present the antigen using class I MHC.

[0113] The term "part" refers to a fraction. With respect to an amino acid sequence or a specific structure such as a protein, the term "part" may specify a continuous or discontinuous fraction of the structure. Preferably, a part of an amino acid sequence contains at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, preferably at least 40%, preferably at least 50%, more preferably at least 60%, more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90% of the amino acids of the amino acid sequence. Preferably, if the part is a discontinuous fraction, the discontinuous fraction consists of two, three, four, five, six, seven, eight, or more parts of the structure, each part being a continuous element of the structure. For example, a discontinuous fraction of the amino acid sequence may consist of 2, 3, 4, 5, 6, 7, 8, or more, preferably 4 or fewer, portions of the amino acid sequence, each portion preferably containing at least 5 consecutive amino acids, at least 10 consecutive amino acids, preferably at least 20 consecutive amino acids, and preferably at least 30 consecutive amino acids.

[0114] The terms “part” and “fragment” are used interchangeably herein and refer to a continuous element. For example, a part of an amino acid sequence or a structure such as a protein means a continuous element of the said structure. A part, part, or fragment of a structure preferably contains one or more functional properties of the said structure. For example, a part, part, or fragment of an epitope, peptide, or protein preferably is immunologically equivalent to the epitope, peptide, or protein from which it originates. In connection with the present invention, a “part” of a structure such as an amino acid sequence preferably comprises, and preferably consists of, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98%, or at least 99% of the entire structure or amino acid sequence.

[0115] In relation to the present invention, the term "immunoreactive cell" refers to a cell that exerts effector function during an immune response. "Immunereactive cells" are preferably cells characterized by the presentation of an antigen, or antigen peptides derived from an antigen, that can bind to and mediate an immune response. For example, such cells secrete cytokines and / or chemokines, secrete antibodies, recognize cancer cells, and optionally eliminate such cells. For example, immunoreactive cells include T cells (cytotoxic T cells, helper T cells, tumor-infiltrating T cells), B cells, natural killer cells, neutrophils, macrophages, and dendritic cells. Preferably, in relation to the present invention, "immunoreactive cell" refers to a T cell, preferably CD4 + and / or CD8 + These are T cells.

[0116] Preferably, "immunoreactive cells" recognize antigens or antigen-derived antigen peptides with a certain degree of specificity, particularly when presented on the surface of antigen-presenting cells or disease cells such as cancer cells, in the context of MHC molecules. Preferably, such recognition enables the cells that recognize the antigen or antigen-derived antigen peptides to become responsive or reactive. The cells are helper T cells (CD4) that possess receptors that recognize antigens or antigen-derived antigen peptides in the context of MHC class II molecules. + If it is a T cell, such responsiveness or reactivity is due to the release of cytokines and / or CD8 + This may include activation of lymphocytes (CTLs) and / or B cells. If the cell is a CTL, such responsiveness or reactivity may include, for example, the elimination of cells presented in the context of MHC class I molecules, i.e., cells characterized by antigen presentation using class I MHCs, by apoptosis or perforin-mediated cytolysis. CTL responsiveness may include sustained calcium flow, cell division, production of cytokines such as IFN-γ and TNF-α, upregulation of activation markers such as CD44 and CD69, and specific cytolytic death of antigen-expressing target cells. CTL responsiveness may also be determined using an artificial reporter that accurately exhibits CTL responsiveness. CTLs that recognize an antigen or antigen-derived antigen peptide and are responsive or reactive are also referred to herein as “antigen-responsive CTLs.” If the cell is a B cell, such responsiveness may include the release of immunoglobulins.

[0117] The terms “T cell” and “T lymphocyte” are used interchangeably herein and include T helper cells (CD4+ T cells) and cytotoxic T cells (CTLs, CD8+ T cells), including cytolytic T cells.

[0118] T cells belong to a group of white blood cells known as lymphocytes and play a central role in cell-mediated immunity. They can be distinguished from other lymphocyte types, such as B cells and natural killer cells, by the presence of a special receptor called the T cell receptor (TCR) on their cell surface. The thymus is the main organ responsible for the maturation of T cells. Several different subgroups of T cells have been discovered, each with distinctly different functions.

[0119] T helper cells assist other leukocytes in immune processes, including the maturation of B cells into plasma cells and the activation of cytotoxic T cells and macrophages, among other functions. These cells are also known as CD4+ T cells because they express the CD4 protein on their surface. Helper T cells are activated when peptide antigens are presented by MHC class II molecules expressed on the surface of antigen-presenting cells (APCs). After activation, they rapidly divide and secrete small proteins called cytokines that modulate or assist the active immune response.

[0120] Cytotoxic T cells destroy virus-infected cells and tumor cells and are also associated with graft rejection. These cells express the CD8 glycoprotein on their surface and are therefore also known as CD8+ T cells. These cells recognize their targets by binding to antigens associated with MHC class I, which are present on the surface of almost all cells in the body.

[0121] In most T cells, the T cell receptor (TCR) exists as a complex of several proteins. The actual T cell receptor consists of two separate peptide chains produced from independent T cell receptor alpha and beta (TCRα and TCRβ) genes, and are called α- and β-TCR chains. Gamma delta T cells (γδT cells) represent a small subgroup of T cells that possess a distinctly different T cell receptor (TCR) on their surface. However, in γδT cells, the TCR consists of one γ-chain and one δ-chain. This group of T cells is far rarer than αβT cells (2% of all T cells).

[0122] According to the present invention, the term "antigen receptor" includes naturally occurring receptors such as T cell receptors and engineered receptors that confer any specificity, such as the specificity of monoclonal antibodies, to immune effector cells such as T cells. In this way, a large number of antigen-specific T cells can be generated for adoptive cell transfer. Thus, antigen receptors according to the present invention may be present on T cells, for example, in place of or in addition to the T cell receptor of the T cell itself. Such T cells do not necessarily require the processing and presentation of antigens for recognition of target cells, but rather can recognize any antigen present on target cells, preferably by specificity. Preferably, the antigen receptor is expressed on the surface of the cell. For the purposes of the present invention, T cells containing antigen receptors are included by the term "T cell" as used herein. In particular, according to the present invention, the term "antigen receptor" includes an artificial receptor comprising a single molecule or a complex of molecules that recognizes a target structure (e.g., an antigen) on target cells such as cancer cells, that is, binds to it (for example, by binding an antigen-binding site or antigen-binding domain to an antigen expressed on the surface of the target cell), and can confer specificity to immune effector cells such as T cells that express the antigen receptor on the cell surface. Preferably, the recognition of the target structure by the antigen receptor leads to the activation of immune effector cells that express the antigen receptor. The antigen receptor may comprise one or more protein units, and the protein units comprise one or more domains as described herein. According to the present invention, the "antigen receptor" may also be a "chimeric antigen receptor (CAR)", a "chimeric T cell receptor", or an "artificial T cell receptor".

[0123] Antigens can be recognized by antigen receptors via antigen-recognition domains (hereinafter also simply referred to as "domains") that can form antigen-binding sites, such as via the antigen-binding portion of antibodies and T cell receptors, which may reside on the same or different peptide chains. In one embodiment, the two domains forming the antigen-binding site are derived from immunoglobulin. In one embodiment, the two domains forming the antigen-binding site are derived from T cell receptors. Antibody-variable domains, such as monoclonal antibody-derived single-chain variable fragments (scFv) and T cell receptor-variable domains, particularly TCR alpha and beta single-chains, are especially preferred. In fact, almost anything that binds to a given target with high affinity can be used as an antigen-recognition domain.

[0124] The primary signal in T cell activation is brought about by the binding of T cell receptors to short peptides presented by major histocompatibility complexes (MHCs) on other cells. This ensures that only T cells with a TCR specific to the peptide are activated. The partner cells are typically professional antigen-presenting cells (APCs), usually dendritic cells in the case of a naive response, although B cells and macrophages can also be important APCs. Peptides presented to CD8+ T cells by MHC class I molecules are typically 8-10 amino acids long; peptides presented to CD4+ T cells by MHC class II molecules are typically longer because the ends of the binding gaps of MHC class II molecules are open.

[0125] According to the present invention, the molecule has a significant affinity for the target in a standard assay and can bind to the predetermined target when it binds to the predetermined target. "Affinity" or "binding affinity" is often expressed as the equilibrium dissociation constant (K). D ) is measured by the standard assay. If a molecule does not have significant affinity for the target in the standard assay and does not bind significantly to the target, it is considered that the molecule cannot (substantially) bind to the target.

[0126] Cytotoxic T lymphocytes can be produced in vivo by introducing an antigen or antigen peptide into antigen-presenting cells in vivo. The antigen or antigen peptide can be represented as a protein, as DNA (e.g., in a vector), or as RNA. While the antigen may be processed to produce peptide partners for MHC molecules, its fragments can be presented without further processing. The latter is particularly true if they can bind to MHC molecules. Generally, administration to patients is possible via intradermal injection. However, intranodular injection into lymph nodes is also possible (Maloy et al., 2001, Proc Natl Acad Sci USA 98:3299-303). The resulting cells present the target complex, are recognized by autologous cytotoxic T lymphocytes, and subsequently proliferate.

[0127] Specific activation of CD4+ or CD8+ T cells can be detected by various methods. Methods for detecting specific T cell activation include detecting T cell proliferation, cytokine (e.g., lymphokine) production, or the occurrence of cytolytic activity. In CD4+ T cells, a preferred method for detecting specific T cell activation is the detection of T cell proliferation. In CD8+ T cells, a preferred method for detecting specific T cell activation is the detection of the occurrence of cytolytic activity.

[0128] The terms "cells characterized by antigen presentation" or "cells that present antigens," or similar expressions, in the context of MHC molecules, particularly MHC class I molecules, refer to disease cells, such as cancer cells or antigen-presenting cells, that present the antigen they express or a fragment derived from said antigen, for example, by antigen processing. Similarly, the term "diseases characterized by antigen presentation" refers to diseases that include cells characterized by antigen presentation, particularly using class I MHC. Antigen presentation by cells can be achieved by transfusing nucleic acids, such as RNA, that encode the antigen into the cell.

[0129] "Presented antigen fragment" or similar expression means a fragment that can be presented by MHC class I or class II, preferably MHC class I, when added directly to antigen-presenting cells, for example. In one embodiment, the fragment is a fragment naturally presented by cells expressing the antigen.

[0130] The term "immunologically homogeneous" means that immunologically homogeneous molecules, such as immunologically homogeneous amino acid sequences, exhibit the same or essentially the same immunological properties and / or exert the same or essentially the same immunological effects, for example, in terms of the type of immunological effect, such as induction of humoral and / or cellular immune responses, the intensity and / or duration of the induced immune response, or the specificity of the induced immune response. In the context of the present invention, the term "immunologically homogeneous" is preferably used in relation to the immunological effects or properties of peptides used for immunization. For example, if an amino acid sequence induces an immune response that has specificity to react with a reference amino acid sequence when exposed to the target immune system, then the amino acid sequence is immunologically homogeneous with the reference amino acid sequence.

[0131] In the context of the present invention, the term "immune effector function" includes any function mediated by components of the immune system that results in inhibition of tumor growth and / or tumor development, including, for example, the death of tumor cells or the inhibition of intratumoral and metastatic metastasis. Preferably, the immune effector function in the context of the present invention is an effector function mediated by T cells. Such a function is mediated by helper T cells (CD4 + In the case of T cells, recognition of antigens or antigen-derived antigen peptides by T cell receptors in the context of MHC class II molecules, cytokine release, and / or CD8 +This includes activation of lymphocytes (CTLs) and / or B cells, in the case of CTLs, recognition of antigens or antigen-derived antigen peptides by T cell receptors in the context of MHC class I molecules, elimination of cells presented in the context of MHC class I molecules, i.e., cells characterized by antigen presentation using class I MHCs, by apoptosis or perforin-mediated cytolysis, production of cytokines such as IFN-γ and TNF-α, and specific cytolytic death of antigen-expressing target cells.

[0132] The term "major histocompatibility complex" and its abbreviation "MHC" refer to a complex of genes that occur in all vertebrates and include MHC class I and MHC class II molecules. MHC proteins or molecules are important for signaling between lymphocytes and antigen-presenting or infected cells in immune responses. MHC proteins or molecules bind to peptides and present them for recognition by T cell receptors. Proteins encoded by MHC are expressed on the cell surface and display both self-antigens (peptide fragments derived from the cell itself) and non-self-antigens (e.g., fragments of invading microorganisms) to T cells.

[0133] The MHC region is divided into three subgroups: Class I, Class II, and Class III. MHC Class I proteins contain the α chain and β2 microglobulin (not the MHC portion encoded by chromosome 15). These present antigen fragments to cytotoxic T cells. In most immune system cells, particularly antigen-presenting cells, MHC Class II proteins contain the α and β chains, which present antigen fragments to T helper cells. The MHC Class III region encodes other immune components, such as complement components, some of which encode cytokines.

[0134] MHC is both polygenic (there are several MHC class I and MHC class II genes) and polymorphic (there are multiple alleles for each gene).

[0135] As used herein, the term “haplotype” refers to an HLA allele found on a single chromosome and the protein it encodes. A haplotype can also refer to an allele located at any single locus within the MHC. Each class of MHC is represented by several loci: for example, Class I includes HLA-A (human leukocyte antigen-A), HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-P, and HLA-V, while Class II includes HLA-DRA, HLA-DRB1-9, HLA-DQA1, HLA-DQB1, HLA-DPA1, HLA-DPB1, HLA-DMA, HLA-DMB, HLA-DOA, and HLA-DOB. The terms “HLA allele” and “MHC allele” are used interchangeably herein.

[0136] MHC exhibits extreme polymorphism. Within the human population, numerous haplotypes exist at each locus, each containing distinct alleles. Different polymorphic MHC alleles, both Class I and Class II, have different peptide specificities, in that each allele encodes a protein that binds to peptides exhibiting a specific sequence pattern.

[0137] In the context of the present invention, the MHC molecule is preferably an HLA molecule.

[0138] In the context of the present invention, the term "MHC-binding peptide" includes MHC class I and / or class II binding peptides, or peptides that can be processed to produce MHC class I and / or class II binding peptides. In the case of class I MHC / peptide complexes, the binding peptide is typically 8 to 12, preferably 8 to 10 amino acids long, although longer or shorter peptides may also be effective. In the case of class II MHC / peptide complexes, the binding peptide is typically 9 to 30, preferably 10 to 25 amino acids long, particularly 13 to 18 amino acids long, although longer and shorter peptides may also be effective.

[0139] "Antigen peptide" preferably relates to a portion or fragment of an antigen that can stimulate an immune response, preferably a cellular response, against an antigen, or a diseased cell, particularly a cancer cell, or other cell characterized by the expression, preferably the presentation, of the antigen. Preferably, the antigen peptide can stimulate a cellular response against a cell characterized by the presentation of an antigen having class I MHC, and preferably stimulate antigen-responsive cytotoxic T lymphocytes (CTLs). Preferably, the antigen peptide is an MHC class I and / or class II presenting peptide, or can be processed to produce an MHC class I and / or class II presenting peptide. Preferably, the antigen peptide comprises an amino acid sequence substantially corresponding to the amino acid sequence of the antigen fragment. Preferably, the aforementioned fragment of the antigen is an MHC class I and / or class II presenting peptide. Preferably, the antigen peptide comprises an amino acid sequence substantially corresponding to the amino acid sequence of such fragment and is processed to produce an MHC class I and / or class II presenting peptide derived from such fragment, i.e., the antigen.

[0140] When the peptide needs to be presented directly, i.e., without processing, and especially without cleavage, it should have a length suitable for binding to MHC molecules, particularly class I MHC molecules, preferably 7 to 20 amino acids long, more preferably 7 to 12 amino acids long, more preferably 8 to 11 amino acids long, and particularly 9 or 10 amino acids long.

[0141] If the peptide is a portion of a larger entity containing an additional sequence, such as a vaccine sequence or polypeptide, and needs to be presented after processing, particularly after cleavage, the peptide produced by processing has a length suitable for binding to MHC molecules, particularly class I MHC molecules, preferably 7 to 20 amino acids long, more preferably 7 to 12 amino acids long, more preferably 8 to 11 amino acids long, particularly 9 or 10 amino acids long. Preferably, the sequence of the peptide that needs to be presented after processing is derived from the amino acid sequence of the antigen, i.e., its sequence substantially corresponds to a fragment of the antigen, and preferably is exactly identical thereto. Thus, the MHC-binding peptide contains a sequence that substantially corresponds to a fragment of the antigen, and preferably is exactly identical thereto.

[0142] Peptides having an amino acid sequence substantially corresponding to the sequence of a peptide presented by a class I MHC may differ by one or more residues that are not essential for TCR recognition of the peptide presented by the class I MHC, or for the binding of the peptide to the MHC. Such substantially corresponding peptides can also stimulate antigen-responsive CTLs and can be considered immunologically equivalent. Peptides having a different amino acid sequence from the presented peptide, with residues that do not affect TCR recognition but improve the stability of binding to the MHC, can improve the immunogenicity of the antigen peptide and may be referred to herein as "optimized peptides." A reasonable approach for designing substantially corresponding peptides can be used using existing knowledge regarding which of these residues is likely to affect binding to either the MHC or the TCR. The functionally resulting peptide is considered an antigen peptide.

[0143] Antigen peptides, when presented by MHC molecules, should be recognizable by T cell receptors. Preferably, when recognized by T cell receptors, antigen peptides can induce the presence of appropriate co-stimulatory signals and clonal growth of T cells possessing T cell receptors that specifically recognize the antigen peptide. Preferably, when presented in the context of MHC molecules, antigen peptides can stimulate an immune response, preferably a cellular response, against cells characterized by the antigen from which they originate, or by the expression of the antigen, and preferably by the presentation of the antigen. Preferably, antigen peptides can stimulate a cellular response against cells characterized by the presentation of an antigen having class I MHC molecules, and preferably, they can stimulate antigen-responsive CTLs. Such cells are preferably target cells.

[0144] The term "genome" refers to the total amount of genetic information contained in the chromosomes of an organism or cell.

[0145] The term "exome" refers to a portion of an organism's genome formed by exons, which are the coding regions of expressed genes. Exomes provide the genetic blueprint used in the synthesis of proteins and other functional gene products. This is the most functionally relevant part of the genome and therefore most likely to contribute to an organism's phenotype. The human genome's exome is estimated to constitute 1.5% of the entire genome (Ng et al., 2008, PLoS Gen., 4(8):1-15).

[0146] The term “transcriptome” refers to any set of RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNAs, produced in a single cell or a population of cells. In the context of the present invention, transcriptome or RNA sequence means any set of RNA molecules produced in a single cell, a population of cells, preferably a population of cancer cells, or any cell of a given individual at a given point in time.

[0147] "Nucleic acid" is preferably deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), more preferably RNA, most preferably in vitro transcription RNA (IVT RNA) or synthetic RNA. Nucleic acids include genomic DNA, cDNA, mRNA, recombinant production, and chemosynthetic molecules. Nucleic acids can exist as single-stranded or double-stranded linear or covalent cyclic closed molecules. Nucleic acids can be isolated. The term "isolated nucleic acid" means that the nucleic acid has been (i) amplified in vitro, e.g., by polymerase chain reaction (PCR), (ii) recombinantly produced by cloning, (iii) purified, e.g., by separation by cleavage and gel electrophoresis, or (iv) synthesized, e.g., by chemosynthesis. Nucleic acids can be used for introduction into cells, i.e., translocation, in the form of RNA, which can be prepared particularly by in vitro transcription from a DNA template. RNA can be further modified before application by sequence stabilization, capping, and polyadenylation.

[0148] The term "genetic material" refers to an isolated nucleic acid, either DNA or RNA, a segment of a double helix, a segment of a chromosome, or the entire genome of an organism or cell, particularly its exome or transcriptome.

[0149] The term "mutation" refers to a change or difference (nucleotide substitution, addition, or deletion) in the nucleic acid sequence of an affected genome compared to a reference, preferably matched, normal genome. "Somatic mutations" can occur in any cell of the body other than germ cells (sperm and egg cells) and are therefore not transmitted to offspring. These changes may (but not necessarily) cause cancer or other diseases. Preferably, mutations are nonsynonymous mutations. The term "nonsynonymous mutation" refers to a mutation that results in an amino acid change, such as an amino acid substitution in the translation product, preferably a nucleotide substitution, which preferably leads to the formation of a neoepitope.

[0150] The term "single nucleotide variant / variation (SNV)" refers to a difference in nucleic acid sequence at a specific site (allele) when comparing the genome of a diseased cell, such as a tumor cell, with a preferably matched (corresponding) genome of a normal, non-disease cell, or a reference genome. As used herein, the term mutation preferably encompasses SNVs.

[0151] Copy number variation (CNV) events in the affected (tumor) genome are somatic copy number variation events that occur only in affected cells and are defined as changes in the number of maternal and / or paternal alleles in a region of the affected (tumor) genome relative to a matched normal genome. This alteration preferably affects a region of the genome that is approximately 1 kb or larger.

[0152] The term "mutation" includes point mutations, indels, fusions, chromothripsis, and RNA editing.

[0153] The term "indel" describes a specific class of mutations defined as those resulting in coexisting insertions and deletions, as well as a net increase or decrease in nucleotides. In the coding regions of the genome, indels, unless their length is a multiple of 3, result in frameshift mutations. Indels can be contrasted with point mutations. While indels insert and delete nucleotides from a sequence, point mutations are a form of substitution that replaces one of the nucleotides.

[0154] Fusion can produce a hybrid gene formed from two previously separate genes. This can result from translocation, intermediate deletion, or chromosomal inversion. Often, fusion genes are oncogenes. Oncogenic fusion genes can yield gene products with different functions from the new or two fusion partners. Alternatively, a proto-oncogene may be fused with a strong promoter, and thus the oncogenic function begins to function through upregulation caused by the strong promoter of the upstream fusion partner. Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events.

[0155] The term "chromothripsis" refers to a genetic phenomenon in which a specific region of the genome is fragmented by a single disruption event and then subsequently reassembled.

[0156] The term "RNA editing" or "to perform RNA editing" refers to a molecular process in which the information content of an RNA molecule is altered by chemical changes in its base composition. RNA editing includes nucleoside modification, such as cytidine (C) to uridine (U) and adenosine (A) to inosine (I), deamination, and non-template nucleotide addition and insertion. RNA editing of mRNA effectively alters the amino acid sequence of the encoded protein to differ from that predicted by the genomic DNA sequence.

[0157] The term "cancer mutation signature" refers to a set of mutations present in cancer cells when compared to non-cancerous reference cells.

[0158] In the context of this invention, “reference” can be used to correlate and compare results obtained from tumor specimens. Typically, “reference” can be obtained based on one or more normal specimens, particularly specimens unaffected by cancer, obtained from either a patient or one or more different individuals, preferably healthy individuals, especially individuals of the same species. “Reference” can be determined empirically by testing a sufficiently large number of normal specimens.

[0159] The term "reference genome" refers to a genome that provides a coordinate system for normal and diseased genomes. The reference genome is used for mapping read data and providing a coordinate system for normal and tumor genomes, which enables the provision of chromosome numbers, nucleotide positions within chromosomes, and the orientation of read data. The reference genome may be based on the genome of one or more members of the same species as the subject providing the diseased sample, or it may be based on the normal genome (matched genome) of the subject.

[0160] Any suitable sequencing method can be used in the context of this invention to identify disease-specific mutations, with next-generation sequencing (NGS) technology being preferred and optionally combined with SNP arrays to obtain absolute copy number information. To increase the speed of the sequencing steps of the method, third-generation sequencing methods may replace NGS technology in the future. For the purpose of clarification, the terms “next-generation sequencing” or “NGS” in the context of this invention mean all novel high-throughput sequencing technologies as opposed to “conventional” sequencing methods known as Sanger chemistry, which read nucleic acid templates randomly and in parallel along the entire genome by cutting the entire genome into small pieces. Such NGS technologies (also known as massively parallel sequencing technologies) can deliver nucleic acid sequence information of the whole genome, exome, transcriptome (all transcribed sequences of the genome), or methylome (all methylated sequences of the genome) in a very short period of time, e.g., within 1-2 weeks, preferably within 1-7 days, or most preferably within less than 24 hours, enabling single-cell sequencing methods in principle. Several NGS platforms, either commercially available or described in detail in the literature, for example, Zhang et al., 2011, The impact of next-generation sequencing on genomics. J. Genet Genomics 38(3): pp. 95-109, or Voelkerding et al., 2009, Next generation sequencing: From basic research to diagnostics. Clinical chemistry 55: pp. 641-658, can be used in the context of this invention. Non-limiting examples of such NGS technologies / platforms are as follows: 1) For example, sequencing by a synthetic technique known as pyrosequencing, implemented in the GS-FLX 454 Genome Sequencer™ of Roche-affiliated 454 Life Sciences (Branford, Connecticut), as first described by Ronaghi et al., 1998, A sequencing method based on real-time pyrophosphate, Science 281: pp. 363-365. This technique uses emulsion PCR, in which single-stranded DNA-binding beads are encapsulated by vigorous vortex stirring in aqueous micelles containing PCR reactants surrounded by oil for emulsion PCR amplification. During the pyrosequencing process, light emitted from phosphate molecules during nucleotide incorporation as polymerase synthesizes the DNA strand is recorded. 2) Synthetic sequencing based on a reversible dye-terminator, developed by Solexa (now part of Illumina Inc., San Diego, California), and implemented in, for example, the Illumina / Solexa Genome Analyzer™ and the Illumina HiSeq 2000 Genome Analyzer™. In this technique, all four nucleotides are simultaneously added to an oligo-initial stimulated cluster fragment in a flow cell channel along with DNA polymerase. Crosslinking amplification extends the cluster chain containing all four fluorescently labeled nucleotides for sequencing. 3) Sequencing using the ligation method, as implemented in the SOLid® platform from Applied Biosystems (now Life Technologies Corporation, Carlsbad, California). In this technique, a pool of all possible fixed-length oligonucleotides is labeled according to the sequencing location. The oligonucleotides are annealed and ligated, and preferential ligation of matched sequences by DNA ligase yields a signal that provides information about the nucleotide at that location. Before sequencing, the DNA is amplified by emulsion PCR. The resulting beads, each containing only one copy of the same DNA molecule, are placed on a glass slide. As a second example, the Polonator® G.007 platform from Dover Systems (Salem, New Hampshire) also uses ligation sequencing by amplified DNA fragments for parallel sequencing using emulsion PCR based on randomly arrayed beads. 4) Single-molecule sequencing technologies, such as those implemented in the PacBio RS system from Pacific Biosciences (Menlo Park, California) or the HeliScope® platform from Helicos Biosciences (Cambridge, Massachusetts). A key feature of this technology is its ability to sequence a single DNA or RNA molecule without amplification, which is defined as single-molecule real-time (SMRT) DNA sequencing. For example, HeliScope uses a highly sensitive fluorescence detection system to directly detect each nucleotide as it is being synthesized. A similar technique based on fluorescence resonance energy transfer (FRET) has been developed by Visigen Biotechnology (Houston, Texas). Other fluorescence-based single-molecule techniques are from USGenomics (GeneEngine®) and Genovoxx (AnyGene®). 5) Nanotechnology for single-molecule sequencing, for example, using various nanostructures placed on a chip to monitor the movement of polymerase molecules on a single strand during replication. Non-limiting examples of nanotechnology-based methods include the GridON® platform from Oxford Nanopore Technologies (Oxford, UK), the Hybridization-Assisted Nanopore Sequencing (HANS®) platform developed by Nabsys (Providence, Rhode Island), and a trademarked ligase-based DNA sequencing platform using DNA nanoball (DNB) technology called Combinatorial Probe Anchor Ligation (cPAL®). 6) Techniques based on electron microscopy observation for single-molecule sequencing, such as those developed by LightSpeed ​​Genomics (Sunnyvale, California) and Halcyon Molecular (Redwood City, California). 7) Ion semiconductor sequencing based on the detection of hydrogen ions released during DNA polymerization. For example, Ion Torrent Systems (San Francisco, California) uses a high-density array of micro-measurement wells to perform this biochemical process in a massively parallel manner. Each well has a different DNA template. Below the wells is an ion-sensitive layer, and below that are trademarked ion sensors.

[0161] In one embodiment, as disclosed in the international PCT patent application titled "Highly Accurate Mutation Detection, In Particular for Personalized Therapeutics," filed on the same date as this specification and incorporated herein by reference in its entirety, whether or not a disease-specific mutation has occurred can be determined by a method relating to determining whether a site in the normal genome matches a homozygous genotype and an ideal noise distribution reflected by a normal allele and three noise alleles; if the corresponding site in the tumor genome does not match a homozygous genotype and an ideal noise distribution, it is determined to be a mutation; and if the read data maps to each of the noise alleles with a probability of one-third of the error rate per nucleotide, the read data matches an ideal noise distribution.

[0162] Preferably, DNA and RNA preparations serve as starting materials for NGS. Such nucleic acids can be readily obtained from samples of biological materials, for example, from fresh, flash-frozen, or formalin-fixed paraffin-embedded tumor tissue (FFPE), from newly isolated cells, or from cytotoxic tumor cells (CTCs) present in the peripheral blood of patients. Normal, unmutated genomic DNA or RNA can be extracted from normal somatic tissue, but germline cells are preferred in the context of this invention. Germline DNA or RNA is extracted from peripheral blood mononuclear cells (PBMCs) in patients with non-hematological malignancies. Nucleic acids extracted from FFPE tissue or freshly isolated single cells are highly fragmented, but they are suitable for NGS applications.

[0163] Several targeted NGS methods for exome sequencing are described in the literature (see, for example, Teer and Mullikin, 2010, Human Mol Genet 19(2): R145-51), and all of these can be used in conjunction with the present invention. Many of these methods (described, for example, as genome capture, genome distribution, genome enrichment, etc.) use hybridization techniques, including array-based (e.g., Hodges et al., 2007, Nat. Genet. 39: 1522-1527) and liquid-based (e.g., Choi et al., 2009, Proc. Natl. Acad. Sci USA 106: 19096-19101) hybridization methods. Commercially available kits for DNA sample preparation and subsequent exome capture are also available. For example, Illumina Inc. (San Diego, California) offers the TruSeq® DNA Sample Preparation Kit and the TruSeq® Exome Enrichment Kit.

[0164] For example, when comparing the sequence of a tumor sample with the sequence of a reference sample, such as a germline sample, it is preferable to determine the sequences during replication of one or both of these sample types in order to reduce the number of false positives in detecting cancer-specific somatic mutations or sequence differences. Therefore, it is preferable to determine the sequence of the reference sample, such as a germline sample, two, three, or more times. Alternatively, in addition to that, the sequence of the tumor sample may be determined two, three, or more times. Furthermore, the sequences of the reference sample, such as a germline sample, and / or the tumor sample may also be determined multiple times by determining the sequence in the genomic DNA at least once and the sequence in the RNA of the reference sample and / or the tumor sample at least once. For example, by determining the mutations between replications of a reference sample, such as a germline sample, the expected false positive rate (FDR) of somatic mutations can be estimated as a statistical quantity. Technical replication of a single sample should produce the same result, and all mutations detected during this "comparison of identical samples" are false positives. In particular, to determine the false detection rate of somatic mutations in tumor samples relative to a reference sample, technical replicates of the reference sample can be used as a reference to estimate the number of false positives. Furthermore, various quality-related metrics (e.g., coverage or SNP quality) can be combined into a single quality score using machine learning techniques. Optionally, for a given somatic mutation, all other mutations with a higher quality score may be counted, thereby allowing for the ranking of all mutations in the dataset.

[0165] In the context of the present invention, the term “RNA” refers to a molecule comprising at least one ribonucleotide residue, preferably composed entirely or substantially of ribonucleotide residues. “Ribonucleotide” refers to a nucleotide having a hydroxyl group at the 2' position of a β-D-ribofuranosyl group. The term “RNA” includes isolated RNA such as double-stranded RNA, single-stranded RNA, partially or completely purified RNA, essentially pure RNA, synthetic RNA, and recombinant-produced RNA such as modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution, and / or alteration of one or more nucleotides. Such alterations may include, for example, the addition of non-nucleotide substances to or within the RNA, for example, at one or more nucleotides of the RNA. Furthermore, nucleotides in an RNA molecule may also include non-standard nucleotides such as naturally occurring nucleotides, chemically synthesized nucleotides, or deoxyribonucleotides. These modified RNAs may be called analogs or analogs of naturally occurring RNA.

[0166] The term "RNA" includes, and preferably relates to, "mRNA." The term "mRNA" means "messenger RNA" and refers to a "transcript" that is produced using a DNA template and encodes a peptide or polypeptide. Typically, mRNA includes a 5'-UTR, a protein-coding region, and a 3'-UTR. mRNA has a limited half-life in cells and in vitro. In the context of this invention, mRNA can be produced from a DNA template by in vitro transcription. In vitro transcription methods are known to those skilled in the art. For example, various in vitro transcription kits are commercially available.

[0167] RNA stability and translation efficiency can be modified as needed. For example, RNA can be stabilized and its translation increased by one or more modifications that have a stabilizing effect on RNA and / or increase translation efficiency. Such modifications are described, for example, in PCT / EP2006 / 009448, which is incorporated herein by reference. To increase the expression of RNA used in embodiments of the present invention, the GC content can be increased, thereby increasing mRNA stability, performing codon optimization, and thus enhancing intracellular translation, by modifying the coding region, i.e., the sequence encoding the peptide or protein to be expressed, preferably without altering the sequence of the peptide or protein to be expressed.

[0168] In the context of RNA as used in this invention, the term "modification" includes any modification of the RNA that is not naturally present in the RNA.

[0169] In one embodiment of the present invention, the RNA used in accordance with the present invention does not have an uncapped 5'-triphosphate. Such removal of the uncapped 5'-triphosphate can be achieved by treating the RNA with a phosphatase.

[0170] The RNA according to the present invention may have modified ribonucleotides to increase its stability and / or reduce its cytotoxicity. For example, in one embodiment, cytidine is partially or completely, preferably completely, substituted with 5-methylcytidine in the RNA used according to the present invention. Or, in addition, in one embodiment, uridine is partially or completely, preferably completely, substituted with pseudouridine in the RNA used according to the present invention.

[0171] In one embodiment, the term "modification" refers to providing RNA with a 5'-cap or a 5'-cap analogue. The term "5'-cap" refers to a cap structure found on the 5' end of an mRNA molecule, and generally consists of a guanosine nucleotide linked to mRNA via an unusual 5'-to-5' triphosphate bond. In one embodiment, this guanosine is methylated at position 7. The term "conventional 5'-cap" refers to a naturally occurring RNA 5'-cap, preferably a 7-methylguanosine cap (m 7 G) refers to the above. In the context of the present invention, the term "5'-cap" includes 5'-cap analogues that are similar to the RNAcap structure and are modified to have the ability to stabilize RNA and / or enhance RNA translation when attached to it in vivo and / or intracellularly.

[0172] Providing a 5'-cap or 5'-cap analogue to RNA may be achieved by in vitro transcription of a DNA template in the presence of the 5'-cap or 5'-cap analogue, wherein the 5'-cap is co-transcribed into the prepared RNA strand, or the RNA may be prepared, for example, by in vitro transcription, and the 5'-cap may be attached to the RNA after transcription using a capping enzyme, such as a vaccinia virus capping enzyme.

[0173] The RNA may undergo further modifications. For example, further modifications of the RNA used in the present invention may include elongation or cleavage of a naturally occurring poly(A) tail, or introduction of UTRs unrelated to the coding region of the RNA, such as alterations of the 5'- or 3'-untranslated region (UTR), including replacement or insertion of an existing 3'-UTR with one or more, preferably two copies, of 3'-UTRs derived from a globin gene, such as alpha-2-globin, alpha-1-globin, beta-globin, preferably beta-globin, more preferably human beta-globin.

[0174] RNA with unmasked poly(A) sequences is translated more efficiently than RNA with masked poly(A) sequences. The term "poly(A) tail" or "poly(A) sequence" refers to a sequence of adenyl (A) residues typically located at the 3' end of an RNA molecule. "Unmasked poly(A) sequence" means that the poly(A) sequence at the 3' end of the RNA molecule ends with an A and is not followed by any nucleotides other than A at the 3' end, i.e., downstream. Furthermore, a long poly(A) sequence of approximately 120 base pairs provides optimal transcriptional stability and translational efficiency for RNA.

[0175] Therefore, in order to increase the stability and / or expression of RNA used according to the present invention, it may be modified to be present with a poly-A sequence having a length of preferably 10 to 500, more preferably 30 to 300, even more preferably 65 to 200, and particularly 100 to 150 adenosine residues. In a particularly preferred embodiment, the poly-A sequence has a length of approximately 120 adenosine residues. To further increase the stability and / or expression of RNA used according to the present invention, the poly-A sequence may not be masked.

[0176] Furthermore, the incorporation of a 3'-untranslated region (UTR) into the 3'-untranslated region of an RNA molecule can lead to improved translation efficiency. Synergistic effects can be achieved by incorporating two or more such 3'-untranslated regions. The 3'-untranslated regions can be self- or heterologous to the RNA into which they are introduced. In a particular embodiment, the 3'-untranslated region is derived from the human β-globin gene.

[0177] The aforementioned combinations of modifications, namely the incorporation of poly-A sequences, demasking of poly-A sequences, and incorporation of one or more 3'-untranslated regions, have a synergistic effect on increasing RNA stability and translation efficiency.

[0178] The term "stability" of RNA relates to its "half-life." Half-life refers to the time required to eliminate half of the activity, quantity, or number of a molecule. In the context of this invention, the half-life of RNA is an indicator of its stability. The half-life of RNA can affect the "duration of expression" of RNA. RNA with a long half-life can be expected to be expressed for a long period of time.

[0179] Of course, if it is desirable to reduce RNA stability and / or translation efficiency, it is possible to modify the RNA in such a way as to interfere with the function of the aforementioned elements that increase RNA stability and / or translation efficiency.

[0180] The term "expression" is used in its most common sense and includes, for example, the production of RNA and / or peptides or polypeptides by transcription and / or translation. With respect to RNA, the terms "expression" or "translation" particularly refer to the production of peptides or polypeptides. This also includes the partial expression of nucleic acids. Furthermore, expression can be transient or stable.

[0181] The term expression also includes “abnormal expression” or “abnormal expression.” “Abnormal expression” or “abnormal expression” means that the expression is altered, preferably increased, compared to a reference, e.g., abnormal expression of a particular protein, e.g., tumor antigen, or a state of a subject not suffering from the disease associated with the abnormal expression. Increased expression means an increase of at least 10%, particularly at least 20%, at least 50%, or at least 100%, or more. In one embodiment, the expression is found only in affected tissue, and expression in healthy tissue is suppressed.

[0182] The term "specifically expressed" means that a protein is expressed essentially only in specific tissues or organs. For example, a tumor antigen specifically expressed in the gastric mucosa means that the protein is primarily expressed in the gastric mucosa and not expressed in other tissues or to a significant degree in other tissues or organ types. Therefore, a protein that is exclusively expressed in gastric mucosal cells and expressed to a significantly lesser degree in any other tissue, such as the testes, is specifically expressed in gastric mucosal cells. In some embodiments, a tumor antigen may also be specifically expressed in multiple tissue types or organs, for example, two or three tissue types or organs, but preferably three or fewer different tissue or organ types, under normal conditions. In this case, the tumor antigen is specifically expressed in these organs. For example, if a tumor antigen is preferably expressed to approximately the same degree in the lungs and stomach under normal conditions, the tumor antigen is specifically expressed in the lungs and stomach.

[0183] In the context of this invention, the term “transcription” refers to the process by which the genetic code of a DNA sequence is transcribed into RNA. The RNA can then be translated into a protein. According to this invention, the term “transcription” includes “in vitro transcription,” which refers to the process by which RNA, particularly mRNA, is synthesized in vitro in a cell-free system, preferably using a suitable cell extract. Preferably, a cloning vector is used to produce the transcript. These cloning vectors are generally named transcription vectors and are encompassed by the term “vector.” The RNA used in this invention is preferably in vitro transcription RNA (IVT-RNA), which can be obtained by in vitro transcription of a suitable DNA template. The promoter for regulating transcription can be any promoter of any RNA polymerase. Specific examples of RNA polymerases are T7, T3, and SP6 RNA polymerases. Preferably, in vitro transcription is regulated by the T7 or SP6 promoter. A DNA template for in vitro transcription can be obtained by cloning a nucleic acid, particularly cDNA, and introducing it into a suitable vector for in vitro transcription. cDNA can be obtained by reverse transcription of RNA.

[0184] The term "translation" refers to the process within a cell's ribosome in which a chain of messenger RNA directs the assembly of amino acid sequences to produce peptides or polypeptides.

[0185] In the context of the present invention, an expression regulatory sequence or regulatory sequence that can be functionally linked to a nucleic acid may be homologous or heterologous with respect to the nucleic acid. The coding sequence and the regulatory sequence are "functionally" linked together if they are covalently bonded together such that the transcription or translation of the coding sequence is under the control or influence of the regulatory sequence. When a coding sequence is translated into a functional protein using functional linkage between a regulatory sequence and a coding sequence, the induction of the regulatory sequence results in the transcription of the coding sequence without causing a shift in the reading frame of the coding sequence or a failure of the coding sequence to be translated into the desired protein or peptide.

[0186] In relation to the present invention, the terms “expression regulatory sequence” or “regulatory sequence” include promoters, ribosome-binding sequences, and other regulatory elements that control the transcription of nucleic acids or the translation of induced RNA. In certain embodiments, regulatory sequences can be controlled. The exact structure of regulatory sequences may vary depending on the species or cell type, but generally include 5'-untranscribed and 5'- and 3'-untranslated sequences involved in the initiation of transcription or translation, such as TATA boxes, capping sequences, and CAAT sequences. In particular, 5'-untranscribed regulatory sequences include promoter regions containing promoter sequences that regulate the transcription of functionally bound genes. Regulatory sequences may also include enhancer sequences or upstream activating sequences.

[0187] Preferably, the RNA to be expressed in the cell is introduced into the cell. In one embodiment of the method according to the present invention, the RNA to be introduced into the cell is obtained by in vitro transcription of a suitable DNA template.

[0188] Terms such as “expressible RNA” and “coding RNA” are used interchangeably herein, and with respect to a particular peptide or polypeptide, RNA can be expressed to produce the peptide or polypeptide when present in a suitable environment, preferably within a cell. Preferably, RNA can interact with the cell’s translation mechanism to provide the peptide or polypeptide it can express.

[0189] Terms such as “transfer,” “introduce,” or “transfer” are used interchangeably herein and relate to the introduction of nucleic acids, particularly exogenous or heterologous nucleic acids, especially RNA, into cells. According to the present invention, cells can form organs, tissues, and / or parts of organisms. According to the present invention, the administration of nucleic acids is achieved either as naked nucleic acids or in combination with an administration reagent. Preferably, the administration of nucleic acids is in the form of naked nucleic acids. Preferably, RNA is administered in combination with a stabilizing substance such as an RNase inhibitor. The present invention also envisions the repeated introduction of nucleic acids into cells to enable long-term sustained expression.

[0190] Cells can be translocated using any carrier that can associate with RNA, for example, by forming a complex with RNA or by forming vesicles that encapsulate or encapsulate RNA, resulting in increased RNA stability compared to naked RNA. Useful carriers include, for example, cationic lipids, liposomes, especially cationic liposomes, and lipid-containing carriers such as micelles, as well as nanoparticles. Cationic lipids can form complexes with nucleic acids of load. Any cationic lipid can be used.

[0191] Preferably, the introduction of RNA encoding a peptide or polypeptide into cells, particularly into cells present in vivo, results in the expression of the peptide or polypeptide in the cells. In certain embodiments, it is preferable to target the nucleic acid to specific cells. In such embodiments, the carrier applied to administer the nucleic acid to cells (e.g., a retrovirus or liposome) represents the targeting molecule. For example, molecules such as antibodies specific to surface membrane proteins on target cells or ligands of receptors on target cells may be incorporated into or bound to the nucleic acid carrier. When nucleic acids are administered by liposomes, proteins that bind to surface membrane proteins involved in endocytosis may be incorporated into the liposome formulation to enable targeting and / or uptake. Such proteins include capsid proteins or fragments thereof that are specific to specific cell types, antibodies against proteins that are internally distributed, proteins that target intracellular locations, etc.

[0192] The terms “cell” or “host cell” preferably refer to an untreated cell, i.e., a cell with an untreated membrane that does not release its normal intracellular components such as enzymes, organelles, or genetic material. An untreated cell is preferably a living cell, i.e., a living cell capable of performing any normal metabolic function. Preferably, the terms refer to any cell that can be transformed or transmigrated using exogenous nucleic acids. The term “cell” includes prokaryotic cells (e.g., Escherichia coli (E. coli)) or eukaryotic cells (e.g., dendritic cells, B cells, CHO cells, COS cells, K562 cells, HEK293 cells, HELA cells, yeast cells, and insect cells). Exogenous nucleic acids may be found (i) freely dispersed within the cell, (ii) incorporated into a recombinant vector, or (iii) incorporated into the host cell genome or mitochondrial DNA. Mammalian cells, such as those from humans, mice, hamsters, pigs, goats, and primates, are particularly preferred. The cells may originate from a number of tissue types, including primary cells and cell lines. Specific examples include keratinocytes, peripheral blood leukocytes, bone marrow stem cells, and embryonic stem cells. In further embodiments, the cells are antigen-presenting cells, particularly dendritic cells, monocytes, or macrophages.

[0193] Cells containing nucleic acid molecules preferably express peptides or polypeptides encoded by the nucleic acid.

[0194] The term "clonal proliferation" refers to the process by which a specific entity increases in number. In the context of the present invention, this term is preferably used in relation to an immunological response in which lymphocytes are stimulated by an antigen, proliferate, and specific lymphocytes that recognize the antigen are amplified. Preferably, clonal proliferation results in the differentiation of lymphocytes.

[0195] Terms such as “reduce” or “inhibit” relate to the ability to cause an overall level reduction of preferably 5% or more, 10% or more, 20% or more, more preferably 50% or more, and most preferably 75% or more. The term “inhibit” or similar phrases include complete or essentially complete inhibition, i.e., a reduction to zero or essentially zero.

[0196] Terms such as “increase,” “enhance,” “promote,” or “extend” preferably relate to an increase, enhancement, promotion, or extension of at least about 10%, preferably at least 20%, preferably at least 30%, preferably at least 40%, preferably at least 50%, preferably at least 80%, preferably at least 100%, preferably at least 200%, and especially at least 300%. These terms may also relate to an increase, enhancement, promotion, or extension from zero or an unmeasurable or undetectable level to a level higher than zero or a measurable or detectable level.

[0197] According to the present invention, the term "peptide" refers to a substance containing two or more, preferably three or more, preferably four or more, preferably six or more, preferably eight or more, preferably ten or more, preferably thirteen or more, preferably sixteen or more, preferably twenty-one or more, and preferably up to eight, ten, twenty, thirty, forty, or fifty, and particularly up to 100 amino acids, covalently bonded by peptide bonds. The term "polypeptide" or "protein" refers to a large peptide, preferably a peptide having more than 100 amino acid residues, but generally the terms "peptide," "polypeptide," and "protein" are synonymous and are used interchangeably in this specification. According to the present invention, the terms "modification" or "sequence change" relating to peptides, polypeptides, or proteins refer to sequence changes in peptides, polypeptides, or proteins compared to a parent sequence, such as the sequence of a wild-type peptide, polypeptide, or protein. This term includes amino acid insertion mutants, amino acid addition mutants, amino acid deletion mutants, and amino acid substitution mutants, preferably amino acid substitution mutants. All of these sequence changes according to the present invention can potentially generate new epitopes.

[0198] Amino acid insertion mutants contain the insertion of one, two, or more amino acids into a specific amino acid sequence.

[0199] Amino acid addition mutants include amino and / or carboxyl terminological fusions of one or more amino acids, such as one, two, three, four, five, or more amino acids.

[0200] Amino acid deletion mutants are characterized by the removal of one or more amino acids from a sequence, for example, the removal of one, two, three, four, five, or more amino acids.

[0201] Amino acid substitution mutants are characterized by the removal of at least one residue in a sequence and the insertion of another residue in its place.

[0202] According to the present invention, the modified or modified peptide used in the test in the method of the present invention may be derived from a modified protein.

[0203] According to the present invention, the term “derived from” means that a particular entity, in particular a particular peptide sequence, is present in the object from which it originates. In the case of an amino acid sequence, in particular a particular sequence region, “derived from” means, in particular, that the relevant amino acid sequence originates from the amino acid sequence in which it exists.

[0204] The agents, compositions, and methods described herein can be used to treat subjects having diseases, for example, diseases characterized by the presence of affected cells that express antigens and present antigenic peptides. Cancer is a particularly preferred disease. The agents, compositions, and methods described herein can also be used for immunization or vaccination to prevent the diseases described herein.

[0205] One such drug is a vaccine, such as a cancer vaccine designed based on a suitable neoepitope that resists immune evasion identified by the method of the present invention.

[0206] According to the present invention, the term "vaccine" refers to a pharmaceutical preparation (pharmaceutical composition) or product that, upon administration, induces an immune response, particularly a cellular immune response, that recognizes and attacks pathogens or disease cells such as cancer cells. Vaccines can be used for the prevention or treatment of disease. The terms "personalized cancer vaccine" or "individualized cancer vaccine" refer to specific cancer patients and mean that the cancer vaccine is adapted to the needs or special circumstances of individual cancer patients.

[0207] The cancer vaccines provided according to the present invention provide one or more T cell epitopes for stimulating, priming, and / or amplifying T cells specific to the patient's tumor when administered to the patient. The T cells are preferably directed to cells expressing the antigen from which the T cell epitope is derived. Thus, the vaccines described herein can induce or promote a cellular response to cancerous disease characterized by the presentation of one or more tumor-associated neoantigens by class I MHC, preferably cytotoxic T cell activity. Since the vaccines provided herein target cancer-specific mutations, they are specific to the patient's tumor.

[0208] In the context of the present invention, the vaccine preferably provides one or more T cell epitopes (neoepitopes, suitable neoepitopes, combinations of suitable neoepitopes as identified herein), such as two or more, five or more, ten or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, preferably up to sixty, up to five five, up to five ten, up to fourty, up to fourty, up to fourty, up to three fifteen or up to thirty T cell epitopes, which, when administered to a patient, incorporate amino acid modifications or modified peptides that are predicted to be suitable epitopes. The presentation of these epitopes by the patient's cells, particularly antigen-presenting cells, preferably results in T cells targeting the epitopes upon binding to MHC, and thus the patient's tumor, preferably a primary tumor and tumor metastases, expressing antigens derived from the MHC-presented T cell epitopes and presenting the same epitopes on the surface of tumor cells.

[0209] The method of the present invention may include further steps to determine the usefulness of a modified peptide comprising an identified amino acid modification or a suitable neoepitope identified herein for cancer vaccination. Thus, the further steps may include one or more of the following: (i) evaluation of whether the modification is located in a known or predicted epitope presented on the MHC; (ii) in vitro and / or in silico testing of whether the modification is located in a MHC-presented epitope, e.g., testing whether the modification is part of a peptide sequence that is processed into and / or presented as an MHC-presented epitope; and (iii) in vitro testing of whether the assumed modified epitope can stimulate T cells, such as patient T cells, with desired specificity, particularly if it exists in its native sequence configuration, e.g., if it is adjacent to an amino acid sequence also adjacent to the epitope in a naturally occurring protein, and if expressed in antigen-presenting cells. Such adjacent sequences may each contain 3 or more, 5 or more, 10 or more, 15 or more, 20 or more, preferably up to 50, up to 45, up to 40, up to 35, or up to 30 amino acids, and may be adjacent to the epitope sequence at the N-terminus and / or C-terminus.

[0210] Modified peptides determined according to the present invention can be ranked for their usefulness as epitopes for cancer vaccines. Accordingly, in one embodiment, the method of the present invention includes a manual or computer-based analytical process in which the identified modified peptides are analyzed and selected for their usefulness in each vaccine to be offered. In a preferred embodiment, the analytical process is a computer algorithm-based process. Preferably, the analytical process includes one or more of the following steps, preferably the analytical method, which includes the step of determining and / or ranking epitopes according to a prediction of their immunogenicity.

[0211] The epitopes identified in accordance with the present invention and provided in the vaccine preferably exist in the form of a polypeptide containing the epitope (neoepitope, suitable neoepitope, neoepitope found in combinations of suitable neoepitopes identified herein), such as a polyepitope polypeptide or a nucleic acid encoding the polypeptide, particularly RNA. Furthermore, the epitope may exist in the polypeptide in the form of a vaccine sequence, i.e., in the context of its natural sequence, for example, adjacent to an amino acid sequence similarly adjacent to the epitope in a naturally occurring protein. Such adjacent sequences may each contain 5 or more, 10 or more, 15 or more, 20 or more, preferably up to 50, up to 45, up to 40, up to 35, or up to 30 amino acids, and may be adjacent to the epitope sequence at the N-terminus and / or C-terminus. Therefore, the vaccine sequence may contain 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, preferably up to 50, up to 45, up to 40, up to 35, or up to 30 amino acids. In one embodiment, the epitope and / or vaccine sequence are arranged head-to-tail in the polypeptide.

[0212] In one embodiment, an epitope / suitable neoepitope and / or vaccine sequence, as identified herein, is separated by a linker, particularly a neutral linker. As used in connection with the present invention, the term "linker" refers to a peptide that is added between two peptide domains, such as an epitope or vaccine sequence, to connect the peptide domains. There are no particular restrictions on the linker sequence. However, it is preferable that the linker sequence reduces steric hindrance between the two peptide domains, is well translated, and assists or allows the processing of the epitope. Furthermore, the linker should have no or only a few immunogenic sequence elements. Preferably, the linker should not create non-endogenous epitopes, such as those resulting from junctional sutures between adjacent epitopes, which may produce an undesirable immune response. Therefore, polyepitope vaccines should preferably contain a linker sequence that can reduce the number of undesirable MHC-binding junctional epitopes. Hoyt et al. (EMBO J. 25(8), pp. 1720-179, 2006) and Zhang et al. (J. Biol. Chem., 279(10), pp. 8635-8641, 2004) showed that glycine-rich sequences impair proteasome processing, and therefore, the use of glycine-rich linker sequences minimizes the number of peptides contained in linkers that are capable of proteasome processing. Furthermore, it was observed that glycine inhibits strong binding at the MHC binding groove site (Abastado et al., 1993, J. Immunol. 151(7): pp. 3569-3575). Schlessinger et al., 2005, Proteins, 61(1): pp. 115-126, found that the inclusion of the amino acids glycine and serine in the amino acid sequence results in a more flexible protein that is more efficiently translated and processed by the proteasome, allowing for better access to the encoded epitope. The linker may contain 3 or more, 6 or more, 9 or more, 10 or more, 15 or more, 20 or more, preferably up to 50, up to 45, up to 40, up to 35, or up to 30 amino acids. Preferably, the linker is rich in glycine and / or serine amino acids.Preferably, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the linker's amino acids are glycine and / or serine. In one preferred embodiment, the linker is substantially composed of the amino acids glycine and serine. In one embodiment, the linker is an amino acid sequence (GGS). a (GSS) b (GGG) c (SSG) d (GSG) e The formula includes, where a, b, c, d, and e are independently digits selected from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, and a+b+c+d+e is different from 0, preferably 2 or greater, 3 or greater, 4 or greater, or 5 or greater. In one embodiment, the linker includes the linker sequences described herein, including the sequence GGSGGGGSG, etc.

[0213] In one particularly preferred embodiment, a polypeptide incorporating one or more suitable neoepitopes identified by the methods herein, such as a polyepitope polypeptide, is administered to the patient in the form of nucleic acids, preferably RNA, such as in vitro transcribed or synthetic RNA, that can be expressed in the patient's cells, such as antigen-presenting cells, to produce polypeptides. The present invention also envisions the administration of one or more multiepitope polypeptides, as defined by the term "polyepitope polypeptide" for the purposes of the present invention, preferably in the form of nucleic acids, preferably RNA, such as in vitro transcribed or synthetic RNA, that can be expressed in the patient's cells, such as antigen-presenting cells, to produce one or more polypeptides. When multiple multiepitope polypeptides are administered, the suitable neoepitopes provided by different multiepitope polypeptides may be different or partially overlapping. Once present in the patient's cells, such as antigen-presenting cells, the polypeptides according to the present invention are processed to produce suitable neoepitopes identified according to the present invention. Administration of a vaccine provided according to the present invention may preferably provide an epitope presented in MHC class II that can induce a CD4+ helper T cell response against cells expressing an antigen derived from the epitope presented in the MHC. Alternatively or, moreover, administration of a vaccine provided according to the present invention may provide a neoepitope presented in MHC class I that can induce a CD8+ T cell response against cells expressing an antigen derived from the neoepitope presented in the MHC. Furthermore, administration of a vaccine provided according to the present invention may provide one or more neoepitopes (including known neoepitopes and suitable neoepitopes identified according to the present invention), as well as one or more epitopes that do not contain cancer-specific somatic mutations but are expressed by cancer cells and preferably induce an immune response against cancer cells, preferably a cancer-specific immune response.In one embodiment, administration of a vaccine provided according to the present invention provides a neoepitope that can induce a CD4+ helper T cell response to cells expressing an antigen that is an epitope presented in MHC class II and / or derived from an epitope presented in MHC, and an epitope that does not contain cancer-specific somatic mutations and can induce a CD8+ T cell response to cells expressing an antigen that is an epitope presented in MHC class I and / or derived from an epitope presented in MHC. In one embodiment, the epitope that does not contain cancer-specific somatic mutations is derived from a tumor antigen. In one embodiment, the neoepitope and epitope that do not contain cancer-specific somatic mutations have a synergistic effect in the treatment of cancer. Preferably, the vaccine provided according to the present invention is useful for polyepitope stimulation of cytotoxicity and / or helper T cell response.

[0214] The vaccine provided in accordance with the present invention may be a recombinant vaccine.

[0215] Another type of agent is an immune cell, such as a T cell expressing a T cell receptor, or a T cell expressing a T cell receptor by recombination, or an artificial or chimeric T cell receptor (CAR), the receptor being targeted to an antigen, for example, a suitable neoepitope identified by the method of the present invention as a suitable disease-specific target, preferably such a neoepitope being expressed on the cell surface in complex with an MHC molecule. Preferably, when the immune cell recognizes the antigen by receptor-antigen binding, the immune cell (immunoreactive cell) is stimulated, primed and / or amplified, or exerts the effector function of the immunoreactive cell described above.

[0216] The term “antigen-specific T cell” or similar terms refers to T cells that recognize an antigen complexed within an MHC class I molecule, such as a suitable neoepitope, and, after binding to the antigen, preferably exert the effector function of the T cell described above. T cells and other lymphoid cells are considered antigen-specific if they kill target cells expressing the antigen. T cell specificity can be evaluated, for example, using any of the various standard techniques in chromium release assays or proliferation assays. Alternatively, the synthesis of lymphokines (such as interferon-γ) can be measured.

[0217] T cell receptors and other antigen receptors are described above. The term “CAR” (or “chimeric antigen receptor”) relates to an artificial receptor comprising a single molecule or a complex of molecules that recognizes, i.e., binds to, a target structure (e.g., an antigen) on target cells such as cancer cells (e.g., by binding an antigen-binding domain to the antigen expressed on the surface of the target cell), and can confer specificity to immune effector cells such as T cells that express the CAR on the cell surface. Preferably, recognition of the target structure by the CAR results in activation of immune effector cells expressing the CAR. A CAR may comprise one or more protein units, the protein units comprising one or more domains as described herein. The term “CAR” does not include T cell receptors.

[0218] In one embodiment, a single-chain variable fragment (scFv) derived from a monoclonal antibody is fused to the CD3-zeta transmembrane and endodomain. Such a molecule results in zeta signaling in response to the recognition of its antigen target by the scFv in target cells and the death of target cells expressing the target antigen. Antigen recognition domains that can be used similarly include, among others, T cell receptor (TCR) alpha and beta single chains. In fact, almost anything that binds to a given target with high affinity can be used as an antigen recognition domain.

[0219] After antigen recognition, receptors cluster together, and signals are transmitted to the cell. In this respect, the "T cell signaling domain" is the domain that transmits an activation signal to the T cell after antigen binding, preferably an endodomain. The most commonly used endodomain component is CD3-zeta.

[0220] Because CAR-modified T cells can be engineered to target virtually any antigen expressed on affected cells, such as tumor antigens, adoptive cell transfer therapy using CAR-modified T cells expressing chimeric antigen receptors is a promising mode of therapy. Preferably, the tumor antigen is a neoepitope resulting from tumor-specific mutations identified by the method of the present invention as a suitable tumor-specific target. For example, the patient's T cells can be genetically engineered (genetically modified) to express a CAR specifically created against a tumor-specific neoepitope that has complexed with MHC molecules on the surface of the patient's tumor cells, and then returned to the patient for injection.

[0221] CARs can replace the function of T cell receptors, and in particular, can confer reactivity, such as cytolytic activity, to cells such as T cells. However, in contrast to the binding of T cell receptors to the antigen peptide-MHC complex described above, CARs can also bind to antigens, especially when expressed on the cell surface.

[0222] According to the present invention, a CAR generally contains three domains. The first domain is a binding domain that recognizes and binds to the antigen. The second domain is a co-stimulatory domain. The co-stimulatory domain functions to enhance the proliferation and survival of cytotoxic lymphocytes after the binding of the CAR to the targeting site. The identity of the co-stimulatory domain is limited only to the fact that it has the ability to enhance cellular proliferation and survival after the binding of the targeting site by the CAR. Suitable co-stimulatory domains include CD28, CD137 (4-1BB), a member of the tumor necrosis factor (TNF) receptor family, CD134 (OX40), a member of the TNFR receptor superfamily, and CD278 (ICOS), a CD28 superfamily co-stimulatory molecule expressed on activated T cells. The third domain is an activation signaling domain (or T cell signaling domain). The activation signaling domain functions to activate cytotoxic lymphocytes after the binding of the CAR to the antigen. The identity of the activation signaling domain is limited solely to its ability to induce activation of selected cytotoxic lymphocytes after antigen binding by CAR. Suitable activation signaling domains include T cell CD3 [zeta] chains and Fc receptor [gamma].

[0223] A CAR can contain three domains collectively in the form of a fusion protein. Such a fusion protein will generally contain a binding domain, one or more co-stimulatory domains, and an activation signaling domain, linked from the N-terminus to the C-terminus. However, CARs are not limited to this arrangement, and other arrangements are permitted, which include a binding domain, an activation signaling domain, and one or more co-stimulatory domains. Since the binding domain must freely bind to the antigen, it will be understood that the placement of the binding domain in a fusion protein will generally result in the display of the region outside the cell. In the same manner, since the co-stimulatory and activation signaling domains function in inducing the activity and proliferation of cytotoxic lymphocytes, the fusion protein will generally display these two domains inside the cell. A CAR may contain additional elements such as a signal peptide to ensure proper transport of the fusion protein to the cell surface, a transmembrane domain to ensure the fusion protein is maintained as an endogenous membrane protein, and a hinge domain (or spacer region) to confer mobility to the binding domain and enable strong binding to the antigen.

[0224] The cells used in association with CARs and other artificial antigen receptors are preferably T cells, particularly cytotoxic lymphocytes, and are preferably selected from cytotoxic T cells, natural killer (NK) cells, and lymphokine-activated killer (LAK) cells. After activation, each of these cytotoxic lymphocytes induces the destruction of target cells. For example, cytotoxic T cells induce the destruction of target cells by one or both of the following means: Firstly, after activation, T cells release cytotoxic substances such as perforin, granzyme, and granulysin. Perforin and granulysin create pores in target cells, and granzyme enters the cells and induces a caspase cascade in the cytoplasm, which induces apoptosis (programmed cell death) of the cells. Secondly, apoptosis may be induced via Fas-Fas ligand interaction between T cells and target cells. Cytotoxic lymphocytes will preferably be autologous cells, but xenogeneic cells or allogeneic intergenerational cells may also be used.

[0225] The antigen-binding domain that may be present within the CAR has the ability to bind to (target) the antigen, that is, the ability to bind to (target) an epitope present on the antigen, preferably a neoepitope identified by the method of the present invention as a suitable disease-specific target, the neoepitope being presented in the context of MHC on the cell surface. Preferably, the antigen-binding domain is antigen-specific.

[0226] Another type of agent is immune cells, such as lymphoid cells loaded with peptides containing suitable neoepitopes identified by the method of the present invention. In preferred embodiments, the lymphoid cells are dendritic cells. In the context of the present invention, lymphoid cells, preferably isolated from a patient to be treated, are incubated with an antigen to be targeted, and the incubated cells are then administered to the patient to induce an immune response against cells expressing the antigen. Thus, peptides containing suitable epitopes can be incubated with dendritic cells, and the incubated cells can be administered to induce an immune response against cells expressing the suitable neoepitopes.

[0227] In relation to the present invention, the term "recombinant" means "produced by genetic manipulation." Preferably, "recombinant entities" such as recombinant polypeptides in relation to the present invention do not exist in nature and are preferably the result of combinations of entities such as amino acids or nucleic acid sequences that do not combine in nature. For example, a recombinant polypeptide in relation to the present invention may contain several amino acid sequences, such as a neoepitope or a vaccine sequence derived from different proteins or different parts of the same protein fused together, for example, by peptide bonds or an appropriate linker.

[0228] As used herein, the term "naturally occurring" means that a substance can be found in nature. For example, peptides or nucleic acids that are present in living organisms (including viruses), can be isolated from natural sources, and have not been intentionally modified by humans in a laboratory are considered naturally occurring.

[0229] According to the present invention, the term "disease" means cancerous disease, and in particular any pathological condition including the forms of cancerous disease described herein.

[0230] The term "normal" refers to a healthy state or a healthy subject or tissue, i.e., a state in a non-pathological condition, and "healthy" preferably means non-cancerous.

[0231] "A disease involving cells expressing an antigen" means that the expression of an antigen is detected in cells of an affected tissue or organ. The expression in cells of an affected tissue or organ may be increased compared to the state of a healthy tissue or organ. An increase means at least 10%, particularly at least 20%, at least 50%, at least 100%, at least 200%, at least 500%, at least 1000%, at least 10000%, or even greater. In one embodiment, the expression is found only in the affected tissue, and the expression in healthy tissue is suppressed. According to the present invention, diseases involving or related to cells expressing an antigen include cancerous diseases.

[0232] Cancer (medically known as malignant neoplasm) is a class of diseases characterized by uncontrolled growth (division beyond normal limits), invasion (invasion and destruction of adjacent tissues), and, in some cases, metastasis (spread to other locations within the body via the lymphatic system or bloodstream). These three malignant characteristics of cancer distinguish it from benign tumors, which are self-contained and do not invade or metastasize. Most cancers form tumors, but some, such as leukemia, do not.

[0233] Malignant tumors are essentially synonymous with cancer. Malignant diseases, malignant neoplasms, and malignant tumors are essentially synonymous with cancer.

[0234] According to the present invention, the terms “tumor” or “tumor disease” preferably mean the abnormal growth of cells (referred to as neoplastic cells, oncoplastic cells, or tumor cells) that form swelling or lesions. “Tumor cells” means abnormal cells that grow by rapid, unregulated cell proliferation and continue to grow after the stimulus that initiated new growth has ended. Tumors exhibit a partial or complete lack of structural organization and functional coordination with normal tissue, and usually form a distinct tissue mass that can be benign, premalignant, or malignant.

[0235] A benign tumor is a tumor that lacks all three malignant characteristics of cancer. Therefore, by definition, a benign tumor does not grow in an unrestricted, invasive manner, does not invade surrounding tissues, and does not spread (metastasize) to non-adjacent tissues.

[0236] A neoplasm is an abnormal tissue mass resulting from neoplasia. Neoplasia (from the Greek word for new growth) is the abnormal proliferation of cells. The growth of cells exceeds the growth of the surrounding normal tissue and is not coordinated. The growth continues in the same excessive manner even after the stimulus has ceased. This usually results in a mass or tumor. Neoplasms can be benign, premalignant, or malignant.

[0237] In relation to the present invention, "tumor growth" or "tumor development" refers to the tendency for a tumor to increase in size and / or the tendency for tumor cells to proliferate.

[0238] For the purposes of this invention, the terms "cancer" and "cancer disease" are used interchangeably with the terms "tumor" and "tumor disease."

[0239] Cancers are classified according to the type of cells that resemble tumors, and therefore the type of tissue from which the tumor is presumed to originate. These are histology and location, respectively.

[0240] The term "cancer" according to the present invention includes leukemia, seminomas, melanoma, teratomas, lymphomas, neuroblastomas, gliomas, rectal cancer, endometrial cancer, kidney cancer, adrenal cancer, thyroid cancer, hematological cancer, skin cancer, brain cancer, cervical cancer, intestinal cancer, liver cancer, colon cancer, gastric cancer, intestinal cancer, head and neck cancer, gastrointestinal cancer, lymph node cancer, esophageal cancer, colorectal cancer, pancreatic cancer, ear, nose and pharynx (ENT) cancers, breast cancer, prostate cancer, uterine cancer, ovarian cancer and lung cancer, and their metastases. Examples include lung carcinoma, breast carcinoma, prostate carcinoma, colon carcinoma, renal cell carcinoma, cervical carcinoma, or metastases of the above-mentioned cancers or tumors. Furthermore, according to the present invention, the term cancer also includes cancer metastasis and cancer relapse.

[0241] "Metastasis" refers to the spread of cancer cells from their original site to another part of the body. The formation of metastasis is a very complex process and depends on the detachment of malignant cells from the primary tumor, invasion of the extracellular matrix, penetration of the endothelial basement membrane to enter the body cavity and blood vessels, and then invasion of the target organ after being transported by the blood. Finally, the growth of a new tumor at the target site, i.e., a secondary tumor or metastatic tumor, depends on angiogenesis. Tumor metastasis often occurs even after the removal of the primary tumor, because tumor cells or components may remain and develop metastatic potential. In one embodiment, the term "metastasis" according to the present invention refers to "distant metastasis" that is a metastasis away from the primary tumor and the regional lymph node system.

[0242] The cells of a secondary or metastatic tumor resemble those of the original tumor. This means, for example, if ovarian cancer metastasizes to the liver, the secondary tumor will consist of abnormal ovarian cells rather than abnormal liver cells. In this case, the tumor in the liver is called metastatic ovarian cancer, not liver cancer.

[0243] The term “circulating tumor cells” or “CTC” refers to cells that detach from a primary tumor or tumor metastasis and circulate in the bloodstream. CTCs can constitute the seeds for the subsequent growth of further tumors (metastases) in various tissues. Circulating tumor cells are found at a frequency of approximately 1 to 10 CTCs per mL of whole blood in patients with metastatic disease. Research methods for isolating CTCs have been developed. Several research methods for isolating CTCs in the art have been described, for example, techniques that utilize the fact that epithelial cells generally express the cell adhesion protein EpCAM, which is not present in normal blood cells. Immunomagnetic bead-based capture involves treating a blood sample with an antibody against EpCAM conjugated with magnetic particles, and then separating the tagged cells in a magnetic field. Rare CTCs are then identified from contaminated leukocytes by staining the isolated cells with antibodies against another epithelial marker, cytokeratin, and the common leukocyte marker CD45. This robust, semi-automated method identifies CTCs with an average yield of approximately 1 CTC / mL and a purity of 0.1% (Allard et al., 2004, Clin Cancer Res 10: pp. 6897-6904). A second method for isolating CTCs uses a microfluidic-based CTC capture device, which involves flowing whole blood through a chamber embedded with 80,000 microposts functionalized by coating with an antibody against EpCAM. The CTCs are then stained with a secondary antibody against either cytokeratin or a tissue-specific marker such as PSA in prostate cancer or HER2 in breast cancer, and visualized by automated scanning of the microposts along a three-dimensional coordinate system across multiple planes. The CTC-chip can identify cytokeratin-positive circulating tumor cells in patients with a median yield of 50 cells / mL and a purity range of 1-80% (Nagrath et al., 2007, Nature 450: pp. 1235-1239). Another possibility for isolating CTCs is to use the CellSearch™ Circulating Tumor Cell (CTC) test from Veridex, LLC (Raritan, New Jersey), which captures, identifies, and counts CTCs in the blood vessels.The CellSearch™ system is a U.S. Food and Drug Administration (FDA) approved method for counting circulating tumor cells (CTCs) in whole blood, based on a combination of immunomagnetic labeling and automated digital microscopy. Other methods for isolating CTCs are described in the literature and can all be used in conjunction with the present invention.

[0244] A relapse or recurrence occurs when a person is affected again by a condition that previously affected them. For example, if a patient suffers from a tumor, receives successful treatment for the disease, and then develops the disease again, the newly developed disease may be considered a relapse or recurrence. However, according to the present invention, a relapse or recurrence of a tumor may, but not necessarily, occur at the site of the original tumor. Therefore, for example, if a patient has an ovarian tumor and receives successful treatment, a relapse or recurrence could be the development of an ovarian tumor or a tumor at a site other than the ovary. Furthermore, tumor relapse or recurrence includes situations in which a tumor develops at a site different from the site of the original tumor, as well as situations in which a tumor develops at the site of the original tumor. Preferably, the original tumor that the patient received treatment for is a primary tumor, and the tumor at a site different from the site of the original tumor is a secondary or metastatic tumor.

[0245] The term "immune response" refers to the immune system's reaction to immunogenic organisms such as bacteria or viruses, cells or substances. The term "immune response" includes innate and adaptive immune responses. Preferably, the immune response involves the activation of immune cells, induction of cytokine biosynthesis, and / or antibody production.

[0246] The immune response induced by the composition of the present invention preferably includes the steps of activation of antigen-presenting cells such as dendritic cells and / or macrophages, presentation of an antigen or fragment thereof by the antigen-presenting cells, and activation of cytotoxic T cells by this presentation.

[0247] The term "immune cell" refers to cells of the immune system that are involved in the defense of an individual's body. The term "immune cell" encompasses specific types of immune cells and their precursors, including leukocytes such as macrophages, monocytes (precursors of macrophages), neutrophils, eosinophils, and basophils, dendritic cells, mast cells, and lymphocytes such as B cells, T cells, and natural killer (NK) cells. Macrophages, monocytes (precursors of macrophages), neutrophils, dendritic cells, and mast cells are phagocytic cells.

[0248] The term "immunotherapy" relates to the treatment of a disease or condition by inducing, enhancing, or suppressing an immune response. Immunotherapies designed to induce or amplify an immune response are classified as activation immunotherapies, while immunotherapies that reduce or suppress an immune response are classified as suppression immunotherapies. The term "immunotherapy" includes antigen immunization or antigen vaccination or tumor immunization or tumor vaccination. The term "immunotherapy" also relates to the manipulation of an immune response such that an inappropriate immune response is modulated to a more appropriate response in the context of autoimmune diseases such as rheumatoid arthritis, allergy, diabetes, or multiple sclerosis.

[0249] The term "immunization" or "vaccination" represents the process of administering an antigen to an individual for the purpose of inducing an immune response, for example for therapeutic or prophylactic reasons.

[0250] "Treating" means administering to a subject a compound or composition described herein to prevent or eliminate a disease, including reducing the size or number of tumors in the subject, stopping or delaying the disease in the subject, inhibiting or delaying the onset of a new disease in the subject, reducing the frequency or severity of symptoms and / or recurrence in a subject currently or previously suffering from the disease, and / or prolonging, i.e., increasing, the lifespan of the subject. In particular, the term "treatment of a disease" includes curing the disease or its symptoms, shortening the duration, remission, prevention, delay or inhibition of progression or worsening, or prevention or delay thereof.

[0251] "At risk" means a subject, i.e., a patient, identified as having a higher than normal likelihood of developing a disease, particularly cancer, compared to the general population. Further, a subject who has had, or currently has, a disease, particularly cancer, is a subject with an increased risk of developing the disease because such a subject can continue to develop the disease. Also, a subject who currently has, or has had, cancer also has an increased risk of cancer metastasis.

[0252] Preventive administration of immunotherapy, e.g., preventive administration of the compositions of the present invention, preferably protects the recipient from the onset of the disease. Therapeutic administration of immunotherapy, e.g., therapeutic administration of the compositions of the present invention, can result in inhibition of disease progression / growth. This preferably includes deceleration of disease progression / growth, particularly disruption of disease progression, which results in elimination of the disease.

[0253] Immunotherapy can be performed using any of a variety of techniques in which the drugs provided herein function to remove disease cells from the patient. Such removal can occur as a result of enhancing or inducing an immune response specific to an antigen or cells expressing an antigen in the patient.

[0254] In certain embodiments, the immunotherapy can be active immunotherapy, and the treatment depends on in vivo stimulation of the host's innate immune system using administration of immune response modifiers (such as polypeptides and nucleic acids provided herein) that react to disease cells.

[0255] The drugs and compositions provided herein can be used alone or in combination with conventional treatment regimens such as surgery, irradiation, chemotherapy, and / or bone marrow transplantation (autologous, syngeneic, allogeneic or unrelated).

[0256] The term "in vivo" relates to the situation within a subject.

[0257] The terms “subject,” “individual,” “organism,” or “patient” refer to vertebrates, particularly mammals. For example, in the context of this invention, mammals include humans, non-human primates, domesticated mammals such as dogs, cats, sheep, cattle, goats, pigs, and horses, laboratory animals such as mice, rats, rabbits, and guinea pigs, and animals kept in cages, such as zoo animals. These terms also refer to non-mammalian vertebrates such as birds (particularly domesticated birds such as chickens, ducks, geese, and turkeys) and fish (particularly farmed fish, such as salmon or catfish). Furthermore, the term “animal” as used herein also includes humans.

[0258] The term "autologous" is used to describe all things that originate from the same source. For example, "autotransplantation" refers to the transplantation of tissue or organs from the same source. Such procedures are advantageous in overcoming immunological barriers that would otherwise lead to rejection.

[0259] The term "xenotransplant" is used to describe something consisting of multiple different elements. For example, transferring bone marrow from one individual into another individual constitutes xenotransplantation. Xenogenes are genes that originate from a source other than the target organism.

[0260] Preferably, one or more of the agents described herein are administered as part of a composition for immunization or vaccination, together with one or more adjuvants for inducing or enhancing an immune response. The term “adjuvant” refers to a compound that prolongs, enhances, or accelerates an immune response. The compositions of the present invention preferably exert their effects without the addition of adjuvants. Nevertheless, the compositions of this application may contain any known adjuvants. Adjuvants include a heterogeneous group of compounds such as oil emulsions (e.g., Freund's adjuvant), inorganic compounds (e.g., alum), bacterial products (e.g., Bordetella pertussis toxin), liposomes, and immunostimulatory complexes. Examples of adjuvants include saponins such as monophosphoryl-lipid-A (MPL SmithKline Beecham), QS21 (SmithKline Beecham), DQS21 (SmithKline Beecham, WO96 / 33739), QS7, QS17, QS18, and QS-L1 (So et al., 1997, Mol. Cells 7: pp. 178-186), incomplete Freund's adjuvant, complete Freund's adjuvant, vitamin E, montanid, alum, CpG oligonucleotides (Krieg et al., 1995, Nature 374: pp. 546-549), and various aqueous emulsions in oil prepared from biodegradable oils such as squalene and / or tocopherol.

[0261] Other substances that stimulate the patient's immune response may also be administered. For example, cytokines can be used during vaccination due to their regulatory properties on lymphocytes. Such cytokines include, for example, interleukin-12 (IL-12) (see Hall, 1995, IL-12 at the crossroads, Science 268: pp. 1432-1434), GM-CSF, and IL-18, which have been shown to enhance the protective effects of vaccines.

[0262] There are several compounds that enhance the immune response and can therefore be used in vaccination. These compounds include co-stimulating molecules provided in the form of proteins or nucleic acids such as B7-1 and B7-2 (CD80 and CD86, respectively).

[0263] According to the present invention, a “tumor specimen” is a body specimen containing tumor or cancer cells such as circulating tumor cells (CTCs), in particular a tissue specimen containing body fluids and / or a cellular specimen. According to the present invention, a “non-tumor specimen” is a body specimen that does not contain tumor or cancer cells such as circulating tumor cells (CTCs), in particular a tissue specimen containing body fluids and / or a cellular specimen. Such body specimens can be obtained by tissue biopsy, including punch biopsy, and by conventional methods such as collecting blood, bronchial aspirate, sputum, urine, feces, or other body fluids. According to the present invention, the term “specimen” also includes processed specimens such as fractions or isolates of biological specimens, for example, nucleic acid or cell isolates.

[0264] The therapeutic agents, vaccines, and compositions described herein may be administered by any conventional route, including by injection or infusion. Administration may be, for example, orally, intravenously, intraperitoneally, intramuscularly, subcutaneously, or percutaneously. In one embodiment, administration is performed within a nodule, such as by injection into a lymph node. Other forms of administration envision in vitro transfusion of antigen-presenting cells, such as dendritic cells, using nucleic acids described herein, followed by administration of the antigen-presenting cells.

[0265] The agents described herein should be administered in an effective dose. “Effective dose” means the amount, alone or in combination with additional doses, that achieves the desired response or effect. In the case of treating a particular disease or condition, the desired response preferably relates to inhibiting the course of the disease. This includes delaying the progression of the disease, particularly interrupting or reversing its progression. Alternatively, the desired response in the treatment of a disease or condition may be a delay in the onset of the disease or condition or a prevention of its onset.

[0266] An effective amount of the agent described herein depends on the condition being treated, the severity of the disease, the age, physiological state, size and weight of the patient, including individual parameters, the duration of treatment, the type of concomitant treatment (if any), the particular route of administration and similar factors. Accordingly, the dosage of the agent described herein can vary depending on such parameters. If the response in the patient is insufficient at the initial dosage, higher dosages (or effectively higher dosages achieved by different, more local routes of administration) can be used.

[0267] The term "pharmaceutically acceptable" refers to the non-toxicity of materials that do not interact with the action of the active ingredient of a pharmaceutical composition.

[0268] The pharmaceutical compositions of the present invention can contain salts, buffers, preservatives, carriers and optionally other therapeutic agents. Preferably, the pharmaceutical compositions of the present invention include one or more pharmaceutically acceptable carriers, diluents and / or excipients.

[0269] The term "excipient" is intended to refer to any substance in a pharmaceutical composition that is not an active ingredient, such as binders, lubricants, thickeners, surfactants, preservatives, emulsifiers, buffers, flavorings or colorants.

[0270] The term "diluent" relates to agents for dilution and / or attenuation. Further, the term "diluent" includes any one or more of fluids, liquids or solid suspensions and / or mixed media.

[0271] The term "carrier" relates to one or more compatible solid or liquid fillers or diluents suitable for administration to humans. The term "carrier" relates to natural or synthetic organic or inorganic components that are combined with the active ingredient to facilitate the application of the active ingredient. Preferably, the carrier component is a sterile liquid such as water or oil, including peanut oil, soybean oil, sesame oil, sunflower oil, etc., mineral oil, those derived from animals or plants. Salt solutions and aqueous dextrose and glycerol solutions can also be used as aqueous carrier compounds.

[0272] Pharmaceutically acceptable carriers or diluents for therapeutic use are well known in the pharmaceutical technology field and are described, for example, in Remington's Pharmaceutical Sciences, Mack Publishing Co. (edited by A. R. Gennaro, 1985). Examples of suitable carriers include, for example, magnesium carbonate, magnesium stearate, talc, sugars, lactose, pectin, dextrin, starch, gelatin, tragacanth, methylcellulose, carboxymethylcellulose sodium, low-melting-point waxes, cocoa butter, and others. Examples of suitable diluents include ethanol, glycerol, and water.

[0273] Pharmaceutical carriers, excipients, or diluents can be selected with respect to the intended route of administration and standard pharmaceutical practices. The pharmaceutical compositions of the present invention may contain, or in addition to, suitable binders, lubricants, suspending agents, coating agents, and / or solubilizers as carriers, excipients, or diluents. Examples of suitable binders include natural sugars such as starch, gelatin, glucose, anhydrous lactose, free-flow lactose, beta-lactose, and corn sweeteners, natural and synthetic gums such as acacia and tragacanth, or sodium alginate, carboxymethylcellulose, and polyethylene glycol. Examples of suitable lubricants include sodium oleate, sodium stearate, magnesium stearate, sodium benzoate, sodium acetate, sodium chloride, and others. Preservatives, stabilizers, colorants, and even flavorings may be provided in the pharmaceutical composition. Examples of preservatives include esters of sodium benzoate, sorbic acid, and p-hydroxybenzoic acid. Antioxidants and suspending agents may be used.

[0274] In one embodiment, the composition is an aqueous composition. The aqueous composition may optionally contain a solute, such as a salt. In one embodiment, the composition is in the form of a freeze-dried composition. The freeze-dried composition can be obtained by freeze-drying each aqueous composition.

[0275] The agents and compositions provided herein may be used alone or in combination with other treatment regimens such as surgery, radiation therapy, chemotherapy and / or bone marrow transplantation (autologous, allogeneic, allogeneic, or unrelated).

[0276] The present invention is described in detail with reference to the drawings and examples, which are for illustrative purposes only and are not intended to limit the invention. Thanks to the description and examples, further embodiments similarly included in the present invention are available to those skilled in the art. [Brief explanation of the drawing]

[0277] [Figure 1a] Glioblastoma sample showing high localized amplification of the epidermal growth factor receptor gene (EGFR). Graph showing localized genes surrounding EGFR on chromosome 7. [Figure 1b] Glioblastoma specimens exhibiting high localized amplification of the epidermal growth factor receptor gene (EGFR). A list of 12 single nucleotide polymorphisms with the highest absolute copy number. [Figure 2] A list of numerous genes containing disease-specific mutations in melanoma samples selected by conjugation status (Cx). [Figure 3] This is a list of numerous genes in a melanoma sample where all copies have disease-specific mutations (conjugation rate equal to 1), of which three genes are essential genes in humans, or are presumed to be essential genes in humans. Abbreviations in the figure: chr_pos, chromosome position; CN, absolute copy number; Cx, conjugation status; EC, absolute copy number error correction; VAF, variant allele frequency; rho, estimated percentage of tumor cells containing mutated alleles (SNVs); FLRT+u, confidence score in mutation; site classification is the confidence class of the mutation; essential gene, Y, if the gene is found to be essential. [Example 1]

[0278] Targeting disease-specific mutations in genes with high copy number Genomic information from glioblastoma samples (Chin et al., 2008, Comprehensive genomic characterization defines human glioblastoma genes and core pathways, Nature 455: pp. 1061-1068) was analyzed by searching for genes with high copy numbers and at least one copy containing disease-specific mutations. Quality analysis showed high fidelity in copy number assignments for the 11,574 individual genes analyzed, and the ploidy of the sample genome was determined to be 1.95. Figure 1a shows a graphical representation of local genes around the epidermal growth factor receptor (EGFR) on chromosome 7, which were known driver genes and targets for treatment. EGFR in this genome was shown to have 76 error-corrected absolute copy numbers, of which 13 copies contained disease-specific single nucleotide polymorphisms. Figure 1b provides a list of genes in this genomic sample with the highest absolute copy numbers. Four additional genes with absolute copy numbers greater than 2 exist. In fact, EGFR amplification is a known genetic feature of primary glioblastoma (Benito et al., 2009, Neuropathology 30 (4): pp. 392-400), and this gene has been considered as a target for treatment (Taylor, 2012, Curr Cancer Drug Targets. Mar; 12(3): pp. 197-209). [Example 2]

[0279] Targeting disease-specific mutations with high zygote status We searched for genes in which at least one copy has a disease-specific mutation and analyzed exomes obtained from samples of human tumor-derived melanoma cells by focusing on the number of copies of the gene with disease-specific mutations, the total number of gene copies, and whether or not the gene has a mutation. Figure 2 provides a list of genes selected by the conjugation status in which disease-specific mutations are found in multiple copies of the gene. For example, the disease-specific mutation in the OXGR1 gene has the best conjugation status (4), and in particular, since there are a total of 5 copies of the OXGR1 gene, of which 4 copies contain disease-specific mutations, it also has the best conjugation rate of 4 / 5 or 0.8. The list provides 10 additional genes in which 3 out of a total of 4 copies of the gene have mutations, and the disease-specific mutations in these genes have a conjugation rate of 3 / 4 or 0.75. The remaining listed genes have mutations in which 2 out of a total of 3 copies have mutations, resulting in a conjugation rate of 2 / 3 or 0.66. [Example 3]

[0280] Targeting disease-specific mutations present in all copies of essential genes. We analyzed exomes obtained from human tumor-derived melanoma cell samples by searching for genes with identical disease-specific mutations in all copies and determining which of these genes are essential. These genes were determined to be essential by inferring their degree of essentiality in humans based on their known essentiality in mice (Georgi et al., 2013, From mouse to human: evolutionary genomics analysis of human orthologs of essential genes, PLoS Genetics 9 (5):e1003484; Liao et al., 2007, Mouse duplicate genes are as essential as singletons, Trends Genet. 23:378-381). Figure 3 lists numerous genes with identical disease-specific mutations in all copies. Furthermore, three highlighted genes were determined to be essential by inferring their degree of essentiality from mouse data.

Claims

1. A method for ranking the suitability of two or more neoepitopes as disease-specific targets, which result from disease-specific mutations in alleles (mutated alleles) in two or more genes, comprising the steps of: determining the copy number of each of the two or more mutated alleles encoding the two or more neoepitopes in affected cells or a population of affected cells; and ranking the two or more neoepitopes according to their suitability as disease-specific targets. A method including, (a) The copy number of the mutated allele encoding the neoepitope is greater than 2, or (b) The number of copies of the mutated allele encoding the neoepitope is greater than 0.5 of the total number of copies of the nucleotide site to which the mutation is mapped. However, this method demonstrates the suitability of neoepitopes as disease-specific targets.

2. The method according to claim 1, wherein the copy number of the mutated allele encoding each neoepitope is 1 relative to the total copy number of the nucleotide sites to which the mutation is mapped.

3. The method according to claim 1 or 2, wherein the copy number is determined using a heterozygous segment containing at least one heterozygous SNP, and the segment is a reference genome or a predetermined region of the genome based on a fluorescence-based method such as FACS or FISH, spectral karyotyping (SKY), or digital PCR.

4. The method according to claim 3, wherein the heterozygous segment comprises an equal number of each version of heterozygous SNP.

5. The method according to claim 3 or 4, wherein the heterozygous SNP is used to determine the allele-specific copy number of the segment.

6. The method according to any one of claims 1 to 5, wherein the number of copies is relative or absolute, preferably absolute.

7. The method according to claim 6, wherein the absolute copy number is determined using next-generation sequencing (NGS) in combination with a single nucleotide polymorphism (SNP) array.

8. The method according to claim 6 or 7, wherein the absolute number of copies is an error-corrected absolute number of copies.

9. The method according to any one of claims 6 to 8, wherein the absolute copy number or error-corrected absolute copy number is normalized with respect to ploidy, preferably with respect to ploidy of the genome of the affected cell, or chromosomes, or the ploidy of the portion of a chromosome in the affected cell where the mutation or mutated gene is located.

10. A computer-based analytical processing method for analyzing and selecting one or more neoepitopes resulting from disease-specific mutations (mutated alleles) in one or more gene alleles for use in the provision of vaccines, wherein the computer-based analytical processing method (a) A step of identifying one or more neoepitopes that have been determined to be suitable as disease-specific targets by the method described in any one of claims 1 to 5, and (b) A step of ranking one or more neoepitopes according to their suitability as disease-specific targets, Methods that include...

11. The method according to claim 10, wherein the copy number of the mutated allele encoding each neoepitope is 1 relative to the total copy number of the nucleotide site on which the mutation is mapped.