A method for selecting new antigens for the development of personalized anti-cancer vaccines.

JP2026529676APending Publication Date: 2026-09-01LG CHEM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026510153
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-23
Filing Date
2024-08-23
Publication Date
2026-09-01

AI Technical Summary

Benefits of technology

【0140】 上記の発明の特徴が適用された場合、患者に対するNGSデータプロセッシングにより効率的にHLA-typeIとMutation Variantに対するProcessingが可能である。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026529676000001_ABST
    Figure 2026529676000001_ABST
Patent Text Reader

Abstract

This invention provides a method for selecting tumor-specific novel antigens (immunogenic peptides), and applications for the selected tumor-specific novel antigens in the production of personalized anti-cancer vaccines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross-Reference to Related Applications] The present application claims the benefit of priority based on Korean Patent Application No. 10-2023-0110887 filed on August 23, 2023, and all contents disclosed in the document of said Korean patent application are incorporated as a part of the present specification.

[0002] The present application provides a method for selecting neoantigens (tumor-specific immunogenic peptides) and the use of the selected tumor-associated neoantigens for producing personalized anti-cancer vaccines.

Background Art

[0003] As used herein, the term "Neoantigen" refers to an antigen generated by cancer-specific mutations, more specifically, a peptide (surface protein) that contains an antigenic determinant and can be recognized as an antigen as a mutant peptide expressed on the cell surface specifically in cancer. Since it is expressed only in cancer cells, it can be applied to the development of targeted anti-cancer agents. As used herein, a neoantigen is a type of immunogenic peptide, and can also be expressed as "tumor-specific neoantigen", "tumor-specific antigen", or "tumor-specific immunogenic peptide".

[0004] "Personalized Cancer Vaccine" is prepared in the form of protein or mRNA by selecting neoantigens having immunopotentiation ability from among cancer-specific antigens expressed in each individual patient, and then administered to the patient using delivery vehicles such as lipid-based nanoparticles including classifications such as Lipid Nanoparticles (including LNP, Solid Lipid Nanoparticle (SLN), Nanostructured Lipid Carrier (NLC), etc.), liposome-nucleic acid complexes (e.g., lipoplex, liposomes, etc.), it refers to a therapeutic agent that activates the patient's immune system in vivo to eliminate cancer cells, or a treatment method using said therapeutic agent.

[0005] For personalized anti-cancer vaccines to be effective, a key factor is the discovery of new antigens unique to each patient and the prioritization and selection of immunogenic new antigens. Immunogenic new antigens must have binding affinity to human leukocyte antigens (HLA), which function in the immune response to antigen stimulation in patients and in regulating cellular and humoral immunity. When the T cell receptor on T cells in the body can recognize a specific new antigen sequence bound to HLA, immune activation against the new antigen occurs.

[0006] However, because cancer-specific mutations that occur in each patient and the HLA types expressed differ from person to person, it is difficult to discover new immunogenic antigens that can be applied to the development of personalized anti-cancer vaccines. Furthermore, there is a lack of established methods for collecting and selecting data for training new antigen AI prediction models that can predict immunogenic new antigens, and for predicting and selecting patient-applicable new antigens during the development of personalized anti-cancer vaccines, making it difficult to use new antigen prediction models efficiently. Conventionally, in order to select new antigens that can be expected to have anti-cancer efficacy when applied to anti-cancer vaccines, AI prediction models have been developed that measure the binding affinity between immunogenic peptides, including mutations in the patient's cancer cells, and HLA. However, even if the binding affinity between the patient's epitope and HLA is high, it does not necessarily mean that a high immune response is observed in T cells, making it difficult to use the derived immunogenic peptides as vaccines. Furthermore, there was a technique to predict the immunogenicity of T cells by comparing the physicochemical properties of amino acids between peptides containing mutations from cancer cells and peptides containing non-mutations from non-cancer cells. However, amino acid sequence information alone cannot reflect the three-dimensional protein structure, and since the T cell receptor comprehensively recognizes the HLA-peptide complex bound to a new antigen peptide containing mutations from cancer cells to generate an immune response, the binding affinity between HLA and the new antigen alone cannot encompass the diverse immune responses, resulting in a low success rate in predicting immune responses. [Overview of the project] [Problems that the invention aims to solve]

[0007] One example of this application is the provision of a method for selecting tumor-specific immunogenic peptides (novel antigens, tumor-specific antigens) (prediction method, selection method, confirmation method, or method for providing information for selection).

[0008] The method for selecting tumor-specific immunogenic peptides may include (I) a step of constructing an immunogenic peptide prediction model, and (II) a step of predicting and / or selecting tumor-specific immunogenic peptides.

[0009] The above step (I) of constructing an immunogenicity peptide prediction model may include the following steps (a) to (d): (a) Obtaining peptide information from a peptide database, in which case the peptide information may include one or more pieces of information selected from the following: sequence information of peptides derived from tumor cells or tumor tissue and amino acid sequence information of peptides derived from normal cells or normal tissue, class I HLA type information to which the peptides bind, binding affinity information between the peptides and the class I HLA and / or immunogenicity information of the peptides; (b) Obtaining protein sequence information of the class I HLA from a biological sequence database; (c) A step in which the information obtained in steps (a) and (b) above is used to train a model capable of predicting the number of binding affinity scores (BA) for HLA to construct a binding force prediction model and obtain the number of binding points; and (d) The step of learning the information obtained in steps (a) and (b) above into a model capable of predicting immunogenicity score (IMM) to construct an immunogenicity prediction model and obtain an immunogenicity score.

[0010] The peptide database available in step (I)(a) above may be one or more peptide databases selected from the group consisting of IEDB (Immune Epitope Database and Analysis Resource), MHCBN (Comprehensive Database of MHC Binding and Non-binding Peptides), etc., but is not limited to these. Any database containing peptide information, such as amino acid sequence information and HLA-related binding information (e.g., IC50, KD value, etc.) and / or immunogenicity-related information of human-derived peptides (from normal cells / tissues and tumor cells / tissues), may be used without special restrictions.

[0011] The biological sequence databases available in step (I)(b) above may be one or more protein databases selected from the group consisting of NCBI (National Center for Biotechnology Information), Protein Database, MHC Motif Atlas, etc., but are not limited to these. Any database that provides the amino acid sequence of an HLA protein can be used for learning without any special restrictions.

[0012] The information obtained from the aforementioned peptide database and / or biological sequence database can be used to train the predictive models in stages (c) and (d). The (II) tumor-specific immunogenicity peptide prediction and / or selection step may include the following steps (1) to (7): (1) The step of obtaining NGS sequence information from patient-derived tumor cells or tumor tissue and normal cells or normal tissue; (2) A step of obtaining sequence information of tumor-specific peptides having one or more amino acid mutations in the NGS sequence of patient-derived tumor cells or tumor tissue, with a length of seven or more amino acids, by comparing it with the NGS sequence of normal cells or normal tissue; (3) measuring the expression level (TPM; Transcripts per Million) of said tumor-specific peptide or a gene encoding said tumor-specific peptide; (4) applying said tumor-specific peptide sequence information to the binding force prediction model constructed in step (c) to measure (confirm, predict) the binding affinity score (BA) for an HLA molecule; (5) applying said tumor-specific peptide sequence information to the immunogenicity prediction model constructed in step (d) to measure (confirm, predict) the immunogenicity score (IMM); (6) a prioritization step of determining the ranking of tumor-specific peptides based on a numerical value calculated by the following formula: Neoantigen prioritization score=min(TPM, 1)*(W1*BA+W2*IMM) / (W1+W2) [0<W1<1 (or 0.1 ≤ W1 ≤ 0.9) and 0<W2<1 (or 0.1 ≤ W2 ≤ 0.9), which are each independently set; "min(TPM, 1)" means that when 0<TPM ≤ 1, said value is used as it is, and when TPM ≥ 1, calculation is performed with TPM=1 (when TPM is 1 or more (e.g., 100, 1000), the Score may have a large bias, so this is to reduce errors in result analysis caused thereby); and (7) a step of selecting two or more tumor-specific immunogenic peptides based on the determined priority order of said tumor-specific peptides.

[0013] The NGS (next-generation sequencing) sequence information in said steps (1) and (2) may be nucleic acid (gene) sequences and / or sequence information obtained by converting the same into amino acid (protein) sequences. In one example, said step (1) may additionally comprise, after the step of obtaining NGS sequence information, a step of converting said obtained NGS sequence into an amino acid sequence.

[0014] The tumor cells or tumor tissue and normal cells or normal tissue in the preceding steps (1) and / or (2) may, but are not limited to, originate from the same patient.

[0015] The TPM measurement step in step (3) may be performed on tumor cells and / or tumor tissue, and the tumor cells and / or tumor tissue may be obtained from a patient-derived sample or from an available tumor (cancer) cell line. If the TPM measurement in step (3) is performed on patient-derived tumor cells or tumor tissue, the tumor cells or tumor tissue in steps (1) and / or (2) may be from the same patient as the normal cells or normal tissue, but are not limited to this.

[0016] The TPM measurement step (3), the binding site measurement step (4), and the immunogenicity score measurement step (5) described above do not have any particular restrictions on the order in which they are performed. For example, they may be performed sequentially regardless of the order, or two or more may be performed simultaneously, but are not limited to these. The two or more selected tumor-specific immunogenic peptides can be used as candidate substances for the final tumor-specific immunogenic peptide (novel antigen).

[0017] In one example, the method for selecting tumor-specific immunogenic peptides provided herein may include the previously described (II) prediction and / or selection steps for tumor-specific immunogenic peptides (i.e., steps (1) to (7)).

[0018] In other examples, the method for selecting tumor-specific immunogenic peptides may additionally include (I) a step of constructing an immunogenic peptide prediction model (i.e., steps (a) to (d)) in addition to the (II) step of predicting and / or selecting tumor-specific immunogenic peptides. In this case, the additionally included (I) step of constructing an immunogenic peptide prediction model may be included before the (II) step of predicting and / or selecting tumor-specific immunogenic peptides.

[0019] For the peptide database used for constructing the prediction model in the step (I) of the selection method, data obtained through one or more selection processes selected from the following i) to iv) may be used: i) deleting data that do not correspond to HLA-A, HLA-B, and HLA-C, ii) deleting data having a mutation in an HLA molecule, iii) deleting data in which a peptide has a length of 7 or less or 15 or more, and iv) removing those including amino acid residues other than 20 types of human-derived amino acids in an amino acid sequence.

[0020] The step (I) (one or more of steps (a) to (d)) and step (II) (one or more of steps (1) to (7)) may be performed by an electronic data processing device such as a computer, or an electronic processing apparatus.

[0021] The peptide database usable for training the binding force prediction model in the step (I) may be one or more peptide databases selected from the group consisting of IEDB (Immune Epitope Database and Analysis Resource), MHCBN (Comprehensive Database of MHC Binding and Non-binding Peptides), and the like, but is not limited thereto. Any database including peptide information, for example, amino acid sequence information of human-derived peptides, binding information related to HLA (e.g., IC50, KD value, etc.) and / or immunogenicity-related information can be used for training without particular limitation.

[0022] In one specific example, the step of confirming (measuring, determining) an immunogenicity score (IMM) includes two or more immunogenicity model training data related to an immune reaction when a peptide-MHC (HLA) complex binds to a T cell receptor (TCR).

[0023] The immunogenicity model training data may consist of two or more selected from a group comprising all immunogenicity-related information for the sequences of candidate peptides obtained from peptide databases or relevant literature, such as T cell-derived cytokine secretion levels including INF-g, TNF-a, IL-2, IL-1b, etc., T cell proliferation due to immunogenicity activation, chemokine levels including CCL4, CXCL9, and values ​​of activated T cell cytotoxicity such as Granzyme B. Specifically, the peptide database may be one or more peptide databases selected from the group comprising IEDB (Immune Epitope Database and Analysis Resource) and MHCBN (Comprehensive Database of MHC Binding and Non-binding Peptides), but is not limited to this. Any database containing peptide information, such as amino acid sequence information and immunogenicity measurements of human-derived peptides, can be used for training without any special restrictions.

[0024] The immunogenicity measurements mentioned above are values ​​obtained by methods including ELISPOT, ELISA, flow cytometry intracellular staining, binding assay, proliferation assay (e.g., T cell proliferation assay), and cytotoxicity assay (e.g., in vitro / ex vivo / in vivo cytotoxicity assay).

[0025] The immunogenicity model training data includes all qualitatively and quantitatively measured values, of which the qualitatively measured values ​​are included as data after being statistically quantified, such as by beta distribution transformation, using the negative control group comparison value.

[0026] Step (II) may be performed using NGS sequence information and / or information obtained by converting it into peptide information from samples (tumor cells / tissues and normal cells / tissues) derived from the individualized patient (mammal; for example, human).

[0027] The tumor-specific immunogenic peptide may be a peptide derived from cancer cell mutations. For example, the tumor-specific immunogenic peptide may be derived from a sequence obtained by genetic sequencing analysis (e.g., Next Generation Sequencing (NGS)) of a patient-derived sample (e.g., cancer cells) (e.g., a nucleic acid sequence obtained from NGS or an amino acid sequence converted therefrom). More specifically, the tumor-specific immunogenic peptide may be derived from a sequence (nucleic acid sequence or an amino acid sequence converted therefrom) obtained by genetic sequencing analysis (e.g., Next Generation Sequencing (NGS)) of a patient-derived sample (e.g., cancer cells) that differs from that of normal cells (e.g., a sequence in which a mutation has occurred compared to normal).

[0028] The HLA molecule may be one or more selected from the group consisting of major histocompatibility complex (MHC) molecules of human origin. The HLA molecule may contain different HLA molecular types and / or different HLA alleles. In one example, the HLA molecule may be an HLA type I molecule.

[0029] In the stage of constructing the immunogenicity peptide prediction model (stage (I)), the number of binding points in stage (c) may be obtained by combining (ensemble) predicted values ​​obtained by training using two or more methods with the amino acid sequence information of the immunogenicity peptide and the HLA molecule. The amino acid sequence information of the HLA molecule may be obtained from the candidate peptide database described above, but is not limited to this.

[0030] The immunogenicity score in step (d) above may be obtained by the amount of cytokine secreted when the immunogenic peptide is applied to a sample containing immune cells. The sample containing immune cells is selected from the group consisting of blood, leukocytes, and peripheral blood mononuclear cells (PBMCs), and the cytokine may be one or more selected from the group consisting of IFN-g, IL-2, TNF-α, and granzyme B, but is not limited thereto. In one example, the immunogenicity score in step (2) above may be determined by the amount of cytokine secreted when the candidate peptide is applied to immune cells (e.g., PBMCs, etc.) or a sample containing them. The amount of cytokine secreted can be obtained from a database selected from the group consisting of, for example, IEDB, MHCBN, etc., but is not limited thereto.

[0031] The immunogenicity score can be indirectly measured as the ratio of positive experiments to the total number of experiments. In one example, the IEDB database provides qualitative measurements, the number of experiments (tested), the number of responses (responded), and the frequency as data instead of quantitative measurements of cytokine secretion, and this can be used as training data for quantifying the immunogenicity score. Actual experimental values ​​can be converted numerically to reflect the ratio of the number of responses to the number of experiments and used as training data when retraining the immunogenicity prediction model. This quantification may be performed by one or more methods selected from the group consisting of mean, median, and beta distribution transformation.

[0032] The immunogenicity score in step (d) above may be obtained by combining (ensemble) predicted values ​​obtained by learning using two or more methods with the peptide amino acid sequence information of the immunogenicity learning data, the amino acid sequence corresponding to the HLA molecule of the immune cell, and the cytokine secretion amount information when the peptide is processed by the immune cell.

[0033] The combination (ensemble) of steps (c) and (d) may involve selecting two or more predictive models, each trained with two or more different learning methods, that have the smallest prediction error on a validation dataset different from the training data, and combining the predicted values. Specifically, the validation dataset may be a portion of the model training data that has been separated in advance.

[0034] In step (c) (binding prediction modeling and scoring) and / or step (d) (immunogenicity prediction modeling and scoring), - The aforementioned learning may be performed using one or more models selected from commonly used AI (artificial intelligence) learning models such as GRU (Gated Recurrent Unit), Transformer, FFN (Feed Forward Neural-network), and GNN (Graph Neural Network). -The learning in stages (c) (measuring the number of binding points) and (d) (measuring the immunogenicity score) may be performed by combining one learning method selected from among the application and non-application of learning methods commonly used in the relevant field with the AI ​​model. - The number of binding points and / or immunogenicity points may be the final predicted value obtained by combining (ensemble) the predicted values ​​obtained from each of the prediction models that were trained using one or more AI models and one or more learning methods using a normal ensemble model. For example, the final predicted value may be obtained as [(mean of predicted values) - (variance of predicted values)].

[0035] In the tumor-specific immunogenicity peptide prediction and / or selection step (step (II)), step (1) may be performed by a bioinformatics analysis process between the NGS sequences of patient-derived tumor cells or tumor tissue and the NGS sequences of patient-derived normal cells or normal tissue. The tumor cells or tumor tissue and the normal cells or normal tissue may originate from the same patient. The bioinformatics analysis process may include, but is not limited to, steps i) through vii) below, and further steps may be added or removed.

[0036] i) Verification of the integrity of NGS data; for example, software such as FastQC can be used, but is not limited to this. ii) Removal of low-quality data; for example, software such as Cutadapt can be used, but is not limited to this. iii) Sequence alignment to a reference genome; for example, the hg19 reference genome and BWA software can be used, but are not limited to these. iv) Quality verification and correction of sequence alignment results; for example, indel realignment, BQSR (Base Quality Score Recalibration), etc. can be applied using software such as GATK, but are not limited to this. v) Tumor-specific mutation detection; for example, software such as GATK, Strelka, and Mutect2 may be used, but are not limited to these; vi) Classification of tumor-specific mutations; for example, mutations can be classified using software such as SnpEff, and only missense, insertion, and deletion mutations that affect the peptide sequence can be selected, but this is not limited to this. vii) Extraction of peptide sequences of length 7 or longer that can be generated containing tumor-specific mutations

[0037] The TPM (transcriptome per million) in step (3) above is an expression value that measures the degree of expression in the process of gene expression, which is the process in which a gene composed of base sequences is transcribed and a transcript is formed. Methods for measuring the expression value include measuring mRNA concentration and measuring the concentration of the produced transcript, but are not limited to these, and the expression value can be measured by a variety of methods. In one embodiment of the present invention, the expression value is the number of reads obtained from the NGS results mapped to the expression body (transcriptome), and can be calculated in TPM units, but is not limited to this, and the expression value can be shown by calculating it in a variety of units. The step of filtering RNA base sequences with a TPM (transcriptome per million) of 1 or more from mutant base sequences or whole RNA base sequences may be further included.

[0038] The expression level of the tumor-specific peptide or its coding gene in step (3) above (e.g., TPM, Transcripts Per Million; Li et al. (2010); normalized based on the proportion of each transcript (number of mapped NGS reads) in the library, where the sum of all TPM values ​​in one library equals 1,000,000, allowing for comparison between different samples) may represent the likelihood that the tumor-specific peptide is present in an actual patient sample (e.g., cancer cells) (e.g., the likelihood, frequency, or extent of gene mutations in cancer cells being produced into proteins through transcription and translation). The expression level of the tumor-specific peptide or its coding gene in step (3) above may be obtained by sequence analysis of patient sample data targeted for personalized therapy, or more specifically, calculated by processing microarray or RNA-seq data. In one example, patient sample RNA-seq data can be analyzed to calculate the TPM of the tumor-specific peptide or its coding gene, and priority can be assigned to the tumor-specific peptide with a TPM value of 1 or greater. For example, since TPM represents the expression level of the gene per million RNA-seq data normalized by gene length and data production, by using TPM ≥ 1 for assigning priority, priority can be assigned to candidate peptides that account for at least one millionth of the total RNA-seq data and are likely to be present in actual patient samples.

[0039] The priority determination step (6) above is performed using the following formula: Neoantigen prioritization score=min(TPM, 1)*(W1*BA+W2*IMM) / (W1+W2). In the above formula, W1 and W2 each represent the weighting of binding score (BA) and immunogenicity score (IMM), respectively, and can be independently set within the ranges of 0<W1<1 (or 0.1≤W1≤0.9) and 0<W2<1 (or 0.1≤W2≤0.9), respectively. In one example, W1+W2=1, but this is not limitative.

[0040] In one example, the calculation of the above formula may be performed on candidate peptides that satisfy the conditions of TPM>1, 0.6>binding score (BA)≥0.1, and immunogenicity score (IMM)≥0.25, but this is not limitative.

[0041] The step of selecting tumor-specific immunogenic peptides (neoantigens) in said step (7) may additionally comprise the step of sorting the results of said step (6) in descending order, and selecting one or more tumor-specific immunogenic peptides from the top 2 to 100, 2 to 75, 2 to 50, 2 to 40, 2 to 30, 2 to 20, 2 to 10, 5 to 100, 5 to 75, 5 to 50, 5 to 40, 5 to 30, 5 to 20, 5 to 10, 10 to 100, 10 to 75, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 20 to 100, 20 to 75, 20 to 50, 20 to 40, or 20 to 30 peptides.

[0042] When the neoantigen (tumor-specific immunogenic peptide) prediction and individualized anti-cancer vaccine provided herein is prepared from mRNA, it may not be necessary to particularly consider hydrophobicity; however, when an anti-cancer vaccine is prepared with peptides or when an immunogenicity test is performed, if the hydrophobicity of the peptide is high (e.g., the content of hydrophobic amino acids is high), production cannot be achieved, so it may be necessary to measure the hydrophobicity of the peptide. In one specific embodiment, said hydrophobicity can be measured by calculating the Grand average of hydropathy (Gravy) score, but this is not limited thereto.

[0043] The tumor-related immunogenic peptides selected by the tumor-specific immunogenic peptide selection method provided herein may be peptides (e.g., mutant peptides) that are specifically present in cancer (tumor) samples and not present in normal samples without cancer (tumor). Therefore, the tumor-specific immunogenic peptides can be usefully used in the production of pharmaceutical compositions for the prevention and / or treatment of cancer, such as anti-cancer vaccines, in particular, personalized anti-cancer vaccines for patients used in the selection method described herein.

[0044] Another example is an anti-cancer vaccine composition comprising a tumor-specific immunogenic peptide (novel antigen) selected by the tumor-associated immunogenic peptide selection method and / or a gene encoding it (mRNA, DNA, etc.). The anti-cancer vaccine composition may further contain a pharmaceutically acceptable carrier and / or excipient and / or one or more adjuvants, stabilizers, etc.

[0045] Another example provides a method for producing an anti-cancer vaccine, comprising the step of mixing a tumor-specific immunogenic peptide selected by the tumor-specific immunogenic peptide selection method and / or a gene encoding it (mRNA, DNA, etc.) with a pharmaceutically acceptable carrier and / or one or more adjuvants, stabilizers, etc.

[0046] The carrier and / or excipient and / or one or more adjuvants, stabilizers, etc., may be selected from all pharmaceutically acceptable substances commonly used in vaccine formulation.

[0047] The vaccine may be in the form of a therapeutic vaccine or a prophylactic vaccine.

[0048] Another example provides a method for preventing and / or treating cancer, comprising the step of administering a tumor-specific immunogenic peptide selected by the tumor-specific immunogenic peptide selection method or the vaccine to a patient in need of cancer prevention and / or treatment.

[0049] The aforementioned treatment of cancer may be combined with other anticancer agents and / or with surgical resection and / or radiotherapy and / or conventional chemotherapy. The aforementioned other anticancer agents can be selected from all known anticancer agents (chemical drugs, antibodies, etc.) and may, but are not limited to, immunosuppressants (immune checkpoint inhibitors (e.g., PD-1 inhibitors and / or PD-L1 inhibitors, etc.)). [Means for solving the problem]

[0050] To discover new antigens with strong immunogenicity among cancer-specific sequence mutations, we developed a novel antigen prediction model using AI learning technology that combines a Binding Affinity model, which predicts the binding affinity between HLA and new antigen sequences, and an Immunogenicity model, which predicts the immunogenicity of new antigen sequences. The new antigen prediction model was trained using algorithms from the Immune Epitope Database and Analysis Resource (IEDB) database, and in silico improvements were made through data updates.

[0051] By training an AI-based new antigen prediction model, we enabled the prediction of immunogenic new antigen candidates. Using this prediction model, we applied NGS data from cancer patients to derive a group of new antigen candidates. Thresholds were defined for the new antigen candidate groups predicted by the Binding Affinity and Immunogenicity models used when applying the prediction model. Peptides were then prepared for the new antigens, and Binding Affinity and immunogenicity tests (e.g., IFNγ ELISPOT) were conducted to validate the new antigen prediction rate of the model and select new antigens applicable to personalized anti-cancer vaccines.

[0052] One aspect of the structure of this application, as an example, is as follows: New antigen prediction model (Transformer-based Binding Affinity Model + Immunogenity Model); Selection of input data for a new antigen prediction model; Learning method for new antigen prediction models; Data processing methods for patient datasets: WES (whole exosome sequencing; mutation selection) and RNA-seq (expression level) data. Neoantigen Prioritization (Checking Neoantigen Ranking Score using TPM, BA Score, and IMM Score); Validating the applicability of predicted novel antigens for personalized anti-cancer vaccines by conducting binding affinity and immunogenicity testing.

[0053] Definition of Terms In this invention, the term "peptide" means a substance containing two or more, three or more, four or more, six or more, eight or more, ten or more, thirteen or more, sixteen or more, twenty-one or more, and up to eight, ten, twenty-two, twenty-one or more amino acids covalently linked by peptide bonds. The terms "polypeptide" or "protein" may mean a substance in which more amino acid residues than peptides are covalently linked by peptide bonds, but in this specification, "peptide," "polypeptide," and "protein" are interchangeable.

[0054] In the present invention, the term "modification" or "mutation" in relation to peptides, polypeptides, or proteins can mean that a mutation has been introduced into the parental sequence of a wild-type peptide, polypeptide, or protein, and such mutation may be selected from the group consisting of amino acid insertion mutations, amino acid addition mutations, amino acid deletion mutations, and amino acid substitution mutations, for example, an amino acid substitution mutation. In this specification, all sequence changes can potentially create new epitopes.

[0055] In this specification, the term "immunogenic peptide" means a peptide that contains an antigenic determinant (epitope) and can induce an immune response.

[0056] In this specification, the term "neoantigen" may refer to a peptide or protein containing one or more amino acid modifications compared to a parental peptide or protein, and such amino acid modifications may be due to tumor-specific mutations. In this specification, a neoantigen may be a type of immunogenic peptide containing tumor (cancer)-specific mutations that are not present or detectable in normal cells. The neoantigen may also be expressed as a tumor-specific neoantigen (tumor-specific antigen) or tumor-specific immunogenic peptide.

[0057] The term "immune response" refers to an integrated bodily reaction to a target such as an antigen, and can mean, for example, a cellular immune response or a cellular and humoral immune response. An immune response may be protective / preventive / protective and / or therapeutic.

[0058] "Inducing an immune response" can mean that there was no immune response before induction, but it can also mean that there was a certain level of immune response before induction, and that the immune response was enhanced after induction. Therefore, "inducing an immune response" also includes the meaning of "enhancing an immune response." For example, after inducing an immune response in a given subject (patient), the subject may be protected from the onset of a disease such as cancer, or the state of the disease may be improved, cured, and / or alleviated by the induction of the immune response.

[0059] The terms “cellular immune response” and “cellular response” or similar terms refer to a cell-directed immune response characterized by the presentation of antigens by class I or class II MHC associated with T cells or T lymphocytes acting as “helper” or “killer” cells. Helper T cells (also called CD4+ T cells) play a central role in regulating the immune response, while killer cells (also called CD8+ T cells, cytotoxic T cells, cytolytic T cells, CD8+ T cells, or CTLs) kill disease cells such as cancer cells and prevent the development of diseased cells. In preferred specific examples, the present invention may include stimulating an antitumor CTL response against tumor cells expressing one or more tumor-expressing antigens, and / or presenting such tumor-expressing antigens by HLA type I or class I MHC.

[0060] "Antigen" encompasses all substances, such as peptides or proteins, that are targets of and / or induce immune responses, such as specific reactions with antibodies or T lymphocytes (T cells). For example, an antigen includes at least one epitope, such as a T cell epitope. In this specification, an antigen may optionally be a molecule (including a cell expressing the antigen) that, after processing, induces an immune response specific to that antigen. An antigen or its T cell epitope may be presented in terms of MHC molecules by an antigen-presenting cell, such as a disease cell, particularly a cancer cell, to provoke an immune response against the antigen (including a cell expressing the antigen).

[0061] In this specification, the terms "tumor-specific antigen," "tumor-specific expression antigen," "cancer-specific antigen," and "cancer-specific expression antigen" can be used interchangeably with the term "novel antigen."

[0062] The term "immunogenicity" can refer to the ability to induce a therapeutic immune response, such as in the treatment of cancer (e.g., antibody formation).

[0063] The term "major histocompatibility complex (MHC)" refers to a gene complex that includes MHC class I and MHC class II molecules and occurs in all vertebrates. MHC proteins or molecules are important in signaling between lymphocytes and antigen-presenting or diseased cells in immune responses, where MHC proteins or molecules bind to peptides and present them for recognition by T cell receptors. Proteins encoded by MHC are expressed on the cell surface to present both self-antigens (peptide fragments from the cell itself) and non-self-antigens (e.g., fragments from invading microorganisms) to T cells.

[0064] The MHC region is divided into three subgroups: Class I, Class II, and Class III. MHC Class I proteins contain α-chains and β2-microglobulin (not part of the MHC encoded by chromosome 15). These present antigen fragments to cytotoxic T cells. On most immune system cells, particularly antigen-presenting cells, MHC Class II proteins contain α and β-chains, which present antigen fragments to T-helper cells. The MHC Class III region encodes other immune components, such as complement components and cytokines.

[0065] MHC is polygenic (having several MHC class I and MHC class II genes) and polymorphic (each gene has multiple alleles).

[0066] In this invention, the term "haplotype" refers to the HLA alleles found on a single chromosome and the protein it encodes. A haplotype may also refer to an allele located at a specific position within the MHC.

[0067] Each class of MHC can be indicated by several positions: for example, - For Class I, HLA (Human Leukocyte Antigen)-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-P, HLA-V, etc.; and - In the case of Class II, these include HLA-DR (HLA-DRA, HLA-DRB1-9, etc.), HLA-DQ (HLA-DQA1, HLA-DQB1, etc.), HLA-DP (HLA-DPA1, HLA-DPB1, etc.), HLA-DM (HLA-DMA, HLA-DMB, etc.), HLA-DO (HLA-DOA, HLA-DOB, etc.), etc.

[0068] In this specification, the terms "HLA or HLA allele" and "MHC or MHC allele" are interchangeable.

[0069] MHCs exhibit very strong pleomorphism. In human populations, there are numerous haplotypes at each gene site, each containing distinct alleles. Different pleomorphic MHC alleles in both Class I and Class II have different peptide specificities: each allele encodes a protein that binds to peptides exhibiting a specific sequence pattern.

[0070] In one specific example, the MHC molecule may be an HLA molecule. When expressing the HLA molecule by its serological type based on the type of HLA DNA, it may be expressed by referring to the "HLA dictionary 2008" (John Wiley & Sons A / S; Tissue Antigens 73, 95-170; included herein by reference).

[0071] The term “MHC-binding peptide” may include MHC class I and / or class II binding peptides or peptides that can be processed to produce MHC class I and / or class II binding peptides. In the case of a class I MHC / peptide complex, the peptide (binding peptide) may generally be 7-12, 7-10, 8-12, or 8-10 amino acid lengths, but is not limited thereto. In the case of a class II MHC / peptide complex, the peptide (binding peptide) may generally be 9-30 or 10-25 amino acid lengths, for example, 13-18 amino acid lengths, but is not limited thereto.

[0072] In one example, if the peptide is presented directly, that is, without processing, particularly without cleavage, its length may be a length suitable for binding to MHC molecules, particularly class I MHC molecules, for example, 7-30, 7-20, 7-12, 8-11, 9, or 10 amino acid lengths. In another example, if the peptide is part of a larger entity containing an additional sequence, such as a vaccine sequence or polypeptide, and is presented after processing, particularly cleavage, the length of the peptide produced by processing may be a length suitable for binding to MHC molecules, particularly class I MHC molecules, for example, 7-30, 7-20, 7-12, 8-11, 9, or 10 amino acid lengths.

[0073] Tumor-specific immunogenic peptides selected by the methods provided herein may bind to class I MHC (HLA) molecules.

[0074] The term "epitope" refers to an antigenic determinant within a molecule, such as an antigen, that is, a part or fragment of a molecule recognized by the immune system. An epitope of a protein, such as a tumor antigen, comprises a continuous or discontinuous portion of the protein, and its length may be 5–100, 5–50, 8–30, or 10–25 amino acids. In this specification, an epitope may also be a T-cell epitope.

[0075] In this specification, the epitope may be an "MHC-binding peptide" that can bind to MHC molecules on the cell surface.

[0076] In this specification, the term “neo-epitope” means an epitope that is not present in a reference, such as a normal non-cancerous or germ cell, but is found in cancer cells. This includes, in particular, situations in which a corresponding epitope is found in a normal non-cancerous or germ cell, but one or more mutations in the cancer cell alter the sequence of the epitope, leading to the neo-epitope.

[0077] In this invention, the term "T cell epitope" refers to a peptide that binds to an MHC molecule in a sequence (configuration) recognized by the T cell receptor. Generally, T cell epitopes are located on the surface of antigen-presenting cells.

[0078] A T cell epitope may contain an amino acid sequence substantially corresponding to the amino acid sequence of a certain antigen fragment. For example, the antigen fragment may be a peptide presented as MHC class I and / or class II.

[0079] The T cell epitopes according to the present invention may be related to a part or fragment of an antigen that can stimulate an immune response, or to a part or fragment of an antigen that can stimulate a cellular response to an antigen or cell characterized by the expression of that antigen and, preferably, the presentation of the antigen, such as disease cells, particularly cancer cells. For example, the T cell epitope may be able to stimulate a cellular response to a cell characterized by the presentation of an antigen by class I MHC, or it may be able to stimulate antigen-reactive cytotoxic T lymphocytes (CTLs).

[0080] "Antigen processing" or "processing" means that a peptide, polypeptide, or protein is broken down (for example, a polypeptide is broken down into peptides) and one or more of these fragments associate with an MHC molecule (for example, by binding) so that they are presented to a cell, e.g., an antigen-presenting cell, on a specific T cell.

[0081] The terms “T cell” and “T lymphocyte” are used interchangeably herein and include cytotoxic T cells (CTLs, CD8+ T cells), including cytolytic T cells, and T helper cells (CD4+ T cells).

[0082] T cells belong to a group of white blood cells known as lymphocytes and play a central role in cell-mediated immunity. They can be distinguished from other lymphocytes, such as B cells and natural killer cells, by the presence of a special receptor called the T cell receptor (TCR) on their cell surface. The thymus is the primary organ responsible for the maturation of T cells. Several different T cell subsets have been discovered, each with a unique function.

[0083] T helper cells play a role in assisting other leukocytes in immunological processes, particularly the maturation of B cells into plasma cells and the activation of cytotoxic T cells and macrophages. These cells are also known as CD4+ T cells because they express the CD4 protein on their surface. Helper T cells are activated when presented with peptide antigens by MHC class II molecules expressed on the surface of antigen-presenting cells (APCs). Once activated, they rapidly divide and secrete small proteins called cytokines that regulate or assist the active immune response.

[0084] Cytotoxic T cells destroy virus-infected cells and tumor cells and are also involved in transplant rejection. These cells are also known as CD8+ T cells because they express the CD8 glycoprotein on their surface. These cells recognize their targets by binding to antigens associated with MHC class I, which are present on the surface of almost all cells in the body.

[0085] The vast majority of T cells possess a T cell receptor (TCR), which exists as a complex of several proteins. The actual T cell receptor is generated from independent T cell receptor α and β (TCRα and TCRβ) genes and consists of two distinct peptide chains, referred to as the α and β-TCR chains. γδT cells (γδT cells) represent a small subset of T cells that possess a unique T cell receptor (TCR) on their surface. However, in γδT cells, the TCR consists of one γ-chain and one δ-chain. This group of T cells is far rarer than αβT cells (2% of all T cells).

[0086] The initial signal for T cell activation is provided when the T cell receptor binds to a short peptide presented by MHC on another cell. This ensures that only T cells with a TCR specific to that peptide are activated. Partner cells are generally antigen-presenting cells, such as professional antigen-presenting cells (APCs), and in the case of a naive response, generally dendritic cells, but B cells and macrophages can also be important APCs.

[0087] In this specification, a molecule is said to be able to bind to a given target if it binds to that target in a standard analysis with a significant affinity for that target. "Affinity" or "binding affinity" can be measured by the equilibrium dissociation constant (KD). If a molecule does not have a significant affinity for the target and does not bind to the target significantly in a standard analysis, then the molecule is said to be unable to bind to the target (substantially).

[0088] Specific activation of CD4+ or CD8+ T cells can be confirmed in various ways. Methods for confirming specific T cell activation include examining T cell proliferation, cytokine (e.g., lymphokine) production, or the generation of cytolytic activity. In the case of CD4+ T cells, specific T cell activation can be confirmed by T cell proliferation. In the case of CD8+ T cells, specific T cell activation can be confirmed by cytolytic activity.

[0089] In this specification, the term “score” generally refers to the result of a test or examination expressed numerically.

[0090] In the present invention, confirming the number of binding sites of a predetermined peptide to an MHC molecule includes detecting the possibility of the peptide binding to the MHC molecule.

[0091] The number of peptide binding sites to the aforementioned MHC molecule can be determined using a peptide:MHC binding prediction tool. For example, the Immunoepitope Database Analysis Resource (IEDB-AR: http: / / tools.iedb.org) can be used.

[0092] The amino acid deformations whose immunogenicity is sought by this invention may be caused by mutations in cellular nucleic acids. Such mutations may be identified by known sequencing techniques.

[0093] In one specific example, cancer-specific situated mutations or sequence differences can be detected in the exome of a tumor sample, for example, the whole exome. Therefore, the present invention may include identifying cancer mutation indicators in the exome, preferably the whole exome, of one or more cancer cells. In one specific example, the step of identifying cancer-specific situated mutations in a tumor sample from a cancer patient includes identifying an exome-wide cancer mutation profile.

[0094] In one specific example, cancer-specific situated mutations or sequence differences can be detected in the transcriptome of a tumor sample, e.g., the whole transcriptome. Therefore, the present invention may include identifying cancer mutation indicators in the transcriptome of one or more cancer cells, e.g., the whole transcriptome. In one specific example, the step of identifying cancer-specific situated mutations in a tumor sample from a cancer patient includes identifying a transcriptome-wide cancer mutation profile.

[0095] In one specific example, the step of identifying cancer-specific somatic mutations or identifying sequence differences includes single-cell sequencing of one or more cancer cells, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more. Therefore, the present invention may include identifying cancer mutation signs in one or more cancer cells. In one specific example, the cancer cells are circulating tumor cells. Cancer cells such as circulating tumor cells can be isolated before single-cell sequencing.

[0096] In one specific example, the step of identifying cancer-specific somatic mutations or the step of identifying sequence differences may be performed using next-generation sequencing (NGS).

[0097] In one specific example, the step of identifying cancer-specific somatic mutations or sequence differences may include sequencing the genetic DNA and / or RNA of the tumor sample.

[0098] To identify cancer-specific somatic mutations or sequence differences, it is preferable to compare sequence information obtained from tumor samples with reference information, such as sequence information obtained by sequencing nucleic acids like DNA or RNA from normal non-cancerous cells, such as germ cells, obtained from the patient or other individuals. In one specific example, the normal hereditary germ cell DNA may be obtained from peripheral blood mononuclear cells (PBMCs).

[0099] The term "genome" refers to the total amount of genetic information contained within the chromosomes of an organism or cell.

[0100] The term "exome" refers to a portion of an organism's genome formed by exons, which are the coding parts of expressed genes. Exomes provide a blueprint for the genes used to synthesize proteins and other functional gene products. Because it is the most functionally relevant part of a gene, it is likely to contribute most significantly to an organism's phenotype. The human genome's exome is estimated to account for 1.5% of the entire genome.

[0101] The term “transcriptome” refers to the set of all RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNAs, produced in a single cell or a population of cells. In this specification, the transcriptome can mean the set of all RNA molecules produced in a single cell, a population of cells, such as a cancer cell population, or all cells of a given individual at a particular point in time.

[0102] In the present invention, the term "nucleic acid" may mean deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Nucleic acids include hereditary DNA, cDNA, mRNA, recombinantly produced molecules, and chemically synthesized molecules according to the present invention. In the present invention, nucleic acids may exist as single-stranded or double-stranded molecules and linear or covalently converted molecules.

[0103] The term "mutation" refers to a change or difference in the nucleic acid sequence compared to a reference (nucleotide substitution, addition, or deletion) and / or a change in the amino acid sequence of the peptide (protein) encoded by such a change in the nucleic acid sequence. "Somatic mutations" can occur in all somatic cells except germ cells (sperm and egg cells) and are therefore not transmitted to children. Such changes can (but not always) induce cancer or other diseases. Preferably, the mutation is a nonsynonymous mutation. The term "nonsynonymous mutation" refers to a mutation that causes a change in amino acids, such as an amino acid substitution, in the translation product, preferably a nucleotide substitution.

[0104] In the present invention, the term "mutation" may include point mutations, indels, fusions, chromothripsis, base editing, and / or RNA edits.

[0105] The term "cancer mutation signature" refers to the set of mutations present in cancer cells when compared to non-cancer reference cells.

[0106] The aforementioned "reference" is used to relate and compare the results obtained from tumor samples using the method of the present invention. Generally, the "reference" is obtained based on one or more normal samples, for example, samples obtained from individuals of the same species that were not affected by cancer.

[0107] All suitable sequencing methods can be used for mutation detection, for example, next-generation sequencing (NGS). To increase the speed of the sequencing process, in the future, NGS technology should be replaceable by third-generation sequencing methods. The term “next-generation sequencing” or “NGS” refers to all novel, high-performance sequencing technologies, in contrast to “conventional” sequencing methodologies known as Sanger chemistry, which read nucleic acid templates randomly and equilibrium by the whole genome by decomposing the whole genome into smaller fragments. Such NGS technologies (also known as massively parallel sequencing technologies) can transmit nucleic acid sequence information of the whole genome, exome, transcriptome (all transcribed sequences of the genome), or methylome (all methylated sequences of the genome) within a very short period, e.g., 1-2 weeks, 1-7 days, or within 24 hours, essentially enabling a single-cell sequencing approach. Commercially available or literature-referenced multiplex NGS platforms can be used.

[0108] The term “RNA” refers to a molecule comprising at least one ribonucleotide residue, preferably consisting solely of or substantially composed of ribonucleotide residues. “Ribonucleotide” refers to a nucleotide having a hydroxyl group at the 2' position of the β-D-ribofuranosyl group. The term “RNA” includes double-stranded RNA, single-stranded RNA, isolated RNA, e.g., partially or completely purified RNA, essentially pure RNA, synthetic RNA, and recombinantly produced RNA, e.g., modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution, and / or alteration of one or more nucleotides. Such alterations may include the addition of non-nucleotide substances to the terminal(s) or internally of RNA, e.g., to one or more nucleotides of RNA. Nucleotides within an RNA molecule may also include non-standard nucleotides, e.g., non-spontaneous nucleotides or chemically synthesized nucleotides or deoxynucleotides. These modified RNAs may also be referred to as analogs or analogs of naturally occurring RNA.

[0109] The term "RNA" may include "mRNA." The term "mRNA" refers to "messenger RNA" that is produced using a DNA template and codes for a peptide or polypeptide; it can also be expressed as a "transcript." Generally, mRNA contains a 5'-UTR, a protein-coding region, and a 3'-UTR. mRNA has only a restrictive half-life in cells and in vitro. mRNA can be produced from a DNA template by in vitro transcription. In vitro transcription methodologies are generally known to technicians.

[0110] Terms such as “decrease” or “suppress” relate to the ability to cause an overall decrease, for example, an overall decrease of 5%, 10%, 20%, 50%, or 75% or more. The term “suppress” or similar expressions include complete or nearly complete suppression, i.e., a decrease to zero, or a decrease to essentially zero.

[0111] Terms such as “increase,” “boost,” “promote,” or “extend” relate to increases, boosts, promotes, or extensions of levels of, for example, at least about 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 80%, at least 100%, at least 200%, or at least 300%. These terms also relate to increases, boosts, promotes, or extensions from levels of 0 or unmeasurable or undetectable levels to levels greater than 0 or measurable or detectable levels.

[0112] The term "vaccine" relates to a pharmaceutical preparation (pharmaceutical composition) or product that, when administered, induces an immune response, particularly a cellular immune response, that recognizes and attacks disease cells or pathogens, such as cancer cells. Vaccines can be used to prevent or treat diseases. The terms "personalized cancer vaccines" or "individualized cancer vaccines" relate to specific cancer patients and refer to cancer vaccines being tailored as needed to suit the individual cancer patient's specific circumstances.

[0113] In one specific example, the vaccine provided by the present invention may include one or more peptides selected as immunogenic peptides by the method of the present invention, or nucleic acids (e.g., RNA) encoding said peptides.

[0114] The anti-cancer vaccine provided by the present invention provides one or more T cell epitopes suitable for stimulating, priming, and / or expanding T cells specific to the patient's tumor upon administration to the patient. The T cells may be directed to cells expressing the antigen from which the T cell epitope originates. Thus, the vaccine described in the present invention may be able to induce or enhance a cellular response, such as the activity of cytotoxic T cells, against cancerous diseases characterized by the presentation of one or more tumor-associated neogenic antigens by class I MHC. Since the vaccine provided by the present invention targets cancer-specific mutations, it may be specific to the tumor of the patient in question.

[0115] The vaccines provided in the present invention, when administered to a patient, provide one or more T cell epitopes associated with peptides selected to be immunogenic by the method of the present invention, for example, 2 or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, and 60 or less, 55 or less, 50 or less, 45 or less, 40 or less, 35 or less, or 30 or less. Such T cell epitopes are also referred to in the present invention as "neo-epitopes". Presentation of these epitopes by the patient's cells, particularly antigen-presenting cells, preferably induces T cells that target the epitope upon binding to MHC, and thus target the patient's tumor, preferably not only the primary tumor but also the metastases of the tumor, expressing the antigen from which the T cell epitope originated and presenting the same epitope on the surface of the tumor cells.

[0116] The term “tumor” or “neoplastic disease” preferably means the abnormal growth of cells (referred to as neoplastic cells, neoplastic cells, or tumor cells) that form swelling or lesions. “Tumor cells” means abnormal cells that grow by rapid, uncontrolled proliferation and continue to grow after the stimulus that initiated their new growth has ceased. Tumors are characterized by a partial or complete lack of functional coordination and structural organization with normal tissue and generally form a distinct tissue mass that can be benign, premalignant, or malignant.

[0117] Cancer is a class of diseases characterized by unregulated growth (division beyond normal limits), invasion (invasion and destruction of adjacent tissues), and occasional metastasis (spread to other parts of the body through the lymphatic system or bloodstream). These three malignant characteristics of cancer distinguish it from benign tumors, which are self-limiting and do not invasive or metastasize. Most cancers form tumors, but some, like leukemia, do not. Malignant, malignant neoplasm, and malignant tumor are essentially synonymous with cancer.

[0118] For the purposes of this invention, the terms "cancer" and "cancer disease" can be used interchangeably with the terms "tumor" and "tumor disease."

[0119] The term "cancer" may include carcinoma, adenocarcinoma, progenitor cell tumor, leukemia, seminomas, melanoma, teratoma, lymphoma, neuroblastoma, glioma, rectal cancer, endometrial cancer, kidney cancer, adrenal cancer, thyroid cancer, hematological cancer, skin cancer, brain cancer, cervical cancer, intestinal cancer, liver cancer, colon cancer, gastric cancer, intestinal cancer, head and neck cancer, gastrointestinal cancer, lymph node cancer, esophageal cancer, colorectal cancer, pancreatic cancer, ear, nose and throat (ENT) cancer, breast cancer, prostate cancer, uterine cancer, ovarian cancer, lung cancer, etc. The cancer may be primary or metastatic.

[0120] "Treatment" can mean reducing the size or number of tumors in a subject; stopping or slowing the progression of a disease in a subject; suppressing or slowing the development of a new disease in a subject; reducing the frequency or severity of symptoms and / or recurrences in a subject who is currently or has previously been ill; and / or extending the lifespan of a subject.

[0121] "Immunotherapy" refers to treatment methods that involve the activation of specific immune responses. From the perspective of the present invention, terms such as "protective," "preventive," "preventive," "preventing," and "protective" are expressions relating to the prevention or treatment, or both, of the onset and / or transmission of a disease in a subject, and in particular, expressions relating to minimizing the opportunity for disease progression or delaying disease progression in a subject. For example, an individual at risk of developing cancer would be a candidate for cancer preventive therapy.

[0122] Prophylactic administration of immunotherapy, for example, prophylactic administration of the vaccine of the present invention, preferably protects the recipient from disease development. Therapeutic administration of immunotherapy, for example, therapeutic administration of the vaccine of the present invention, can lead to the suppression of disease progression / growth. This includes slowing the progression / growth of the disease, in particular disrupting the progression of the disease, and, for example, leading to the elimination of the disease.

[0123] Immunotherapy may be carried out using a variety of techniques in which the formulations provided in the present invention function to remove diseased cells from a patient. Such removal may occur as a result of inducing or enhancing an immune response in the patient that is specific to an antigen or cells expressing the antigen.

[0124] The formulations and compositions provided in the present invention may be used alone or in combination with therapies such as surgery, radiotherapy, chemotherapy and / or bone marrow transplantation (autologous, allogeneic, allogeneic or unrelated).

[0125] The terms "immunization" or "vaccination" describe the process of treating a patient with the aim of inducing an immune response for therapeutic or preventive reasons.

[0126] The term "in vivo" refers to the condition of a subject's body.

[0127] The terms “subject,” “individual,” or “patient” are used interchangeably and relate to vertebrates, such as mammals. For example, in the context of this invention, mammals include not only humans, non-human primates, domesticated animals such as dogs, cats, sheep, cattle, goats, pigs, and horses, and laboratory animals such as mice, rats, rabbits, and guinea pigs, but also animals captured in zoos. In this invention, the term “animal” may also include humans. The term “subject” includes patients, i.e., animals, preferably diseases, preferably humans suffering from the diseases described in this invention, or biological samples isolated therefrom (cancer cells, blood, etc.).

[0128] The vaccine compositions provided herein may additionally contain adjuvants. The term “adjuvant” refers to compounds that prolong, enhance, or accelerate an immune response. The compositions of the present invention preferably exert their effects without the addition of adjuvants. However, the compositions of the present invention may contain known adjuvants. Adjuvants include a heterogeneous group of compounds such as oil emulsions (e.g., Freund's adjuvants), mineral compounds (e.g., alum), bacterial products (e.g., Bordetella pertussis toxin), liposomes, and immune-stimulating complexes. An example of an adjuvant is monophosphoryl-lipid-A (MPL SmithKline Beecham). Examples include saponins, such as QS21 (SmithKline Beecham), DQS21 (SmithKline Beecham; WO96 / 33739), QS7, QS17, QS18, and QS-L1, incomplete Freund's adjuvants, complete Freund's adjuvants, vitamin E, montanides, alum, CpG oligonucleotides, and a variety of water-in-oil emulsions produced from biodegradable oils such as squalene and / or tocopherol.

[0129] The vaccine composition may further contain pharmaceutically acceptable substances that stimulate the patient's immune response. For example, such substances may be cytokines. Examples of such cytokines include interleukin-12 (IL-12), GM-CSF, and IL-18, which increase the protective activity of the vaccine.

[0130] The vaccine composition may further contain compounds that enhance the immune response. These compounds may include co-stimulatory molecules provided in the form of nucleic acids or proteins, such as B7-1 and B7-2 (CD80 and CD86, respectively).

[0131] The vaccine compositions provided herein may be administered by conventional routes, including injection or infusion. Administration may be, for example, orally, intravenously, intraperitoneally, intramuscularly, subcutaneously, or percutaneously. In one specific example, administration can be performed intranodally, such as by injection into a lymph node.

[0132] The vaccine composition may be administered in a pharmaceutically effective dose. A "pharmaceutically effective dose" means an amount that, alone or in combination with additional doses, is capable of producing the desired response or effect. The pharmaceutically effective dose may vary depending on the condition to be treated, the severity of the disease, individual patient variables such as age, physiological state, size and weight, duration of treatment, type of accompanying treatment (if any), specific route of administration and similar factors.

[0133] The vaccine composition may be sterile and may contain an effective amount of therapeutically active substance to produce the desired reaction or effect.

[0134] The term "pharmaceutically acceptable" means non-toxic substances that do not interact with the action of the active ingredient in a pharmaceutical composition, and may include, for example, co-immunoenhancing substances such as salts, buffers, preservatives, carriers, and adjuvants, such as CpG oligonucleotides, cytokines, chemokines, saponins, GM-CSF and / or RNA, and, where appropriate, other therapeutically active compounds.

[0135] "Carrier" refers to a natural or synthetic organic or inorganic component that is combined with an active ingredient to facilitate, enhance, or enable the application of the active ingredient. In this invention, the term "pharmaceutically acceptable carrier" includes one or more suitable solid or liquid fillers, diluents, or encapsulating materials suitable for administration to a patient.

[0136] The vaccine composition may contain appropriate buffering agents such as acetates, citrates, borates, and phosphates.

[0137] The vaccine composition may contain suitable preservatives such as benzalkonium chloride, chlorobutanol, parabens, and thimerosal.

[0138] The vaccine compositions are generally provided in unit dose form and can be manufactured by known methods. The vaccine compositions of the present invention may be in the form of injections, capsules, tablets, lozenges, solutions, suspensions, syrups, elixirs, or, for example, emulsions.

[0139] Compositions suitable for parenteral administration may generally include sterile aqueous or non-aqueous formulations of the active compound that are isotonic with the recipient's blood. Examples of suitable carriers and solvents include Ringer's solution and isotonic sodium chloride solution. In addition, sterile fixative oils can generally be used as the solution or suspension medium. [Effects of the Invention]

[0140] When the features of the above invention are applied, NGS data processing of patients enables efficient processing of HLA-type I and mutation variants.

[0141] This invention presents a method for selecting new antigens applicable to personalized anti-cancer vaccines by using AI technology to perform database learning, selecting a new antigen candidate group using a Binding Affinity model and an Immunogenicity model, and then applying and verifying a novel Neoantigen prioritization process. [Brief explanation of the drawing]

[0142] [Figure 1] The distribution of binding sites (BA) and immunogenicity sites (IMM) is shown. [Figure 2A] The measurement process for immunogenicity scores according to Example 4 is schematically shown. [Figure 2B] Example 4 shows the results of confirming immunogenicity by INFg ELISPOT for each HLA-I type (A*02:01). [Figure 2C] Example 4 shows the results of confirming immunogenicity by INFg ELISPOT for each HLA-I type (A*24:02). [Figure 2D] Example 4 shows the results of confirming immunogenicity by INFg ELISPOT for each HLA-I type (A*11:01). [Figure 3] The test process for Example 5.1 is schematically shown. [Figure 4] This graph shows the results of splenocyte IFNg ELISPOT obtained in Example 5.1. [Figure 5] The test process for Example 5.2 is schematically shown. [Figure 6] This graph shows the results of splenocyte IFNg ELISPOT obtained in Example 5.2. [Figure 7] This graph shows the anticancer effect (in vivo) of the new antigen confirmed in Example 6. [Modes for carrying out the invention]

[0143] The present invention will be described more specifically below with reference to the following examples. However, these are merely illustrative examples of the present invention, and the scope of the present invention is not limited by these examples.

[0144] Example 1. Construction of a new antigen prediction model (model learning) (New Antigen Prediction Model I) We prepared an artificial intelligence model to predict the binding affinity and immunogenicity of a new antigen candidate peptide by learning from HLA (Human Leukocyte Antigen) type I HLA IA, B, and C, along with the amino acid sequence information of the peptide, as found in the IEDB (Immune Epitope Database and Analysis Resource) (https: / / www.iedb.org). Specifically, the binding affinity and immunogenicity prediction models are ensemble models consisting of 150 different models, each trained with a combination of different model structures, training data partitioning, whether or not transfer learning is included, training optimization methods, and training interruption points. For the model structure, we selected either a GRU (Gated Recurrent Unit) or a Transformer. For the training data partitioning, we applied 5-Fold Cross Validation. For transfer learning, in the case of the binding affinity prediction model, we used a classifier trained on qualitative data, and in the case of the immunogenicity prediction model, we utilized the binding affinity prediction model with the smallest prediction error on the validation data. For the learning optimization method, we selected one of AdamW or SWA (Stochastic Weights Averaging). The five points at which the prediction error on the validation data was lowest during the entire learning process were selected as the stopping points for learning.

[0145] The following describes in more detail the preparation process for the new antigen prediction model: 1.1. Selection of Input Data for the New Antigen Prediction Model Data necessary for training HLA type I was selected from the IEDB data, converted for use in training the new antigen prediction model, and applied to the Binding Affinity model and the Immunogenicity model, respectively. For example, the data selection may involve applying one or more operations such as deleting data that does not correspond to HLA-A, HLA-B, or HLA-C, deleting data with mutations in the HLA molecule, deleting peptides with a length of 7 or less or 15 or more, or removing amino acid sequences that contain sequences other than the 20 amino acids derived from humans. For example, the data transformation may involve applying one or more operations such as normalizing binding strength measurements with a large range of data values ​​(e.g., IC50) to the interval of 0 to 1, or converting immunogenicity score data provided as [number of experiments, number of reactions] instead of quantitative measurements into a beta distribution mean (L. Guangyuan, et al. "DeepImmuno: deep learning-empowered prediction and generation of immunogenic peptides for T-cell immunity." Briefings in bioinformatics 22.6 (2021)).

[0146] During the training of the immunogenicity prediction model, we selected and proceeded with training data from the diverse experimental results of the aforementioned IEDB data, focusing on immunogenicity-related test methods such as cytokine secretion levels (IFNg, IL-2, TNFa, Granzyme A, etc.). These data included T cell-derived cytokines such as INF-g, TNFa, IL-2, and IL-1b during PBMC treatment of each peptide, T cell proliferation due to immunogenic activation, and data corresponding to activated T cell cytotoxicity, including chemokines such as CCL4 and CXCL9, and Granzyme B.

[0147] 1.2. Learning method of the new antigen prediction model Using the data selected in Example 1.1, the following was performed: For the binding affinity prediction model, one of two model structures (GRU (J. Chung et al., https: / / arxiv.org / abs / 1412.3555, 2014) or transformer (Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017.)) was selected, one of three learning methods was selected, and one of five-fold cross-validation data partitioning methods was selected. A total of 150 models were trained using a method in which the model was saved at the five training points where the prediction error on the validation data was lowest (2 × 3 × 5 × 5 = 150). The mean and variance of the predicted values ​​of the 150 trained models were calculated to predict the final binding affinity (predicted value 1).

[0148] The immunogenicity prediction model was trained using a method that selected one of two model structures (GRU, transformer), one of three training methods (including training the model with the lowest prediction error on the validation data among the coupling strength prediction models), and one of five-fold cross-validation data splitting methods. A total of 150 models were trained (2 × 3 × 5 × 5 = 150). The mean and variance of the prediction values ​​of the 150 trained models were calculated to predict the final immunogenicity (predicted value 2).

[0149] In calculating Prediction Value 1 and Prediction Value 2, the final prediction value was calculated as [(average of 150 model predictions) - (variance of 150 model predictions)], respectively. This is because a smaller deviation between the predictions of 150 models using different data and learning methods indicates a more stable prediction.

[0150] Example 2. Prediction and selection of new antigens A new antigen was selected from the new antigen prediction model constructed in Example 1, and the selection of the new antigen was carried out in the following steps: 1) Collecting NGS data of cancer patients from publicly available databases; 2) Using NGS data from the aforementioned cancer patients, determine HLA type and peptides derived from cancer cells, and calculate TPM; 3) Calculation of predicted score using the new antigen prediction model described in Example 1; 4) Prioritization of new antigens and selection of target peptides for experimentation using the formula in Example 3 below; 5) Binding affinity and immunogenicity experiments using PBMCs identical to the HLA type identified in 2) above.

[0151] The publicly available data used to validate the new antigen prediction model were from four melanoma patients (HLA type: A*02:01; see "John Wiley & Sons A / S, Tissue Antigens 73, 95-170") who showed a complete response to pembrolizumab treatment (hereinafter referred to as "patient sample data" or "patient sample dataset"). [Hugo, Willy, et al. "Genomic and transcriptomic features of response to anti-PD-1 therapy in metastatic melanoma." Cell 165.1 (2016): 35-44.]

[0152] 2.1. Data preparation for new antigen selection (Whole exome sequencing and RNA-seq data processing of patient-derived samples) For the aforementioned patient sample dataset, we configured the system to combine codes for data pre-processing, variant discovery, and callset refinement to derive information such as mutation sequences and HLA-types. More specifically,

[0153] 1) DNA mutation information was converted into peptide mutation information, and RNA abundance data and HLA type data were concatenated. Furthermore, a process was conducted to confirm whether peptides obtained by conversion using known Homo sapience protein sequences could be found in Uniprot's SwissProt.

[0154] 2) Results were obtained using Mutect2 from Broad Institute, which is included in the GATK package, a program for searching for cancer-specific genetic mutations, and Strelka, developed by Illumina. Then, after processing the DNA mutation information and organizing the information obtained from the tools for protein-level missense mutations and in-frame deletion data, duplicate peptides obtained with Mutect2 and Strelka were unified (Koboldt, DC et al., VarScan2: Somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Research 22, 568-576 (2012); Cai, L., Yuan, W., Zhang, Z., He, L. & Chou, KC, In-depth comparison of somatic point mutation callers based on different tumor next-generation sequencing depth data. Scientific Reports 6, 9, https: / / doi.org / 10.1038 / srep36540 (2016); Kim, S. et al., Strelka2: fast and accurate calling of germline and somatic variants. Nature Methods15,591-594,https: / / doi.org / 10.1038 / s41592-018-0051-x(2018);McKenna,A.et al.,The Genome Analysis Toolkit:a MapReduce framework for analyzing next-generation DNA sequencing data.Genome Res.20, 1297-1303, https: / / doi.org / 10.1101 / gr.107524.110(2010);Cibulskis, K., Lawrence, MS, Carter, SL, Sivachenko, A., Jaffe, D., Sougnez, C., & Getz, G. (2013).Sensitive detection of somatic point mutations in impure and heterogeneous cancer samples.Nature Biotechnology,31(3),213-219.doi:10.1038 / nbt.2514). .

[0155] 3) Hydrophobicity for peptides was calculated to obtain data for use in the production of new antigens thereafter. When new antigens are predicted and personalized anti-cancer vaccines are produced using mRNA, consideration of hydrophobicity may not be necessary. However, when producing anti-cancer vaccines with peptides or during immunogenicity testing, if the hydrophobicity of the peptide is high (for example, if it has a high content of hydrophobic amino acids), production may be impossible, so it may be necessary to calculate the hydrophobicity of the peptide. In this example, the hydrophobicity was determined by calculating the Gravy (Grand average of hydropathy) score.

[0156] 2.2. Selection of new antigens using a new antigen prediction model (binding affinity and immunogenicity) (confirmation of applicability to personalized anti-cancer vaccines) Using the data prepared in Example 2.1, binding strength and immunogenicity thresholds were set, and the novel antigen prediction model was applied to validate the selected novel antigen: - The binding affinity (BA) test compares the binding affinity of the predicted new antigen to that of NY-ESO-1 (New York esophageal squamous cell carcinoma 1 peptide (157-165, SLLMWITQC; synthesized by AnyGen), for which the IC50 value is known, to verify its binding affinity. - The immunogenicity test was conducted using the IFNγ ELISPOT Immunogenicity assay. -HLA-A*02:01 type normal PBMCs were rested for 24 hours, then treated with candidate peptides at 50 μM IL-2 (10 IU / ml) for pulsing (72 hours, 37°C, 5% CO2). Subsequently, the PBMCs were washed and transferred to an ELISPOT assay plate for re-stimulation with the same peptides (50 μM, 24 hours, 37°C, 5% CO2). The ELISPOT assay was then performed according to Immunospot's ELISPOT protocol. The assay was repeated with PBMCs from three other donor types (to confirm reproducibility and statistical significance).

[0157] The aforementioned HLA-A*02:01 uses a portion of the alpha chain sequence, and the sequence used is as follows: SHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYWDGETRKVKAHSQTHRVDLGTLRGYYNQSEAGSHTVQRMYGCDVGSDWRFLRGYHQYAYDGKDYIALKEDLRSWTAADMAAQTTKHKWEAAHVAEQLRAYLEGTCVEWLRRYLENGKETLQRT(Sequence ID 1)

[0158] Since the test was conducted using cells, there was a deviation in the immunogenicity test. Therefore, we applied Min-Max Normalization and Weight Value (1 lot detection: x1, 2 lots detection: x2, 3 lots detection: x3) to rank the new antigens and established a method for selecting new antigens applicable to personalized anti-cancer vaccines.

[0159] Figure 1 shows the results obtained by applying the information of 118,057 mutant peptides identified from patient-derived samples (four melanoma patients (HLA type: A*02:01) who showed a complete response to pembrolizumab treatment in Example 2.1) to the binding affinity (BA) prediction model and immunogenicity (IMM) prediction model of Example 1.

[0160] Furthermore, for the BA high group (G1: 0.15 ≤ IMM Score < 0.25 and 0.6 ≤ BA Score) and the IMM high group (G2: IMM Score ≥ 0.25 and 0.6 > BA Score ≥ 0.1) from the results in Figure 1, the actual binding affinity and immunogenicity were measured using the method described above and are shown in Table 1:

[0161] [Table 1]

[0162] As can be seen from the results in Table 1 above, When the threshold is set to 0.15 ≤ IMM Score < 0.25 and 0.6 ≤ BA Score (G1), the prediction rate of new antigens with positive binding affinity is 100%. When the threshold was set to IMM Score ≥ 0.25 and 0.6 > BA Score ≥ 0.1 (G2), we confirmed that the prediction rate for new immunogenic positive antigens was 23.6%.

[0163] Example 2.3. Measurement of RNA expression level (TPM) as an additional variable. As an additional variable for evaluating the predicted new antigens, the TPM (RNA expression level; transcripts per million) variable was measured. TPM represents the expression rate of RNA sequences of the new antigen candidate group, and can indicate the possibility that mutations in the patient's cancer cells will be generated in the form of actual peptides. However, since the IEDB does not provide the TPM data, it can be used for evaluation (prioritization) of model prediction results that are not model learning. In order to develop evaluation criteria utilizing the TPM, publicly available NGS data from cancer patients (TESLA dataset, Wells, Daniel K., et al. "Key parameters of tumor epitope immunogenicity revealed through a consortium approach improve neoantigen prediction." Cell 183.3 (2020):818-834.) were collected to select a group of new antigen candidates (peptide candidates) derived from cancer, and the predicted values ​​1 and 2 for the new antigen candidate group were predicted using the prediction model (see Example 1.2). The RNA-seq data included in the publicly available NGS data from cancer patients was processed to calculate the expression rate (TPM) of RNA sequences of the new antigen candidate group.

[0164] Example 3. Evaluation of selected new antigens (Neoantigen Prioritization) For the predicted values ​​1 and 2 obtained in Example 1.2, the predictive model variables were optimized using a Binding Affinity Model and an Immunogenicity Model through multivariate linear regression analysis. Finally, TPM (RNA expression level), Binding affinity prediction score (BA; predicted value 1 in Example 3), and Immunogenicity prediction score (IMM; predicted value 2 in Example 3) were determined as variables, and the ranking of the new antigens was selected.

[0165] The aforementioned TESLA dataset includes HLA types and NGS data from eight patients, peptide sequences of 845 new antigen candidates derived from these patients, and results of immunogenicity experiments. Of these 845 candidates, 38 peptides were verified to be immunogenic. To compare the excellence of the new antigen prediction criteria, TPM was applied to the new antigen candidate group, and 20 new antigen candidates were selected per patient (160 in total) by varying the weighting of the Binding affinity prediction score (BA) and Immunogenicity prediction score (IMM). Table 2 below shows how many of the 38 immunogenic new antigen candidates were included in this group (Threshold: TPM (transcripts per million) > 1):

[0166] [Table 2]

[0167] As can be seen from Table 2 above, when the Binding affinity prediction score (BA) and Immunogenicity prediction score (IMM) were weighted identically (1:1), we were able to find the new antigen with the highest immunogenicity.

[0168] Based on the results described above, the following formula was established for ranking the new antigens: Prioritization scoring scheme for the new antigen: [min(TPM, 1)*(BA+IMM) / 2] The number of new antigens detected before and after the application of the Prioritization procedure is shown in Table 3 below:

[0169] [Table 3]

[0170] As shown in Table 3, it can be confirmed that applying Prioritization increases the number of new antigens that can be selected and enhances the new antigen prediction rate.

[0171] Example 4. Selection of a new antigen prediction model II and a new antigen. Referring to the methods of Examples 1-3, melanoma patients possessing three HLA types (A*02:01, A*24:02, and A*11:01) were selected, and 20 new antigens were predicted using a new antigen prediction model. The presence or absence of immunogenicity to the predicted new antigens was confirmed using IFNg ELISPOT.

[0172] More specifically, the selection criteria for the patient group data are as follows: -High Mutation Burden Tumor→Melanoma; -Complete WES (whole exosome sequencing; mutation selection), RNA-seq (expression level) dataset (converted to peptide sequence for use); -ICI (Immune Checkpoint Inhibitor) treatment (anti-PD-1 therapy: Pembrolizumab, Nivolumab) → CR (Complete Response), PR (Partial Response), PD (Progressive Disease); -HLA type:A*02:01, A*24:02, A*11:01([HLA-A*02:01]: SHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYWDGETRKVKAHSQTHRVDLGTLRGYYNQSEAGSHTVQRMYGCDVGSDWRFLRGYHQYAYDGKDYIALKEDLRSWTAADMAAQTTKHKWEAAHVAEQLRAYLEGTCVEWLRRYLENGKETLQRT(Sequence ID 1), [HLA-A*11:01]: SHSMRYFYTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYWDQETRNVKAQSQTDRVDLGTLRGYYNQSEDGSHTIQIMYGCDVGPDGRFLRGYRQDAYDGKDYIALNEDLRSWTAADMAAQITKRKWEAAHAAEQQRAYLEGRCVEWLRRYLENGKETLQRT (Sequence No. 2), [HLA-A*24:02]: SHSMRYFSTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYWDEETGKVKAHSQTDRENLRIALRYYNQSEAGSHTLQMMFGCDVGSDGRFLRGYHQYAYDGKDYIALKEDLRSWTAADMAAQITKRKWEAAHVAEQQRAYLEGTCVDGLRRYLENGKETLQRT (Sequence No. 3); -Source: Hugo et al., 2016, Cell 165, 35-44, Riaz et al., 2017, Cell 171, 934-949 (included herein by reference).

[0173] The novel antigen prediction model applied to the selected patient group data is as follows: -Binding Affinity Score (BA) = (Mean value of Binding Affinity score for the selected patient group) - (Variance of Binding Affinity Score for the selected patient group); -Immunogenicity Score (IMM) = (Mean Immunogenicity Score of the selected patient group) - (Variance of Immunogenicity Score of the selected patient group)

[0174] The BA score is obtained from the binding results of each patient's new antigen to the patient's specific A*02:01, A*24:02, and A*11:01, and is a value obtained using a binding strength prediction model that learned predicted value 1 from Example 1.2 (see Example 2.1). The IMM score is obtained from the predicted immunogenicity results of each patient's new antigen to the patient's specific HLA type, and is a value obtained using an immunogenicity prediction model that learned predicted value 2 (Example 1.2) (see Example 2.1).

[0175] The Neoantigen Prioritization Score was calculated by applying the TPM (RNA expression level; transcripts per million) score (TPM threshold: 1; see Example 3): Prioritization Score=min(TPM, 1)*[(BA+IMM) / 2] Min(TPM, 1): This means that if TPM > 1, replace it with TPM = 1, and if TPM < 1, use the TPM Value as is. The ranking was advanced using the obtained Prioritization Score, and the top 20 peptides were predicted as new antigens (Tables 4-6).

[0176] [Table 4]

[0177] [Table 5]

[0178] [Table 6]

[0179] To validate the predicted new antigen, Peripheral Blood Mononuclear Cells (PBMCs) were obtained from three adult donor patients with available HLA type information, and immunogenicity was evaluated using IFNg ELISPOT (each of the three donor patients had a different HLA type, and the obtained PBMCs were A*02:01, A*24:02, or A*11:01, respectively). The test process is schematically shown in Figure 2.

[0180] The ELISPOT results obtained above were analyzed by performing normalization using the following formula:

number

[0181] (Normalization is performed to maintain consistency in the results for each test.) For eight (A*02:01), seven (A*24:02), and five (A*11:01) HLA type-matched patients, the Top 20 and Top 10 new antigens were predicted using the novel antigen prediction models described in Examples 1 and 2. Immunogenicity evaluation of the 20 predicted new antigens was then performed using the INFg ELISPOT as described above, according to the test methods corresponding to Examples 3 and 4, to confirm the number of new antigens among the 20 new antigens in each patient that were judged to be immunogenic.

[0182] The results obtained are shown in Table 7 and Figures 2B-2D:

[0183] [Table 7]

[0184] When immunogenicity was experimentally evaluated in 8 individuals with A*02:01, 7 individuals with A*24:02, and 5 individuals with A*11:01, the predicted immunogenicity rates for the new antigens were confirmed to be 26.8%, 20.7%, and 21.0%, respectively.

[0185] Example 5. Ex vivo Immunogenicity 5.1.Human Cell Line (MCF7 cell line)Derived Xenograft Model The MCF7 breast cancer cell line is a human cancer cell line that expresses HLA-A*02:01. Referring to Example 4, whole exosome sequencing (Mutation Selection) and RNA-Seq (Express level) NGS sequencing were performed on MCF7 cell line (ATCC, cat#HTB-22). Using a new antigen prediction model (BA, IMM, TPM), the top 20 antigens (Top 20) were selected, resulting in 20 new antigens (mRNA). These were then synthesized into mRNA-lipoplex anti-cancer vaccines (see Example A(1) of KR10-2022-0149461A and KR10-2022-0149460A), administered to A02:01 transgenic mice (C57BL / 6 mice expressing HLA-A2.1; Taconic biosciences, cat#8906-F) (60ug x 3 times QW, tail-vein IV injection), and immunogenicity was confirmed by splenocyte IFNg ELISPOT.

[0186] The immunogenicity test was performed as follows: A 96-well ELISPOT plate coated with anti-IFNg antibody was prepared. After washing the plate once with PBS, 1.0 × 10⁶ spleen cells were collected from the immunized mice per well. 6 Cells were loaded one at a time. For each immunization group, cells were advanced in 3 wells and re-stimulated with 10 μM peptide or medium alone in a total volume of 200 μL. The plates were then incubated at 37°C for 18 hours. After washing three times with PBST, each well was treated with biotin-labeled mouse IFNg detection antibody, and the plates were incubated at room temperature for 2 hours. The cells were washed again, and AP-conjugated streptavidin was added to each well, and the plates were incubated at room temperature for 30 minutes. The AP reaction was developed using a substrate reagent set (Immunospot). The number of spot-forming units (SFUs) per well was automatically counted using an ELISPOT reader (CTL ELISPOT reader).

[0187] Figure 3 schematically shows the test process, including the administration of the aforementioned vaccine. The 20 new antigens selected are shown in Table 8:

[0188] [Table 8]

[0189] The results of the splenocyte IFNg ELISPOT obtained are shown in Figure 4. In Figure 4, the control group is the untreated group that was not treated with the new antigen peptide, and "MCF7 NEO" is the group that was administered the mRNA-lipoplex anti-cancer vaccine containing the 20 selected new antigens. A threshold of more than three times the average ELISPOT of the untreated group (control) was set, and ELISPOT results above the threshold were judged to have anti-cancer efficacy. As a result, 25% of the 20 new antigens tested were evaluated as having expected anti-cancer efficacy (indicated by * in Figure 4; 5 / 20).

[0190] 5.2.B16F10(Melanoma)AI Model Predicted Neoantigens Referring to Example 4, NGS was performed on the B16F10 (Mouse Melanoma; ATCC, cat#CRL-6475) cell line (converted to amino acid sequence), and the Top 50 and Top 20 new antigens were predicted using new antigen prediction models (BA, IMM, TPM). The predicted new antigens were prepared as mRNA-lipoplex anticancer vaccines and administered to C57BL / 6 mice (60ug x 2 QW IV injections). Immunogenicity was confirmed by splenocyte IFNg ELISPOT (see Figure 5). As a positive control, a mouse model administered with mTRP2 (Tyrosinase-related protein 2) (Melanoma Tumor-associated Antigen; SVYDFFVWL (SEQ ID NO: 134)) in the same manner as described above was used, and the immunogenicity of the predicted new antigens was determined by setting a threshold.

[0191] Table 9 lists the Top 50 novel antigen sequences used in this embodiment:

[0192] [Table 9-1] [Table 9-2]

[0193] The results obtained are shown in Figure 6. As shown in Figure 6, immunogenicity was confirmed for 45% (9 / 20) of the top 20 predicted new antigens (see upper part of Figure 6). Of the 12 new antigens with high immunogenicity among the top 50 new antigens, 9 were distributed within the top 20, confirming that the top 20 are highly likely to contain the target new antigen during the search for new antigens (see Table 10).

[0194] [Table 10] (Immunogenicity Threshold:>3x Media ELISPOT Avg.)

[0195] Example 6: In vivo Efficacy: Predicted anticancer efficacy of the novel antigen B16F10 tumor cells 5×10 4 The individual cells were transplanted into C57BL / 6 mice using a subcutaneous transplantation method, and the resulting 50mm 3 Group separation was performed based on tumor size. The test groups consisted of a group receiving mRNA-lipoplex containing the top 10 of the 12 B16F10 predicted novel antigens secured in Example 5.2, a positive control receiving mRNA-lipoplex containing mTRP2 (Tyrosinase-related protein 2) (Melanoma Tumor-associated Antigen) (SEQ ID NO: 134), and a negative control receiving only lipoplex without novel antigen-containing mRNA. mRNA-lipoplex was administered at a dose of 60 ug every 5 days for 3 doses, and tumor size was recorded over 3 weeks.

[0196] The results obtained (tumor size) are shown in Figure 7 (Neoantigen Cancer Vaccine: group administered with 10 new antigens, Positive control: group administered with mTRP2, control: group administered with lipoplex only (negative control)). As shown in Figure 7, a 63% reduction in tumor size was observed in the mRNA-lipoplex group that received 10 new antigens predicted by B16F10.

[0197] From the above description, those skilled in the art in which the present invention pertains will understand that the present invention can be implemented in other specific forms without altering its technical idea or essential features. In this regard, it should be understood that the embodiments described above are illustrative in all respects and not limiting. The scope of the present invention should be interpreted as encompassing all modified or altered forms derived from the meaning and scope of the claims, which are described below in more detail than the above, and their equivalent concepts.

Claims

1. A method for selecting patient-derived tumor-specific immunogenic peptides, including the following steps (1) to (7): (1) The step of obtaining NGS sequence information from patient-derived tumor cells or tumor tissue and normal cells or normal tissue; (2) A step of obtaining tumor-specific peptide sequence information having one or more amino acid mutations in the NGS sequence of tumor cells or tumor tissue, and having a length of seven or more amino acids, by comparing it with the NGS sequence of normal cells or normal tissue; (3) A step of measuring the expression level (TPM; Transcripts per Million) of tumor-specific peptides or their coding genes within tumor cells or tumor tissue; (4) The step of applying the tumor-specific peptide sequence information to a binding affinity prediction model to confirm the number of binding sites (binding affinity score; BA) for HLA molecules; (5) The step of applying the tumor-specific peptide sequence information to an immunogenicity prediction model to confirm the immunogenicity score (IMM); (6) Prioritization step to determine the ranking of tumor-related peptides based on the values ​​calculated using the following formula: Neoantigen prioritization score=min(TPM, 1)*(W1*BA+W2*IMM) / (W1+W2) [Set within the range of 0 < W1 < 1, 0 < W2 < 1]; and (7) A step of selecting two or more tumor-specific immunogenic peptides based on the priority order of the tumor-specific peptides determined above.

2. A method for selecting a patient-derived tumor-specific immunogenic peptide according to claim 1, further comprising the following steps (a) to (d): (a) The step of obtaining peptide sequence information derived from tumor cells or tumor tissue and peptide sequence information derived from normal cells or normal tissue from a peptide database. (b) Steps to obtain class I HLA peptide sequence information from the biological sequence database, (c) A step of training a model capable of predicting the binding affinity score (BA) for an HLA with the data obtained in steps (a) and (b) above to construct a binding strength prediction model and obtain the number of binding points; and (d) A step in which the data obtained in steps (a) and (b) above is used to train a model capable of predicting immunogenicity score (IMM) to construct an immunogenicity prediction model and obtain an immunogenicity score.

3. The method for selecting patient-derived tumor-specific immunogenic peptides according to claim 2, characterized in that the peptide database used to construct the predictive model in step (a) undergoes one or more of the following selection processes i) to iv): i) Deletion of data that does not fall under HLA-A, HLA-B, or HLA-C. ii) Deletion of data containing mutations in the HLA molecule. iii) Deletion of data where the peptide length is 7 or less or 15 or more, and iv) Removal of amino acid residues other than the 20 amino acids of human origin from the amino acid sequence.

4. The method for selecting a patient-derived tumor-specific immunogenic peptide according to claims 1 to 3, wherein the HLA comprises different HLA types or different HLA alleles.

5. The learning described in (c) and (d) above is carried out by combining one or more artificial intelligence models and one or more learning methods. A method for selecting a patient-derived tumor-specific immunogenic peptide according to any one of claims 1 to 3.

6. The method for selecting a patient-derived tumor-specific immunogenic peptide according to any one of claims 1 to 3, wherein the number of binding sites is determined by the IC50 value or Kd value related to binding to HLA.

7. A method for selecting a patient-derived tumor-specific immunogenic peptide according to any one of claims 1 to 3, wherein the immunogenicity score is determined by one or more pieces of information selected from the group consisting of the amount of cytokine secretion from T cells when the candidate peptide is treated with immune cells or a sample containing the same, T cell favorization due to immunogenic activation, chemokine secretion, and information related to the cytotoxicity of activated T cells.

8. The method for selecting a patient-derived tumor-specific immunogenic peptide according to claim 7, wherein the sample containing the immune cells is selected from the group consisting of blood, leukocytes, and peripheral blood mononuclear cells (PBMCs), and the information is two or more selected from the group consisting of IFN-g, IL-2, TNF-a, CCL4, CXCL9, and granzyme B.

9. The method for selecting a patient-derived tumor-specific immunogenic peptide according to any one of claims 1 to 3, wherein the expression level of the candidate peptide or its coding gene in the patient-derived sample in step (3) is obtained from the patient sample data by RNA sequencing.

10. A method for selecting a patient-derived tumor-specific immunogenic peptide according to any one of claims 1 to 3, wherein the patient-derived tumor-specific immunogenic peptide is for use in the manufacture of a personalized anti-cancer vaccine for the patient.

11. A tumor-specific immunogenic peptide selected by the method of any one of claims 1 to 3, or a nucleic acid molecule encoding the peptide, and Pharmacologically acceptable carriers An anti-cancer vaccine composition containing [a specific ingredient].

12. A tumor-specific immunogenic peptide selected by the method of any one of claims 1 to 3, or a nucleic acid molecule encoding the peptide, and Pharmacologically acceptable carriers A method for manufacturing an anti-cancer vaccine, including a step of mixing the following.