Method for predicting immunogenic epitopes and device using the same
The method integrates multiple biological processes and characteristics using AI to predict immunogenic epitopes, addressing the limitations of existing methods and improving the accuracy and efficacy of neoantigen-based vaccines.
Patent Information
- Application Number
- JP2025526531
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2023-11-07
- Publication Date
- 2025-12-09
AI Technical Summary
Existing methods for predicting immunogenic epitopes focus on individual factors or a limited combination, failing to account for the complex interplay of multiple biological processes and characteristics that influence epitope immunogenicity.
A method and device that integrate various factors and characteristics, including antigen processing, presentation, immune response, and tumor microenvironment, using a pre-trained artificial intelligence model to predict immunogenicity, incorporating algorithms and tools to calculate and standardize these factors.
Accurately predicts immunogenic epitopes with high accuracy, enhancing the efficacy of neoantigen-based vaccines and reducing side effects, achieving high accuracy and efficacy in the development of neoantigen-based vaccines.
Smart Images

Figure 2025539731000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for predicting immunogenic epitopes and an apparatus using the same, and more specifically to a method for predicting immunogenic epitopes by integrating various factors involved in multiple biological processes and the characteristics of the epitopes, and an apparatus using the same. [Background technology]
[0002] T cell epitope recognition plays a key role in the immune response to pathogens and tumors. Neoantigens, expressed through cancer mutations, can activate immune cells in the body and treat cancer. Neoantigens can convert tumors that cannot be recognized by T cells (cold tumors) into tumors that are easily infiltrated by T cells (hot tumors), thereby solving the problem of low responsiveness to anti-cancer immunotherapeutic agents. In other words, neoantigens, when expressed specifically in cancer cells through mutation, can selectively induce the proliferation of T cells that recognize neoantigens / neoepitopes and / or attack only tumor cells.
[0003] The development of immunogenic epitopes may involve factors in biological processors associated with immune responses, such as MHC class I / II binding affinity, proteasomal cleavage, TAP transporter efficiency, peptide-MHC (pMHC) stability, and T-cell receptor (TCR)-epitope interactions. Furthermore, the characteristics of the epitope peptide sequence, such as post-translational modifications, sequence similarity to known epitopes, dissimilarity to the self-proteome, the extent of functional alterations due to mutations, the anti-inflammatory / inflammatory properties of the epitope, and the tumor microenvironment (TME), may also influence the immunogenicity of the epitope.
[0004] To predict epitopes that induce immune responses in silico, various factors involved in various biological processes, such as MHC binding affinity, proteasomal cleavage, TAP transporter efficiency, peptide-MHC (pMHC) stability, and T-cell receptor (TCR) interactions, as well as the characteristics of epitope sequences, have been studied. For example, the PRIME tool predicts epitopes that induce immune responses by integrating MHC I binding affinity at the antigen presentation stage and TCR recognition at the immunogenicity stage. The MHCflurry 2.0 tool predicts epitopes that induce immune responses by integrating factors involved in the antigen processing stage and MHC I binding affinity at the antigen presentation stage. Recently published research results from the TESLA Consortium have presented criteria / parameters for predicting immunogenic epitopes that take into account MHC I binding affinity, peptide-MHC (pMHC) binding stability, agretopicity (the ratio of mutant binding affinity to wild-type binding affinity), and foreignness (homology to known pathogenic peptides). Summary of the Invention [Problem to be solved by the invention]
[0005] Previously, methods have been developed to predict epitopes that induce immune responses by individually using various factors and / or epitope characteristics involved in numerous biological processes, or by combining a small number of factors and / or characteristics. However, determining the immunogenicity of an epitope can be influenced by multiple factors and / or characteristics, either together or sequentially. The present invention proposes additional various factors (biomarkers) that can be considered to accurately predict the immunogenicity of an epitope, and proposes a method and device using the same for integrating and analyzing factors and / or characteristics involved / affecting immunogenicity. [Means for solving the problem]
[0006] The method for predicting an immunogenic epitope according to the present invention includes the steps of calculating the degree of influence of factors involved in the biological process of a tumor cell for an epitope and characteristics of the epitope on the immunogenicity of the epitope, performing at least one of standardization and normalization on the calculated value, and inputting the value after performing at least one of standardization and normalization into a pre-trained artificial intelligence model to predict the immunogenicity of the epitope, wherein the biological process comprises an antigen processing stage, an antigen presentation stage, an immune stage, and a tumor microenvironment, and the factors involved in the tumor microenvironment may be at least one of an inflammatory response, a B-cell linear epitope, and a B-cell conformational epitope.
[0007] In the method for predicting an immunogenic epitope according to the present invention, the step of calculating the degree of influence of each of the factors involved in the biological process for the epitope and the characteristics of the epitope on the immunogenicity of the epitope may be a step of performing the calculation using an algorithm, program, or tool.
[0008] In the method for predicting immunogenic epitopes according to the present invention, the characteristics of the epitope may be at least one of functional effects of mutation, phosphorylation, hydrophobicity, similarity to known epitopes, dissimilarity to the self-proteome with reference to humans, and peptide stability.
[0009] In the method for predicting immunogenic epitopes according to the present invention, the factor involved in the antigen processing step among the biological processes may be at least one of proteasome cleavage and TAP transporter efficiency.
[0010] In the method for predicting immunogenic epitopes according to the present invention, the factor involved in the antigen presentation step among the biological processes may be at least one of MHC I binding affinity, pMHC stability, average binding affinity to MHC I alleles, cross-reactivity to MHC I alleles, average binding affinity to MHC II alleles, and cross-reactivity to MHC II alleles.
[0011] In the method for predicting immunogenic epitopes according to the present invention, the average binding affinity for the MHC I / MHC II alleles is calculated using the following formula:
[0012]
number
[0013] The computer-readable recording medium for predicting immunogenic epitopes according to the present invention may have a computer program recorded thereon for performing at least one of the above-described methods.
[0014] The device for predicting immunogenic epitopes according to the present invention includes an input / output device that receives input of an epitope or outputs predicted results; a storage device that stores an AI model trained to predict the immunogenicity of a learned epitope based on factors involved in the biological process of the tumor cell epitope and at least some of the characteristics of the epitope; and a calculation device that calculates the degree of influence of each of the factors involved in the biological process of the tumor cell epitope and the characteristics of the epitope on the immunogenicity of the epitope, performs at least one of standardization and normalization on the calculated value, and inputs the value after the standardization and normalization into the trained AI model to predict the immunogenicity of the epitope. The biological process may consist of an antigen processing stage, an antigen presentation stage, an immune stage, and a tumor microenvironment, and the factors involved in the tumor microenvironment may be at least one of an inflammatory response, a B-cell linear epitope, and a B-cell conformational epitope. [Effects of the Invention]
[0015] According to the present invention, immunogenic epitopes can be predicted with high accuracy by simultaneously utilizing various factors involved in the various biological processes leading to the epitope's immunogenicity and the characteristics of the epitope. Highly predicted immunogenic epitopes are expected to have higher efficacy and fewer side effects in the development of neoantigen-based vaccines. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 illustrates an example of a biological process in tumor cells. [Figure 2]FIG. 1 illustrates a number of biological process elements and epitope properties that can affect the immunogenicity of an epitope. [Figure 3] FIG. 1 shows an example of a process for learning immunogenic epitopes using an artificial intelligence model. [Figure 4] FIG. 1 shows an example of a process for predicting immunogenic epitopes using an artificial intelligence model. [Figure 5] FIG. 1 shows an example of a configuration diagram of a device that predicts immunogenic epitopes using an artificial intelligence model. DETAILED DESCRIPTION OF THE INVENTION
[0017] The following embodiments are merely for the purpose of explaining the present invention in more detail, and in light of the gist of the present invention, it will be obvious to those skilled in the art that the scope of the present invention is not limited to these embodiments. It should be understood that the present invention includes all modifications, equivalents, and alternatives that fall within the technical idea and technical scope described below.
[0018] As used herein, singular terms should be construed as including plural terms unless the context clearly dictates otherwise, and terms such as "comprises" should be understood to mean the presence of stated features, numbers, steps, operations, components, parts, or combinations thereof, but not to exclude the presence or additional possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. Also, the term "and / or" means the inclusion of a combination of two or more associated stated items or any two or more associated stated items.
[0019] In performing a method or method of operation, the processes / steps constituting the method may be performed in an order different from that described, unless the context clearly dictates a particular order. That is, the processes / steps may be performed in the order described, substantially simultaneously, or in the reverse order. FIG. 1 shows an example of a biological process in tumor cells.
[0020] Referring to Figure 1, the biological process of tumor cells can be divided into three stages: antigen processing, antigen presentation, and immunogenicity. The tumor microenvironment (TME) surrounding tumor cells can also affect the biological process and can therefore be considered part of the tumor cell biological process. While the biological process of tumor cells is the same as or similar to that of general cells, various factors can affect the generation and / or characteristics of epitopes at each stage.
[0021] The antigen processing stage occurs when DNA in the tumor cell nucleus is transcribed into mRNA and released into the cytoplasm. The transcribed mRNA is then converted into peptides and binds to MHC class I. Tumor cell DNA can be altered due to mutations and other factors. After mRNA is translated and synthesized into protein, it can be cleaved into peptides by the proteasome. The cleaved peptides are transported into the rough endoplasmic reticulum by TAP transporters and bind to MHC class I.
[0022] In the antigen presentation step, peptides generated in the antigen processing step bind to MHC class I and are presented on the cell surface as peptide-MHC complexes.
[0023] The immunogenic stage is the stage at which peptide-MHC complexes displayed on the cell surface are recognized by T cells.
[0024] The tumor microenvironment (TME) is the environment surrounding cells, which may be surrounded by blood vessels, immune cells, fibroblasts, bone marrow-derived inflammatory cells, lymphocytes, signaling molecules, and extracellular matrix. Tumor cells can affect the tumor microenvironment by releasing extracellular signals, promoting neovascularization, and inducing peripheral immune tolerance, which in turn can affect tumor cell growth and / or metastasis. For example, when the disease inducing an immune response is cancer, the clonality of the sample may affect the immunogenicity of the epitope. The distribution / ratio of immune cells, tumor-infiltrating lymphocytes, RNA expression, and other factors may also affect the immunogenicity of the generated epitope.
[0025] FIG. 2 is a diagram illustrating the properties of epitopes and elements of many biological processes that can affect the immunogenicity of an epitope.
[0026] In addition to the biological processes described in Figure 1, the immunogenicity of an epitope can also be influenced by its properties. For example, negative factors that affect the human body, such as the toxicity or allergenicity of an epitope, can affect the immunogenicity of an epitope. Factors that affect the immunogenicity of an epitope can occur one at a time, or multiple factors can affect the epitope simultaneously and / or sequentially.
[0027] Below we detail several factors in epitope properties and biological processes that influence the immunogenicity of epitopes.
[0028] Among the epitope characteristics, factors that affect immunogenicity include functional impact by mutations, phosphorylation (PTM), hydrophobicity (hydrophobicity or physicochemical hydrophobicity), similarity to known epitopes, dissimilarity to self (human reference)-proteome, and peptide stability.
[0029] Below, specific programs / algorithms / tools, etc. will be described as examples for determining the degree of influence of each factor on the immunogenicity of an epitope, but are not limited to these.
[0030] The functional impact of a mutation is a factor that indicates the impact of tumor-causing somatic mutations on protein function. Epitopes (peptides) containing mutations with a large functional impact are naturally eliminated by existing immune cells, while epitopes containing mutations with a relatively small functional impact may exhibit immunogenicity at a higher frequency. To determine the extent to which the functional impact of a mutation affects the immunogenicity of an epitope, for example, programs such as PROVEAN or PolyPhen-2 can be used.
[0031] Phosphorylation is a factor that indicates the degree to which the MHC class I ligand of the epitope whose immunogenicity is to be predicted is phosphorylated. To confirm the degree to which phosphorylation affects the immunogenicity of an epitope, for example, the NetMHCphosPan program can be used.
[0032] Hydrophobicity is a factor that indicates whether the epitope whose immunogenicity is to be predicted is hydrophobic or not. To confirm the degree of influence of hydrophobicity on the immunogenicity of an epitope, a program called ProtFP can be used.
[0033] Similarity to known epitopes is a factor that indicates whether the epitope whose immunogenicity is to be predicted is similar to an epitope known to be immunogenic (or non-immunogenic). To confirm the extent to which similarity to known epitopes affects the immunogenicity of an epitope, for example, the blastP program can be used.
[0034] The dissimilarity to the self-proteome with a human reference is a factor indicating whether the epitope for which immunogenicity is to be predicted is dissimilar to the self-proteome. Here, the self-proteome can be the self-proteome with a human reference. To confirm the extent to which the dissimilarity to the self-proteome with a human reference affects the immunogenicity of the epitope, for example, the pairwise2 module of the Python Bio package can be used.
[0035] Epitope (peptide) stability is a factor that indicates the stability of the epitope (peptide) itself. Traditionally, peptide-MHC binding stability has been used, but epitope (peptide) stability has not been used. If the epitope stability is low, the opportunity to bind / interact with MHC molecules decreases, and the possibility of immunogenicity decreases. To confirm the extent to which epitope (peptide) stability affects the immunogenicity of an epitope, for example, ProtParam can be used.
[0036] During antigen processing, proteasomal cleavage and / or TAP transporter efficiency can be used to predict the immunogenicity of an epitope. For example, the NetCTLpan program can be used to determine the extent to which proteasomal cleavage and TAP transporter efficiency affect the immunogenicity of an epitope.
[0037] At the antigen presentation stage of the biological process, at least some of the following can be used to predict epitopes that will elicit an immune response: MHC I binding affinity, pMHC stability, average binding affinity to MHC I alleles, cross-reactivity to MHC I alleles, average binding affinity to MHC II alleles, and cross-reactivity to MHC II alleles.
[0038] MHC class I binding affinity is a factor that indicates the degree to which an epitope whose immunogenicity is to be predicted binds to MHC class I. To confirm the degree to which MHC class I binding affinity affects the immunogenicity of an epitope, programs such as NetMHCpan and / or MHCflurry can be used.
[0039] pMHC stability is a factor that indicates the peptide-MHC stability of an epitope. To confirm the extent to which pMHC stability affects the immunogenicity of an epitope, for example, the NetMHCstabpan program can be used.
[0040] Traditionally, binding affinity for a single epitope-MHC pair has been utilized. However, epitopes and MHCs can have many-to-many interactions. Average binding affinity for MHC I alleles is a factor indicating the degree to which an epitope is recognized by multiple MHC I alleles by measuring the average binding strength of a single epitope to multiple MHC I alleles. Cross-reactivity for MHC I alleles is a factor indicating how many MHC I alleles can bind to a given epitope, calculated based on a specific reference binding affinity value.
[0041] The average binding affinity for MHC I alleles can be expressed as a value obtained by collecting MHC I alleles occurring in a population as shown in [Equation 1], calculating the binding affinity of the epitope to be calculated for all or a specific number of MHC I alleles, and using the frequency of the allele as a weight.
[0042]
number
[0043] Cross-reactivity to MHC I alleles can be calculated by collecting MHC I alleles occurring in a population, calculating the binding affinity of the epitope to be calculated for all or a specific number of MHC I alleles, and calculating the number of MHC I alleles with which the epitope reacts, where the binding affinity is lower than a reference value (e.g., 50 nM).
[0044] The average binding affinity for MHC I alleles and cross-reactivity for MHC I alleles for MHC I can also be applied to MHC II. An epitope that binds to both an MHC I allele that induces a CD8+ (cytotoxic) T-cell response and an MHC II allele that induces a CD4+ (helper) T-cell response may elicit a stronger immune response. The average binding affinity for MHC II alleles is a factor that indicates the degree to which an epitope is recognized by multiple MHC II alleles by measuring the average binding strength of a single epitope to multiple MHC II alleles. Cross-reactivity for MHC II alleles is a factor that indicates how many MHC II alleles can bind to a given epitope, calculated based on a specific reference binding affinity value.
[0045] The average binding affinity for MHC II alleles can be expressed as a value obtained by collecting MHC II alleles occurring in a population using [Equation 2], calculating the binding affinity of the epitope to be calculated for all or a specific number of MHC II alleles, and calculating the average weighted by the frequency of the alleles.
[0046]
number
[0047] Cross-reactivity to MHC II alleles can be calculated by collecting MHC II alleles occurring in a population, calculating the binding affinity of the epitope to be calculated for all or a specific number of MHC II alleles, and calculating the number of MHC II alleles with which the epitope reacts, where the binding affinity is lower than a reference value (e.g., adjusted consensus percentile rank 10%).
[0048] In the immunogenicity stage of the biological process, immunogenicity is used to predict epitopes that will induce an immune response. To confirm the extent to which immunogenicity affects the immunogenicity of an epitope, at least one of IEDB, PRIME, DeepHLApan, and DeepImmuno can be used.
[0049] In the tumor microenvironment, at least one of the following factors can be considered as predictors of epitopes that induce an immune response: inflammatory response, B-cell linear epitopes, and B-cell conformational epitopes.
[0050] Immune cells can regulate immune responses by secreting inflammatory substances such as cytokines. The ability of an epitope to induce such cytokines can be assessed, and if the epitope induces too much or too little cytokines, the immune response can be suppressed. To determine the extent to which the inflammatory response affects the immunogenicity of an epitope, at least one of PIP-EL, AIPpred, and IFNepitope can be used.
[0051] Tumor-causing somatic mutations can be recognized by B-cells. B-cells can recognize both linear epitopes in the form of peptides and conformational epitopes in protein structures. B-cell linear epitopes are elements that indicate potential linear epitopes, and B-cell conformational epitopes are elements that indicate potential conformational epitopes. To determine the extent to which B-cell linear epitopes affect the immunogenicity of an epitope, for example, the Bepipred program can be used. To determine the extent to which B-cell conformational epitopes affect the immunogenicity of an epitope, for example, Discope can be used.
[0052] FIG. 3 is a diagram showing an example of a process for learning immunogenic epitopes using an artificial intelligence model.
[0053] Referring to FIG. 3, an artificial intelligence model is used to predict immunogenic epitopes. The artificial intelligence model may be a variety of learning models, such as machine learning, neural networks, deep learning, logistic regression, AdaBoost, LogitBoost, XgBoost, support vector machines, or random forest methods. Ensemble machine learning methods (e.g., bagging, stacking, voting, boosting, etc.), which integrate various machine learning algorithms and statistical methods, may also be used. The input to the artificial intelligence model may be at least some of the various factors that may affect the immunogenicity of an epitope described in FIG. 2, and the output may be a prediction result of whether or not an epitope is immunogenic. The various factors that may affect the immunogenicity of an epitope may be values calculated using publicly available programs / algorithms / tools or values calculated using the method proposed in the present invention, and these values may be input to the artificial intelligence model. According to one embodiment, values derived using a program / algorithm / tool may have different value ranges and resolutions depending on each element or the applied program / algorithm / tool, and may be used after standardization and / or normalization.
[0054] The AI model can be trained using both immunogenic and non-immunogenic epitopes. The AI model can be trained by integrating various elements, or by each element individually. The data used to train the AI model is publicly available data stored in databases such as the Immune Epitope Database (IEDB).
[0055] According to one embodiment, immunogenic / non-immunogenic epitopes can be predicted for each element. Based on the results calculated for each element, statistical methods such as AUC (area under the curve), prAUC (precision-recal AUC), and PPV (positive predictive value) can be used to evaluate and add or remove programs / algorithms / tools, or to remove elements that affect the immunogenicity of epitopes. In this case, the addition or subtraction of programs / algorithms / tools and / or individual elements may vary depending on the disease being studied. For example, different combinations of elements and / or programs / algorithms / tools may be created depending on the type of cancer.
[0056] The result of the artificial intelligence model is a predicted value of immunogenicity for the epitope. The predicted value of immunogenicity for the epitope indicates the probability that the epitope is immunogenic or non-immunogenic, and can be a value between 0 and 1. For example, if the predicted value of immunogenicity for the epitope is close to 1, it can be determined to be immunogenic, and if it is close to 0, it can be determined to be non-immunogenic.
[0057] According to another embodiment, the result of the AI model may be a determination of whether an epitope is immunogenic. The AI model may compare the predicted immunogenicity value for the epitope with a reference value to determine whether the epitope is immunogenic or non-immunogenic. The reference value may vary depending on any one or a combination of tumor type, purpose, and situation.
[0058] FIG. 4 shows an example of a process for predicting immunogenic epitopes using an artificial intelligence model.
[0059] 4, a method for predicting the immunogenicity of an epitope according to one embodiment of the present invention may first include calculating the degree to which factors involved in the biological process of a tumor cell for the epitope and characteristics of the epitope affect the immunogenicity of the epitope (S410). The biological process may include an antigen processing stage, an antigen presentation stage, an immune stage, and the tumor microenvironment, and the factors involved in the tumor microenvironment may be at least one of an inflammatory response, a B-cell linear epitope, and a B-cell conformational epitope.
[0060] The method for predicting the immunogenicity of an epitope may include a step (S420) of performing at least one of standardization and normalization on the calculated value. Because the calculated values may have different ranges and resolutions, standardization and / or normalization may be necessary. Note that if the range and resolution of the degree of impact on the immunogenicity of the epitope derived by the calculation are the same, standardization and normalization are not necessary.
[0061] The method for predicting the immunogenicity of an epitope may include a step (S430) of inputting the value obtained by at least one of the standardization and normalization into a pre-trained artificial intelligence model to predict the immunogenicity of the epitope.
[0062] FIG. 5 shows an example of the configuration of a device for predicting immunogenic epitopes using an artificial intelligence model.
[0063] Referring to Figure 5, the device for predicting immunogenic epitopes using an artificial intelligence model may be composed of an input / output device 510, a storage device 520, and a computing device 530. The input / output device 510 (or input / output device) may be configured to receive input of an epitope or output predicted results. The input / output device 510 may be configured to receive input of an epitope from a user or from another device. Although the input / output device 510 is shown as a single device here, it may also be configured as separate devices. For example, the input device may be a mouse and / or a keyboard, and the output device may be a display such as a monitor or a speaker.
[0064] The storage device 520 may store an AI model trained to predict the immunogenicity of a learned epitope based on at least some of the factors involved in the biological process of the epitope in tumor cells and the characteristics of the epitope. The storage device 520 may store programs, algorithms, and tools necessary for data processing in addition to the AI model.
[0065] The computing device 530 calculates the degree of influence of factors involved in the biological process of tumor cells for an epitope and characteristics of the epitope on the immunogenicity of the epitope, performs at least one of standardization and normalization on the calculated value, and inputs the value after the standardization and normalization into the trained artificial intelligence model to predict the immunogenicity of the epitope. The biological process includes an antigen processing stage, an antigen presentation stage, an immune stage, and a tumor microenvironment, and the factors involved in the tumor microenvironment may be at least one of an inflammatory response, a B-cell linear epitope, and a B-cell conformational epitope.
[0066] Here, the device for predicting immunogenic epitopes using an AI model is described as comprising an input / output device 510, a storage device 520, and a calculation device 530, but multiple components may be combined into one component, and one component may be combined into multiple components. In addition, the device for predicting immunogenic epitopes using an AI model may further include a communication device, etc.
[0067] The following describes one embodiment of predicting immunogenic epitopes using the above-described method.
[0068] To learn and evaluate the immunogenicity of epitopes, data on immunogenic and non-immunogenic epitopes against neoepitopes generated in tumors were collected from public databases and / or papers.
[0069] Table 1 below shows the number of training datasets (training datasets) and evaluation datasets (evaluation datasets).
[0070] [Table 1] In this study, a total of 315 immunogenic epitopes (positive) and 4,067 non-immunogenic epitopes (negative) (i.e., epitopes and MHC I alleles that bind to them) collected from dbPepNeo, NEPdb, PRIME data, and INeo-Epp data (Table 1) were used as the training set, and 79 immunogenic epitopes (positive) and 3,125 non-immunogenic epitopes (negative) collected from McPAS-TCR, VDJdb, IEDB t-cell db, and TESLA consortium data were used as the independent set. To predict the immunogenicity of epitopes, the characteristics of each biological process and epitope were calculated using the programs and methods presented in Table 2.
[0071] [Table 2] The average binding affinity for MHC I alleles, cross-reactivity for MHC I alleles, average binding affinity for MHC II alleles, and cross-reactivity for MHC II alleles proposed in this invention were calculated using the methods described above. Here, binding affinity for MHC I alleles was calculated using NetMHCpan, and binding affinity for MHC II alleles was calculated using an IEDB-recommended program (http: / / tools.iedb.org / mhcii / ). For B-cell conformational epitopes, 3D protein structure information is required. Therefore, the epitope source protein was compared with human protein sequences in the AlphaFold Protein Structure database using blastP to identify the 3D protein structure. The 3D protein structure information containing the corresponding epitope was generated using the modeller program, followed by Discotope calculations. The ability of the calculated biological process and epitope characteristics to distinguish immunogenic and non-immunogenic epitopes was confirmed using AUC, prAUC, and P-value (Wilcoxon rank-sum test, P<0.05). All biological process and epitope characteristics except NetCTLpan (TAP efficiency) statistically significantly distinguished immunogenic / non-immunogenic epitopes.
[0072] [Table 3] A total of 22 factors statistically significant for the epitopes in the training dataset and evaluation dataset were organized into a two-dimensional matrix (rows: epitope-HLA allele, columns: 22 factors) and subjected to machine learning analysis. The 22 factors represent factors involved in immunogenicity-related biological processes and epitope characteristics, and factors or characteristics measured using multiple programs / algorithms / tools are treated as multiple factors. For example, immunogenicity can be measured using PRIME and IEDB, resulting in two factors. In this embodiment, multiple machine learning algorithm analyses were performed using the Weka program, and a final model was generated by integrating and selecting multiple machine learning algorithms. In this embodiment, the final model was selected based on the average probability of model results from nine machine learning algorithms: LogitBoost, BayesNet, CSForest, AdaBoost, Logistic, PART, NaiveBayes, RandomForest, and SMO. To compare and evaluate the performance improvement of the present invention, the performance of publicly available immunogenicity-related programs was compared using the same data. Table 4 below shows the evaluation results of the epitope immunogenicity prediction performance for the training dataset, and Table 5 shows the evaluation results of the epitope immunogenicity prediction performance for the evaluation dataset.
[0073] [Table 4]
[0074] [Table 5] In the case of the present invention, these are the results of 10-fold cross-validation, while in the case of existing publicly available programs and algorithms, they are the results of evaluation on the entire data. Referring to [Table 4], it can be seen that the method according to the present invention exhibits superior prediction accuracy in terms of AUC, prAUC, F-measure, etc. compared to existing publicly available programs and algorithms. Referring to [Table 5], it can be seen that the method according to the present invention exhibits superior prediction accuracy in terms of AUC, prAUC, F-measure, etc. compared to existing publicly available programs and algorithms, even for the evaluation dataset not used for training. Table 6 below compares the results of evaluating the epitope immunogenicity prediction performance using only the factors previously used for CD8+ T cell immunogenicity, and evaluating the epitope immunogenicity prediction performance taking into account the newly proposed factors of the present invention (functional effect of mutation, peptide stability, average binding affinity to MHC I alleles, cross-reactivity to MHC I alleles, average binding affinity to MHC II alleles, cross-reactivity to MHC II alleles, inflammatory response, B-cell linear epitopes, and B-cell conformational epitopes).
[0075] [Table 6] Referring to Table 6, when the newly proposed features of the present invention are taken into consideration, it can be seen that higher accuracy can be obtained compared to the prAUC criterion, which is suitable for evaluating imbalanced data.
[0076] Although specific aspects of the present invention have been described in detail above, it will be apparent to those skilled in the art that these specific technical details are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the true scope of the present invention is to be defined by the appended claims and their equivalents.
Claims
1. calculating the degree to which each of the factors involved in the biological process of tumor cells for the epitope and the characteristics of the epitope influences the immunogenicity of the epitope; performing at least one of standardization and normalization on the calculated values; and inputting the standardized and / or normalized values into a pre-trained artificial intelligence model to predict the immunogenicity of the epitope; The biological process is It consists of the antigen processing stage, antigen presentation stage, immune stage, and tumor microenvironment. The method for predicting the immunogenicity of an epitope, wherein the factor involved in the tumor microenvironment is at least one of an inflammatory response, a B-cell linear epitope, and a B-cell conformational epitope.
2. The step of calculating the degree of influence of each of the factors involved in the biological process for the epitope and the characteristics of the epitope on the immunogenicity of the epitope comprises: A method for predicting the immunogenicity of an epitope according to claim 1, comprising the step of calculating using an algorithm, program or tool.
3. The characteristics of the epitope include: The method for predicting the immunogenicity of an epitope according to claim 1, wherein the method is based on at least one of the following: functional impact of mutation, phosphorylation, hydrophobicity, similarity to known epitopes, dissimilarity to the self-proteome with reference to humans, and peptide stability.
4. Among the biological processes, the elements involved in the antigen processing stage are: The method for predicting the immunogenicity of the epitope of claim 1, which is at least one of proteasomal cleavage and TAP transporter efficiency.
5. Among the biological processes, the elements involved in the antigen presentation stage are:
2. The method for predicting the immunogenicity of an epitope according to claim 1, wherein the immunogenicity is at least one of MHC I binding affinity, pMHC stability, average binding affinity to MHC I alleles, cross-reactivity to MHC I alleles, average binding affinity to MHC II alleles, and cross-reactivity to MHC II alleles.
6. The average binding affinity for the MHC I / MHC II alleles is It can be calculated according to [Equation 3], [Formula 3] [Equation 1] where n is the type / number of MHC I / MHC II alleles, F is the frequency of occurrence of the corresponding MHC I / MHC II alleles in all or a specific number of alleles, and B is the binding affinity of the corresponding MHC I / MHC II allele to the epitope; The cross-reactivity to the MHC I / MHC II alleles is 6. The method for predicting the immunogenicity of an epitope according to claim 5, wherein the method comprises collecting MHC I / MHC II alleles occurring in a population, calculating the binding affinity of the epitope to be calculated for all or a specific number of MHC I / MHC II alleles, and calculating the number of MHC I / MHC II alleles with which the epitope reacts by calculating the case where the binding affinity is smaller than a predetermined reference value.
7. A computer-readable recording medium having recorded thereon a computer program for carrying out the method according to any one of claims 1 to 6.
8. an input / output device that receives an input of an epitope or outputs the result of predicting the immunogenicity of the epitope; a storage device for storing an artificial intelligence model trained to predict the immunogenicity of an epitope based on at least some of the factors involved in the biological process of tumor cells for the epitope and the characteristics of the epitope; a computing device that calculates the degree of influence of each of factors involved in the biological process of a tumor cell epitope and characteristics of the epitope on the immunogenicity of the epitope, performs at least one of standardization and normalization on the calculated value, and inputs the value after performing at least one of standardization and normalization into the trained artificial intelligence model to predict the immunogenicity of the epitope; The biological process is It consists of the antigen processing stage, antigen presentation stage, immune stage, and tumor microenvironment. The device for predicting the immunogenicity of an epitope, wherein the factor involved in the tumor microenvironment is at least one of an inflammatory response, a B-cell linear epitope, and a B-cell conformational epitope.