Methods, systems, and computer program products for determining the immunogenicity of peptides.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-15
- Publication Date
- 2026-08-14
AI Technical Summary
[0009]然而,与本发明的预测器系统相比,两种预测工具的性能都明显不足
[0027]本发明是有利的,因为机器学习分类器可以通过记录集训练,从而学习基于其氨基酸序列识别肽的免疫原性。这例如显著地简化了疫苗开发过程,因为可以基于充分了解的免疫原性指征来选择肽。此外,与本领域已知的其它免疫原性肽分类器相比,基准化测试(benchmarking)显示本发明的机器学习分类器的性能显著改善。
Smart Images

Figure CN116368570B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to methods, systems, and computer program products for determining the immunogenicity of peptides, and their applications in cancer treatment, viral inoculation, etc. Background Technology
[0002] T-cell epitope prediction plays an important role in the design of immune experiments and vaccine preparation.
[0003] Currently, most epitope prediction studies focus on peptide processing and presentation, such as proteasome cleavage, antigen-associated transporters (TAPs), and binding to the major histocompatibility complex (MHC).
[0004] However, the mechanisms of epitope immunogenicity remain unclear to date. It is generally believed that T cell immunogenicity may be influenced to varying degrees by the foreignness, accessibility, molecular weight and structure, molecular conformation, chemical properties, and physical properties of the target peptide.
[0005] For vaccine development (i.e., cancer vaccines, viral vaccines), it is crucial to identify peptides that can elicit an immune response. Vaccine efficacy is highly dependent on the immunogenicity of epitopes, i.e., the ability of potential target peptides to trigger an immune response. Current technology for immunogenicity testing involves labor- and time-intensive assays and requires safety measures in a laboratory setting.
[0006] Current strategies for identifying peptides that can elicit an immune response use the predicted binding affinity (IC50) to MHC molecules. Examples of this strategy are discussed in US 2013330335, US 2016132631, and US 2019346442. However, not all peptides with high (predicted) binding affinity also elicit an immune response.
[0007] Using this binding prediction model to predict whether a peptide is likely to be presented on the MHC, and assuming that those selected peptides are also immunogenic neoantigens, leads to a large number of false positive predictions. As few as 5% of the neoantigens identified by this model for predicting peptide binding / presentation on the MHC are capable of evoking an immune response, highlighting the need for better identification of truly immunogenic peptides.
[0008] IEDB Immunogenicity Predictor and Ineo-Epp are two publicly available immunogenicity prediction tools. IEDB is a simple model that evaluates peptides based on the position and properties of their amino acids. It is based on positional enrichment of amino acids in immunogenic and non-immunogenic peptides. The resulting peptide immunogenicity score is the sum of its individual amino acid scores. Ineo-Epp is a random forest classifier trained on HLA-I immunogenic peptides, in which several features are extracted, such as amino acid physicochemical properties and eluted ligand probability percentile (EL rank (%)) scores. Ineo-Epp uses a range of peptides along with HLA alleles as input. For antigen prediction, nine major HLA-I supertypes (A1, A2, A3, A24, B7, B27, B44, B58, B62) can be used.
[0009] However, both prediction tools are significantly inferior in performance compared to the predictor system of the present invention. The present invention aims to provide an improved immunogenicity prediction tool that surpasses the prior art. Summary of the Invention
[0010] In a first aspect, the present invention relates to a computer-executed method for determining the immunogenicity of peptides, preferably cell surface-presented peptides, comprising the following steps: - Obtain the amino acid sequence of the peptide, wherein the peptide comprises protein-derived and / or non-protein-derived amino acids; - For each amino acid source of a protein, multiple numerical indices are obtained, each associated with a specific physicochemical property; - Obtain a training dataset containing positive and negative datasets, wherein the positive dataset includes (numerical) data associated with multiple amino acid sequences of immunogenic peptides, and wherein the negative dataset includes (numerical) data associated with multiple amino acid sequences of non-immunogenic peptides. - Train a mathematical classification model on the training dataset; - Determine the likelihood that the peptide will trigger an immune response using a trained classification model; The method is characterized by comprising the following steps: - Perform principal component analysis on the numerical index of the amino acid source of each protein to obtain the principal components of the amino acid source of each analyzed protein; - Obtain the feature vectors of the amino acid sequences of the training dataset and the amino acid sequences of the peptides, wherein the feature vectors of the amino acid sequences are obtained by replacing each amino acid of the amino acid sequence with one or more corresponding principal components; The classification model is trained on the feature vectors of the training dataset; The likelihood that a peptide is immunogenic is determined by analyzing its feature vector.
[0011] According to the method in the first aspect, the negative dataset comprising (numerical) data related to multiple amino acid sequences of non-immunogenic peptides is obtained by: - Obtain the amino acid sequence of the peptide that can bind to the major histocompatibility complex (MHC) and / or presented on the major histocompatibility complex (MHC), preferably identified by binding assay and / or mass spectrometry. - Obtain the amino acid sequence corresponding to the housekeeping protein; - Compare the amino acid sequence of the peptide capable of binding to MHC and / or presented on MHC with the amino acid sequence corresponding to the housekeeping protein to determine the match between them; The negative dataset contains the matching amino acid sequence.
[0012] According to the method described in the first aspect, the negative dataset containing (numerical) data associated with multiple amino acid sequences of non-immunogenic peptides is obtained by: - Obtain the amino acid sequence that can bind to MHC and / or present peptides on MHC; - Obtain the amino acid sequence corresponding to the protein histidine; - Compare the amino acid sequences of the peptides that can bind to MHC and / or present on MHC with the amino acid sequences of the proteome to determine the match between them; The negative dataset includes amino acid sequences of peptides that can bind to MHC and / or present on MHC, said peptides being closely related to but not identical to peptides in the proteome, particularly having one or more amino acid mismatches compared to peptides in the proteome, preferably having one, two or three amino acid mismatches compared to peptides in the proteome.
[0013] According to the method described in the first aspect, the obtained MHC-presented amino acid sequence is a linear sequence.
[0014] According to the method in the first aspect, the immunogenic peptides of the positive dataset are obtained by: - Obtain the amino acid sequence of the peptide that can induce a T cell response; - Obtain the amino acid sequence corresponding to the protein histidine; - Compare the amino acid sequences that induce T cell responses with the amino acid sequences of the proteome to determine the matches between them; Positive datasets include amino acid sequences that induce T cell responses, in addition to the matched amino acid sequences.
[0015] According to the method in the first aspect, the immunogenic peptides of the positive dataset are obtained by: - Obtain the amino acid sequence of the peptide that can induce a T cell response; - Obtain the amino acid sequence corresponding to the protein histidine; - Compare the amino acid sequences that induce T cell responses with the amino acid sequences of the proteome to determine the matches between them; Positive datasets include amino acid sequences that are closely related to but not identical to peptides in the proteome and induce T cell responses, wherein the amino acid sequences inducing T cell responses have one or more amino acid mismatches compared to peptides in the proteome, preferably one, two or three amino acid mismatches compared to peptides in the proteome.
[0016] According to the method in the first aspect, the amino acid sequence that induces the T cell response is a linear sequence.
[0017] According to the method in the first aspect, the training dataset includes a positive dataset and a negative dataset, wherein the negative dataset includes substantially more records compared to the positive dataset.
[0018] According to the method in the first aspect, a feature vector of the amino acid sequence is obtained by replacing each amino acid of the amino acid sequence with at least two corresponding principal components, preferably two to ten corresponding principal components, and most preferably three corresponding principal components.
[0019] According to the method in the first aspect, the plurality of numerical indices are transformed into z-values prior to the principal component analysis.
[0020] According to the method in the first aspect, wherein the classification model is a supervised classification machine learning algorithm, more preferably, wherein the machine learning algorithm is one or more of a multilayer perceptron classifier, a Gaussian Naive Bayes classifier, a linear support vector machine, a kernel support vector machine, a K nearest neighbor classifier, or a random forest classifier, and most preferably, wherein the machine learning algorithm is a random forest classifier.
[0021] In a second aspect, the present invention relates to a computer system for determining the immunogenicity of peptides, said computer system being configured to perform a computer-executed method according to any one of the first aspects.
[0022] In a third aspect, the present invention relates to a computer program product for determining the immunogenicity of a peptide, the computer program product comprising instructions that, when executed by a computer, cause the computer to perform a computer execution method according to any one of the first aspects.
[0023] In a fourth aspect, the present invention relates to the application of the computer execution method described in the first aspect and / or the computer system described in the second aspect and / or the computer program product described in the third aspect for determining cancer treatment of a subject, particularly by determining the immunogenicity of a neoantigen presented by the subject's tumor cells.
[0024] In a fifth aspect, the present invention relates to the application of the computer execution method described in the first aspect and / or the computer system described in the second aspect and / or the computer program product described in the third aspect for screening viruses or bacteria on vaccine targets, particularly by measuring the immunogenicity of epitopes from viral or bacterial proteins.
[0025] In a sixth aspect, the present invention relates to the application of the computer execution method described in the first aspect and / or the computer system described in the second aspect and / or the computer program product described in the third aspect for characterizing the autoimmune response of a subject, particularly by measuring the immunogenicity of self-antigens presented by the subject's cells.
[0026] In a seventh aspect, the present invention relates to the use of the computer execution method described in the first aspect and / or the computer system described in the second aspect and / or the computer program product described in the third aspect for evaluating changes in the immunogenicity of a peptide when altering one or more amino acids of the amino acid sequence of the peptide, particularly simulating amino acid changes that may increase or decrease the immunogenicity of a particular peptide.
[0027] This invention is advantageous because the machine learning classifier can be trained on a set of records to learn how to recognize the immunogenicity of peptides based on their amino acid sequences. This significantly simplifies the vaccine development process, for example, as peptides can be selected based on well-understood immunogenicity indicators. Furthermore, benchmarking shows a significant performance improvement in the machine learning classifier of this invention compared to other immunogenic peptide classifiers known in the art.
[0028] Preferred embodiments of the invention are discussed in claims 2 to 11, as well as throughout the description, embodiments, and drawings. Attached Figure Description
[0029] Figure 1 An illustrative workflow of the present invention is shown.
[0030] Figures 2 to 5 The receiver operating characteristics (ROC) and precision-recall benchmark tests of different configurations and classifiers performed according to the present invention are illustrated.
[0031] Figure 6 The ROC and precision-recall benchmark tests of the present invention and other immunogenic peptide classifiers known in the art are shown.
[0032] Figure 7 Histograms and density curves showing the immunogenicity probabilities of envelope-derived peptides of SARS-CoV-2, SARS-CoV, and MERS-CoV.
[0033] Figure 8 The peptide sequence markers are displayed, showing the differences in amino acid usage between a 9-amino acid-long peptide in the immunogenic positive dataset (top panel) and a 9-amino acid-long peptide in the non-immunogenic negative dataset (bottom panel). Amino acids are not grouped. Figure 8 A) Grouped by chemical properties Figure 8 B) Grouped by charge ( Figure 8 C) or grouped by size ( Figure 8 D).
[0034] Figure 9 shows the correlation between different parameters used to predict immunogenic responses and ELISpot results for 33 peptides, all of which were predicted to be presented by netMHCpan4.0. Figure 9A The invention demonstrates the predictive potential for the immunogenicity of peptides determined from a library of peptides with positive or negative ELISpot results using the computer-executed method of the present invention, the IEDB immunogenicity predictor, and netMHCpan 4.0 (the latter based on a subject’s specific HLA allele or all HLA alleles). Figure 9B The comparison of recipient operating characteristic (ROC) curves for various immunogenicity predictors is shown, including the computer-executed method of this invention, the IEDB immunogenicity predictor, and netMHCpan 4.0, which is based on a subject’s specific HLA allele or all HLA alleles. Figure 9C Cumulative gain plots are shown for several different immunogenicity predictors, including the computer-executed method of this invention, the IEDB immunogenicity predictor, and netMHCpan 4.0, which is based on a specific HLA allele or all HLA alleles of the subject. The cumulative gain curves illustrate the percentage of true positives identified using the least probable portion of the sample.
[0035] Figure 10 This diagram shows the relationship between the response rate and duration of response to CTLA-4 blockade in melanoma patients. Patients were divided into two groups based on the median mutation (tumor mutational burden, TMB), the number of presented mutations, or the number of immunogenic mutations, and represented as "high" (solid line) or "low" (dashed line). Detailed Implementation
[0036] This invention relates to methods, systems, and computer program products for determining the immunogenicity of peptides. Furthermore, this invention relates to the application of said methods, systems, and / or products in cancer treatment and viral inoculation. The invention will be described in detail below, preferred embodiments will be discussed, and the invention will be illustrated by non-limiting examples.
[0037] Unless otherwise defined, all terms used in disclosing this invention, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. The guidance, including terminology definitions, is provided to better understand the teachings of this invention.
[0038] As used herein, the following terms have the following meanings: As used herein, “a,” “an,” and “the” refer to singular and plural indicators, respectively, unless the context clearly indicates otherwise. For example, “compartment” refers to one or more compartments.
[0039] As used herein, “about” refers to a measurable value, such as a parameter, quantity, and duration, and means a variation covering less than + / -20%, preferably less than + / -10%, more preferably less than + / -5%, even more preferably less than + / -1%, and still more preferably less than + / -0.1% of the specified value, such variations being suitable for implementation in the disclosed invention to date. However, it should be understood that the value referred to by the modifier “about” is itself specifically disclosed.
[0040] As used herein, “comprising,” “including,” and “consisting of” are synonymous with “having” or “containing” and are inclusive or open-ended terms that specify the presence of the following content, such as components, and do not exclude or preclude the presence of additional, unlisted elements, members, steps, etc., known in the art or disclosed herein.
[0041] Furthermore, the terms first, second, third, etc., used in the specification and claims are used to distinguish similar elements and are not necessarily used to describe order or chronological order, unless specifically stated otherwise. It should be understood that such terms are interchangeable where appropriate, and that embodiments of the invention described herein can operate in orders other than those described or shown herein.
[0042] Although the terms “one or more” or “at least one”, such as one or more or at least one member of a group of members, are self-evident, by further examples, the term specifically covers references to any one of the members or to any two or more of the members, such as any ≥3, ≥4, ≥5, ≥6 or ≥7 of the members, and up to all of the members.
[0043] The range of values represented by the endpoints includes all values and ratings contained within that range, as well as the endpoints they represent.
[0044] Unless otherwise defined, all terms used in this disclosure, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. Further guidance, including definitions of terms used in the specification, is provided to better understand the teachings of this invention. The terms or definitions used herein are for illustrative purposes only.
[0045] Throughout this specification, references to "one embodiment" or "implementation" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing throughout this specification do not necessarily refer to the same embodiment, but may refer to the same embodiment. Furthermore, in one or more embodiments, particular features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those skilled in the art from this disclosure. Moreover, while some embodiments described herein include some but not others, features in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments, as will be understood by those skilled in the art. For example, in the following claims, any claimed embodiment may be used in any combination.
[0046] In a first aspect, the present invention relates to a computer-executed method for determining the immunogenicity of a peptide. The method preferably includes the step of obtaining the amino acid sequence of the peptide. The method preferably further includes the step of obtaining a plurality of numerical indices, each associated with a physicochemical property, for each protein-derived amino acid. The method preferably further includes the step of obtaining a training dataset comprising a positive dataset and a negative dataset, wherein the positive dataset comprises (numerical) data associated with a plurality of amino acid sequences of immunogenic peptides, and wherein the negative dataset comprises (numerical) data associated with a plurality of amino acid sequences of non-immunogenic peptides. The method preferably further includes the step of training a mathematical classification model on the training dataset. The method preferably further includes the step of determining the probability that the peptide will elicit an immune response using the trained classification model. The method preferably further includes the step of performing principal component analysis on the numerical indices of each protein-derived amino acid to obtain principal components for each analyzed protein-derived amino acid. The method preferably further includes the step of obtaining feature vectors of the amino acid sequences of the training dataset and the amino acid sequences of the peptide, wherein the feature vectors of the amino acid sequences are obtained by replacing each amino acid of the amino acid sequence with one or more corresponding principal components, preferably one or more major principal components. Preferably, the classification model is trained on the feature vectors of the training dataset. Preferably, the characteristic vector of the peptide determines the likelihood that the peptide is immunogenic.
[0047] In a second aspect, the present invention relates to a computer system for determining the immunogenicity of peptides. The computer system is configured to perform a computer-executed method according to the first aspect.
[0048] In a third aspect, the present invention relates to a computer program product for determining the immunogenicity of peptides, wherein the computer program product includes instructions that, when executed by a computer, cause the computer to perform a computer execution method according to the first aspect.
[0049] In a fourth aspect, the present invention relates to the application of the computer execution method of the first aspect and / or the computer system of the second aspect and / or the computer program product of the third aspect for determining cancer treatment of a subject, particularly by determining the immunogenicity of a neoantigen presented by the subject's tumor cells.
[0050] In a fifth aspect, the present invention relates to the application of the computer-executed method of the first aspect and / or the system of the second aspect and / or the product of the third aspect for screening viruses on inoculation targets, particularly by determining the immunogenicity of epitopes from viral proteins.
[0051] In a sixth aspect, the present invention relates to the application of the computer-executed method of the first aspect and / or the system of the second aspect and / or the product of the third aspect for characterizing the autoimmune response of a subject, particularly by determining the immunogenicity of an autoantigen presented by the subject's cells.
[0052] In a seventh aspect, the present invention relates to the application of the computer-executed method of the first aspect and / or the system of the second aspect and / or the product of the third aspect for screening bacteria on an inoculated target, particularly by measuring the immunogenicity of epitopes from bacterial proteins.
[0053] In an eighth aspect, the present invention relates to the application of the computer-executed method of the first aspect and / or the system of the second aspect and / or the product of the third aspect for evaluating changes in the immunogenicity of a peptide when one or more amino acids of the amino acid sequence of the peptide are altered. In particular, it relates to applications for simulating amino acid changes that may increase or decrease the immunogenicity of a specific peptide.
[0054] This invention provides computer-executed methods, computer systems, and computer program products for determining the immunogenicity of peptides, and applications of any of said methods, systems, or products for cancer treatment, viral inoculation, characterizing the autoimmune response of a subject, screening for bacteria or other pathogens on an inoculation target, and evaluating changes in the immunogenicity of said peptide when its amino acid sequence is altered. Those skilled in the art will understand that the method is implemented in a computer program product and executed using a computer system. Those skilled in the art will also appreciate that immunogenicity prediction can be used for cancer treatment, viral inoculation, characterizing the autoimmune response of a subject, and screening for bacteria on an inoculation target. Therefore, all eight aspects of this invention are addressed together below.
[0055] A simplified embodiment of the invention preferably provides the amino acid sequence of the peptide to which immunogenicity is to be determined. Preferably, the peptide is a cell surface-presented peptide. More preferably, it is a major histocompatibility complex (MHC)-binding peptide presented on the cell surface. The peptide preferably binds to MHC class I and elicits a CD8 T cell response, or binds to MHC class II and elicits a CD4 T cell response. Therefore, the invention is applicable to MHC class I alleles in some embodiments and to MHC class II alleles in others. Preferably, the peptide sequence (query sequence) used to predict immunogenicity using the invention should be known to be MHC-presented. This has the advantage that the query peptide will more closely resemble peptides from which the training dataset (see below) is derived and which are known to be MHC-presented. Therefore, the characteristic relating to immunogenicity rather than MHC presentation is the most significant differentiating factor between positive and negative peptides (see below).
[0056] The peptide can be obtained from any subject. As used herein, “subject” refers to a term known in the art and should preferably be understood as a human or animal, most preferably a human. As used herein, “animal” preferably refers to a vertebrate, more preferably birds and mammals, and even more preferably mammals. As used herein, “subject in need” should be understood as a subject who will benefit from (preventive) treatment, such as cancer or viral vaccination. Furthermore, the peptide thus determined to be immunogenic comprises an amino acid (aa) sequence and preferably has a length of 9-11 aa, or 8-12, or 13-25.
[0057] A preferred embodiment of the invention provides multiple numerical indices for each protein amino acid, wherein each numerical index is associated with the physicochemical properties of the corresponding protein amino acid. As used herein, “protein-derived amino acids” are amino acids that are biosynthetically incorporated into proteins during translation. “Protein-derived” means “producing protein.” In known life, there are 22 genetically encoded (protein-derived) amino acids, 20 of which are in the standard genetic code, and the other two can be incorporated through specific translational mechanisms. Conversely, “non-protein-derived amino acids” are amino acids that are not incorporated into proteins (such as GABA, L-DOPA, or triiodothyronine), do not substitute for genetically encoded amino acids, or are not directly produced and isolated through standard cellular mechanisms (such as hydroxyproline). The latter are typically produced by post-translational modifications of proteins. By obtaining information about protein-derived amino acids, information about the “building blocks” of each peptide in the subject is obtained. Therefore, peptides may contain protein-derived amino acids, non-protein-derived amino acids, or a combination of both.
[0058] Physicochemical properties are preferably obtained from databases known in the art, such as the AA index database on the GenomeNet website operated by the Bioinformatics Center at Kyoto University. The AA index is a database of numerical indices representing various physicochemical and biochemical properties of amino acids and amino acid pairs. The AA index currently consists of three segments. The segment of interest is AAindex1, which contains an amino acid index of 20 values. Examples of physicochemical properties include, but are not limited to, hydrophobicity, free energy, and amino acid distribution.
[0059] A simplified embodiment of the invention preferably provides training a classification model on a training dataset and determining the likelihood that the peptide will elicit an immune response using the trained classification model. In the context of training mathematical models such as classification algorithms, the following terminology is used, and is thus further explained with the aid of guidance.
[0060] A "training set" is a set of data observations (also called "records," "examples," or "instances") used to train or learn a model. The analytical model has parameters that need to be estimated to make good predictions. This translates to finding the optimal parameter values for the analytical model. For this, we use a training set to find or estimate these optimal parameter values. Once we have a trained model, we can use it for predictions. In supervised classification tasks, class labels (e.g., "immunogenic," "non-immunogenic") are also attached to each observation to estimate the optimal parameter values. This allows the algorithm to be trained on patterns that help identify fraud cases.
[0061] A "validation set" is used for models with parameters that cannot be directly estimated from the data. However, to find optimal values for these parameters (called hyperparameters), a validation set is used. Typically, a set of candidate values for the hyperparameters can be identified. One candidate value is chosen, the model is trained on the training set, and the predictive performance is evaluated on the validation set. Then, the next candidate value is chosen, and this process continues in a similar manner until all candidate values have been tried. Finally, for each candidate value, a corresponding estimate of the predictive performance is obtained. Based on the performance estimated on the validation set, a candidate value corresponding to the best performance can be chosen. The training and validation sets are preferably strictly separated throughout the process to obtain reliable performance estimates. That is, observations in the validation set cannot be found in the training set (or, for that matter, the test set). Alternatively, the training and validation sets are not strictly separated. This is, for example, the case of cross-validation.
[0062] The "test set," also known as the "holdout," is a set of data observations used to test whether the trained model makes good predictions. In other words, during model evaluation, the true values of the test observations are known, and one can check how many predictions are correct by comparing them to the true values. It's important to note that class labels here are only used to evaluate the predictive performance (e.g., accuracy) of the classification model. That is, observations in the test set cannot be from the training or validation sets. This strict separation is crucial because one expects the model to predict observations not used during training. Only when this is guaranteed and the model shows good performance can one be certain that the model will also perform well on new, previously unseen data.
[0063] The "retention strategy" or "single training-test split strategy" refers to the simplest split, where the data is divided into two subsets: one for training and one for testing. One can train the model using the former and then test it using the latter. Note that the training and testing process is performed only once. This data split is randomized, meaning observations are randomly assigned to either the training or test set. Typically, for a set of candidate models, performance is evaluated on the test set, and the best model is selected. Furthermore, to account for overfitting, it is often necessary to compare the trained and validated models. Some models have parameters that cannot be directly estimated from the data. These are called hyperparameters. One can rely on a validation set to find the best model. Here, the data can be split into three subsets: one for training, one for validation, and one for testing. This split is also randomized. With the help of the validation set, one can find the model with the best hyperparameter values (i.e., model selection) and ultimately evaluate the best model on the test set. Note that the selection of the best predictive model from a set of various candidate models is based on performance measured on the test set; for example, one might need to determine whether a logistic regression model, a decision tree, or a random forest is the best performing model. The performance of the test set is crucial for making this determination. Once the final predictive model is found, it can be implemented in the operating system for predicting new, previously unseen data.
[0064] The term "k-fold cross-validation strategy" refers to an alternative to the simple training-test split. It corresponds to repeated training-test splits, whereby the test set is systematically shifted. The performance on the resulting test set is then averaged. The advantage of this strategy is that each observation will be considered once in the test set. However, more importantly, the estimated predictive performance becomes more reliable, providing a better characterization of the model's generalization performance.
[0065] The training dataset of this invention includes at least a positive dataset containing multiple amino acid sequences of immunogenic peptides. As used herein, "immunogenicity" or "immune responsiveness" arises from biological materials detected by the body's immune system as foreign objects. Immunoreactive biological materials, such as peptides, are detected through antigenic responses on cells. A biochemical cascade then occurs, thereby causing T helper cells to migrate toward the biological material. Therefore, a classification model is trained to identify immunogenicity.
[0066] According to a preferred embodiment, the immunogenic peptides in the positive dataset are obtained by: obtaining the amino acid sequence of a peptide capable of inducing a T cell response; obtaining the amino acid sequence corresponding to a proteomic peptide; comparing the amino acid sequence inducing a T cell response with the proteomic amino acid sequence to determine a match between them; the positive dataset includes amino acid sequences inducing a T cell response other than the matched amino acid sequence. Preferably, the obtained amino acid sequence inducing a T cell response is a linear sequence.
[0067] According to an embodiment for determining the immunogenicity of autoimmune peptides, the immunogenic peptides in the positive dataset are obtained by: obtaining the amino acid sequence of a peptide capable of inducing a T cell response; obtaining the amino acid sequence corresponding to a proteomic peptide; comparing the amino acid sequence inducing a T cell response with the proteomic amino acid sequence to determine a match between them; the positive dataset includes the matching amino acid sequence inducing a T cell response. Preferably, the obtained amino acid sequence inducing a T cell response is a linear sequence.
[0068] According to an embodiment for determining the immunogenicity of peptides resulting from mutations, immunogenic peptides in a positive dataset are obtained by: obtaining the amino acid sequence of a peptide capable of inducing a T-cell response; obtaining the amino acid sequence corresponding to a proteomic peptide; comparing the amino acid sequence inducing a T-cell response with the proteomic amino acid sequence to determine a match between them; the positive dataset includes amino acid sequences that are closely related to, but not identical to, peptides in the (e.g., human) proteome and induce a T-cell response. Preferably, "closely related but not identical" means allowing for 1, 2, or 3 amino acid mismatches between the peptides. Those skilled in the art will know methods for comparing a query peptide (the peptide to be evaluated) with peptides in the human proteome, and how to identify the number of mismatches. The advantage is that by filtering out peptides identical to those in the human proteome, the peptides in the positive dataset are more closely similar to the peptide of interest, which is a peptide with a modified sequence due to mutation.
[0069] According to an embodiment, the immunogenic peptides in the positive dataset are obtained by: obtaining the amino acid sequence of a peptide capable of inducing a T cell response; obtaining the amino acid sequence corresponding to a proteomic peptide; comparing the amino acid sequence inducing a T cell response with the proteomic amino acid sequence to determine a match between them; the positive dataset includes amino acid sequences inducing a T cell response that have one or more, two or more, three or more, four or more, five or more… mismatches with peptides in the (e.g., human) proteome. Those skilled in the art know methods for comparing a query peptide (the peptide to be evaluated) with peptides in the (human) proteome and how to identify the number of mismatches. This positive dataset can be used when a model specific to a foreign antigen (e.g., non-human or non-mammal, such as bacteria or viruses, compared to a human proteome sequence) is required. The positive dataset obtained with respect to peptides having multiple mismatches with peptides in the (human) proteome can be similar to a positive dataset obtained from the foreign (e.g., bacterial or viral) antigen.
[0070] Peptides capable of inducing T cell responses can be experimentally characterized using human T cell assays. T cell assays can measure cytokine release, cytotoxicity, or qualitative T cell binding to antigen-presenting cells (APCs). Cytokines used in T cell assays can be selected from interferon-γ (IFNγ), tumor necrosis factor-α (TNFα), interleukin-2 (IL-2), interleukin-4 (IL-4), interleukin-5 (IL-5), interleukin-6 (IL-6), interleukin-8 (IL-8), interleukin-10 (IL-10), interleukin-17 (IL-17), interleukin-21 (IL-21), interleukin-22 (IL-22), granzyme A, and granzyme B. In some variants, qualitative T cell binding can be measured by MHC multimer staining. In some variants, T cell assays can be performed in vitro or ex vivo. Experimentally characterized peptides recognized by T cells can be selected from the Immunoepitope Database (IEDB).
[0071] According to a preferred embodiment, the training dataset also includes a negative dataset.
[0072] The negative dataset can be obtained by: obtaining the amino acid sequences of peptides capable of binding to and / or presenting on the major histocompatibility complex (MHC), preferably identified by means such as binding assays and / or mass spectrometry; obtaining the amino acid sequences corresponding to housekeeping protein histidine peptides; comparing the MHC-presented / MHC-binding amino acid sequences with the housekeeping protein histidine amino acid sequences to determine a match between them; the negative dataset includes the match between them, i.e., the MHC-presented amino acid sequences corresponding to housekeeping proteins. The binding assays and / or mass spectrometry used to identify peptides capable of binding to and / or presenting on MHC can be any type of binding assay and / or mass spectrometry known in the art. The type of binding assay and / or mass spectrometry is known to those skilled in the art. Preferably, the obtained MHC-presented amino acid sequences are linear sequences.
[0073] In this article, "housekeeper proteome" refers to the proteome that participates in the basic functions of cells or cell groups in an organism.
[0074] According to an implementation method for determining the immunogenicity of peptides resulting from mutations, a negative dataset is obtained by: obtaining the amino acid sequences of peptides capable of binding to and / or presenting on major histocompatibility complexes; obtaining the amino acid sequences corresponding to proteomic peptides; comparing the MHC-presented / MHC-bound amino acid sequences with the proteomic amino acid sequences to determine a match between them; the negative dataset includes MHC-presented / MHC-bound amino acid sequences that are closely related to but not identical to peptides in the human proteome. Preferably, "closely related but not identical" means allowing for 1, 2, or 3 amino acid mismatches between the peptides. Those skilled in the art will know methods for comparing the query peptide (the peptide to be evaluated) with peptides in the (e.g., human) proteome and how to identify the number of mismatches. The advantage is that by filtering out peptides identical to those in the (human) proteome, the peptides in the negative dataset more closely resemble peptides with altered sequences due to mutations. By selecting this negative dataset, training a classification model to distinguish between mutated and non-mutated peptides, rather than immunogenic and non-immunogenic peptides, is avoided. Another or related advantage is that the negative dataset will be very similar to the peptides to be tested by the model, thereby increasing the model's specificity.
[0075] According to one embodiment, the negative dataset is obtained by: obtaining the amino acid sequence of a peptide capable of presenting a major histocompatibility complex; obtaining the amino acid sequence corresponding to a proteomic peptide; comparing the MHC-presented amino acid sequence with the proteomic amino acid sequence to determine a match between them; the negative dataset includes MHC-presented amino acid sequences that have one or more, two or more, three or more, four or more, five or more… mismatches with peptides of the (e.g., human) proteome. Those skilled in the art will know methods for comparing a query peptide (the peptide to be evaluated) with peptides of the (human) proteome and how to identify the number of mismatches. This negative dataset can be used when a model specific to a foreign (non-human or non-mammal, such as bacterial or viral) antigen is required. The negative dataset obtained from peptides having multiple mismatches with peptides of the human proteome can be similar to the negative dataset obtained from the foreign (e.g., non-human or non-mammal, such as bacterial or viral, if compared with human proteome sequences) antigen.
[0076] According to a preferred embodiment, the training dataset includes positive and negative datasets, wherein the negative dataset contains substantially more records than the positive dataset. Preferably, the ratio is at least 3:1. According to another embodiment, the training dataset includes a positive dataset and a negative dataset, wherein the negative dataset contains a similar number of records to the positive dataset, or fewer records than the positive dataset.
[0077] According to a preferred embodiment, the classification model is a classification machine learning algorithm, preferably a supervised classification machine learning algorithm, and more preferably, the machine learning algorithm is one or more of the following: multilayer perceptron classifier, decision tree classifier, Gaussian Naive Bayes classifier, Gaussian process classifier, stochastic gradient descent classifier, linear support vector machine, kernel support vector machine, K nearest neighbor classifier, or random forest classifier. After benchmarking, the random forest classifier shows the best results. Furthermore, the inventors noted that the random forest classifier algorithm requires minimal effort in feature engineering and parameter tuning. Therefore, the most preferred classification model according to the present invention is the random forest classifier.
[0078] A preferred embodiment of the invention involves performing principal component analysis (PCA) on the numerical index of each protein amino acid before training the classification model to obtain several principal components for each analyzed protein amino acid. The “principal component analysis” or simply “PCA” used herein is used to reconstruct the most distinctive feature subspace, which is then used as input in representation-based classification for prediction. PCA improves the predictability and tractability of the data, but also eliminates noise if limited to a finite number of principal components. The transformation of amino acid properties in the principal components reduces dimensionality and focuses on the most unique properties among amino acids.
[0079] A preferred embodiment of the present invention provides that, before training a classification model, feature vectors of the amino acid sequences of a training dataset and the amino acid sequences of the peptide are obtained, wherein the feature vectors of the amino acid sequences are obtained by replacing each amino acid of the amino acid sequence with one or more major corresponding principal components; wherein the classification model is trained on the feature vectors of the training dataset; and wherein the likelihood of the peptide having immunogenicity is determined for the feature vectors of the peptide. Translating the peptide into a feature vector preserves the positional information of its physicochemical properties.
[0080] Preferably, one or more principal components should be understood as one or more first principal components, more preferably 10 first principal components, even more preferably 9 first principal components, even more preferably 8 first principal components, even more preferably 7 first principal components, even more preferably 6 first principal components, even more preferably 5 first principal components, even more preferably 4 first principal components, and most preferably 3 first principal components.
[0081] Therefore, the feature vector of the amino acid sequence is obtained by replacing each amino acid in the amino acid sequence with 2 to 10 corresponding principal components, or any range therebetween. Most preferably, the feature vector of the amino acid sequence is obtained by replacing each amino acid in the amino acid sequence with 3 corresponding principal components. The inventors have noted that, in their experience, limiting the feature vector to only the first three principal components prevents excessive processing requirements while preserving sufficient information for very accurate predictions.
[0082] In addition, prior to principal component analysis, it is preferable to convert multiple numerical indices into z-values, which represent the probability that each feature will conform to the normal value.
[0083] According to one embodiment, the present invention also provides training of an unsupervised training algorithm for further defining the characteristics of immunogenic and / or non-immunogenic peptides.
[0084] According to one embodiment, the present invention also provides for obtaining peptide-level properties, such as peptide solubility, molecular weight, peptide stability, peptide three-dimensional structure, global localization of the peptide in the genome (e.g., encoded as distance from the nuclear center), similarity to non-mutant peptides, etc., which can be used to further train classification models.
[0085] As will be apparent to those skilled in the art, the computer-executed method, computer system, and computer program product of the present invention can be used to identify immunogenic peptides (immunogenic epitopes) within a group of MHC-presented peptides (epitaxies) and to obtain identification of therapeutic targets. Furthermore, it can be used to assess changes in the immunogenicity of a peptide when one or more amino acids of the peptide's amino acid sequence are altered, particularly mimicking amino acid changes that can increase or decrease the immunogenicity of a specific peptide.
[0086] Description of embodiments and accompanying drawings
[0087] The invention is further described by way of the following non-limiting embodiments, which further illustrate the invention and are not intended to, nor should they be construed as, limiting the scope of the invention.
[0088] Example 1: Preferred Implementation
[0089] This embodiment relates to a preferred implementation of predicting the likelihood that peptides presented on the cell surface will trigger an immune response.
[0090] This embodiment accepts input peptides of 9-11 amino acids in length presented on the cell surface (determined by prediction algorithms or experiments).
[0091] The output of this embodiment is a probability score between 0 and 1, describing the likelihood that the peptide is immunogenic.
[0092] Preprocessing
[0093] The input is preprocessed as follows: exist GenomeNet website operated by Kyoto University Bioinformatics Center Physicochemical properties of 20 protein-derived amino acids were downloaded from the AAindex database (date 13.2.2017). AAindex is a database of numerical indices representing various physicochemical and biochemical properties of amino acids and amino acid pairs. AAindex is currently composed of three sections: AAindex1 represents the amino acid index with 20 numerical values, AAindex2 represents the amino acid mutation matrix, and AAindex3 represents the statistical protein contact potential. All data are from published literature. The values from AAindex1 were converted to z-values (mean set to 0, standard deviation to 1). Principal component analysis was performed on the scaled physicochemical properties using the sklearn PCA module to reduce the dimensionality of the data. Then, each amino acid in the peptide was translated into the values of the three first principal components of the corresponding peptide, resulting in a length of (9-11). The feature vector is 3. If the peptide is shorter than 11 amino acids, the missing amino acids are encoded as three zeros at their respective N-termini (peptide initiation). Therefore, the feature vector is always 11. 3.
[0094] Supervised classifier: This example uses the Random Forest classifier module from the sklearn python package, specifically: RandomForestClassifier(n_estimators=1000,criterion='entropy',max_depth=None,min_samples_split=3,min_samples_leaf=1,max_features=2,bootstrap=False).
[0095] Training classifier
[0096] The training dataset includes peptides of lengths 9, 10, and 11 amino acids, and three times more negative peptides than positive peptides. The training dataset is transformed into a matrix. Each peptide in the dataset is translated into a feature vector as described above. Therefore, each row in the matrix corresponds to the feature vector of one peptide in the dataset. This matrix is used to train the random forest model described above and to store the trained model.
[0097] Training data
[0098] The training data used to train the supervised classifier is constructed as follows: Negative data (non-immunogenic peptides): Peptides derived from endogenous housekeeping genes are considered non-immunogenic because they are frequently present on all healthy cells. Therefore, the negative training dataset is constructed from peptides derived from said housekeeping genes, which have been shown to present on MHC molecules. MHC presentation prediction is not the goal of this module. Furthermore, we do not want it to be a factor in this module. Therefore, positive and negative peptides are MHC-binding peptides. IEBDB data on MHC-binding peptides were imported from the Immune Epitope Database & Tools website (version Feb. 19, 2020). Peptides from Homo sapiens were also filtered for linear sequences, producing at least one positive assay with a length of 9, 10, or 11 amino acids. The resulting peptide list contains only unique peptides.
[0099] In addition, housekeeping genes were obtained from the Human Protein Atlas website. These were filtered to exclude genes annotated as "involved in disease" or "RNA cancer specific." Fast sequences of housekeeping gene proteins were retrieved from Uniprot based on the gene IDs from the previous step, and a blast database was constructed from these sequences. The IEDB MHC-binding peptides were then queried against the housekeeping gene blast database. All perfectly matching peptides (peptides produced by these housekeeping genes) were retained as negative datasets.
[0100] Positive dataset (immunogenic peptides): Positive data (i.e., immunogenic peptides) were obtained from the IEDB database (dated February 19, 2020): peptides with relevant T-cell assay data. Furthermore, peptides were filtered to obtain linear sequences, tested on humans, yielding at least one positive assay, and having a length of 9 amino acids. The resulting peptide list contained only unique peptides. After filtering, peptides were ensured to have no overlap with the human proteome by filtering for complete matches to human proteins. This was achieved by filtering for peptides that perfectly matched a blast database prepared from the human proteome.
[0101] Figure 1 The workflow of this embodiment as described above is illustrated. The training dataset 1 includes a positive dataset 102 and a negative dataset 103. The negative dataset, containing the matched amino acid sequences, is obtained by comparing MHC-presented amino acid sequences / peptide sequences capable of binding to and / or presented on MHC 104 with housekeeping amino acid sequences 105 to determine matches 106. The positive dataset is obtained by comparing T-cell-inducing amino acid sequences 107 with proteomic amino acid sequences 108 to determine matches 109, wherein the positive dataset includes T-cell-inducing amino acid sequences other than the matched amino acid sequences. In the training dataset, the negative and positive datasets are compiled at a 3:1 ratio 110.
[0102] Figure 1Preprocessing according to this embodiment is also illustrated. First, multiple numerical indices 111 related to physicochemical properties are obtained for each protein amino acid. Second, principal component analysis 112 is performed on the numerical indices for each protein-derived amino acid. Thus, the principal numerical components 113 of each protein-derived amino acid are determined. These can then be used to train peptides in dataset 114 and peptides of interest 115, 116 for peptide-to-feature translation. The resulting feature matrix 117 is used to train 118 a random forest classifier 119 to obtain a trained classifier 120, which can be used to predict 122 the immunogenicity of the peptide of interest using the feature vector 121 thus determined. The predictions obtained according to this embodiment are preferably displayed as scores 123.
[0103] Figure 2 The receiver operating characteristics (ROC) of this example are shown. Therefore, the true positive rate (the proportion of actual immunogenic epitopes, referred to as immunogenicity, i.e., sensitivity), is shown as a function of the false positive rate (the proportion of non-immunogenic epitopes, referred to as immunogenicity, i.e., 1-specificity). The precision-recall curve shows the trade-off between precision and recall for different thresholds. The area under the curve (AUC) of the ROC curve is 0.88.
[0104] Figure 4 The ROC results are shown when training with different positive to negative data ratios while validating at a 1:3 ratio. The results are improved by increasing the ratio of negative to positive data.
[0105] Figure 5 The metrics for training and validation using 1, 2, 3, 5, 10, or 20 principal components are shown. Performance increases when using 2 or more principal components. However, processing speed increases significantly when using more than 10 principal components.
[0106] Example 2: Benchmark Classifier Benchmark Classifier )
[0107] This example illustrates the performance of supervised classification algorithms available in scikit-learn.
[0108] Figure 3 Receiver operating characteristics (ROC) and precision-recall curve benchmarks for different classifiers known in the art are shown. The workflow characteristics of Example 1 are followed for the implementation of each classifier.
[0109] The classification algorithms tested included: Multilayer Perceptron (MLP) classifiers. Figure 3 ;1 1) Random Forest Classifier Figure 3 ;1 2), K-Nearest Neighbor Classifier Figure 3 ;1 3) Gaussian Naive Bayes classifier ( Figure 3 ;2 1) Gaussian process classifier ( Figure 3 ;2 2), Support Vector Classifier ( Figure 3 ;2 3). The corresponding AUC values are: 0.77; 0.88; 0.73; 0.74; 0.82; 0.82.
[0110] Although all classifiers are effective for the current workflow, the Random Forest classifier significantly outperforms the classifiers tested on other tests.
[0111] Example 3: Benchmark Prior Art
[0112] This embodiment relates to the benchmarking test of the workflow of Example 1 relative to other immunogenic peptide classifiers known in the art.
[0113] In the benchmark study, the workflow according to Example 1 will be compared with two other publicly available immunogenicity prediction tools, in particular IEDB immunogenicity prediction and Ineo-Epp.
[0114] IEDB immunogenicity prediction values are models that evaluate peptides based on the position and properties of their amino acids. Ineo-Epp is a random forest classifier trained on HLA-I immunogenic peptides, in which several features are extracted, such as amino acid physicochemical properties and eluted ligand probability percentile (EL rank (%)) scores. Ineo-Epp uses a range of peptides and HLA alleles as inputs. For antigen prediction, nine major HLA-I supertypes (A1, A2, A3, A24, B7, B27, B44, B58, B62) can be used.
[0115] For this benchmark study, immunogenic peptides were retrieved from the Immunoeptope Database (IEDB) (accessed 15.06.2020). To avoid bias in the benchmark data for epitopes present in the training data of any tools, antigenic epitope data were retrieved from the most recent 2020 publications. The data was further filtered (linear epitopes, human, positive T cell assays) and epitopes were removed from those already preserved in earlier versions of the database. This reduced the number of positive peptides from 2213 to 64 unique epitopes, all of which are 9 amino acids in length. As negative data, non-immunogenic peptides published in “Chowell, D. et al., TCR contact residue hydrophobicity is a hallmark of immunogenic CD8+ T cell epitopes. Proc. Natl. Acad. Sci. USA 112, E1754–E1762 (2015)” were used. The non-immunogenic peptides in the dataset are ligand-eluting MHC-I presented self-peptides that have been processed by antigens and bound to MHC. For amino acid lengths of 9-11, filtering yields 4254 unique non-episodes.
[0116] Since all peptides in the benchmark dataset are MHC conjugates, they can all be used as inputs for the workflow of Example 1 and for IEDB immunogenicity prediction. For Ineo-Epp, predictions were made for 9 HLA-I supertypes, considering only the scores of the predicted conjugates.
[0117] Based on the predicted scores of the three tools, a receiver operator characteristic (ROC) curve was constructed, and... Figure 6 The curves shown in the figure illustrate the true and false positive rates for predictions with lower cutoff values.
[0118] In the comparison of the three tools, workflow 41 according to Example 1 outperformed IEDB immunogenicity predictor 43 and Ineo-Epp 42. Example 1 had an AUC of 0.84, which is close to the internal validation AUC of 0.88. The IEDB immunogenicity predictor had an AUC of 0.57. Ineo-Epp also had lower performance (AUC of 0.65).
[0119] Example 4: Cancer Vaccine
[0120] To demonstrate the effectiveness of the immunogenicity scoring framework outlined in Example 1, the framework was applied to data from published studies on breast cancer vaccines.
[0121] This study, titled "Dillon, PM et al. A pilot study of the immunogenicity of a9-peptide breast cancer vaccine plus poly-ICLC in early stage breast cancer. J. Immunother. Cancer 5, 1–10 (2017)," evaluated the immune response of 12 breast cancer patients to a vaccine composed of nine MHC class I restricted breast cancer-related peptides. CD8+ T cell responses to the vaccine were evaluated using direct and stimulated interferon-gamma ELISpot assays. This ELISpot assay quantitatively assesses the frequency of cytokine secretion by cells in response to stimuli, and is therefore a suitable method for evaluating the immunogenicity of the vaccine.
[0122] In stimulated ELISpot assays, peptide-specific CD8+ T cell responses were detected in 4 of 11 evaluable patients. Two ELISpot responses were observed in response to the modified HLA-A2 CEA571-579 peptide (YLS-D antigen), and two responses were observed in response to HLA-A3 CEA27-35 (HLF antigen). One patient showed a borderline response to HLA-A3 MAGE-A196-104.
[0123] The framework of this embodiment is used to score the immunogenicity of each of the nine peptides in the vaccine (Table 1). As shown in Table 1, the two antigens (YLS-D and HLF antigens) that produced positive ELISpot assays had the highest backbone / immunogenicity scores.
[0124]
[0125] Table 1. Vaccine peptides with ELISpot response and frame score.
[0126] Example 5: Viral Immunogenicity
[0127] Immunogenicity prediction of epitopes derived from viral proteins allows for rapid screening of novel viruses on vaccine targets.
[0128] The immunogenicity classifier outlined in Example 1 is trained on immunogenic peptides of already tested viral strains and can be used to pre-screen a complete list of potential epitopes for new viral strains. It reduces the number of in vitro assays required for a particular strain, effectively minimizing the associated timeframe and cost. Additionally, it helps optimize target peptides in vaccines that are likely to elicit an immune response in a larger segment of the population.
[0129] To demonstrate the applicability of this algorithm to viral peptides, it was used to compare the immunogenic epitopes of SARS-CoV-2 with those of SARS-CoV and MERS-CoV. These three viruses are closely related but can cause symptoms of vastly different orders of magnitude. In fact, SARS-CoV symptoms (approximately 10% mortality) are more severe than those of current coronaviruses (estimated at 1.3%), while MERS-CoV is even more deadly than both (20.4%–69.2% mortality).
[0130] After collecting the envelope protein sequences of SARS-CoV-2, as well as SARS-CoV and MERS-CoV, all possible 9-mers were extracted. These were then run using the neoMS algorithm to determine which peptides were likely to be presented on the surface of infected cells. Subsequently, the algorithm of this invention categorized the presented peptides according to their predicted immunogenicity. The resulting immunogenicity probability distribution of the presented peptides in this reduced set was examined, as shown below. Figure 7 As shown, SARS-CoV-2 exhibits fewer potential immunogenic peptides on average than SARS-CoV, while SARS-CoV has fewer potential immunogenic peptides than MERS-CoV. This implies that the immunogenicity load of these three viruses appears to be related to their pathogenicity, thus confirming the hypothesis of "Mcmanus, LM & Mitchell, RN Pathobiology of human disease: a dynamic encyclopedia of disease mechanisms. in Pathobiology of human disease: a dynamic encyclopedia of diseasemechanisms. 1036 (Elsevier, 2014)". Figure 7 Histograms and correlation density curves showing the immunogenicity probabilities of envelope-derived cell surface-presented peptides of SARS-CoV-2 (red), SARS-CoV (blue), and MERS-CoV (green).
[0131] Example 6: The differential amino acid composition of the peptides on which the positive and negative datasets of this invention are based.
[0132] As mentioned above, the positive and negative datasets are based on peptides that are assumed to elicit an immune response (positive dataset) or not (negative dataset).
[0133] Differential amino acid group usage analysis (dagLogo: Jianhong Ou, Haibo Liu, Niraj K. Nirala, Alexey Stukalov, Usha Acharya, Michael R. Green, Lihua Julie Zhu, dagLogo: An R / Bioconductor package for identifying and visualizing differential amino acid group usage in proteomics data, PLOS ONE, Published: November 6, 2020) assessed the presence of conserved sequence patterns in peptides of positive and negative sets, and whether these patterns differed between the two groups. For this purpose, the positive data group was based on immunogenic peptides that did not match the human proteome, while the negative data group was based on MHC-presenting peptides that did not match the human proteome.
[0134] Figure 8 Displaying peptide sequence markers, showing the difference in amino acid usage between a 9-amino acid-long peptide in the immunogenic positive dataset (top) and a 9-amino acid-long peptide in the non-immunogenic negative dataset (bottom). Amino acids or ungrouped ( Figure 8 A), or according to chemical properties ( Figure 8 B, acidity, basicity, amide group, hydroxyl group, sulfur-containing, aromatic or aliphatic), charge grouping ( Figure 8 C; positive, negative or neutral) or magnitude ( Figure 8 D; tiny, small, medium, large, or huge.
[0135] The figure clearly shows the physicochemical differences between the two groups of peptides used in the construction and training models. Furthermore, it reveals that the properties of certain amino acids, or amino acids at certain positions, are related to immunogenicity.
[0136] Example 7: Test the efficiency of the computer execution method of the present invention on the newly generated dataset and compare it with... He has available methods for comparison.
[0137] Experimental immunogenicity assays of 33 peptides, measured by ELISpot, were performed on healthy donor T cells. Assays were performed using autologous dendritic cells loaded with peptides, with two rounds of stimulation. All tests were performed in duplicate or triplicate.
[0138] The netMHCpan4.0 tool predicts the presentation of all peptides on at least two donor HLA alleles. netMHCpan4.0 is a tool that uses an artificial neural network to predict peptide binding to MHC molecules. This tool is trained with naturally eluted ligand and binding affinity data. It returns two properties: the probability that the peptide will become a natural ligand or the predicted binding affinity.
[0139] A positive ELISpot result is defined as every 5 × 10 4 Each cell had at least 25 spots, an increase of at least 2-fold compared to the control. Of the 33 peptides predicted to be presented, 8 peptides (23.5%) were positive by ELISpot assay.
[0140] The immunogenicity of the same peptide was determined using the computer-executed method of the present invention. The positive data set was based on immunogenic peptides that did not match the human proteome, while the negative data set was based on MHC-presenting peptides that did not match the human proteome.
[0141] Different rating systems: - Scoring of the method of the present invention: The immunogenicity score of each peptide is calculated using the computer execution method of the present invention.
[0142] -IEDB Immunogenicity Score: Calculates a score for each peptide using IEDB immunogenicity prediction values.
[0143] -netMHCpan rating score (HLA-A) 02:01): Using netMHCpan 4.0, calculate HLA-A 02:01 Net MHCpan grade score for each peptide of the allele.
[0144] -netMHCpan rank score (all donor HLA): Using netMHCpan 4.0, calculate the netMHCpan rank score for each peptide of all donor HLA alleles, taking into account the lowest score.
[0145] Figure 9A The predictive power of the immunogenicity of peptide libraries with positive or negative ELISpot results (p=0.019) determined using the computer-executed method of the present invention is shown. These results clearly demonstrate that peptides with positive ELISpot results (n=8) have significantly higher immunogenicity scores using the method of the present invention compared to peptides with negative results (n=25). Therefore, the computer-executed method of the present invention is able to distinguish between immunogenic and non-immunogenic peptides (p=0.019). No significant differences were found between the IEDB and netMHCpan tools.
[0146] Figure 9BA comparison of receiver operating characteristic (ROC) curves for several different immunogenicity predictors is shown, including the computer-executed method of this invention, the IEDB immunogenicity predictor, and netMHCpan 4.0, the latter based on a specific HLA allele or all HLA alleles of the subject. The area under the curve (AUC) of the ROC curves is:
[0147] The results show that the computer-executed method of the present invention is more accurate than other methods such as the IEDB immunogenicity predictor and netMHCpan4.0 in distinguishing between immunogenic and non-immunogenic peptides, which only predict binding affinity rather than immunogenicity.
[0148] In subsequent steps, the peptides obtained by the method of this invention are sorted from highest to lowest immunogenicity score and represented as an ascending curve. All positive peptides as determined by ELISpot are present in the top 60% of peptides, allowing 40% of all peptides to be discarded without loss of any true positives, such as... Figure 9C As shown, Figure 9C The percentage of true positives in the population is shown (as a cumulative gain). The cumulative gain curve illustrates the percentage of true positives identified using the minimum possible portion of the sample. Compared to the IEDB immunogenicity predictor and netMHCpan 4.0, the method of this invention results in a significant increase in specificity from 23% to 40% without loss of sensitivity, and a significant reduction in the number of candidate tumor neoantigens.
[0149] In conclusion, the computer-executed method of the present invention allows for better prioritization of presented novel epitopes by their immunogenic potential, significantly improving tumor neoepithelial prediction and reducing false positives. Prioritizing tumor neoepithels with high predictive immunogenicity using the computer-executed method of the present invention increases the likelihood of selecting actionable tumor neoantigens. Therefore, the computer-executed method of the present invention is a powerful tool for improving target selection for personalized immunotherapy and for providing immunogenic peptide loads as biomarkers for improved patient survival and treatment response.
[0150] Example 8: Identifying patients with low tumor mutation burden who will still benefit from immune checkpoint inhibitor therapy
[0151] A common biomarker for selecting patients for immune checkpoint inhibitor therapy (ICI) is tumor mutational burden (TMB) with >100 (ns=nonsynonymous) mutations, a form of cancer therapy designed to reactivate a patient's immune system. High TMB has been shown to be associated with the clinical benefit of ICI. However, some patients with low TMB will still benefit from ICI, and identifying these patients is crucial.
[0152] Three potential biomarkers for response to treatment were investigated in a cohort of low-TMB melanoma patients receiving ICI.
[0153] The first biomarker was TMB (total number of nonsynonymous mutations). Based on TMB, the group was divided into two equal-sized groups: "high" containing 50% of patients with the highest number of mutations and "low" containing 50% of patients with the lowest number of mutations.
[0154] Figure 10 The Kaplan-Meier curve is shown, illustrating the probability of an event at corresponding time intervals. It can be seen that TMB ( Figure 10 (Left figure) is not a good biomarker for identifying patients who respond to treatment.
[0155] The experiment then examined the number of mutations presented by each patient. For all possible peptides of 9–11 amino acids in length generated by each mutation, the number of presented mutations was obtained by running netMHCpan 4.1. A mutation was considered present if at least one peptide with a grade score <2 was identified. Based on the number of presented mutations, the group was divided into two equal-sized groups: "high" containing 50% of patients with the highest number of presented mutations, and "low" containing 50% of patients with the lowest number of presented mutations.
[0156] Figure 10 The intermediate figure shows that the number of presented mutations, similar to TMB, is not a good biomarker for identifying patients showing a response to treatment.
[0157] Finally, the number of immunogenic mutations in each patient was observed. The number of immunogenic mutations was obtained by running netMHCpan4.1 on all possible peptides of 9-11 amino acid length generated by each mutation. For all peptides considered to be presented (rank score <2), an immunogenicity score was calculated using the computer-executed method of this invention. If at least one peptide with an immunogenicity score >0.75 was identified, the mutation was considered immunogenic. Based on the number of immunogenic mutations, the group was divided into two equal-sized groups: "high" containing 50% of patients with the highest number of immunogenic mutations and "low" containing 50% of patients with the lowest number of immunogenic mutations.
[0158] Figure 10 The right figure shows that the number of immunogenic mutations is a better biomarker for identifying patients who show a response to treatment than the number of TMB or presented mutations.
[0159] This indicates that the response to treatment is driven not only by the absolute number of mutations in the tumor (especially in patients with a low number of mutations), but also by the immunogenic potential of these mutations. Therefore, the computer-executed method of the present invention, based on the immunogenic potential of mutations, can thus be used to identify patients with low tumor mutational burden who would still benefit from immune checkpoint inhibitor therapy.
[0160] This invention is by no means limited to the embodiments described and / or shown in the accompanying drawings. Rather, the method according to the invention can be implemented in many different ways without departing from the scope of the invention.
Claims
1. A computer-executed method for determining the immunogenicity of peptides, comprising the following steps: Obtain the amino acid sequence of the peptide, wherein the peptide comprises protein-derived and / or non-protein-derived amino acids; For each amino acid source of a protein, multiple numerical indices related to its physicochemical properties were obtained. Obtain a training dataset containing a positive dataset and a negative dataset, wherein the positive dataset includes data associated with multiple amino acid sequences of immunogenic peptides, and wherein the negative dataset includes data associated with multiple amino acid sequences of non-immunogenic peptides. Train a mathematical classification model on the training dataset; The likelihood of the peptide triggering an immune response is determined by a trained classification model. The method includes the following steps: Principal component analysis was performed on the numerical index of the amino acid source of each protein to obtain the principal components of the amino acid source of each analyzed protein; The feature vectors of the amino acid sequences of the training dataset and the amino acid sequences of the peptides are obtained, wherein the feature vectors of the amino acid sequences are obtained by replacing each amino acid of the amino acid sequence with one or more corresponding principal components; The classification model is trained on the feature vectors of the training dataset; The likelihood that a peptide is immunogenic is determined based on its feature vector. The negative dataset, characterized in that it contains data related to multiple amino acid sequences of non-immunogenic peptides, is obtained through the following method: Obtain the amino acid sequence of the peptide that can bind to the major histocompatibility complex and / or be presented on the major histocompatibility complex; Obtain the amino acid sequence corresponding to the protein histidine; The amino acid sequences of the peptides capable of binding to and / or presented on the major histocompatibility complex are compared with the amino acid sequences of the proteome to determine the match between them. The negative dataset includes amino acid sequences capable of binding to the major histocompatibility complex and / or presenting peptides on the major histocompatibility complex, said peptides being closely related to but not identical to peptides in the proteome, said closely related amino acid sequences having 1, 2 or 3 amino acid mismatches compared to peptides in the proteome.
2. The method according to claim 1, wherein, The peptide is a peptide presented on the cell surface.
3. The method according to claim 1, wherein, The negative dataset, which includes data related to multiple amino acid sequences of non-immunogenic peptides, was obtained by: Obtain the amino acid sequence of the peptide that can bind to the major histocompatibility complex and / or be presented on the major histocompatibility complex; Obtain the amino acid sequence corresponding to the housekeeping protein; The amino acid sequences of the peptides capable of binding to and / or presented on the major histocompatibility complex are compared with the amino acid sequences corresponding to the housekeeping proteins to determine the match between them. The negative dataset contains the matching amino acid sequence.
4. The method according to claim 3, wherein, The amino acid sequences of peptides capable of binding to and / or presenting on the major histocompatibility complex are identified by binding assays and / or mass spectrometry.
5. The method according to claim 1 or 3, wherein, The data associated with the multiple amino acid sequences of the non-immunogenic peptides are digital data.
6. The method according to claim 1 or 3, wherein, The amino acid sequence presented by the obtained major histocompatibility complex is a linear sequence.
7. The method according to claim 1, wherein, The immunogenic peptides in the positive dataset were obtained through the following methods: Obtain the amino acid sequence of the peptide that can induce a T cell response; Obtain the amino acid sequence corresponding to the protein histidine; The amino acid sequences that induce T cell responses are compared with the amino acid sequences of the proteome to determine the matches between them; Positive datasets include amino acid sequences that induce T cell responses, in addition to the matched amino acid sequences.
8. The method according to claim 1, wherein, The immunogenic peptides in the positive dataset were obtained through the following methods: Obtain the amino acid sequence of the peptide that can induce a T cell response; Obtain the amino acid sequence corresponding to the protein histidine; The amino acid sequences that induce T cell responses are compared with the amino acid sequences of the proteome to determine the matches between them; Positive datasets include amino acid sequences that are closely related to but not identical to peptides in the proteome and induce T cell responses, wherein the closely related amino acid sequences have 1, 2 or 3 amino acid mismatches compared to peptides in the proteome.
9. The method according to any one of claims 7 or 8, wherein, The amino acid sequence that induces the T cell response is a linear sequence.
10. The method according to claim 1 or 3, wherein, The training dataset includes a positive dataset and a negative dataset, wherein the negative dataset includes substantially more records than the positive dataset.
11. The method according to claim 1 or 3, wherein, The feature vector of the amino acid sequence is obtained by replacing each amino acid in the amino acid sequence with at least two corresponding principal components.
12. The method according to claim 1 or 3, wherein, The feature vector of the amino acid sequence is obtained by replacing each amino acid in the amino acid sequence with 2 to 10 corresponding principal components.
13. The method according to claim 1 or 3, wherein, The feature vector of the amino acid sequence is obtained by replacing each amino acid in the amino acid sequence with three corresponding principal components.
14. The method according to claim 1 or 3, wherein, Prior to the principal component analysis, the plurality of numerical indices are transformed into z-values.
15. The method according to claim 1 or 3, wherein, The classification model is a supervised classification machine learning algorithm.
16. The method according to claim 15, wherein, The machine learning algorithm is one or more of the following: multilayer perceptron classifier, Gaussian Naive Bayes classifier, linear support vector machine, kernel support vector machine, K nearest neighbor classifier, or random forest classifier.
17. A computer system for determining the immunogenicity of a peptide, said computer system being configured to perform a computer-executed method according to any one of claims 1 to 16.
18. A computer program product for determining the immunogenicity of a peptide, the computer program product comprising instructions that, when executed by a computer, cause the computer to perform the computer-executed method according to any one of claims 1 to 16.
19. The use of the computer-executed method according to any one of claims 1 to 16 and / or the computer system according to claim 17 and / or the computer program product according to claim 18 in the preparation of a kit for determining cancer treatment of a subject by determining the immunogenicity of a tumor neoantigen presented by the subject's tumor cells.
20. The application of the computer execution method according to any one of claims 1 to 16 and / or the computer system according to claim 17 and / or the computer program product according to claim 18, for screening viruses or bacteria on vaccine targets by measuring the immunogenicity of epitopes from viral or bacterial proteins.
21. An application of a computer-executed method according to any one of claims 1 to 16 and / or a computer system according to claim 17 and / or a computer program product according to claim 18, for characterizing an autoimmune response of a subject by determining the immunogenicity of an autoantigen presented by the subject's cells.
22. An application of a computer-executed method according to any one of claims 1 to 16 and / or a computer system according to claim 17 and / or a computer program product according to claim 18, for evaluating an immunogenicity change of a peptide when one or more amino acids of the amino acid sequence of the peptide are altered.
Citation Information
Patent Citations
Bioinformatic processes for determination of peptide binding
US20130330335A1
Bioinformatic processes for determination of peptide binding
US20160132631A1
Improved HLA epitope prediction
US20190346442A1
Personalized cancer vaccines and adoptive immune cell therapies
CN104662171A
Peptide vaccines for cancers expressing MPHOSPH1 or DEPDC1 polypeptides
CN105693843A