A method for predicting the affinity of a polypeptide and an hla type

By combining multi-dimensional feature statistical analysis of HLA typing and peptides with the LightGBM regression model, the iNeo-PRED model was constructed, which solved the problem of low accuracy in peptide-HLA affinity prediction and achieved higher prediction accuracy and stability.

CN115188418BActive Publication Date: 2026-07-24HANGZHOU XINYUANLI BIOTECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU XINYUANLI BIOTECHNOLOGY CO LTD
Filing Date
2022-06-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

The accuracy of peptide-HLA affinity prediction in existing technologies is not high, and the false positive rate is high, mainly due to the scarcity of HLA typing and internal features of peptide binding, as well as insufficient data.

Method used

By performing multi-dimensional feature statistical analysis on HLA typing and peptides, the iNeo-PRED model was constructed. The LightGBM regression model was used, combined with BLOSUM encoding, unique thermal encoding and physicochemical property encoding, to generate high-dimensional feature data. HLA-A, HLA-B and HLA-C models were trained respectively to improve prediction accuracy.

Benefits of technology

It improves the accuracy of peptide and HLA typing affinity prediction, outperforming existing models such as CNN, NetMHCpan4.0 and ACME, and is particularly stable when data is scarce.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115188418B_ABST
    Figure CN115188418B_ABST
Patent Text Reader

Abstract

The application discloses a kind of polypeptide and HLA typing affinity prediction method, including the following contents: one, obtains the characteristic relationship of polypeptide and HLA typing interaction;Two, based on the analysis result of obtained characteristic relationship, carry out the multidimensional internal feature mining of HLA typing and polypeptide, and carry out feature screening, to generate multidimensional training data for constructing machine learning model;Three, using machine learning algorithm, multidimensional internal statistical feature data mined is used as model input data after coding processing and feature screening, to carry out the classification of biological category to typing, respectively establish the training of complete independent model, complete the affinity prediction of polypeptide and HLA typing.The application is by multidimensional internal feature mining statistics to HLA typing and polypeptide respectively, analyze the association between original features of training data set, to extract new features from implicit information.Extract new features affinity prediction method is used to improve and optimize, improve the accuracy of prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to a method for predicting peptide and HLA typing affinity. Background Technology

[0002] Human leukocyte antigen (HLA) is the major histocompatibility complex (MHC) in humans, the most complex polymorphic system known to exist in the human body. The binding of HLA to antigenic peptides plays a crucial role in activating the body's T-cell immune response. Exogenous or endogenous antigens undergo a series of steps within antigen-presenting cells (APCs), ultimately being hydrolyzed or enzymatically broken down into short peptides carrying antigenic epitopes. These peptides then bind to MHC molecules intracellularly, forming a stable antigen-peptide-MHC complex (pMHC), which is then transported to the cell membrane surface. T cells can recognize and bind to immunogenic pMHC on the surface of APCs via T-cell receptors (TCRs), thereby activating T cells. Therefore, for immunotherapy, especially personalized tumor immunotherapy, accurately predicting the affinity between highly polymorphic HLA and different antigens is crucial for therapeutic efficacy.

[0003] Current research utilizes machine learning models to predict the binding affinity between HLA and antigenic peptides. Key methods for predicting peptide-HLA affinity include NetMHC, NetMHCpan (see Jurtz, V. et al. (2017) NetMHCpan-4.0: Improved Peptide–MHC Class I Interaction Predictions Integrating Eluted Ligand and Peptide Binding Affinity Data. J. Immunol., ji1700893), ConvMHC (Han and Kim, 2017), and ACME. However, existing techniques suffer from inconsistent peptide fragment lengths, resulting in a scarcity of internal HLA typing and peptide binding characteristics involved in model construction. Furthermore, some HLA typing and peptide binding data are limited or nonexistent, leading to low accuracy and high false-positive rates in predicting peptide-HLA typing affinity.

[0004] To address the shortcomings of existing technologies, the market needs a model for predicting peptide-HLA affinity that can improve the accuracy of predictions. This invention solves this problem. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to propose a method for predicting the affinity of peptides and HLA typing. This invention obtains high-dimensional features for model training through multi-dimensional statistical analysis of HLA typing and peptide characteristics. This is then combined with the optimization of the iNeo-PRED model (including HLA-A, HLA-B, and HLA-C models) constructed based on the LightGBM regression model in machine learning algorithms, thereby completing the prediction of peptide and HLA typing affinity and improving the accuracy of the prediction results.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for predicting peptide and HLA typing affinity includes the following:

[0008] Step 1, Feature Relationship:

[0009] Data on the interaction between peptides and HLA typing were obtained, and the data were enhanced after data processing and analysis to obtain the characteristic relationship between the two.

[0010] Step 2, Feature Engineering:

[0011] Analyze the feature relationships obtained in step one, perform HLA typing and peptide feature mining, and then perform feature screening to generate multi-dimensional training data for building machine learning models.

[0012] Step 3, Model Building:

[0013] Using machine learning algorithms, the multi-dimensional internal statistical feature data is encoded and filtered before being used as input data for the model. The subtypes are classified at the biological function level, and complete and independent models are built and trained for each category to complete the affinity prediction of peptides and HLA subtypes.

[0014] As a further explanation of the peptide and HLA typing affinity prediction method, the characteristic relationship is obtained in step one as follows:

[0015] S1, Data Collection:

[0016] Training data were obtained from the IEDB database, experiments, and relevant literature, including IC50 data and mass spectrometry data matching peptide and HLA typing.

[0017] S2, Data Analysis:

[0018] 1) Analyze affinity values ​​using IC50 and mass spectrometry data, including the distribution of affinity value A from IC50 data, the distribution of affinity value B from mass spectrometry data, and the ratio of positive to negative data.

[0019] 2) The lengths of the peptide sequences in the collected data included four types: 8, 9, 10, and 11 amino acids. The hypothetical sequences corresponding to HLA typing originated from amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing was less than 10 angstroms. Based on the supertype to which each HLA typing belonged, three supertypes were obtained: HLA-A, HLA-B, and HLA-C. Then, three training datasets were selected based on the supertypes, and three models were trained based on the training data of each supertype.

[0020] 3) Analyze the data collected in step S1. The analysis includes: the proportion distribution of data binding to different HLA types for the same peptide, the overall data distribution of binding to different HLA types in IC50, the probability distribution of positive data binding to different HLA types in IC50, the probability distribution of negative data binding to different HLA types in IC50, the overall data probability distribution of binding to different HLA types in mass spectrometry, the probability distribution of positive data binding to different HLA types in mass spectrometry, and the probability distribution of negative data binding to different HLA types in mass spectrometry.

[0021] The distribution of data proportions of the same HLA type and different peptide binding degrees, the overall data distribution of the same type and different peptide binding degrees in IC50, the probability distribution of positive data of the same type and different peptide binding degrees in IC50, the probability distribution of negative data of the same type and different peptide binding degrees in IC50, the overall data probability distribution of the same type and different peptide binding degrees in mass spectrometry, the probability distribution of positive data of the same type and different peptide binding degrees in mass spectrometry, and the probability distribution of negative data of the same type and different peptide binding degrees in mass spectrometry.

[0022] 4) The frequency of different amino acids appearing at the same position in all HLA typing peptides, the probability distribution of positive data of HLA typing and peptide binding degree under each amino acid position of the peptide sequence, the probability distribution of negative data of HLA typing and peptide binding degree, and the analysis of amino acid motif distribution in HLA-A supertype and its corresponding peptide data, HLA-B supertype and its corresponding peptide data, and HLA-C supertype and its corresponding peptide data;

[0023] S3, Data Preprocessing:

[0024] By performing length analysis on the training data, peptides and HLA typing data with a length of 8-11 amino acids were selected as the model training dataset. The training dataset was then classified into peptides and HLA typings based on affinity values, categorizing them as affinity-free or affinity-free. The specific process included:

[0025] 1) Remove data that does not have affinity values ​​and remove data from non-human species;

[0026] 2) Based on the distribution pattern of mass spectrometry data, the affinity value B of the mass spectrometry data is transformed into the affinity value C through a specific conversion. The affinity value C and the affinity value A of the IC50 data are combined and then converted to generate the affinity value D.

[0027] 3) By combining BLOSUM coding, unique heat coding, and physicochemical property coding, amino acid coding is performed on HLA types and peptides; peptides with a length of 8-11 amino acids are fixed into amino acid sequences of uniform length by filling in specific characters; hypothetical sequences corresponding to different HLA types are fixed into uniform length within the same supertype.

[0028] As a further explanation of the affinity prediction method for peptides and HLA typing, in step S2, during data analysis, the distribution of affinity values ​​A in IC50 data and B in mass spectrometry data are analyzed. Specifically, the affinity value A in IC50 data is a continuous value greater than 0, while the affinity value B in mass spectrometry data takes two discrete values, 0 and 1. There is a difference between the distributions of affinity values ​​A and B.

[0029] As a further explanation of the peptide and HLA typing affinity prediction method, the hypothetical sequence corresponding to the HLA typing in step S2 is derived from amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing is less than 4 angstroms.

[0030] As a further explanation of the affinity prediction method for peptides and HLA typing, the amino acids in each HLA typing and its corresponding peptide data have different motif distributions, and the position of each amino acid in the peptide is assigned a corresponding weight based on this distribution.

[0031] As a preferred option, BLOSUM is encoded as a BLOSUM50 encoding matrix.

[0032] As a preferred option, one-hot encoding is used.

[0033] Further explanation of this method for predicting peptide and HLA typing affinity;

[0034] In step S3, based on the distribution pattern of the mass spectrometry data, the affinity value B of the mass spectrometry data is transformed into an affinity value C through a specific conversion. The affinity value C and the affinity value A of the IC50 data are then combined and transformed again to generate the affinity value D. The specific method is as follows:

[0035] IC50 data: A threshold is selected for the affinity value A to determine whether there is affinity between the HLA type and the peptide. When the affinity value A is less than the threshold, the binding between the HLA type and the peptide has affinity; when the affinity value A is greater than the threshold, the HLA type and the peptide do not bind, i.e., there is no affinity. A maximum value is set to normalize the data boundaries. When the affinity value A is greater than the maximum value, it is set to the maximum value.

[0036] Mass spectrometry data: When the affinity value B is 1, the HLA type and the peptide bind effectively and have affinity; when the affinity value B is 0, the HLA type and the peptide do not bind, i.e., there is no affinity.

[0037] Effective processing is performed on the affinity values ​​B, which are discrete values ​​of 0 and 1 in the mass spectrometry data: A strong affinity range is defined, truncated between 0 and a threshold. When the affinity value B is 1, an integer is randomly generated within this strong affinity range as the affinity value C. A weak affinity range is also defined, truncated between the threshold and the maximum value. When the affinity value B is 0, an integer is randomly generated within this weak affinity range as the affinity value C.

[0038] The affinity values ​​C of the mass spectrometry data and A of the IC50 data are combined and then converted to generate the affinity value D: According to formula (1), the affinity values ​​A and C are converted into a decimal f(x) between 0 and 1, where x represents the affinity value A or the affinity value C, and f(x) represents the affinity value D;

[0039] f(x)=1-logx / log50000 (1).

[0040] As a further explanation of the peptide and HLA typing affinity prediction method, the features in step two include one or more of the following: peptide sequence distribution features, HLA typing features, features of correlation between peptide and HLA typing, or amino acid position and distribution features.

[0041] As a further explanation of the peptide and HLA typing affinity prediction method, the multi-dimensional training data mentioned in step two is at least one of the following: features obtained by data screening of the same peptide, features obtained by data screening of the same typing, or features extracted based on the amino acid positions of the peptide sequence.

[0042] Further explanation of this method for predicting peptide and HLA typing affinity;

[0043] The specific features obtained from data screening of the same polypeptide are as follows:

[0044] The statistics include the number of existing HLA types, the total number of data on the binding degree of each peptide to different HLA types, the number of data with affinity between each peptide and different HLA types, the percentage of data with affinity, and the mean and standard deviation of affinity value D, and the number of data without affinity between each peptide and different HLA types, the percentage of data with affinity, and the mean and standard deviation of affinity value D.

[0045] The specific features obtained by filtering data for the same subtype are as follows:

[0046] The total number of peptides present, the total number of data for each HLA type and different peptide binding degrees, the number of data with affinity for each HLA type and different peptide binding degrees, the percentage of data with affinity, and the mean and standard deviation of affinity value D, and the number of data without affinity for each HLA type and different peptide binding degrees, the percentage of data with affinity, and the mean and standard deviation of affinity value D;

[0047] The specific features obtained from extracting amino acid positions in peptide sequences are as follows: For peptide sequences of 8-11 amino acids in length, a fixed-length sequence is generated by filling in fixed characters. The number, percentage, mean, and standard deviation of the affinity values ​​D for each amino acid at each position with HLA typing are statistically analyzed. The number, percentage, mean, and standard deviation of the affinity values ​​D for each amino acid at each position with HLA typing are statistically analyzed. The frequency characteristics of different amino acids appearing at the same position in peptide sequences corresponding to all HLA typings are analyzed. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in HLA-A supertype and its corresponding peptide data. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in HLA-B supertype and its corresponding peptide data. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in HLA-C supertype and its corresponding peptide data.

[0048] The feature data generated above includes both categorical and numerical data. Numerical features are retained without any additional processing, while categorical features are converted into categorical codes and one-hot codes, transforming the original data into a data format that can be recognized by machine learning models.

[0049] Further explanation of this method for predicting peptide and HLA typing affinity;

[0050] In step three, the iNeo-PRED model, constructed based on the LightGBM regression model of Python machine learning algorithm, uses the mined multi-dimensional internal statistical features as input data after encoding and feature selection. It trains models for the three major supertypes of HLA-A, HLA-B, and HLA-C respectively, and constructs three complete and independent models to complete the affinity prediction of peptides and HLA types.

[0051] The advantages of this invention are:

[0052] This invention extracts new features by performing multi-dimensional internal feature mining and statistical analysis on HLA typing and peptides, respectively, and analyzing the correlations between the original features of the training dataset. The extracted features are then encoded and transformed, and feature selection methods from Scikit-learn are used to filter features. The feature dimensionality of the training and test datasets is expanded to ultimately obtain high-dimensional features.

[0053] This invention transforms discrete affinity values ​​into specific continuous values ​​based on the distribution patterns of mass spectrometry data, better integrating them with IC50 data and increasing training data to facilitate model training and prediction. Based on HLA supertype similarity, the training dataset is divided into three categories: HLA-A, HLA-B, and HLA-C. The LightGBM regression model, which is fast, memory-efficient, and capable of handling massive amounts of data, is selected. The extended multi-dimensional statistical features can be used individually or in combination to complete the construction and optimization of the three models, ultimately achieving the prediction of binding affinity between peptides and HLA subtypes.

[0054] This invention utilizes a high-dimensional feature-optimized iNeo-PRED model, which synergistically improves prediction accuracy. Overall, the iNeo-PRED model outperforms CNN, netMHCpan4.0, and ACME models in prediction performance. It outperforms netMHCpan4.0 in 90.3% of HLA-A test set classifications, in 78.3% of HLA-B test set classifications, and in all HLA-C test set classifications. Furthermore, prediction results were evaluated and compared on test sets with very limited peptide and HLA typing data. The iNeo-PRED model significantly outperformed the other three existing models in both mean and median AUC values, demonstrating more stable prediction results. In the benchmark dataset (43 datasets in total), the iNeo-PRED model outperformed the CNN model in AUC across 36 test sets (76.7%), the netMHCpan4.0 model in AUC across 32 test sets (74.4%), and the ACME model in AUC across 34 test sets (79.1%). Attached Figure Description

[0055] Figure 1 This is a histogram showing the statistical number of HLA typing bindings of the same peptide in the data of this invention; (the horizontal axis represents the interval labels, 0 represents the number of HLA typings of the same peptide falling into the interval of 0-10, 1 represents the number falling into the interval of 10-20, 2 represents the number falling into the interval of 20-30, 3 represents the number falling into the interval of 30-40, and 4 represents the number falling into the interval of 40-50; the vertical axis represents the number of times the data appears, representing the total number of times the number of HLA typings of the same peptide appears in the dataset in the current interval.)

[0056] Figure 2 This is a schematic diagram illustrating the overall data probability distribution of the binding degree of the same peptide and different HLA types in the IC data of this invention; (the horizontal axis seq_amount represents the number of HLA types under the same peptide, and the vertical axis represents the probability value of occurrence in all IC50 datasets.)

[0057] Figure 3 This is a schematic diagram illustrating the probability distribution of positive data for the same polypeptide and different HLA types in the data IC of this invention; (the horizontal axis seq_pos_count represents the number of HLA types with binding affinity in the data of the same polypeptide and different HLA types, and the vertical axis represents the probability value of data with binding affinity appearing in the data of the same polypeptide and different HLA types.)

[0058] Figure 4 This is a schematic diagram illustrating the probability distribution of negative data for the same peptide and different HLA types in the data of this invention; (the horizontal axis seq_neg_count represents the number of HLA types with no binding affinity in the data for the same peptide and different HLA types, and the vertical axis represents the probability value of no binding affinity data in the data for the same peptide and HLA types.)

[0059] Figure 5 This is a schematic diagram showing the number of times the same amino acid appears at the same position in the data of the same polypeptide and different HLA typing binding degrees in the data of this invention; (the horizontal axis represents each position of the 34 amino acids in the hypothetical sequence represented by the HLA typing, and the vertical axis represents the number of times the same amino acid appears at each position in the hypothetical sequence represented by the same polypeptide and the HLA typing)

[0060] Figure 6 This is a comparison of the overall AUC of four models on the HLA-A benchmark test set. (The horizontal axis represents the names of the four models: iNeo-PRED_A (based on feature A), iNeo-PRED_AB (based on features A and B), iNeo-PRED_ (based on all features, including features A, B, and C), and netMHCpan4.0. The vertical axis represents the AUC value of each model on the HLA-A benchmark test set. The triangles in the box plot represent the mean AUC.)

[0061] Figure 7 This is a comparison of the overall AUC of four models on the HLA-B benchmark test set; (the horizontal axis represents the names of the four models: iNeo-PRED_A (based on feature A), iNeo-PRED_AB (based on features A and B), iNeo-PRED_ (based on all features, including features A, B, and C), and netMHCpan4.0; the vertical axis represents the AUC value of each model on the HLA-B benchmark test set; the triangles in the box plot represent the mean AUC.)

[0062] Figure 8This is a comparison of the overall AUC of four models on the HLA-C benchmark test set. (The horizontal axis represents the names of the four models: iNeo-PRED_A (based on feature A), iNeo-PRED_AB (based on features A and B), iNeo-PRED_ (based on all features, including features A, B, and C), and netMHCpan4. The vertical axis represents the AUC value of each model on the HLA-C benchmark test set. The triangles in the box plot represent the mean AUC.)

[0063] Figure 9 This is a comparison of the overall AUC of four models on the HLA-A benchmark test set; (the horizontal axis represents the names of the four models: the iNeo-PRED model built based on all features (including features A, B, and C), the CNN model built based on all features (including features A, B, and C), the netMHCpan4.0 model, and the ACME model; the vertical axis represents the AUC value of each model evaluated on the HLA-A benchmark test set; the triangles in the box plot represent the mean AUC.)

[0064] Figure 10 This is a comparison of the overall AUC of four models on the HLA-B benchmark test set; (the horizontal axis represents the names of the four models: the iNeo-PRED model built based on all features (including features A, B, and C), the CNN model built based on all features (including features A, B, and C), the netMHCpan4.0 model, and the ACME model; the vertical axis represents the AUC value of each model evaluated on the HLA-B benchmark test set; the triangles in the box plot represent the mean AUC.)

[0065] Figure 11 This is a comparison of the overall AUC of four models on the HLA-C benchmark test set; (the horizontal axis represents the names of the four models: the iNeo-PRED model built based on all features (including features A, B, and C), the CNN model built based on all features (including features A, B, and C), the netMHCpan4.0 model, and the ACME model; the vertical axis represents the AUC value of each model evaluated on the HLA-C benchmark test set; the triangles in the box plot represent the mean AUC.)

[0066] Figure 12This is a histogram comparing the overall AUC values ​​of four models on the benchmark HLA-A, HLA-B, and HLA-C test sets when the amount of peptide and HLA typing data is sparse (less than 60 data points). (The horizontal axis represents the names of the four models: the iNeo-PRED model built based on all features (including features A, B, and C), the CNN model built based on all features (including features A, B, and C), the netMHCpan4.0 model, and the ACME model. The vertical axis represents the AUC value of each model's prediction results on the test set with less than 60 data points. The triangles in the box plot represent the mean AUC.) Detailed Implementation

[0067] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0068] Step 1, Feature Relationship:

[0069] Data on the interaction between peptides and HLA typing were obtained, and the data were enhanced after data processing and analysis to obtain the characteristic relationship between the two.

[0070] The method for obtaining feature relationships in step one is as follows:

[0071] Training data, including IC50 and mass spectrometry data matching peptide and HLA typing, were collected from the IEDB database, experimental data, and relevant literature. Taking the IC50 data of HLA-A02:01 and HLA-C04:01 typing, and the mass spectrometry data of HLA-B07:02 and HLA-A11:01 typing as examples, the specific training data format is shown in Table 1.

[0072] Table 1

[0073]

[0074]

[0075] S2, Data Analysis:

[0076] 1) Analyze affinity values ​​using IC50 and mass spectrometry data, including the distribution of affinity value A from IC50 data, the distribution of affinity value B from mass spectrometry data, and the ratio of positive to negative data.

[0077] The distribution of affinity value A in IC50 data and the distribution of affinity value B in mass spectrometry data are analyzed. Specifically, the affinity value A in IC50 data is a continuous value greater than 0, while the affinity value B in mass spectrometry data takes two discrete values, 0 and 1. There is a difference between the distributions of affinity value A and affinity value B.

[0078] 2) The peptide sequences in the collected data contained four lengths: 8, 9, 10, and 11 amino acids. The hypothetical sequence corresponding to the HLA typing originated from amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing was less than 10 angstroms. As a further optimization, the hypothetical sequence corresponding to the HLA typing originated from amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing was less than 4 angstroms. All amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing was less than 4 angstroms were taken as hypothetical sequences, and the length of the hypothetical sequence corresponding to the HLA typing was 34 amino acids. Based on the supertype to which each HLA typing belonged, three supertypes were obtained: HLA-A, HLA-B, and HLA-C. Then, three training datasets were selected based on the supertypes, and three models, HLA-A, HLA-B, and HLA-C, were trained based on their respective training data.

[0079] 3) Analyze the data collected in step S1. The analysis includes: the proportion distribution of data binding to different HLA types for the same peptide, the overall data distribution of binding to different HLA types in IC50, the probability distribution of positive data binding to different HLA types in IC50, the probability distribution of negative data binding to different HLA types in IC50, the overall data probability distribution of binding to different HLA types in mass spectrometry, the probability distribution of positive data binding to different HLA types in mass spectrometry, and the probability distribution of negative data binding to different HLA types in mass spectrometry.

[0080] The distribution of data proportions for the same HLA type and different peptide binding degrees; the overall data distribution of the same type and different peptide binding degrees in IC50; the probability distribution of positive data for the same type and different peptide binding degrees in IC50; the probability distribution of negative data for the same type and different peptide binding degrees in IC50; the overall data probability distribution of the same type and different peptide binding degrees in mass spectrometry; the probability distribution of positive data for the same type and different peptide binding degrees in mass spectrometry; and the probability distribution of negative data for the same type and different peptide binding degrees in mass spectrometry.

[0081] The frequency of different amino acids appearing at the same position in the polypeptide sequences corresponding to all HLA types; the probability distribution of positive data on HLA type and polypeptide binding degree at each amino acid position in the polypeptide sequence; and the probability distribution of negative data on HLA type and polypeptide binding degree.

[0082] 4) The motif distribution of amino acids in HLA-A supertype and its corresponding peptide data, HLA-B supertype and its corresponding peptide data, and HLA-C supertype and its corresponding peptide data were analyzed. The specific methods are as follows:

[0083] WebLogo software (http: / / weblogo.berkeley.edu / ) was used to analyze the amino acid motif distributions in HLA-A supertype and its corresponding peptide data, HLA-B supertype and its corresponding peptide data, and HLA-C supertype and its corresponding peptide data. A motif is a conserved region within a sequence, or a short sequence pattern shared by a group of sequences. More often, it refers to any sequence pattern that may have molecular function, structural properties, or be related to a family member.

[0084] Different HLA types and their corresponding peptide data exhibit different motif distributions of amino acids. Based on this distribution, different weights are assigned to the position of each amino acid in the peptide (generated using WebLogo software (http: / / weblogo.berkeley.edu / )).

[0085] The motif distribution of amino acids in HLA-A supertype and its corresponding peptide data is as follows: they are mainly concentrated in the second and ninth positions, indicating that the amino acids in these two positions are very important. Meanwhile, the amino acids in the first and tenth positions have the next lower weight. The corresponding weights can be assigned according to the concentration of amino acid distribution in each position.

[0086] The motif distribution of amino acids in HLA-B supertype and its corresponding peptide data is as follows: they are mainly concentrated at the second and ninth positions, indicating that the amino acids at these two positions are very important. Meanwhile, the distribution at the first, third, and tenth positions is relatively less concentrated than the remaining positions. Appropriate weights can be assigned according to the degree of concentration of amino acid distribution at each position.

[0087] The motif distribution of amino acids in HLA-C supertype and its corresponding peptide data is as follows: they are mainly concentrated at the first and ninth positions, indicating that the amino acids at these two positions are very important. Meanwhile, the second, third, and eighth positions are relatively less concentrated than the remaining positions. Appropriate weights can be assigned based on the degree of concentration of amino acid distribution at each position.

[0088] S3, Data Preprocessing:

[0089] By performing length analysis on the training data, peptides and HLA typing data with a length of 8-11 amino acids were selected as the model training dataset. The training dataset was then classified into peptides and HLA typings based on affinity values, categorizing them as affinity-free or affinity-free. The specific process included:

[0090] 1) Remove data that does not have affinity values ​​and remove data from non-human species;

[0091] 2) Based on the distribution pattern of mass spectrometry data, the affinity value B of the mass spectrometry data is transformed into the affinity value C through a specific conversion. The affinity value C and the affinity value A of the IC50 data are combined and then converted to generate the affinity value D.

[0092] In step S3, based on the distribution pattern of the mass spectrometry data, the affinity value B of the mass spectrometry data is specifically converted into the affinity value C. The affinity value C and the affinity value A of the IC50 data are then combined and converted again to generate the affinity value D. The specific method is as follows:

[0093] IC50 data: A threshold value A is selected to determine whether there is affinity between the HLA type and the peptide. When the affinity value A is less than the threshold, the binding between the HLA type and the peptide is affinity-based; when the affinity value A is greater than the threshold, the HLA type and the peptide do not bind, i.e., there is no affinity. A maximum value is set to normalize the data boundaries; when the affinity value A is greater than the maximum value, it is set to the maximum value.

[0094] Mass spectrometry data: When the affinity value B is 1, the HLA type and the peptide bind effectively and have affinity; when the affinity value B is 0, the HLA type and the peptide do not bind, i.e., there is no affinity.

[0095] Effective processing is performed on the affinity values ​​B, which are discrete values ​​of 0 and 1 in the mass spectrometry data: A strong affinity value interval is set, which is cut off between 0 and a threshold. When the affinity value B in the mass spectrometry data is 1, an integer is randomly generated within this interval as the affinity value C. A weak affinity value interval is set, which is cut off between the threshold and the maximum value. When the affinity value B in the mass spectrometry data is 0, an integer is randomly generated within this interval as the affinity value C.

[0096] The affinity values ​​C of the mass spectrometry data and A of the IC50 data are combined and then converted to generate the affinity value D: According to formula (1), the affinity values ​​A and C are converted into a decimal f(x) between 0 and 1, where x represents the affinity value A or the affinity value C, and f(x) represents the affinity value D;

[0097] f(x)=1-logx / log50000 (1).

[0098] 3) By combining BLOSUM encoding, one-hot encoding, and physicochemical property encoding, HLA typing and peptides are encoded using amino acid sequences; peptides of 8-11 amino acids in length are fixed to a uniform length amino acid sequence by padding with specific characters; hypothetical sequences corresponding to different HLA typings are fixed to a uniform length within the same supertype. As a preferred embodiment, BLOSUM encoding is a BLOSUM50 encoding matrix. One-hot encoding is used.

[0099] Step 2, Feature Engineering:

[0100] Analyze the feature relationships obtained in step one, perform HLA typing and peptide feature mining, and then perform feature screening to generate multi-dimensional training data for building machine learning models.

[0101] Features include one or more of the following: polypeptide sequence distribution features, HLA typing features, features of correlation between polypeptides and HLA typing, or amino acid position and distribution features.

[0102] The multidimensional training data can be any one of the following: features obtained by data screening of the same polypeptide, features obtained by data screening of the same subtype, or features extracted based on the amino acid positions of the polypeptide sequence.

[0103] The specific features obtained from data screening of the same polypeptide are as follows:

[0104] The statistics include the number of existing HLA types, the total number of data on the binding degree of each peptide to different HLA types, the number of data with affinity between each peptide and different HLA types, the percentage of data with affinity, and the mean and standard deviation of affinity value D, and the number of data without affinity between each peptide and different HLA types, the percentage of data with affinity, and the mean and standard deviation of affinity value D.

[0105] The specific features obtained by filtering data for the same subtype are as follows:

[0106] The total number of peptides present, the total number of data for each HLA type and different peptide binding degrees, the number of data with affinity for each HLA type and different peptide binding degrees, the percentage of data with affinity, and the mean and standard deviation of affinity value D, and the number of data without affinity for each HLA type and different peptide binding degrees, the percentage of data with affinity, and the mean and standard deviation of affinity value D;

[0107] The specific features obtained from extracting amino acid positions in peptide sequences are as follows: For peptide sequences of 8-11 amino acids in length, a fixed-length sequence is generated by filling in fixed characters. The number, percentage, mean, and standard deviation of the affinity values ​​D for each amino acid at each position with HLA typing are statistically analyzed. The number, percentage, mean, and standard deviation of the affinity values ​​D for each amino acid at each position with HLA typing are statistically analyzed. The frequency characteristics of different amino acids appearing at the same position in peptide sequences corresponding to all HLA typings are analyzed. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in HLA-A supertype and its corresponding peptide data. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in HLA-B supertype and its corresponding peptide data. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in HLA-C supertype and its corresponding peptide data.

[0108] The generated feature data includes both categorical and numerical data. Numerical features are retained as their original values ​​without any additional processing. Categorical data is unordered and defined as an object in Pandas (a Python library), such as characters 'A' or 'HLA'. Numerical data is ordered, such as 100, 1.0, etc. Numerical features are not processed further. Categorical features are encoded using the LabelEncoder or OneHotEncoder functions from the Scikit-learn (a Python machine learning library) data processing module, converting the original data into a format that machine learning models can recognize.

[0109] Then, the recursive feature elimination method in Scikit-learn is used to select features by recursively reducing the size of the feature set under consideration. First, the prediction model is trained on the original features, with each feature assigned a weight. Then, features with the smallest absolute weight are removed from the feature set, and this process is repeated recursively until the number of remaining features reaches the required number. Finally, at least 100 features are selected and retained as input data for the final model.

[0110] Step 3, Model Building:

[0111] Using machine learning algorithms, the mined multi-dimensional internal statistical features are encoded and filtered before being used as input data for the models. Biological classification is performed on the typing, and independent models are trained to predict the affinity of peptides and HLA typing. As a further optimization, the machine learning algorithm chosen is the iNeo-PRED model, built using the Python-based LightGBM regression model; the deep learning algorithm chosen is the Python-based Convolutional Neural Network (CNN) framework.

[0112] The specific process is as follows:

[0113] The iNeo-PRED model uses L2 loss as its loss function, and the validation set evaluation metrics are mean squared error and area under the curve (AUC). The key hyperparameters for model tuning are: number of trees (n_estimators), maximum number of leaf nodes (num_leaves), minimum number of samples per leaf node (min_data_leaf), learning rate, feature sampling ratio and data sampling ratio per tree, maximum number of bins for features, and regularization parameters lambda_l1 and lambda_l2.

[0114] First, the built-in `cv` function in LightGBM was used for tuning, and fast cross-validation was performed on the continuous `n_estimators` parameters. The remaining parameters were then subjected to five-fold cross-validation using the `GridSearchCV` function in the model selection module of the Scikit-learn library to obtain the optimal parameters. Based on the importance ranking of all features output by the model with the optimal parameters, 100-500 important features were selected as the input to the final model. Then, `n` important features were extracted from all data of the three supertypes (HLA-A, HLA-B, and HLA-C) as training data to construct three complete and independent models, completing the affinity prediction for peptides and HLA typing.

[0115] This Python-based Convolutional Neural Network (CNN) uses multi-dimensional internal statistical features, after encoding and feature selection, as input data to predict the affinity of peptides and HLA typing. The specific process involves: automatically extracting and fusing the input features through multiple convolutional and max-pooling layers, and finally connecting them using a fully connected layer of size 512. The CNN uses mean squared error as the loss function, Adam as the network optimizer, a batch size of 256, and an initial learning rate of 0.001, which decays during training based on actual conditions. The maximum number of iterations is set to 25; if the loss function stops improving within 10 iterations, the model is forcibly stopped early. A total of five models were generated, each consisting of more than 25 deep neural networks. The average prediction score of all networks is used as the final prediction result output.

[0116] Step 4, Model Evaluation:

[0117] Multiple datasets from the IEDB Benchmark provided by netMHCpan 4.0 were selected as the test set for HLA typing and peptide affinity models, and AUC (Area Under Curve) was used as the evaluation metric for model prediction performance.

[0118] The test set consists of multiple datasets from the IEDB Benchmark provided by netMHCpan 4.0, including:

[0119] 1) The test dataset under the HLA-A supertype includes HLA types such as HLA-A*02:01, HLA-A*02:02, HLA-A*02:03, HLA-A*02:06, HLA-A*03:01, HLA-A*11:01, HLA-A*24:02, HLA-A*30:01, HLA-A*30:02, HLA-A*31:01, HLA-A*68:01, HLA-A*68:01, etc.

[0120] 2) The test dataset under HLA-B supertype includes HLA types such as HLA-B*07:02, HLA-B*15:02, HLA-B27:03, HLA-B*27:04, HLA-B*27:05, HLA-B*27:06, HLA-B*35:01, HLA-B*38:01, HLA-B*39:06, HLA-B*40:01, HLA-B*44:03, HLA-B*55:02, HLA-B*57:01, and HLA-B*58:01.

[0121] 3) The test dataset under the HLA-C supertype includes HLA-C*03:03, HLA-C*04:01, HLA-C*05:01, HLA-C*06:02, HLA-C*07:01, HLA-C*07:02, HLA-C*08:02, HLA-C*12:03, HLA-C*14:02, and HLA-C*15:02.

[0122] The prediction performance of the iNeo-PRED model, using the following specific implementation, is compared with that of the CNN model, the netMHCpan4.0 model, and the ACME model:

[0123] 1. Data Source

[0124] Data were collected from two sources: the IEDB database, experimental data, and relevant literature: IC50 data (data format as shown in Table 2) and mass spectrometry data (data format as shown in Table 3).

[0125] Table 2 IC50 Data

[0126] HLA-A11:01 EVAQRAYR 8 ic50 50843.6 HLA-C14:02 CKNFLKQVY 9 ic50 6530.075544 HLA-A02:01 NYMPYVFTL 9 ic50 78125 HLA-B07:02 LSDDSGLMV 9 ic50 22 HLA-B15:02 ETDQMDTIY 9 ic50 384 HLA-A02:06 ISKIPGGAMY 10 ic50 23810.60368 HLA-A01:01 ATSRTLSYY 9 ic50 4.808301 HLA-C08:01 RPFNNILNL 9 ic50 70422.53521 HLA-A01:01 SSSMRKTDWL 10 ic50 49792.02554 HLA-C03:03 CASSSDWFY 9 ic50 3 HLA-A02:03 ILGAQALPVY 10 ic50 70422.53521 HLA-A68:01 LTKGTLEPEYC 11 ic50 3202.155159 HLA-A24:02 ISAGFSLWIY 10 ic50 52.337642 HLA-B07:02 ATVAYFNMVY 10 ic50 892.776409 HLA-A02:01 RYLALYNKY 9 ic50 6413.917478 HLA-A02:01 QTHFPQFYW 9 ic50 20000 HLA-A01:01 YSDPLALREF 10 ic50 21.306046

[0127] Table 3 Mass Spectrometry Data

[0128]

[0129]

[0130] 2. Data Analysis

[0131] 2.1 Analysis of Affinity Values

[0132] Tables 2 and 3 show that the data fields include: HLA typing, peptide sequence, peptide sequence length, data source, and affinity value (where IC50 data represents affinity value A, and mass spectrometry data represents affinity value B). The affinity value A of the IC50 data is a continuous value greater than 0, while the affinity value B of the mass spectrometry data is two discrete values, 0 and 1; there is a certain difference in the distribution of affinity values ​​A and B.

[0133] 2.2 Analysis of peptides and HLA typing

[0134] The peptide sequences included four lengths: 8, 9, 10, and 11 amino acids. All amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing was less than 4 angstroms were considered as hypothetical sequences. The hypothetical sequence length corresponding to the HLA typing was 34 amino acids. Whether HLA typing and peptide binding is closely related to the sequence itself. Therefore, statistical analysis of existing training data was used to uncover patterns in the existence of HLA typing and peptide affinity, which were then used as input for a specific model to improve its learning and thus increase the accuracy of affinity prediction.

[0135] Characteristics of multidimensional analysis of the binding degree between the same peptide and HLA typing:

[0136] 1) Perform histogram statistics on the number of HLA types for the same peptide, binning by bins = [0, 10, 20, 30, 40, 50] and labels = [0, 1, 2, 3, 4], and count the number within each interval. For example... Figure 1 As shown in the figure, the number of HLA types corresponding to most of the same polypeptides is concentrated in the 0 tag (i.e., the 0-10 range).

[0137] 2) The probability distribution of seq_amount of the overall data for the binding degree of the same peptide to different HLA types in IC data is as follows: Figure 2 As shown, in the data on the binding degree of the same peptide with different HLA types, 9.0% of the data showed complete affinity, 22.5% showed complete non-affinity, and the remaining 68.5% showed both affinity and non-affinity. The probability distribution of positive data on the binding degree of the same peptide with HLA types is shown below. Figure 3 As shown in the figure. The probability distribution of negative data for the binding degree of the same polypeptide and HLA typing is as follows. Figure 4 As shown.

[0138] 3) The hypothetical sequence length for the HLA typing of the same polypeptide in the data is 34 amino acids, where the number of the same amino acid at each position in the sequence is as follows: Figure 5 As shown. Especially when the HLA supertypes are the same (e.g., HLA-A*02:01 and HLA-A*02:02, both belonging to the HLA-A supertype), there is a high degree of similarity between the hypothetical sequences represented by HLA typing, so three categories, HLA-A, HLA-B and HLA-C, are obtained based on the supertype.

[0139] Training datasets were selected based on supertypes, and three machine learning models were trained on these datasets. The trained models were then used to predict the affinity of HLA types and peptides on the IEDB benchmark test dataset. The prediction results show that the AUC (Average Acceptance Value) metric was improved, enhancing prediction accuracy and resolving the issue of unpredictable affinity due to limited or nonexistent training data for some HLA types and peptides. In other words, when the HLA type in the test set is absent or has very limited data in the training dataset, the model trained using the supertype to which the HLA type belongs can effectively address the aforementioned problem of unpredictable HLA types and achieve relatively accurate prediction results.

[0140] 2.3 Motif analysis of amino acid positions in the polypeptide sequence

[0141] In HLA-A supertype and its corresponding peptide data, the motif distribution of amino acids is mainly concentrated at the second and ninth positions, indicating that the amino acids at these two positions are very important. Meanwhile, the amino acids at the first and tenth positions have the next lower weight. Appropriate weights can be assigned based on the concentration of amino acid distribution at each position.

[0142] In HLA-B supertype and corresponding peptide data, the motif distribution of amino acids is mainly concentrated at the second and ninth positions, indicating that the amino acids at these two positions are very important. Meanwhile, the distribution at the first, third, and tenth positions is relatively less concentrated than the remaining positions. Appropriate weights can be assigned based on the degree of concentration of amino acid distribution at each position.

[0143] In HLA-C supertype and its corresponding peptide data, the motif distribution of amino acids is mainly concentrated at the first and ninth positions, indicating that the amino acids at these two positions are very important. Meanwhile, the distribution at the second, third, and eighth positions is relatively less concentrated than the remaining positions. Appropriate weights can be assigned based on the degree of concentration of amino acid distribution at each position.

[0144] 3. Data Preprocessing

[0145] By performing length analysis on the training data, peptides and HLA typing data with a length of 8-11 aa were selected as the affinity model training dataset. The training dataset was then classified into peptides and HLA typings with or without affinity based on their affinity values.

[0146] 3.1 Data Cleaning

[0147] 1) Some data points do not have affinity values ​​and need to be removed.

[0148] 2) In the existing collected data (Table 4), it can be seen that the MHC (major histocompatibility complex) in the field type (representing subtype) not only includes human MHC (i.e. HLA), but also animal subtypes (such as H2-Kb, Ptal-N01:01). Therefore, non-human data needs to be removed from the training dataset.

[0149] Table 4 MHC typing information

[0150] AAFEFINSL H2-Kb 1 mass AARPATSTL HLA-B07:02 1 mass AIMDKNIML HLA-A02:01 1 mass MYIFLHTVD H2-Kb 1 mass VVMPLYQSHW Ptal-N01:01 0 mass KYFDEHYEY HLA-C03:01 1 mass

[0151] 3.2 Data Processing

[0152] IC50 data: When the affinity value A is less than 500, the HLA type and the peptide bind with affinity; when the affinity value A is greater than 500, the HLA type and the peptide do not bind, i.e., there is no affinity. Affinity values ​​A greater than 40000 are set to a fixed value of 40000.

[0153] Mass spectrometry data: When the affinity value B is 1, the HLA type and the peptide bind effectively and have affinity; when the affinity value B is 0, the HLA type and the peptide do not bind, that is, there is no affinity.

[0154] Because the two existing data sources present affinity values ​​in different formats, this study effectively processed the continuous values ​​of IC50 affinity value A and the discrete values ​​of affinity value B in the mass spectrometry data to achieve better model training results and facilitate data fusion. Analysis of the normal distribution and log-transformed normal distribution of the affinity values ​​B in the mass spectrometry data revealed that negative data mainly concentrated between 0 and 0.4, while positive data mainly concentrated between 0.4 and 0.8. Therefore, for mass spectrometry data with an affinity value B of 1, a random integer C is generated from a uniform distribution between 1 and 100; for mass spectrometry data with an affinity value B of 0, a random integer C is generated from a uniform distribution between 1000 and 40000.

[0155] The affinity values ​​C from mass spectrometry data and A from IC50 data are combined. Affinity values ​​A and C are then converted to generate affinity value D. According to formula (1), affinity values ​​A and C are converted into a decimal f(x) between 0 and 1, where x represents the affinity value A or C between the HLA typing and the peptide sequence, and f(x) represents the affinity value D. The affinity value of 500 is used as the critical point for determining whether HLA typing and peptides have affinity; that is, when x = 500, the threshold obtained by calculating f(x) is 0.42562.

[0156] f(x)=1-logx / log50000 (1)

[0157] 3.3 Amino acid encoding

[0158] HLA typing and peptide amino acid encoding methods involve a combination of BLOSUM encoding (as shown in the BLOSUM50 encoding matrix in Table 5), unique heat encoding, and physicochemical property encoding (as shown in Table 6). Peptides of variable length (8-11 amino acids) are padded with an "X" character at the end, encoding them as 12 amino acid sequences (e.g., RAQNSPYDCXXX). Each HLA typing corresponds to a different hypothetical sequence, but all sequences are 34 amino acids in length (as shown in Table 7, which displays some hypothetical sequences for HLA typing).

[0159] Table 5 BLOSUM50 encoding matrix

[0160] R -2 7 -1 -2 -4 1 0 -3 0 -4 -3 3 -2 -3 -3 -1 -1 -3 -1 -3 N -1 -1 7 2 -2 0 0 0 1 -3 -4 0 -2 -4 -2 1 0 -4 -2 -3 D -2 -2 2 8 -4 0 2 -1 -1 -4 -4 -1 -4 -5 -1 0 -1 -5 -3 -4 C -1 -4 -2 -4 13 -3 -3 -3 -3 -2 -2 -3 -2 -2 -4 -1 -1 -5 -3 -1 Q -1 1 0 0 -3 7 2 -2 1 -3 -2 2 0 -4 -1 0 -1 -1 -1 -3 E -1 0 0 2 -3 2 6 -3 0 -4 -3 1 -2 -3 -1 -1 -1 -3 -2 -3 G 0 -3 0 -1 -3 -2 -3 8 -2 -4 -4 -2 -3 -4 -2 0 -2 -3 -3 -4 H -2 0 1 -1 -3 1 0 -2 10 -4 -3 0 -1 -1 -2 -1 -2 -3 2 -4 I -1 -4 -3 -4 -2 -3 -4 -4 -4 5 2 -3 2 0 -3 -3 -1 -3 -1 4 L -2 -3 -4 -4 -2 -2 -3 -4 -3 2 5 -3 3 1 -4 -3 -1 -2 -1 1 K -1 3 0 -1 -3 2 1 -2 0 -3 -3 6 -2 -4 -1 0 -1 -3 -2 -3 M -1 -2 -2 -4 -2 0 -2 -3 -1 2 3 -2 7 0 -3 -2 -1 -1 0 1 F -3 -3 -4 -5 -2 -4 -3 -4 -1 0 1 -4 0 8 -4 -3 -2 1 4 -1 P -1 -3 -2 -1 -4 -1 -1 -2 -2 -3 -4 -1 -3 -4 10 -1 -1 -4 -3 -3 S 1 -1 1 0 -1 0 -1 0 -1 -3 -3 0 -2 -3 -1 5 2 -4 -2 -2 T 0 -1 0 -1 -1 -1 -1 -2 -2 -1 -1 -1 -1 -2 -1 2 5 -3 -2 0 W -3 -3 -4 -5 -5 -1 -3 -3 -3 -3 -2 -3 -1 1 -4 -4 -3 15 2 -3 Y -2 -1 -2 -3 -3 -1 -2 -3 2 -1 -1 -2 0 4 -3 -2 -2 2 8 -1 V 0 -3 -3 -4 -1 -3 -3 -4 -4 4 1 -3 1 -1 -3 -2 0 -3 -1 5 X 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0

[0161] Table 6 Physicochemical Property Encoding Matrix

[0162] R 0.667 0.726 0.321 1 0.905 0.529 0.327 0.169 0.372 1 1 N 0.745 0.39 0.164 0.658 0.51 0.235 0.14 0.313 0.116 0.065 0.33 D 0.745 0.304 0.021 0.793 0.515 0.235 0.14 0.601 0.14 0.956 0 C 0.608 0.314 0.76 0.072 0 0.559 0.14 0.947 0.907 0.028 0.285 Q 0.667 0.531 0.178 0.649 0.608 0.529 0.14 0.416 0.023 0.068 0.36 E 0.667 0.482 0.092 0.883 0.602 0.529 0.14 0.561 0.163 0.96 0.056 G 0 0 0.275 0.189 0.103 0 0 0.24 0.581 0 0.401 H 0.686 0.554 0.326 0.468 0.402 0.529 0.14 0.313 0.581 0.992 0.603 I 1 0.65 1 0 0.083 0.824 0.308 0.424 0.93 0.003 0.407 L 0.961 0.65 0.734 0.081 0.138 0.824 0.308 0.463 0.907 0.003 0.402 K 0.667 0.692 0 0.568 1 0.529 0.327 0.313 0 0.952 0.872 M 0.765 0.612 0.603 0.171 0.206 0.765 0.308 0.405 0.814 0.028 0.372 F 0.686 0.772 0.665 0 0.114 0.853 0.682 0.462 1 0.007 0.339 P 0.353 0.372 0.012 0.198 0.411 0.588 0.271 0 0.302 0.03 0.442 S 0.52 0.172 0.155 0.477 0.303 0.206 0 0.24 0.419 0.032 0.364 T 0.49 0.349 0.256 0.523 0.337 0.235 0.14 0.313 0.419 0.032 0.362 W 0.686 1 0.681 0.207 0.219 1 1 0.537 0.674 0.04 0.39 Y 0.686 0.796 0.591 0.477 0.454 0.853 0.682 1 0.419 0.031 0.362 V 0.745 0.487 0.859 0.036 0.094 0.647 0.234 0.369 0.674 0.003 0.399

[0163] Table 7. Hypothesis sequences consisting of 34 amino acids corresponding to HLA typing.

[0164]

[0165]

[0166] 4. Feature Engineering

[0167] Through the above analysis of the training dataset across various dimensions, it was discovered that there are many internal data patterns that can be mined when determining the affinity between HLA types and peptides. Therefore, this invention performs multi-dimensional internal feature mining and statistics on HLA types and peptides, converting the statistical results into features of the binding degree between the same peptide and different HLA types in the training dataset, features of the binding degree between the same HLA type and different peptides, frequency features of different amino acids appearing at the same position in peptide sequences corresponding to all HLA types, features of amino acids at each position of the peptide sequence binding to the HLA type, and motif features of amino acids in the data of different HLA supertypes (HLA-A, HLA-B, HLA-C) and their corresponding peptides. By extracting latent information in this way, the feature dimensions of the training and test datasets are expanded, and 500-dimensional features are retained after feature encoding and feature filtering.

[0168] Features include one or more of the following: polypeptide sequence distribution features, HLA typing features, features of correlation between polypeptides and HLA typing, or amino acid position and distribution features.

[0169] The multidimensional training data can be any one of the following: features obtained by data screening of the same polypeptide, features obtained by data screening of the same subtype, or features extracted based on the amino acid positions of the polypeptide sequence.

[0170] 4.1 Data screening of the same polypeptide yields feature A:

[0171] The statistics include the number of existing HLA types, the total number of data points on the binding degree of each peptide to different HLA types, the number of data points with affinity between each peptide and different HLA types, the percentage of data points with affinity, and the mean and standard deviation of affinity value D. The statistics also include the number of data points with no affinity between each peptide and different HLA types, the percentage of data points with no affinity, and the mean and standard deviation of affinity value D.

[0172] 4.2 Data filtering for the same subtype yields feature B:

[0173] The statistics include the number of peptides present, the total number of data for each HLA type and different peptide binding degrees, the number of data with affinity for each type and different peptide binding degrees, the percentage of data with affinity D, the mean and standard deviation of affinity D, and the number of data without affinity for each type and different peptide binding degrees, the percentage of data with affinity D, the mean and standard deviation of affinity D.

[0174] 4.3 Feature C is obtained by extracting amino acid positions from the polypeptide sequence:

[0175] Peptide sequences of 8-11 amino acids are transformed into fixed-length sequences of 12 amino acids (e.g., RANDCEAYLNXX) by inserting "X" at the end of the sequence. Amino acids at each position are extracted to generate 12 new fields. These 12 new fields are then grouped and statistically analyzed. The number, percentage, mean, and standard deviation of data showing affinity between the peptide and different HLA types at each position, as well as the number, percentage, mean, and standard deviation of affinity values ​​(D), are calculated. Similarly, the number, percentage, mean, and standard deviation of data showing no affinity between the peptide and different HLA types at each position are also analyzed. Furthermore, different positional weights are assigned to the frequency characteristics of different amino acids appearing at the same position in peptide sequences corresponding to all HLA types, and to the concentration of motif distribution of amino acids in the corresponding peptide data for different HLA supertypes (HLA-A, HLA-B, HLA-C).

[0176] 4.4 Feature Encoding

[0177] The generated feature data includes both categorical and numerical data. Categorical data is unordered and defined as an object in Pandas (a Python library), such as characters 'A' or 'HLA'. Numerical data is ordered, such as numbers like 100 and 1.0. Numerical features are left unprocessed. Categorical features are encoded using the `LabelEncoder` or `OneHotEncoder` functions from the Scikit-learn (a Python machine learning library) data processing module, converting the original data into a format recognizable by machine learning models.

[0178] Then, using feature selection methods from Scikit-learn (univariate feature selection and recursive feature elimination), features are selected by reducing the size of the feature set being examined. First, the prediction model is trained on the original features, with each feature assigned a weight. Then, features with the smallest absolute weight are removed from the feature set, and this process is repeated recursively until the number of remaining features reaches the required number. Finally, 500 features are selected and retained as the input data for the final model.

[0179] 5. Model Selection

[0180] GBDT (Gradient Boosting Decision Tree) is a long-standing and popular model in machine learning. Its main idea is to use weak classifiers (decision trees) for iterative training to obtain the optimal model. This model has advantages such as good training performance and low overfitting risk. Common machine learning algorithms, such as neural networks, can be trained in mini-batch mode, and the size of the training data is not limited by memory. However, GBDT requires traversing the entire training data multiple times in each iteration. Loading the entire training data into memory limits its size; not loading it into memory results in excessively long read and write times. Especially when dealing with massive industrial datasets, the ordinary GBDT algorithm cannot meet the requirements due to these limitations. The main reason for proposing LightGBM (Light Gradient Boosting Machine) is to solve the problems encountered by GBDT in handling massive datasets. The LightGBM framework uses a histogram-based algorithm, supports highly efficient parallel training, and has advantages such as faster training speed, lower memory consumption, higher accuracy, and the ability to quickly process massive datasets.

[0181] 5.1 Basic Principles of LightGBM

[0182] 1) Histogram-based algorithm: First, discretize the continuous floating-point feature values ​​into m integers, and construct a histogram of width m. When traversing the data, accumulate statistics in the histogram using the discretized values ​​as indices. After one data traversal, the histogram has accumulated the required statistics. Then, based on the discrete values ​​of the histogram, find the optimal split point. The advantages are lower memory usage and lower computational cost.

[0183] 2) Histogram Subtraction Speedup: The histogram of a leaf node can be obtained by subtracting the histogram of its parent node from the histogram of its sibling node, which can be twice as fast as a tree model. Normally, constructing a histogram requires traversing all the data in the leaf, but histogram subtraction only requires traversing the k buckets of the histogram. In the actual tree construction process, LightGBM can also first calculate the leaf node with the smaller histogram, and then use histogram subtraction to obtain the leaf node with the larger histogram. This allows the histogram of its sibling leaf to be obtained at a very low cost.

[0184] 3) Leaf-wise algorithm with depth constraints

[0185] LightGBM further optimizes the histogram algorithm. It abandons the level-wise decision tree growth strategy and uses a leaf-wise algorithm with depth constraints. The leaf-wise growth strategy involves finding the leaf with the largest splitting gain from all current leaves and splitting it, repeating this process. Therefore, compared to level-wise, leaf-wise has the advantage of reducing error and achieving better accuracy with the same number of splits; however, it can lead to overfitting by growing a relatively deep decision tree. Therefore, LightGBM adds a maximum depth constraint to the leaf-wise algorithm to prevent overfitting while maintaining high efficiency.

[0186] 4) One-sided gradient sampling algorithm

[0187] One-sided Gradient Sampling (GOSS) is a sample sampling algorithm designed to discard samples that do not contribute to information gain calculation and retain those that do. According to the definition of information gain, samples with larger gradients have a greater impact on information gain. Therefore, GOSS only retains data with larger gradients during data sampling. However, discarding all data with smaller gradients would affect the overall data distribution. Thus, the GOSS algorithm focuses more on under-trained samples without significantly altering the distribution of the original dataset.

[0188] 5) Mutually Exclusive Feature Bundling Algorithm

[0189] High-dimensional data is often sparse, and this sparsity inspires us to design a lossless method to reduce feature dimensionality. Typically, bundled features are mutually exclusive (e.g., one-hot) to prevent information loss when bundling them. If two features are not completely mutually exclusive (in some cases, both features have non-zero values), a metric called the conflict ratio can be used to measure their non-exclusivity. When this value is low, it's preferable to bundle the not-completely mutually exclusive features without affecting the final accuracy. The exclusive feature bundling algorithm suggests that fusing and bundling some features can reduce the number of features.

[0190] 5.2 Model Construction and Optimization of HLA Typing and Peptides

[0191] This invention selects the iNeo-PRED model, constructed from the regression model in the LightGBM machine learning algorithm based on Python. The multi-dimensional internal statistical features, after encoding and feature selection, are used as input data for the model. Three independent models are trained using the three supertypes HLA-A, HLA-B, and HLA-C respectively, to predict the affinity of peptides and HLA types. LightGBM involves many hyperparameters. Due to its leaf-first splitting method, the key hyperparameters to be optimized are: the number of trees (n_estimators), the maximum number of leaf nodes (num_leaves), the minimum number of samples per leaf node (min_data_leaf), the learning rate, the feature sampling ratio and data sampling ratio per tree, the maximum number of feature bins, and the regularization parameters lambda_l1 and lambda_l2. The L2 loss function is chosen for the model, and the validation set evaluation metrics are mean squared error and the area under the ROC curve (AUC).

[0192] The model was optimized using the built-in `cv` function in LightGBM, which is extremely fast. Rapid cross-validation was performed on the `n_estimators` parameter, representing the number of consecutive trees. The remaining parameters were then subjected to five-fold cross-validation using the `GridSearchCV` function from the Scikit-learn model selection module to obtain the optimal parameters and a ranking of all feature importance. The model with the optimal parameters output the top n (1-500) most important features from the ranking of all feature importance as the final input. The model then extracted n important features from the complete datasets of HLA-A, HLA-B, and HLA-C supertypes as training data to train three independent models, enabling affinity prediction for peptides and HLA typing.

[0193] This invention also constructs a Python-based Convolutional Neural Network (CNN) model. The multi-dimensional internal statistical features, after encoding and feature selection, are used as input data to predict the affinity of peptides and HLA typing. Specifically, the input feature information is automatically extracted and fused through multiple convolutional and max-pooling layers, and finally connected using a fully connected layer of size 512. The CNN uses mean squared error as the loss function, Adam as the network optimizer, a batch size of 256, and an initial learning rate of 0.001, which is decayed during model training based on actual conditions. The maximum number of iterations is set to 25; if the loss function stops improving within 10 iterations, the model is forcibly stopped early. A total of five models were generated, each consisting of more than 25 deep neural networks. The average prediction score of all networks is used as the final prediction result output.

[0194] 5.3 Evaluation Test Set for HLA Typing and Peptide Models

[0195] As an example of the combination of features 4.1-4.3, three iNeo-PRED_A models were constructed based on feature A for each of the HLA-A, HLA-B, and HLA-C supertypes; three iNeo-PRED_AB models were constructed based on the combination of feature A and feature B for each of the HLA-A, HLA-B, and HLA-C supertypes; and three iNeo-PRED models were constructed based on all features for each of the HLA-A, HLA-B, and HLA-C supertypes. These models were used to predict HLA typing and affinity of peptides with lengths of 8-11 amino acids. Simultaneously, a deep learning algorithm CNN model and the iNeo-PRED models were constructed based on all input features for comparison. The test set consisted of multiple datasets from the IEDB Benchmark provided by netMHCpan 4.0, and the evaluation metric was the area under the ROC curve (AUC).

[0196] The Benchmark test suite contains the following data:

[0197] 1) The test dataset under the HLA-A supertype includes the following HLA types: HLA-A*02:01, HLA-A*02:02, HLA-A*02:03, HLA-A*02:06, HLA-A*03:01, HLA-A*11:01, HLA-A*24:02, HLA-A*30:01, HLA-A*30:02, HLA-A*31:01, HLA-A*68:01, and HLA-A*68:01.

[0198] 2) The test dataset under HLA-B supertype includes HLA types such as HLA-B*07:02, HLA-B*15:02, HLA-B27:03, HLA-B*27:04, HLA-B*27:05, HLA-B*27:06, HLA-B*35:01, HLA-B*38:01, HLA-B*39:06, HLA-B*40:01, HLA-B*44:03, HLA-B*55:02, HLA-B*57:01, and HLA-B*58:01.

[0199] 3) The test dataset under the HLA-C supertype includes HLA-C*03:03, HLA-C*04:01, HLA-C*05:01, HLA-C*06:02, HLA-C*07:01, HLA-C*07:02, HLA-C*08:02, HLA-C*12:03, HLA-C*14:02, and HLA-C*15:02.

[0200] 5.4 Comparison of results from HLA typing and peptide models based on different characteristics

[0201] The iNeo-PRED model, constructed using this invention and incorporating all features, along with the iNeo-PRED_A model based on feature A, the iNeo-PRED_AB model based on a combination of features A and B, and the currently popular netMHCpan4.0 model, were used to predict HLA typing and peptide affinity test files for Benchmark HLA-A, Benchmark HLA-B, and Benchmark HLA-C in IEDB. The comprehensive results were compared, and the evaluation metric was the area under the ROC curve (AUC). A higher AUC value indicates a more accurate prediction.

[0202] 1) Use Matplotlib box plots to analyze the benchmark HLA-A ( Figure 6 Benchmark HLA-B Figure 7 ) and Benchmark HLA-C ( Figure 8 A general comparison was made among the three categories.

[0203] Box plots display the distribution of data, including upper and lower quartiles, median, mean, etc., and can also be used to reflect whether there are outliers in the data. The horizontal axis in the plot represents the name of the model used, and the vertical axis represents the AUC value of the corresponding model's prediction result.

[0204] from Figure 6-8It can be seen that the iNeo-PRED_AB model, built based on a combination of features A and B, performs better in predicting results than the iNeo-PRED_A model built based on feature A and the netMHCpan4.0 model. Both the mean and median AUCs are higher than the latter two, indicating that the iNeo-PRED_AB model has better predictive performance. However, the iNeo-PRED model, built based on all features (including features A, B, and C), has higher mean and median AUCs than the iNeo-PRED_A, iNeo-PRED_AB, and netMHCpan4.0 models, indicating that the iNeo-PRED model has better predictive performance. The more concentrated bin distribution of the iNeo-PRED model indicates more stable predictive results. These findings demonstrate that modeling based on all features (including features A, B, and C) performs better than modeling based solely on feature A or a combination of features A and B, suggesting that adding more extracted effective features can improve the model's predictive performance.

[0205] Comparison of 5.5 HLA typing and peptide model results

[0206] The iNeo_PRED model constructed using this invention, along with the popular netMHCpan4.0 and ACME models based on the same input (including features A, B, and C), were used to predict HLA typing and peptide affinity. The comprehensive results were compared, and the evaluation metric was the area under the ROC curve (AUC). The higher the AUC value, the more accurate the prediction.

[0207] 2) Use Matplotlib box plots to analyze the benchmark HLA-A ( Figure 9 Benchmark HLA-B Figure 10 ) and Benchmark HLA-C ( Figure 11 A general comparison was made among the three categories.

[0208] Box plots display the distribution of data, including upper and lower quartiles, median, and other information, and can also be used to reflect whether there are outliers in the data. The horizontal axis in the plot represents the name of the model used, and the vertical axis represents the AUC value.

[0209] Figure 9 The p-value of the AUC between the iNeo_PRED model and the netMHCpan4.0 model is 0.00077, the p-value of the AUC between the iNeo_PRED model and the ACME model is 0.00176, and the p-value of the AUC between the iNeo_PRED model and the CNN model is 0.045.

[0210] Figure 10 The p-value of AUC for the iNeo_PRED model and the netMHCpan4.0 model is 0.01468, the p-value of AUC for the iNeo_PRED model and the ACME model is 0.02558, and the p-value of AUC for the iNeo_PRED model and the CNN model is 0.224.

[0211] Figure 11 The p-values ​​of the AUC values ​​of the iNeo_PRED model and the netMHCpan4.0 model are 0.02824, the p-values ​​of the AUC values ​​of the iNeo_PRED model and the ACME model are 2.7929E-07, and the p-values ​​of the AUC values ​​of the iNeo_PRED model and the netMHCpan4.0 model are 0.0186.

[0212] from Figure 9-11 It can be seen that the average and median AUC values ​​of the iNeo_PRED model are higher than those of the other three existing models, indicating that the iNeo_PRED model performs better in prediction; and the box distribution of the iNeo_PRED model is more concentrated, which means that its prediction results are more stable.

[0213] Of the 31 test sets in the HLA-A benchmark, the iNeo-PRED model outperformed the netMHCpan4.0 and ACME models in AUC on 26 test sets, while achieving the same AUC on the same test sets. This indicates that the iNeo-PRED model outperformed or equaled the other two models in 28 / 31 = 90.3% of the test sets. Particularly noteworthy is the iNeo-PRED model's superior AUC compared to the netMHCpan4.0 model in predicting HLA-A*02:02, HLA-A*02:03, HLA-A*02:06, HLA-A*03:01, HLA-A*11:01, and HLA-A*68:02 classifications.

[0214] Of the 31 test sets in the HLA-B benchmark, 17 test sets showed AUC results that were better than or equal to those of the ACME model on the iNeo-PRED model, accounting for 17 / 23 = 73.9%; and 18 test sets showed AUC results that were better than or equal to those of the netMHCpan4.0 model on the iNeo-PRED model, accounting for 18 / 23 = 78.3%. In particular, the iNeo-PRED model significantly outperformed the netMHCpan4.0 model in predicting HLA-B*27:05, HLA-B*27:06, HLA-B*35:01, HLA-B*39:06, HLA-B*40:01, HLA-B*44:03, and HLA-B*44:03 classifications.

[0215] A total of 11 benchmark HLA-C test set files were used. All test sets outperformed the netMHCpan4.0 model and ACME model in AUC results on the iNeo-PRED model, accounting for 100%.

[0216] 5.6 Comparison of evaluation results of HLA typing and peptide models on small datasets

[0217] The iNeo-PRED model constructed in this invention was compared with the CNN model built based on all the same features, the currently popular netMHCpan4.0 model and ACME model in terms of affinity prediction results under the condition that the amount of peptide and HLA typing data is very scarce. The evaluation index is the AUC value (area under the ROC curve).

[0218] Box plots were used to evaluate the models for test sets with limited peptide and HLA typing data (less than 60 data points per test file) in the benchmark, and the AUC results of different models were compared.

[0219] Box plot ( Figure 12 The graph displays the distribution of AUC values ​​in the evaluation results, including the upper and lower quartiles, median, and other information. It can also be used to reflect whether there are any outliers in the data. The horizontal axis in the graph represents the name of the model used, and the vertical axis represents the AUC value corresponding to the current model. Figure 12The p-value for the AUC of the iNeo-PRED model and the netMHCpan4.0 model is 1E-05, while the p-value for the AUC of the iNeo-PRED model and the ACME model is 7.12E-05. As can be seen from the figure, the iNeo-PRED model has the highest average and median AUC, meaning it performs better than the other three existing models; moreover, it has smaller bin compression, indicating more stable prediction results.

[0220] In Benchmark's prediction results for test sets (43 test files in total) with limited peptide and HLA typing data, 33 test sets showed better AUC results on the iNeo-PRED model than CNN models built based on the same features (76.7%), 32 test sets showed better AUC results on the iNeo-PRED model than the netMHCpan4.0 model (74.4%), and 34 test sets showed better AUC results on the iNeo-PRED model than the ACME model (79.1%).

[0221] In summary, the iNeo-PRED model of this invention demonstrates substantial improvement in prediction performance compared to the CNN, netMHCpan4.0, and ACME models. The three iNeo-PRED models constructed through high-dimensional feature optimization exhibit excellent synergistic effects in improving affinity prediction accuracy. When comparing prediction results on test sets with very limited peptide and HLA typing data, the iNeo-PRED model significantly outperforms the other three existing models in both average and median AUC values, demonstrating more stable prediction results.

[0222] The above embodiments illustrate the basic principles, main features, and advantages of the present invention. However, the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. sequence list <110> Hangzhou New Anjin Biotechnology Co., Ltd. <120> A method for predicting peptide and HLA typing affinity <141> 2022-06-29 <160> 13 <170> SIPOSequenceListing 1.0 <210> 1 <211> 34 <212> PRT <213> Artificial Sequence(Artificial Sequence) <400> 1 Tyr Phe Ala Met Tyr Gly Glu Lys Val Ala His Thr His Val Asp Thr 1 5 10 15 Leu Tyr Val Arg Tyr His Tyr Tyr Thr Trp Ala Val Leu Ala Tyr Thr 20 25 30 Trp Tyr <210> 2 <211> 34 <212> PRT <213> Artificial Sequence(Artificial Sequence) <400> 2 Tyr Ser Ala Gly Tyr Arg Glu Lys Tyr Arg Gln Ala Asp Val Asn Lys 1 5 10 15 Leu Tyr Leu Arg Phe Asn Phe Tyr Thr Trp Ala Glu Arg Ala Tyr Thr 20 25 30 Trp Tyr <210> 3 <211> 34 <212> PRT <213> Artificial Sequence(Artificial Sequence) <400> 3 Tyr Tyr Ser Glu Tyr Arg Asn Ile Tyr Ala Gln Thr Asp Glu Ser Asn 1 5 10 15 Leu Tyr Leu Ser Tyr Asp Tyr Tyr Thr Trp Ala Glu Arg Ala Tyr Glu 20 25 30 Trp Tyr <210> 4 <211> 34 <212> PRT <213> Artificial Sequence(Artificial Sequence) <400> 4 Tyr Tyr Ala Met Tyr Gln Glu Asn Val Ala Gln Thr Asp Val Asp Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Ala Ala Gln Ala Tyr Arg 20 25 30 Trp Tyr <210> 5 <211> 34 <212> PRT <213> Artificial Sequence(Artificial Sequence) <400> 5 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Arg 20 25 30 Gly Tyr <210> 6 <211> 34 <212> PRT <213> Artificial Sequence(Artificial Sequence) <400> 6 Tyr Ser Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Arg 20 25 30 Gly Tyr <210> 7 <211> 34 <212> PRT <213> Artificial Sequence <400> 7 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Met Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Arg 20 25 30 Gly Tyr <210> 8 <211> 34 <212> PRT <213> Artificial Sequence <400> 8 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Arg 20 25 30 Gly Tyr <210> 9 <211> 34 <212> PRT <213> Artificial Sequence <400> 9 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Leu Ala Tyr Arg 20 25 30 Gly Tyr <210> 10 <211> 34 <212> PRT <213> Artificial Sequence <400> 10 Tyr Phe Ala Met Tyr Gln Glu Asn Val Ala His Thr Asp Glu Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Arg 20 25 30 Gly Tyr <210> 11 <211> 34 <212> PRT <213> Artificial Sequence <400> 11 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Trp 20 25 30 Gly Tyr <210> 12 <211> 34 <212> PRT <213> Artificial Sequence <400> 12 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Val Ala Arg Val Tyr Arg 20 25 30 Gly Tyr <210> 13 <211> 34 <212> PRT <213> Artificial Sequence <400> 13 Tyr Phe Ala Met Tyr Gln Glu Asn Met Ala His Thr Asp Ala Asn Thr 1 5 10 15 Leu Tyr Ile Ile Tyr Arg Asp Tyr Thr Trp Ala Arg Arg Val Tyr Arg 20 25 30 Gly Tyr

Claims

1. A method for predicting peptide and HLA typing affinity, characterized in that, Includes the following: Step 1, Feature Relationship: Data on the interaction between peptides and HLA typing were obtained, and the data were enhanced after data processing and analysis to obtain the characteristic relationship between the two. The characteristic relationship includes S1, data collection: Training data were obtained from the IEDB database, experiments, and relevant literature, including IC50 data and mass spectrometry data matching peptide and HLA typing. Based on the distribution pattern of mass spectrometry data, the affinity value B of the mass spectrometry data is transformed into an affinity value C through a specific conversion. The affinity value C is then combined with the affinity value A of the IC50 data and transformed again to generate the affinity value D. The specific method is as follows: IC50 data: A threshold is selected for the affinity value A to determine whether there is affinity between the HLA type and the peptide. When the affinity value A is less than the threshold, the binding between the HLA type and the peptide has affinity; when the affinity value A is greater than the threshold, the HLA type and the peptide do not bind, i.e., there is no affinity. A maximum value is set to normalize the data boundaries. When the affinity value A is greater than the maximum value, it is set to the maximum value. Mass spectrometry data: When the affinity value B is 1, the HLA type and the peptide bind effectively and have affinity; when the affinity value B is 0, the HLA type and the peptide do not bind, i.e., there is no affinity. Effective processing is performed on the affinity values ​​B, which are discrete values ​​of 0 and 1 in the mass spectrometry data: A strong affinity range is defined, truncated between 0 and a threshold. When the affinity value B is 1, an integer is randomly generated within this strong affinity range as the affinity value C. A weak affinity range is also defined, truncated between the threshold and the maximum value. When the affinity value B is 0, an integer is randomly generated within this weak affinity range as the affinity value C. The affinity value C from the mass spectrometry data and the affinity value A from the IC50 data are combined and then converted to generate the affinity value D. According to formula (1), the affinity values ​​A and C are converted into decimals f(x) between 0 and 1, where x represents the affinity value A or the affinity value C, and f(x) represents the affinity value D; f(x) = 1 - logx / log50000 (1); Step 2, Feature Engineering: Analyze the feature relationships obtained in step one, perform HLA typing and peptide feature mining, and then perform feature screening to generate multi-dimensional training data for building machine learning models. Step 3, Model Building: The iNeo-PRED model, built using the LightGBM regression model based on Python machine learning algorithm, uses the mined multi-dimensional internal statistical features as input data after encoding and feature selection. It trains models for the three major supertypes HLA-A, HLA-B, and HLA-C respectively, constructing three complete and independent models to predict the affinity of peptides and HLA types.

2. The method for predicting peptide and HLA typing affinity according to claim 1, characterized in that, The method for obtaining the feature relationship in step one is as follows: The feature relationship also includes S2, Data Analysis: 1) Analyze affinity values ​​using IC50 and mass spectrometry data, including the distribution of affinity value A from IC50 data, the distribution of affinity value B from mass spectrometry data, and the ratio of positive to negative data. 2) The lengths of the peptide sequences in the collected data included four types: 8, 9, 10, and 11 amino acids. The hypothetical sequences corresponding to HLA typing originated from amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing was less than 10 angstroms. Based on the supertype to which each HLA typing belonged, three supertypes were obtained: HLA-A, HLA-B, and HLA-C. Then, three training datasets were selected based on the supertypes, and three models were trained based on the training data of each supertype. 3) Analyze the data collected in step S1. The analysis includes: the proportion distribution of binding levels of the same peptide with different HLA types; the overall data distribution of binding levels of the same peptide with different HLA types in IC50; the probability distribution of positive data of binding levels of the same peptide with different HLA types in IC50; the probability distribution of negative data of binding levels of the same peptide with different HLA types in IC50; the overall data probability distribution of binding levels of the same peptide with different HLA types in mass spectrometry; the probability distribution of positive data of binding levels of the same peptide with different HLA types in mass spectrometry; and the probability distribution of negative data of binding levels of the same peptide with different HLA types in mass spectrometry. Probability distribution of negative data with different HLA typing binding levels; proportional distribution of data with different peptide binding levels for the same HLA typing; overall data distribution of the same typing with different peptide binding levels in IC50; probability distribution of positive data with the same typing with different peptide binding levels in IC50; probability distribution of negative data with the same typing with different peptide binding levels in IC50; overall data probability distribution of the same typing with different peptide binding levels in mass spectrometry; probability distribution of positive data with the same typing with different peptide binding levels in mass spectrometry; probability distribution of negative data with the same typing with different peptide binding levels in mass spectrometry. 4) The frequency of different amino acids appearing at the same position in all HLA typing peptides, the probability distribution of positive data of HLA typing and peptide binding degree under each amino acid position of the peptide sequence, the probability distribution of negative data of HLA typing and peptide binding degree, and the analysis of amino acid motif distribution in HLA-A supertype and its corresponding peptide data, HLA-B supertype and its corresponding peptide data, and HLA-C supertype and its corresponding peptide data; S3, Data Preprocessing: By performing length analysis on the training data, peptides and HLA typing data with a length of 8-11 amino acids were selected as the model training dataset. The training dataset was then classified into peptides and HLA typings based on affinity values, categorizing them as affinity-free or affinity-free. The specific process included: 1) Remove data that does not have affinity values ​​and remove data from non-human species; 2) Based on the distribution pattern of mass spectrometry data, the affinity value B of the mass spectrometry data is transformed into the affinity value C through a specific conversion. The affinity value C and the affinity value A of the IC50 data are combined and then converted to generate the affinity value D. 3) By combining BLOSUM coding, unique heat coding, and physicochemical property coding, amino acid coding is performed on HLA types and peptides; peptides with a length of 8-11 amino acids are fixed into amino acid sequences of uniform length by filling in specific characters; hypothetical sequences corresponding to different HLA types are fixed into uniform length within the same supertype.

3. The method for predicting peptide and HLA typing affinity according to claim 2, characterized in that, In step S2, during data analysis, the distribution of affinity value A in IC50 data and the distribution of affinity value B in mass spectrometry data are analyzed. Specifically, the affinity value A in IC50 data is a continuous value greater than 0, while the affinity value B in mass spectrometry data takes two discrete values, 0 and 1. There is a difference between the distributions of affinity value A and affinity value B.

4. The method for predicting peptide and HLA typing affinity according to claim 2, characterized in that, In step S2, the hypothetical sequence corresponding to HLA typing originates from amino acids at positions where the spatial atomic binding distance between the peptide and the HLA typing site is less than 4 angstroms.

5. The method for predicting peptide and HLA typing affinity according to claim 2, characterized in that, The amino acids in each HLA type and its corresponding peptide data have different motif distributions. Based on this distribution, the position of each amino acid in the peptide is assigned a corresponding weight.

6. The method for predicting peptide and HLA typing affinity according to claim 2, characterized in that, The BLOSUM encoding is a BLOSUM50 encoding matrix.

7. The method for predicting peptide and HLA typing affinity according to claim 2, characterized in that, The one-hot encoding is a one-hot encoding.

8. The method for predicting peptide and HLA typing affinity according to claim 1, characterized in that, The features described in step two include one or more of the following: polypeptide sequence distribution features, HLA typing features, features of correlation between polypeptides and HLA typing, or amino acid position and distribution features.

9. The method for predicting peptide and HLA typing affinity according to claim 1, characterized in that, The multi-dimensional training data mentioned in step two refers to at least one of the following: features obtained by data screening of the same polypeptide, features obtained by data screening of the same subtype, or features extracted based on the amino acid positions of the polypeptide sequence.

10. The method for predicting peptide and HLA typing affinity according to claim 9, characterized in that, The specific features obtained by data screening of the same polypeptide are as follows: The statistics include the number of existing HLA types, the total number of data on the binding degree of each peptide to different HLA types, the number of data with affinity between each peptide and different HLA types, the percentage of data with affinity, and the mean and standard deviation of affinity value D, and the number of data without affinity between each peptide and different HLA types, the percentage of data with affinity, and the mean and standard deviation of affinity value D. The specific content of the features obtained by filtering data for the same subtype is as follows: The total number of peptides present, the total number of data for each HLA type and different peptide binding degrees, the number of data with affinity for each HLA type and different peptide binding degrees, the percentage of data with affinity, and the mean and standard deviation of affinity value D, and the number of data without affinity for each HLA type and different peptide binding degrees, the percentage of data with affinity, and the mean and standard deviation of affinity value D; The specific content of the features extracted based on the amino acid positions of the polypeptide sequence is as follows: For polypeptide sequences of 8-11 amino acids in length, a fixed-length sequence is generated by filling in fixed characters. The number, proportion, and average and standard deviation of the affinity values ​​D for amino acids at each position with HLA typing are counted. The number, proportion, and average and standard deviation of the affinity values ​​D for amino acids at each position with HLA typing are counted. The frequency characteristics of different amino acids appearing at the same position in the polypeptide sequences corresponding to all HLA typings are also counted. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in the HLA-A supertype and its corresponding polypeptide data. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in the HLA-B supertype and its corresponding polypeptide data. The weights of amino acids at different positions are assigned according to the motif distribution concentration of amino acids in the HLA-C supertype and its corresponding polypeptide data. The feature data generated above includes both categorical and numerical data. Numerical features are retained without any additional processing, while categorical features are converted into categorical codes and one-hot codes, transforming the original data into a data format that can be recognized by machine learning models.