An xgboost-based prediction method for blood pressure-lowering dipeptide activity
By constructing a regression model based on the XGBoost method and combining dipeptide fragment features, the problem of limited and low-accuracy prediction methods for antihypertensive peptides in existing technologies is solved, and efficient and accurate prediction of antihypertensive dipeptide activity is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2026-04-14
AI Technical Summary
There are few existing methods for predicting blood pressure-lowering peptides, and their accuracy is low. Traditional methods are inefficient and cannot meet the technical requirements of the modern functional food industry.
Using an XGBoost-based approach, combining the sequence characteristics, physicochemical characteristics, and atomic composition characteristics of dipeptide fragments, an XGBoost regression model was constructed through principal component analysis and feature optimization to predict the IC50 value of antihypertensive activity.
It improves the accuracy of predicting the activity of antihypertensive dipeptides, simplifies the screening process, reduces screening costs, and achieves efficient prediction results.
Smart Images

Figure CN115938489B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for predicting the activity of antihypertensive dipeptides based on XGBoost, belonging to the fields of food science and biological technology. Background Technology
[0002] Peptides are compounds formed by α-amino acids linked together by peptide bonds, and are intermediate products of protein hydrolysis. A compound formed by the dehydration condensation of two amino acid molecules is called a dipeptide; similarly, there are tripeptides, tetrapeptides, pentapeptides, etc. A peptide composed of three or more amino acid molecules is called a polypeptide. The amino terminus of a polypeptide molecule is called the N-terminus, and the carboxyl terminus is called the C-terminus. The amino acid sequence of a polypeptide chain—the arrangement and combination of amino acid residues at different positions from the N-terminus to the C-terminus—determines the structure and function of the polypeptide.
[0003] Hypertension is a clinical syndrome characterized primarily by elevated systemic arterial blood pressure, which may be accompanied by functional or organic damage to organs such as the heart, brain, and kidneys. Angiotensin-converting enzyme (ACE) is a core component of the renin-angiotensin system (RAS), which controls blood pressure by regulating the volume of blood within the blood vessels. ACE converts inactive angiotensin I into angiotensin II, which has vasoconstrictive effects, thereby causing vasoconstriction and indirectly increasing blood pressure.
[0004] Angiotensin I is a 10-amino acid polypeptide sequence (Asp-Arg-Val-Tyr-Ile-His-Pro-Phe-His-Leu) formed by the activation of angiotensinogen through renin cleavage. ACE inhibitors convert the inactive decapeptide angiotensin I into the octapeptide angiotensin II by removing the C-terminal dipeptide His-Leu. Angiotensin II is a potent substrate-concentration-dependent vasoconstrictor. It binds to the type 1 angiotensin II receptor (AT1), triggering a series of actions that lead to vasoconstriction and increased blood pressure. Inhibiting ACE activity reduces the production of angiotensin II and the destruction of kinins, thereby treating hypertension. ACE inhibitors are widely used as drugs for treating cardiovascular diseases.
[0005] Blood pressure-lowering peptides are substances prepared from food proteins through acid and enzymatic hydrolysis, resulting in low toxicity and high absorption rates. They have relatively small molecular weights, generally below 2000U, making them easily absorbed and utilized by the body. Furthermore, they exhibit strong targeting effects, showing significant benefits for hypertensive patients without lowering normal blood pressure. In recent years, angiotensin-converting enzyme (ACE) inhibitory peptides with blood pressure-lowering activity have become one of the most popular research directions in the field of blood pressure-lowering peptides. Blood pressure-lowering activity is typically determined by IC50. 50 Value expression. Based on the clarification of this scientific question, research on antihypertensive active peptides in the treatment and prevention of hypertension has attracted great attention.
[0006] In contrast to the abundant information available on the production and characterization of antihypertensive peptides, information on the relationship between the structure and activity of food protein-derived antihypertensive peptides is very limited. Traditional techniques, which involve preparing peptide mixtures by enzymatically hydrolyzing proteins, separating and preparing high-purity peptides, and then verifying their ACE inhibitory activity one by one, are inefficient, highly unreliable, and do not meet the technical requirements of modern functional food industry.
[0007] To improve prediction efficiency, some researchers have proposed prediction schemes based on machine learning models. R. He et al. proposed a prediction method based on the Z3 descriptor, which uses the Z3 descriptor combined with partial least squares for prediction. While this method can predict the activity of antihypertensive dipeptides to some extent, its accuracy is not high (J. WU, RE ALUKO, AND S. NAKAI, Structural Requirements of Angiotensin I-Converting Enzyme Inhibitory Peptides: Quantitative Structure). Activity Relationship Study of Di- andTripeptides, J. Agric. Food. Chem, 2006, 54, 732 738); R. He et al. proposed a prediction method based on the Z3 descriptor, which uses the Z3 descriptor combined with an artificial neural network for prediction. M. Sandberg et al. proposed a prediction method based on the Z5 descriptor, which uses the Z3 descriptor combined with partial least squares for prediction. Although both methods can predict the activity of antihypertensive dipeptides to a certain extent, they still have the problem of low accuracy. (R. He, H.Ma, W. Zhao, W. Qu, J. Zhao, L. Luo, and W. Zhu, Modeling the QSAR of ACE-Inhibitory Peptides with ANN and Its Applied Illustration, International journal of peptides, 2012, 2012, 620609; M. Sandberg, L. Eriksson, J.Jonsson, M. Sjostrom, and S. Wold, New Chemical Descriptors Relevant for theDesign of Biologically Active Peptides. A Multivariate Characterization of 87Amino Acids, J. Med. Chem. 1998, 41, 2481-2491) Summary of the Invention
[0008] To address the limitations and low accuracy of current methods for predicting antihypertensive peptides, this invention provides a method for predicting the activity of antihypertensive dipeptides based on XGBoost, as follows:
[0009] Step 1: Extract input features based on the sequence characteristics, physicochemical characteristics, and atomic composition characteristics of dipeptide segments;
[0010] Step 2: Perform principal component analysis on the input features to obtain principal component features;
[0011] Step 3: Train the XGBoost regression model based on the principal component feature input to predict the IC50 antihypertensive activity characterization. 50 value.
[0012] Optionally, step one extracts 51-dimensional input features, which include: 2-dimensional amino acid sequence features, 1-dimensional peptide sequence features, 20-dimensional amino acid combination features, 12-dimensional physicochemical features, and 16-dimensional atomic composition features. The feature extraction process includes:
[0013] S11: Encodes amino acids 1-20, converting the dipeptide sequence into 2D numerical features based on Formula 1:
[0014] (Formula 1)
[0015] in, P For length is L peptide chain sequence, R i Represents amino acid residues;
[0016] S12: Encode the amino acid sequence of the peptide in base-21 digital form, converting the peptide into a 1-dimensional digital feature;
[0017] S13: Extract amino acid combinatorial features from peptide amino acid sequences, and extract protein sequences. P Using a 20-dimensional vector, the peptide segment is transformed into a 20-dimensional numerical feature, as shown in Formula 2:
[0018] (Formula 2)
[0019] in, Indicates the first One amino acid in the protein sequence P The probability of it appearing in the sequence, where T represents the transpose operation;
[0020] S14: Introduce basic physicochemical properties of amino acids to digitally characterize the sequence, including hydrophilicity, hydrophobicity, molecular weight, PK1 value, PK2 value, and PI value, converting the dipeptide segment into 12-dimensional digital features, as shown in Formula 3:
[0021] (Formula 3)
[0022] in, Represents hydrophobicity. Represents hydrophilicity. Represents the molecular mass of the side chain group. The dissociation constant representing the hydroxyl-terminal hydrogen ion. Represents the dissociation constant of the amino-terminal hydrogen ion. Represents the isoelectric point. Representing the N-terminal amino acid residue, Represents the amino acid residue at the C-terminus;
[0023] S15: Introduce the atomic composition features of amino acids to digitally characterize the sequence, including the number of C, H, N, O, and S atoms, the number of bonds, the number of single bonds, and the number of double bonds, converting the dipeptide segment into a 16-dimensional digital feature, as shown in Formula 4:
[0024] (Formula 4)
[0025] In the formula, Represents the type of atom. These represent the number of bonds, the number of single bonds, and the number of double bonds, respectively.
[0026] S16: Merge S11, S12, S13, S14, and S15 to convert dipeptide segments into 51-dimensional digital features.
[0027] Optionally, the training process of the XGBoost regression model includes:
[0028] S21: Perform principal component analysis on the 51-dimensional feature input features to form a principal component expression matrix;
[0029] S22: Based on the principal component expression matrix, calculate the goodness of fit between the principal components and the XGBoost regression model to obtain the preferred principal component features;
[0030] S23: Combine the preferred principal component features and the XGBoost regression model to construct an XGBoost antihypertensive dipeptide activity prediction model incorporating feature selection;
[0031] S24: Based on the preferred principal component features and the corresponding IC 50 A dataset was constructed, and the 500-time hold-out method was applied to divide it into a training set and a test set. The blood pressure-lowering dipeptide activity prediction model was trained using the training set, and the model accuracy was evaluated using the test set.
[0032] S25: The most accurate prediction model for antihypertensive dipeptide activity and its parameters.
[0033] Optionally, the process of forming the principal component expression matrix in S21 includes:
[0034] S31: Calculate the covariance matrix of the 51-dimensional eigenvector F. and the centered value vector ;
[0035] The covariance matrix Each element The calculation method is shown in Formula 5:
[0036] (Formula 5)
[0037] in, The first characteristic vector F representing each peptide segment i A vector composed of features Calculate the covariance between the two features corresponding to each peptide segment;
[0038] The centered value vector Each element The calculation method is shown in Formula 6:
[0039] (Formula 6)
[0040] in, Representing each peptide segment i The average of the features;
[0041] S32: Calculate the covariance matrix in S31 eigenvalues;
[0042] S33: Calculate the covariance matrix in S31 The eigenvectors corresponding to the eigenvalues;
[0043] S34: Sort the eigenvalues in S32 from largest to smallest. The eigenvectors follow the eigenvalue sorting. The coefficient of the first principal component is equal to the eigenvector corresponding to the first largest eigenvalue of the covariance matrix, and so on, to obtain the coefficients of each principal component. ;
[0044] S35: Combining the feature vector F and the centered value vector Substitute the principal component coefficients into S34 to form the principal components. The expression is shown in Formula 7:
[0045] (Formula 7).
[0046] Optionally, the process of obtaining preferred principal component features in S23 includes:
[0047] S41: Select the first three vectors of the principal component expression matrix as principal components, divide the dataset based on ten-fold cross-validation, input them into the XGBoost regression model, train the model parameters with the training set, and test the goodness of fit of the model with the test set.
[0048] S42: According to the order of the principal component expression matrix, add a principal component vector as a principal component. Based on ten-fold cross-validation, divide the dataset and feed it into the XGBoost regression model. Train the model parameters with the training set and test the goodness of fit of the model with the test set. Determine whether the goodness of fit increases. If the goodness of fit no longer increases, stop adding principal components and determine the number of preferred principal components and the preferred principal component features.
[0049] Optionally, the step of constructing the dataset specifically includes:
[0050] S51: Collection of antihypertensive dipeptide sequences and characterization of their antihypertensive activity (IC) 50 Values are used to construct the dataset;
[0051] S52: Extract 51-dimensional input features from the input sequence based on the sequence features, physicochemical features, and atomic composition features of the blood pressure-lowering dipeptide fragment;
[0052] S53: Collect published blood pressure-lowering peptide sequences and their activity IC. 50 value;
[0053] S54: Construct a dataset containing two columns of data. The first column is the amino acid sequence composition of the dipeptide sequence, and the second column is the corresponding antihypertensive activity IC50. 50 Value-related information; remove duplicate peptides from the dataset.
[0054] S55: Remove IC from data center 50 Peptides with uncertain values;
[0055] S56: Retrieve IC 50 The logarithm of the value is often used as the output of the model mapping;
[0056] S57: The dataset is divided using 10-fold cross-validation and 500 hold-outs respectively for training and evaluation;
[0057] Optionally, the method for predicting the activity of the antihypertensive dipeptide is characterized by further employing mean absolute error (MAE), mean squared error (MSE), explained variance score (EVS), and goodness of fit. As an evaluation index for the predictive model of antihypertensive dipeptide activity.
[0058] The beneficial effects of this invention are:
[0059] The method for predicting the activity of antihypertensive dipeptides of the present invention uses the sequence characteristics, amino acid physicochemical properties, and atomic composition characteristics of the dipeptide as inputs. It optimizes the prediction results of the model from the perspective of feature optimization. Compared with the scheme of using some features such as Z3 descriptors and Z5 descriptors for prediction in the prior art, the prediction accuracy of the present invention is higher.
[0060] Furthermore, to address the issue of excessive computational load that may result from introducing more features, this invention, based on principal component optimization, further optimizes and selects principal component features that match the prediction model. In terms of prediction model construction, it utilizes the XGBoost ensemble learning method. Compared to existing prediction schemes, the optimized input features and the constructed prediction model of this invention have a higher degree of matching, achieving higher prediction accuracy while reducing the computational load and complexity of the method. This invention can efficiently predict ACE inhibitory activity before preparation, providing input for post-preparation ACE activity inhibitory peptides and activity verification schemes, effectively simplifying the screening process for antihypertensive peptides and reducing screening costs. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 A block diagram of a method for predicting the activity of an XGBoost antihypertensive dipeptide in an embodiment of the present invention.
[0063] Figure 2 This is a diagram showing the digital encoding results of peptide segments.
[0064] Figure 3 It is a principal component score histogram of 51-dimensional mixed features.
[0065] Figure 4 This is the result of 500 hold-out tests predicting goodness of fit and explained variance variation.
[0066] Figure 5 This shows the changes in the average absolute error and mean square error of the leave-out method over 500 cycles.
[0067] Figure 6 It consists of the amino acid sequences of 10 predicted dipeptide segments. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0069] Example 1:
[0070] This embodiment provides a method for predicting the activity of antihypertensive dipeptides, including:
[0071] Step 1: Extract input features based on the sequence characteristics, physicochemical characteristics, and atomic composition characteristics of dipeptide segments;
[0072] Step 2: Perform principal component analysis on the input features to obtain principal component features;
[0073] Step 3: By training and selecting N principal component features as inputs to the XGBoost regression model, the IC50 of the antihypertensive activity is predicted. 50 value.
[0074] Example 2:
[0075] This embodiment provides a method for predicting the activity of antihypertensive dipeptides based on XGBoost. The overall method includes the training and evaluation of the XGBoost model, and the prediction of antihypertensive dipeptide activity using the trained XGBoost model. The overall process is as follows:
[0076] S1: Collection of antihypertensive dipeptide sequences and characterization of their antihypertensive activity IC 50 Values are used to construct the dataset;
[0077] S2: Extract 51-dimensional features from the input sequence based on peptide sequence features, physicochemical features, and atomic composition features;
[0078] S3: Perform principal component analysis on the input features and map the input features to principal component components;
[0079] S4: Using mean absolute error (MAE), mean squared error (MSE), explained variance score (EVS), and goodness of fit ( () as an indicator for evaluating prediction accuracy;
[0080] S5: Select an initial principal component count of 3, divide the dataset based on 10-fold cross-validation, and feed it into the XGBoost regression model. Train the model parameters on the training set; test the model's goodness of fit on the test set.
[0081] S6: Increase the number of principal components by 1, divide the dataset based on 10-fold cross-validation, and feed it into the XGBoost regression model. Train the model parameters on the training set; test the model's goodness of fit on the test set, and determine whether the goodness of fit has increased. If the goodness of fit no longer increases, stop increasing the number of principal components and determine the number of principal components.
[0082] S7: Combine the input feature principal component and the XGBoost regression model to construct a blood pressure-lowering dipeptide activity prediction model incorporating feature selection;
[0083] S8: Divide the training and test sets using the 500-round hold-out method, and calculate the average MAE, MSE, and EVS of the test set. Evaluate the accuracy of the model;
[0084] S9: Train model parameters based on existing data to generate a prediction model;
[0085] S10: Generate 20 amino acid pairs, totaling 400 dipeptide sequences; remove published dipeptide sequences from these 400 sequences, and use a model to predict the antihypertensive activity IC50 of the remaining unpublished sequences. 50 value.
[0086] Step S1 in this embodiment further includes:
[0087] S11: Collect published antihypertensive peptide sequences and their activity IC. 50 Values; using information on antihypertensive peptides published on two major bioactive peptide websites and the following three literatures, an initial antihypertensive peptide dipeptide sequence and its activity dataset were constructed. This example contains 140 dipeptide data.
[0088] Document 1: B. Manavalan, S. Basith, TH Shin, L. Wei3, G. Lee, mAHTPred: a sequence-based meta-predictor for improving the prediction of anti-hypertensive peptides using effective feature representation, Bioinformatics, 35(16), 2019, 2757–2765.
[0089] Document 2: R. Kumar, K. Chaudhary, JS Chauhan, G. Nagpal, R. Kumar, M. Sharma, GPS Raghava, An in silico platform for predicting, screening and designing of antihypertensive peptides, Sci. Rep., 2015, 5, 12512
[0090] Document 3: J. WU, RE ALUKO, S. NAKAI, Structural Requirements ofAngiotensin I-Converting Enzyme Inhibitory Peptides: Quantitative Structure Activity Relationship Study of Di- and Tripeptides, J. Agric. Food Chem.2006, 54, 732 738
[0091] S12: Construct a dataset with two columns. The first column is the amino acid sequence composition of the dipeptide sequence, and the second column is the corresponding antihypertensive activity IC50. 50 Value-related information; remove duplicate peptides from the dataset.
[0092] S13: Remove IC from the dataset 50 Peptides with values that are uncertain;
[0093] S14: Retrieve IC 50 The logarithmic value is often used as the output of the model mapping to narrow the range of data values.
[0094] Step S2 in this embodiment further includes the following steps:
[0095] S21: Encodes amino acids 1-20, as shown in Table 1; the antihypertensive peptide sequence is a string of varying lengths generated by arranging and combining 20 English letters, with one of length being... L peptide chain sequence P The representation is shown in Formula 1:
[0096] (Formula 1)
[0097] In the formula, Represents amino acid residues;
[0098] The numerical codes of amino acid sequences are shown in Table 1:
[0099] Table 1. Numerical coding of amino acids
[0100]
[0101] Based on Table 1 and Formula 1, dipeptide sequences are converted into 2D numerical features, such as AA being expressed as 11, AC as 12, and so on.
[0102] S22: The amino acid sequence of the peptide is encoded using a base-21 numerical encoding method, converting the peptide into a 1D numerical feature. The results of the conversion of 140 peptides are as follows: Figure 2 As shown;
[0103] S23: Extract amino acid combinatorial features from peptide amino acid sequences, and extract protein sequence... P Using a 20-dimensional vector, the peptide segment is transformed into a 20-dimensional numerical feature, as shown in Formula 2.
[0104] (Formula 2)
[0105] In the formula, Indicates the first The probability of an amino acid appearing in protein sequence P, where T represents the transpose operation;
[0106] S24: The basic physicochemical properties of amino acids are introduced to digitally characterize the sequence, including hydrophilicity, hydrophobicity, molecular weight, PK1 value, PK2 value, and PI value, as shown in Table 2. Based on Table 2, the dipeptide segments are converted into 12-dimensional digital features, as shown in Formula 3:
[0107] (Formula 3)
[0108] In the formula, Represents hydrophobicity. Represents hydrophilicity. Represents the molecular mass of the side chain group. The dissociation constant representing the hydroxyl-terminal hydrogen ion. Represents the dissociation constant of the amino-terminal hydrogen ion. Represents the isoelectric point. Representing the N-terminal amino acid residue, Represents the amino acid residue at the C-terminus;
[0109] S25: The atomic composition features of amino acids are introduced to digitally characterize the sequence, including the number of atom types (CHNOS), the number of bonds, the number of single bonds, and the number of double bonds, as shown in Table 3. Based on Table 3, the dipeptide segments are converted into 16-dimensional digital features, as shown in Formula 4:
[0110] (Formula 4)
[0111] In the formula, Represents the type of atom. These represent the number of bonds, the number of single bonds, and the number of double bonds, respectively.
[0112] Table 2 Common physicochemical properties of 20 amino acids
[0113]
[0114] Table 3 Atomic composition of amino acids
[0115]
[0116] S26: Combine S21, S22, S23, S24, and S26 to convert the peptide into 51-dimensional digital features. The specific feature order is shown in Table 4.
[0117] Table 4. Physical meaning of the 51-dimensional feature vectors
[0118]
[0119] Step S3 in this embodiment further includes the following steps:
[0120] S31: Calculate the covariance matrix of the 51-dimensional mixed eigenvector F. and the centered value vector Each element of the covariance matrix The calculation method is shown in Formulas 5 and 6, where each element of the centered value vector... The calculation method is shown in Formula 6.
[0121] (Formula 5)
[0122] In the formula, The vector composed of the i-th feature of the feature vector F representing each peptide segment. Calculate the covariance between the two features corresponding to each peptide segment;
[0123] (Formula 6)
[0124] In the formula, Representing each peptide segment i The average value of each feature.
[0125] S32: Calculate the covariance matrix in S31 eigenvalues;
[0126] S33: Calculate the covariance matrix in S31 The eigenvectors corresponding to the eigenvalues;
[0127] S34: Sort the eigenvalues in S32 from largest to smallest. The eigenvectors follow the eigenvalue sorting. The coefficient of the first principal component is equal to the eigenvector corresponding to the first largest eigenvalue of the covariance matrix, and so on, to obtain the coefficients of each principal component. ;
[0128] S35: Combining the numerical characteristic F before change and the centered value vector Substitute the principal component coefficients from S34 to form the principal components. The expression is shown in Formula 7.
[0129] (Formula 7)
[0130] Step S4 of this embodiment further includes the following steps:
[0131] S41: Using mean absolute error (MAE), mean squared error (MSE), explained variance score (EVS), and goodness of fit ( () as an indicator for evaluating prediction accuracy;
[0132] S42: Calculate the mean absolute error (MAE) as shown in Equation 8; calculate the mean squared error (MSE) as shown in Equation 9; calculate the explained variance (EVS) as shown in Equation 10; calculate the goodness of fit. The indicators are shown in Formula 11.
[0133] (Formula 8)
[0134] In the formula, The number of data points in the test set. For the first One test data point, for The predicted data.
[0135] (Formula 9)
[0136] (Formula 10)
[0137] In the formula, This represents the variance.
[0138] (Formula 11)
[0139] The initial number of principal components was set to 3. Based on 10-fold cross-validation, the dataset was divided and fed into the XGBoost regression model. The training set was used to train the model parameters, and the test set was used to test the model's goodness of fit.
[0140] Step S5 of this embodiment further includes the following steps:
[0141] S51: Take the first three vectors of the principal component representation matrix calculated above as the principal components;
[0142] S52: Divide the principal component dataset into 10 mutually exclusive subsets of similar size;
[0143] S53: Use 9 subsets as the training set and the remaining 1 as the test set.
[0144] S54: Save the data sequence numbers corresponding to the 10 splits of the training and test sets.
[0145] S55: Install the XGBoost module on the Anaconda platform and import the XGBoost module into the Python IDE (PyCharm);
[0146] S56: The regressor uses the XGBoost module, with a decision tree as the base learner. The squared error is set as the loss function, the subsample ratio is 0.51, the learning rate is set to 0.1, and the number of decision trees is set to 160.
[0147] S57: Feed the principal components into the XGBoost regressor.
[0148] S58: Use training and test sets divided into ten parts, apply the training set to train the model parameters, and use the test set to test the model's goodness of fit.
[0149] S59: Accumulate ten data points and calculate the goodness of fit.
[0150] Step S6 of this embodiment further includes the following steps:
[0151] S61: Based on S5, add a principal component vector as the principal component according to the order of the principal component expression matrix, repeat S5 to train the XGBoost regressor, and calculate the goodness-of-fit value.
[0152] S62: Determine whether the goodness of fit has increased; if the goodness of fit has increased, repeat S61 and S5; otherwise, select the corresponding principal component vector as the input feature to determine the principal components of the model input.
[0153] Step S7 of this embodiment further includes the following steps:
[0154] S71: Combine the principal component features selected in S62 with the XGBoost regressor to construct a prediction model.
[0155] Step S8 of this embodiment further includes the following steps:
[0156] S81: Combining principal component characteristics and ACE activity IC 50 The dataset was divided into two mutually exclusive subsets in a 7:3 ratio.
[0157] S82: Following the partitioning method of S81, randomly generate 500 partitions of data;
[0158] S83: Train the S7 model using the training set, evaluate the model using the test set, and calculate the average MAE, MSE, and EVS on the test set. Evaluate the accuracy of the model.
[0159] Step S9 of this embodiment further includes the following steps:
[0160] S91: Train model parameters based on existing data to generate a prediction model;
[0161] S92: Store the model and model parameters.
[0162] The specific steps for predicting the activity of antihypertensive dipeptides using a trained XGBoost model are as follows:
[0163] Step S10 of this embodiment further includes the following steps:
[0164] S101: Generates 20 amino acid sequences in two-amino acid arrangements, totaling 400 dipeptide sequences;
[0165] S102: 140 dipeptide sequences were removed from 400 dipeptide sequence data;
[0166] S103: Calculate the principal component features of 260 unknown dipeptide sequences;
[0167] S104: Principal component features are fed into the S9 model to predict the antihypertensive activity IC50 of the remaining unpublished sequences. 50 value.
[0168] S105: Sort the predicted values and identify the active IC. 50 Peptide prediction is achieved by selecting the 10 peptides with the lowest values.
[0169] To further illustrate the beneficial effects achievable by this invention, comparative experiments were conducted, and the results are shown in Table 5. Z3-ANN represents Z3 descriptors combined with artificial neural networks; Z3-PLS represents Z3 descriptors combined with partial least squares; Z5-PLS represents Z5 descriptors combined with partial least squares.
[0170] Table 5. Comparison of prediction accuracy of different methods using ten-fold cross-validation
[0171]
[0172] As shown in Table 5, the MAE MSE index of this invention is smaller than that of Z3-ANN, Z3-PLS, and Z5-PLS, indicating that the error of the proposed algorithm is smaller than that of the Z3-ANN, Z3-PLS, and Z5-PLS methods, and the EVS and... The fact that the index is greater than that of Z3-ANN, Z3-PLS, and Z5-PLS indicates that the regression performance of the method in this invention is superior to those methods. Based on evaluation metrics from various different algorithms, it can be seen that the overall prediction performance of this invention is better than that of Z3-ANN, Z3-PLS, and Z5-PLS.
[0173] Some steps in the embodiments of the present invention can be implemented using software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk.
[0174] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the activity of antihypertensive dipeptides, characterized in that, The method for predicting the activity of the antihypertensive dipeptide includes: Step 1: Extract input features based on the sequence characteristics, physicochemical characteristics, and atomic composition characteristics of dipeptide segments; Step 2: Perform principal component analysis on the input features to obtain principal component features; Step 3: By training N principal component features and inputting them into the XGBoost regression model, predict the IC50 of the antihypertensive activity. 50 value; Step one extracts 51-dimensional input features, which include: 2-dimensional amino acid sequence features, 1-dimensional peptide sequence features, 20-dimensional amino acid combination features, 12-dimensional physicochemical features, and 16-dimensional atomic structure features. The feature extraction process includes: S11: Encodes amino acids 1-20, converting the dipeptide sequence into 2D numerical features based on Formula 1: (Official 1) in, P For length is L peptide chain sequence, R i Represents amino acid residues; S12: Encode the amino acid sequence of the peptide in base-21 digital form, converting the peptide into a 1-dimensional digital feature; S13: Extract amino acid combinatorial features from peptide amino acid sequences, and extract protein sequences. P Using a 20-dimensional vector, the peptide segment is transformed into a 20-dimensional numerical feature, as shown in Formula 2: (Official 2) in, Indicates the first One amino acid in the protein sequence P The probability of it appearing in T represents the transpose operation; S14: Introduce basic physicochemical properties of amino acids to digitally characterize the sequence, including hydrophilicity, hydrophobicity, molecular weight, PK1 value, PK2 value, and PI value, converting the dipeptide segment into 12-dimensional digital features, as shown in Formula 3: (Official 3) in, Represents hydrophobicity. Represents hydrophilicity. Represents the molecular mass of the side chain group. The dissociation constant representing the hydroxyl-terminal hydrogen ion. Represents the dissociation constant of the amino-terminal hydrogen ion. Represents the isoelectric point. Representing the N-terminal amino acid residue, Represents the amino acid residue at the C-terminus; S15: Introduce the atomic composition features of amino acids to digitally characterize the sequence, including the number of C, H, N, O, and S atoms, the number of bonds, the number of single bonds, and the number of double bonds, converting the dipeptide segment into a 16-dimensional digital feature, as shown in Formula 4: (Official 4) In the formula, Represents the type of atom. These represent the number of bonds, the number of single bonds, and the number of double bonds, respectively. S16: Merge S11, S12, S13, S14, and S15 to convert the dipeptide segments into 51-dimensional digital features; The training process of the XGBoost regression model includes: S21: Perform principal component analysis on the 51-dimensional feature input features to form a principal component expression matrix; S22: Based on the principal component expression matrix, calculate the goodness of fit between the principal components and the XGBoost regression model to obtain the preferred principal component features; S23: Combine the preferred principal component features and the XGBoost regression model to construct an XGBoost antihypertensive dipeptide activity prediction model incorporating feature selection; S24: Based on the preferred principal component features and the corresponding IC 50 A dataset was constructed, and the 500-time hold-out method was applied to divide it into a training set and a test set. The blood pressure-lowering dipeptide activity prediction model was trained using the training set, and the model accuracy was evaluated using the test set. S25: The most accurate prediction model for antihypertensive dipeptide activity and its parameters.
2. The method for predicting the activity of antihypertensive dipeptides according to claim 1, characterized in that, The process of forming the principal component expression matrix in S21 includes: S31: Calculate the covariance matrix of the 51-dimensional eigenvector F. and the centered value vector ; The covariance matrix Each element The calculation method is shown in Formula 5: (Official 5) in, The first characteristic vector F representing each peptide segment i A vector composed of features Calculate the covariance between the two features corresponding to each peptide segment; The centered value vector Each element The calculation method is shown in Formula 6: (Official 6) in, Representing each peptide segment i The average value of each feature; S32: Calculate the covariance matrix in S31 eigenvalues; S33: Calculate the covariance matrix in S31 The eigenvectors corresponding to the eigenvalues; S34: Sort the eigenvalues in S32 from largest to smallest. The eigenvectors follow the eigenvalue sorting. The coefficient of the first principal component is equal to the eigenvector corresponding to the first largest eigenvalue of the covariance matrix, and so on, to obtain the coefficients of each principal component. ; S35: Combining the feature vector F and the centered value vector Substitute the principal component coefficients into S34 to form the principal components. The expression is shown in Formula 7: (Official 7).
3. The method for predicting the activity of antihypertensive dipeptides according to claim 1, characterized in that, The process of obtaining the preferred principal component features in S23 includes: S41: Select the first three vectors of the principal component expression matrix as principal components, divide the dataset based on ten-fold cross-validation, input them into the XGBoost regression model, train the model parameters with the training set, and test the goodness of fit of the model with the test set. S42: According to the order of the principal component expression matrix, add a principal component vector as a principal component. Based on ten-fold cross-validation, divide the dataset and feed it into the XGBoost regression model. Train the model parameters with the training set and test the goodness of fit of the model with the test set. Determine whether the goodness of fit increases. If the goodness of fit no longer increases, stop adding principal components and determine the number of preferred principal components and the preferred principal component features.
4. The method for predicting the activity of antihypertensive dipeptides according to claim 1, characterized in that, It also includes the step of constructing a dataset, which specifically includes: S51: Collection of antihypertensive dipeptide sequences and characterization of their antihypertensive activity (IC) 50 Values are used to construct the dataset; S52: Extracting 51-dimensional features of the input sequence based on the sequence features, physicochemical features, and atomic composition features of the antihypertensive dipeptide fragment; S53: Collect published blood pressure-lowering peptide sequences and their activity IC. 50 value; S54: Construct a dataset containing two columns of data. The first column is the amino acid sequence composition of the dipeptide sequence, and the second column is the corresponding antihypertensive activity IC50. 50 Value-related information; remove duplicate peptides from the dataset. S55: Remove IC from the dataset 50 Peptides with values that are uncertain; S56: Retrieve IC 50 The logarithm of the value is often used as the output of the model mapping; S57: The dataset is divided using 10-fold cross-validation and 500 hold-outs respectively for training and evaluation.
5. The method for predicting the activity of antihypertensive dipeptides according to claim 1, characterized in that, It also includes the use of mean absolute error (MAE), mean squared error (MSE), explained variance score (EVS), and goodness of fit. As an evaluation index for the predictive model of antihypertensive dipeptide activity.
Citation Information
Patent Citations
Method for predicting activity of angiotensin converting enzyme inhibitor
CN103207947A
Amino acid sequence prediction method for antihypertensive peptide activity
CN114187967A