A method for predicting umami peptide recognition threshold based on machine learning and amino acid sequence information

By constructing an integrated model based on machine learning and amino acid sequence information, we can predict the umami peptide identification threshold, which solves the problem of low screening efficiency of umami peptides in existing technologies and achieves efficient and accurate screening and quantitative prediction of umami peptides.

CN119479908BActive Publication Date: 2025-10-28SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411600974.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-10-28
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing technologies for screening umami peptides suffer from high labor costs, long time cycles, and low yields. Furthermore, existing models can only perform binary classification judgments on whether something is umami-rich, which cannot meet the needs of rapid pre-screening of large quantities of umami peptides.

Method used

Using machine learning and amino acid sequence information-based methods, umami peptide data are digitally characterized by molecular fingerprints, molecular descriptors, and amino acid matrices. An integrated model is constructed by combining Pearson correlation, F-regression, mutual information, and encapsulation screening to predict the umami peptide identification threshold, and a gradient boosting machine model is used for fitting.

Benefits of technology

It enables rapid screening of umami peptides, with efficient and accurate calculations, providing quantitative prediction references for large-scale umami peptides, reducing costs and improving screening efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479908B_ABST
    Figure CN119479908B_ABST
Patent Text Reader

Abstract

This invention provides a method for predicting the recognition threshold of umami peptides based on machine learning and amino acid sequence information. The method includes: acquiring umami peptide data and preprocessing the data; digitally representing the preprocessed umami peptide data using molecular fingerprints, molecular descriptors, and an amino acid matrix; performing feature screening on the digitally represented umami peptide data using Pearson correlation, F-regression, mutual information, and encapsulation to obtain screened molecular descriptors, amino acid counts, and molecular fingerprints; constructing sub-models using the screened molecular descriptors, amino acid counts, and molecular fingerprints, evaluating them using regression model evaluation metrics, selecting the top six best-performing sub-models for ensemble model construction, fitting the ensemble model using a gradient boosting machine model; and predicting the umami peptide recognition threshold using the constructed ensemble model. This invention can quantitatively predict a large number of umami peptides, thus providing a quantitative reference for large-scale screening of umami peptides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine learning, specifically relating to a method for predicting the threshold of umami peptide recognition based on machine learning and amino acid sequence information. Background Technology

[0002] Umami is one of the five basic tastes, and umami peptides are the most frequently reported umami components due to their excellent nutritional properties and sensory appeal. Common umami substances include nucleotides and their derivatives (IMP, BMP, etc., totaling 30 types), amino acids, polypeptides, organic acids and their salts (succinic acid), among which umami peptides are the most studied and widely distributed. Umami peptides can serve as natural flavor modifiers and nutritional supplements, reducing reliance on synthetic additives. Traditional methods for umami peptide discovery, such as ultrafiltration, gel filtration chromatography, reversed-phase high-performance liquid chromatography, and mass spectrometry combinations, suffer from high labor costs, long processing times, and low yields. With advancements in computer performance and algorithm accuracy, computer-aided rapid screening of umami peptides has become possible. Existing models, which can only perform binary classification to determine umami levels, are insufficient for the rapid pre-screening of large quantities of umami peptides; therefore, a high-performance umami intensity prediction model is urgently needed. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0005] A method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information includes:

[0006] Acquire umami peptide data and preprocess the umami peptide data;

[0007] The pre-processed umami peptide data were digitally characterized using molecular fingerprinting, molecular descriptors, and amino acid matrices.

[0008] Feature screening of umami peptide data after digital characterization was performed using Pearson correlation, F regression, mutual information, and encapsulation to obtain the selected molecular descriptors, amino acid counts, and molecular fingerprints.

[0009] Sub-models were constructed using the selected molecular descriptors, amino acid counts, and molecular fingerprints. These sub-models were evaluated using regression model evaluation metrics. The top six sub-models with the best performance were selected for ensemble model construction. The ensemble model was then fitted using a gradient booster model.

[0010] The threshold for umami peptide identification is predicted using a well-constructed ensemble model.

[0011] Furthermore, the preprocessing specifically includes:

[0012] First, sensory tests and in vitro synthesis verification were performed on the obtained umami peptide data. Then, deduplication was performed, and finally, the recognition threshold of the same peptide was averaged.

[0013] Furthermore, the evaluation metrics for the regression model are the coefficient of determination, the variance explained, and the mean absolute error.

[0014] Furthermore, the coefficient of determination (R) 2 ), variance explained (Explained_variance) and

[0015] The formula for Mean Absolute Error (MAE) is as follows:

[0016]

[0017]

[0018]

[0019] Among them, y i This represents the i-th observation; This represents the model's prediction for the i-th observation; It is all y i The mean; ∑ is the summation symbol, indicating that the following expression is accumulated from the first observation to the nth observation i; Var{y} represents the variance of the observations, representing the total degree of variation in the data; This represents the difference between the actual value y and the predicted value for each observation. The absolute value of the difference between them.

[0020] Furthermore, the specific parameters of the Pearson screening are as follows:

[0021] The Pearson correlation coefficient measures the linear relationship between X-Input and Y-Value. The result is between 1 and 1. A result of 1 or -1 represents a positive linear correlation and a negative linear correlation, respectively. A result of 0 indicates no correlation.

[0022] Furthermore, the specific parameters of the F-regression are as follows:

[0023] The univariate linear regression test (F-regression) uses a fast linear model to test the individual regression effects of multiple regressors in turn, and uses the F-score (F-statistic) and significance (p-values) as the evaluation results.

[0024] Furthermore, the specific parameters of the mutual information are as follows: mutual information is a non-negative value that measures the dependency between continuous variables. It is equal to zero if and only if the two random variables are independent, and a higher value means a higher dependency.

[0025] Furthermore, the specific parameters of the wrap-around filter are as follows: In order to take into account both linear data and nonlinear features, Random Forest Regressor and Linear Regression are selected as the algorithms for the evaluator. If a feature is true in at least one of the two model evaluators, then the feature will be retained.

[0026] Compared with the prior art, the present invention has the following advantages:

[0027] This invention establishes a rapid screening model for umami peptides. This model is computationally efficient and accurate, requiring no complex preprocessing or expensive equipment, and achieves rapid identification based on umami peptide sequences. This method is the first regression model to realize the identification threshold of umami peptides, and can quantitatively predict a large number of umami peptides, thus providing a quantitative reference for large-scale screening of umami peptides. It is worthy of widespread application. Attached Figure Description

[0028] The accompanying drawings illustrate various embodiments generally by way of example rather than limitation, and are used, together with the specification and claims, to explain embodiments of the invention. Where appropriate, the same reference numerals are used in all drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be exhaustive or exclusive embodiments of the apparatus or method.

[0029] Figure 1 This is a schematic diagram illustrating the training process and evaluation results of the model of this invention;

[0030] Figure 2 This is a schematic diagram illustrating the performance of a sub-model of the present invention.

[0031] Figure 3 This is a schematic diagram illustrating the web-based publishing effect of the present invention. Detailed Implementation

[0032] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0033] The following specific embodiments further illustrate the implementation process of this invention.

[0034] A method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information includes:

[0035] Acquire umami peptide data and preprocess the umami peptide data;

[0036] The pre-processed umami peptide data were digitally characterized using molecular fingerprinting, molecular descriptors, and amino acid matrices.

[0037] Feature screening of umami peptide data after digital characterization was performed using Pearson correlation, F regression, mutual information, and encapsulation to obtain the selected molecular descriptors, amino acid counts, and molecular fingerprints.

[0038] Sub-models were constructed using the selected molecular descriptors, amino acid counts, and molecular fingerprints. These sub-models were evaluated using regression model evaluation metrics. The top six sub-models with the best performance were selected for ensemble model construction. The ensemble model was then fitted using a gradient booster model.

[0039] The threshold for umami peptide identification is predicted using a well-constructed ensemble model.

[0040] The preprocessing specifically includes:

[0041] First, sensory tests and in vitro synthesis verification were performed on the obtained umami peptide data. Then, deduplication was performed, and finally, the recognition threshold of the same peptide was averaged.

[0042] The evaluation metrics for the regression model are the coefficient of determination, variance explained rate, and mean absolute error.

[0043] The coefficient of determination (R) 2 The formulas for the explained variance and the mean absolute error (MAE) are as follows:

[0044]

[0045]

[0046]

[0047] (1)coefficient ofdetermination(R 2 (1) Refer to Formula 1, the best score for this indicator is 1.0. If the model always predicts the average y without considering the changes in input features, then R2 will be 0; (2) Explained variance: Refer to Formula 2, the best possible score is 1.0, lower values ​​are worse; (3) Mean absolute error: Refer to Formula 3, the function calculates the mean absolute error. The meanings of the given parameters are as follows: y i This represents the i-th observation; This represents the model's prediction for the i-th observation; It is all y i The mean; ∑ is the summation symbol, indicating that the following expression is accumulated from the first observation to the nth observation i; Var{y} represents the variance of the observations, representing the total degree of variation in the data; This represents the difference between the actual value y and the predicted value for each observation. The absolute value of the difference between them.

[0048] The specific parameters for the Pearson screening are as follows:

[0049] The Pearson correlation coefficient measures the linear relationship between X-Input and Y-Value. The result is between 1 and 1. A result of 1 or -1 represents a positive linear correlation and a negative linear correlation, respectively. A result of 0 indicates no correlation.

[0050] The specific parameters for the F-regression are as follows:

[0051] The univariate linear regression test (F-regression) uses a fast linear model to test the individual regression effects of multiple regressors in turn, and uses the F-score (F-statistic) and significance (p-values) as the evaluation results.

[0052] Furthermore, the specific parameters of the mutual information are as follows: mutual information is a non-negative value that measures the dependency between continuous variables. It is equal to zero if and only if the two random variables are independent, and a higher value means a higher dependency.

[0053] Furthermore, the specific parameters of the wrap-around filter are as follows: In order to take into account both linear data and nonlinear features, Random Forest Regressor and Linear Regression are selected as the algorithms for the evaluator. If a feature is true in at least one of the two model evaluators, then the feature will be retained.

[0054] Table 1 shows the model training data, downloaded from TastePeptidesDB (accessible at http: / / tastepeptides-meta.com / TastePeptidesDB, as of September 20, 2023). A total of 473 umami peptides were collected, of which 426 were validated through sensory testing and in vitro synthesis. After deduplication and averaging the recognition thresholds of identical peptides from multiple studies, a total of 324 peptides with recognition thresholds were obtained. These umami peptides and their recognition thresholds are shown in Table 1.

[0055] Table 1. Umami Peptides and Identification Thresholds

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] Table 2 lists 13 different regressor algorithms used for fitting and training the model. (Using R...) 2 We compared the modeling results of 13 algorithms on various feature data using two metrics: variance explained rate and overall performance. Among them, the decision tree regressor, additional decision tree regressor, random forest regressor, adaptive boosting regressor, and gradient boosting tree regressor showed the best performance. Furthermore, we optimized the parameters of the highest-scoring sub-model using a grid search method. The model parameters and performance graphs are shown below. Figure 2 The test set R for these models 2 Most of the values ​​were around 0.65. The gradient boosting tree algorithm was used to further fit the six models above, ultimately resulting in the ensemble model UmamiIP. UmamiIP demonstrated similar performance on both the test and training sets, with an absolute coefficient R0. 2 The variance was 84.95% / 84.46% (Train / Test set), and the explained variance was 86.82% / 84.47% (Train / Test set). This indicates that the model did not experience significant overfitting. Figure 1 B). In contrast, the average value predicted by these sub-models ( Figure 1 C) did not perform as well as UmamiIP. Since UmamiIP was the first tool for predicting umami intensity, there are almost no comparable models available. Therefore, this invention can only compare UmamiIP with its own sub-models. UmamiIP, however, performed much more rigorously, achieving first place in all three metrics. Figure 1 D). To better explore the predictive effect of UmamiIP, this invention further analyzes the amino acid length distribution ( Figure 1 E) and threshold distribution ( Figure 1 F) The mean absolute error was calculated from two perspectives. From the perspective of amino acid length, short peptides (1-3 amino acids) showed a larger error, presumably because short peptides have higher activity, resulting in a larger error value at the same error rate. UmamiIP performed better in predicting peptides with lengths of 4-10 amino acids, as this length has the highest number of amino acids recorded in TastePeptidesDB. The data diversity used to train UmamiIP increased with the abundance of peptides of the same length; with the help of higher data diversity, UmamiIP could better predict the recognition threshold of peptides of this length. Regarding the threshold, UmamiIP performed better in the region with high threshold density (0-3 mM), and the error then increased with the increase of the threshold value.

[0071] Table 2. Detailed information on hyperparameter search for 13 different regressors.

[0072]

[0073]

[0074]

[0075]

[0076] Example 2

[0077] Figure 3 This document demonstrates the graphical interface and operational logic of a web-based application for predicting umami peptides. Users simply submit the amino acid sequence of the peptide to be predicted in FASTA or CSV file format. The umami assessment results can be downloaded after three processing steps.

[0078] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information, characterized in that, include: Acquire umami peptide data and preprocess the umami peptide data; The pre-processed umami peptide data were digitally characterized using molecular fingerprinting, molecular descriptors, and amino acid matrices. Feature screening of umami peptide data after digital characterization was performed using Pearson correlation, F regression, mutual information, and encapsulation to obtain the selected molecular descriptors, amino acid counts, and molecular fingerprints. Sub-models were constructed using the selected molecular descriptors, amino acid counts, and molecular fingerprints. These sub-models included: RidgeCV, LinearRegression, LinearSVR, Lasso, MLPRegressor, DecisionTreeRegressor, ExtraTreeRegressor, RandomForestRegressor, AdaBoostRegressor, GradientBoostingRegressor, BaggingRegressor, KNeighborsRegressor, and GaussianNB. The model was evaluated using regression model evaluation metrics, and the top 6 best-performing sub-models were selected for ensemble model construction. The ensemble model was then fitted using a gradient booster machine model. The threshold for umami peptide identification is predicted using a well-constructed ensemble model.

2. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 1, characterized in that, The preprocessing specifically includes: First, sensory tests and in vitro synthesis verification were performed on the obtained umami peptide data. Then, deduplication was performed, and finally, the recognition threshold of the same peptide was averaged.

3. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 1, characterized in that, The evaluation metrics for the regression model are the coefficient of determination, variance explained rate, and mean absolute error.

4. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 3, characterized in that, The determination coefficient R 2 The formulas for the variance explained and the mean absolute error (MAE) are as follows: Among them, y i This represents the i-th observation; This represents the model's prediction for the i-th observation; It is all y i The mean; ∑ is the summation symbol, indicating that the following expression is accumulated from the first observation to the nth observation; Var{y} represents the variance of the observations, indicating the total degree of variation in the data; Represents each observation value y i Compared with the predicted value The absolute value of the difference between them.

5. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 1, characterized in that, The specific parameters for the Pearson screening are as follows: The Pearson correlation coefficient measures the linear relationship between X-Input and Y-Value. The result is between -1 and 1. A result of 1 or -1 represents a positive linear correlation and a negative linear correlation, respectively. A result of 0 indicates no correlation.

6. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 1, characterized in that, The specific parameters for the F-regression are as follows: The univariate linear regression test uses a fast linear model to test the individual regression effects of multiple regressors in turn, and uses the F-score and significance as the evaluation results.

7. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 1, characterized in that, The specific parameters of mutual information are as follows: mutual information is a non-negative value that measures the dependency between continuous variables. It is equal to zero if and only if the two random variables are independent, and a higher value means a higher dependency.

8. The method for predicting the umami peptide recognition threshold based on machine learning and amino acid sequence information according to claim 1, characterized in that, The specific parameters of the wrap-around filter are as follows: In order to take into account both linear data and nonlinear features, random forest regressors and linear regressors are selected as the algorithms for the evaluator. If a feature is true in at least one of the two model evaluators, then the feature will be retained.

Citation Information

Patent Citations

  • Method for predicting common physicochemical properties of organic molecules based on multiple learning models

    CN115985415A

  • Aquatic organism acute toxicity multi-classification prediction method based on machine learning and integration method

    CN116403659A