A method and model for screening umami peptides

By combining molecular docking, molecular descriptors and molecular fingerprints, a umami peptide screening model was constructed, which solved the problems of high cost, long cycle and low accuracy of traditional umami peptide screening methods, and achieved rapid and accurate umami peptide screening and identification.

CN116798544BActive Publication Date: 2025-10-03SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211523847.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-10-03
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

Traditional umami peptide screening methods are costly, time-consuming, and have low accuracy, making it difficult to quickly and efficiently screen out umami peptides. Furthermore, they are highly dependent on artificial sensory experiments, and the results are difficult to transfer.

Method used

A method based on molecular docking, molecular descriptors, molecular fingerprints and integrated algorithms was used to construct a screening model for umami peptides. By organizing existing umami peptide data, molecular fingerprint features and intermolecular interaction residue features were constructed, a screening sub-model was established using a machine learning algorithm, and the screening model was integrated using a support vector machine algorithm.

Benefits of technology

It achieves rapid and accurate screening of umami peptides, reduces costs, improves screening efficiency, reduces dependence on manual experience, and has strong model transferability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116798544B_ABST
    Figure CN116798544B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and a screening model for umami peptides. The screening method comprises the following steps: S1: organizing existing umami peptide data and establishing a database; S2: constructing molecular fingerprint feature data of umami peptides based on the structural fragments of the existing umami peptides themselves; S3: constructing intermolecular interaction residue feature data based on the interaction mode between the umami peptides and the umami receptors T1R1 / T1R3 analyzed by molecular docking technology; S4: obtaining molecular descriptor feature data of the physicochemical properties of the umami peptides based on molecular descriptors; S5: using a machine learning algorithm to establish umami peptide screening sub-models for the data obtained in steps S2-S4; S6: integrating the three umami peptide screening sub-models using a support vector machine algorithm to establish an umami peptide screening model; and S7: screening umami peptides using the umami peptide screening model established in step S6. The screening method of the present invention can quickly and accurately screen umami peptides, and the screening method is reusable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of umami peptides, and in particular to an umami peptide screening method and a screening model. Background Art

[0002] In 2002, umami was listed as the fifth basic taste, following sweet, bitter, sour, and salty. Umami may help reduce the risk of chronic diseases in adults by reducing dietary sodium intake. In addition to monosodium L-glutamate (MSG), some bifunctional acids, free L-amino acids, peptides, and their derivatives or reaction products have been found to possess umami. Peptides and their derivatives are particularly important components of seasonings. Therefore, the discovery of umami peptides is of great significance for the development of new umami seasonings or food additives.

[0003] At present, more than 100 umami peptides have been reported, including CM, GCG, EDG, TESSSE, RGENESEEEGAIVT, etc. At present, multidimensional chromatography and ultra-high performance liquid chromatography-electrospray ionization-quadrupole time-of-flight mass spectrometry (UPLC-ESI-QTOFMS / MS) are mainly used to identify umami peptides in protein hydrolysates. However, the traditional umami peptide screening method has the following shortcomings: (1) The traditional umami substance mining work first requires the selection of an umami base material, and then requires the use of ultrafiltration, gel filtration chromatography, reverse phase-high performance liquid chromatography and mass spectrometry to screen potential umami peptides; finally, the screening results need to be identified by artificial sensory experiments. The cost in terms of manpower, time, and economy is high, and the experimental cycle is long. (2) The traditional umami peptide mining process requires a large amount of umami source materials for experimental extraction, which has a high material cost. (3) Traditional umami peptide judgment relies on a team of trained human sensory personnel. Each artificial sensory panel requires 3-6 months of systematic training for a specific umami base. This long training cycle requires training on specific umami sources, resulting in difficulty in transferring results. It is difficult to form an efficient and unified judgment for new sources in a short period of time. (4) The results of traditional umami peptide mining are not accurate, and are prone to omissions, making it difficult to fully mine.

[0004] Computer-assisted screening can be applied to the screening and identification of umami peptides. It is time-saving, has a unified measurement standard, and is low-cost. This series of methods includes molecular docking, molecular fingerprint modeling, and molecular descriptor QSAR modeling. These three methods have been widely used in the field of active molecule mining, but molecular docking is a method of drug design based on the characteristics of receptors and the interaction between receptors and drug molecules. It is a theoretical simulation method that mainly studies the interaction between molecules (such as ligands and receptors) and predicts their binding mode and affinity. Molecular fingerprints are matrices that describe the characteristic structure of molecules through a set of byte strings filled with a specific number of bits (128 / 1024 / 2048) and 0 / 1. It is often used as a method of molecular characterization to convert 3D molecular structure information into 2D information for machine learning. Molecular descriptors describe the molecules to be characterized from multiple dimensions (0-3 dimensions), multiple angles (physicochemical property indicators), and multiple description methods (qualitative / quantitative). They are often used in the study of quantitative structure-activity relationships, such as molecular composition (such as the number of hydrogen bond donors, the number of chemical bonds), physicochemical properties (such as ester-water distribution coefficient) descriptors, molecular field descriptors, and molecular shape descriptors. Summary of the Invention

[0005] In response to the above-mentioned problems, the present invention aims to provide an umami peptide screening method and screening model based on the combination of molecular docking, molecular descriptors, molecular fingerprints and integrated algorithms, which can quickly and accurately screen umami peptides, and the screening method is reusable.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A method for screening umami peptides, characterized by comprising the following steps:

[0008] S1: Organize existing umami peptide data and establish a database;

[0009] S2: Constructing molecular fingerprint feature data of umami peptides based on the existing structural fragments of umami peptides themselves;

[0010] S3: Based on the interaction mode between umami peptides and umami receptors T1R1 / T1R3 analyzed by molecular docking technology, the characteristic data of intermolecular interaction residues were constructed;

[0011] S4: Based on the molecular descriptors, obtain the molecular descriptor feature data of the physicochemical properties of umami peptides;

[0012] S5: Using a machine learning algorithm to establish umami peptide screening sub-models for the data obtained in steps S2-S4;

[0013] S6: Use the support vector machine algorithm to integrate the umami peptide screening sub-models to establish an umami peptide screening model;

[0014] S7: Screening umami peptides using the umami peptide screening model established in step S6.

[0015] Furthermore, the molecular fingerprint feature data described in step S2 include Morgan148, 322, 428, 509, 598, 650, 805, 952, 1150, 1409, 1573, 1687, 1706, 1907 and 2017.

[0016] Furthermore, in step S5, the umami peptide screening sub-model established for the molecular fingerprint feature data obtained in step S2 using a machine learning algorithm includes a stochastic gradient descent discriminant model, a molecular fingerprint feature data logistic regression model, a molecular fingerprint feature data gradient boosting tree model and a Gaussian distribution naive Bayes discriminant model.

[0017] Furthermore, the intermolecular interaction residue feature data described in step S3 include: HI_A_2_LEU, HI_A_3_LEU, HI_A_108_ASP, HI_A_154_THR, HI_A_157_ALA, HI_A_158_LEU, HI_A_161_PRO, HI_A_163_LEU, HI_A_179_LYS, HI_A_181_GLN, HI_A_182_TYR, HI_A_183_PRO, HI_A_218_ASP, HI_A_246_PRO, HI_A_419_TRP, HI_A_19_THR, HI_A_56_ ARG, HI_B_57_PRO, HI_B_106_PRO, HI_B_107_VAL, HI_B_152_VAL, HI_B_155_LYS, HI_B_156_PHE, HI_B_179_THR, HI_B_245_LEU, HdB_A_48_SER , HdB_A_50_CYS, HdB_A_52_GLN, HdB_A_107_SER, HdB_A_109_SER, HdB_A_148_SER, HdB_A_150_ASN, HdB_A_151_ARG, HdB_A_161_PRO, HdB_A_217 _SER, HdB_A_218_ASP, HdB_A_219_ASP, HdB_A_222_GLN, HdB_A_247_PHE, HdB_A_248_SER, HdB_A_249_ALA, HdB_A_276_SER, HdB_A_278_GLN, Hd B_B_15_LEU, HdB_B_17_PRO, HdB_B_56_ARG, HdB_B_57_PRO, HdB_B_58_SER, HdB_B_146_SER, HdB_B_148_GLU, HdB_B_155_LYS, HdB_B_178_GLU, The number of HdB_B_179_THR, HdB_B_215_ASP, HdB_B_217_GLU, HdB_B_221_GLN, pSp_A_247_PHE, SB_A_151_ARG, SB_B_155_LYS, SB_B_220_ARG, SB_B_247_ARG, and SB_B_252_ARG. If no interaction occurs between the umami peptide and the umami receptor T1R1 / T1R3, the corresponding intermolecular interaction residue feature data is 0; where A is T1R1 protein, B is T1R3 protein, HdB is hydrogen bond interaction, HI is hydrophobic interaction, SB is salt bridge, and pSp is Π-stacking.

[0018] Furthermore, in step S5, the umami peptide screening sub-model established using a machine learning algorithm for the intermolecular interaction residue feature data obtained in step S3 is a random forest model.

[0019] Furthermore, the molecular descriptor feature data described in step S4 include BCUT2D_MWLOW, BCUT2D_LOGPHI, SMR_VSA1, MinEStateIndex, VSA_EState5, VSA_EState6, VSA_EState7, MolLogP, the number of times D appears in the peptide sequence, the number of times E appears in the peptide sequence, the sum of the number of times D and E appear in the peptide sequence, the position of the first appearance of D in the peptide sequence, and the position of the first appearance of E in the peptide sequence; if the peptide sequence includes both D and E, the corresponding molecular descriptor feature data is 1.

[0020] Furthermore, in step S5, the umami peptide screening sub-model established for the molecular descriptor feature data obtained in step S4 using a machine learning algorithm includes a molecular descriptor feature data logistic regression model and a molecular descriptor feature data gradient boosting tree model.

[0021] Furthermore, a umami peptide screening model includes the umami peptide screening method described above.

[0022] Furthermore, the screening model includes a web page display system and a back-end calculation and analysis system, and the back-end calculation and analysis system adopts the umami peptide screening method when performing data calculation and analysis.

[0023] The beneficial effects of the present invention are:

[0024] 1. The present invention discloses a method for screening umami peptides. The method comprises constructing molecular fingerprint feature data of umami peptides based on the structural fragments of existing umami peptides themselves. Intermolecular interaction residue feature data are constructed based on the interaction mode between umami peptides and umami receptors T1R1 / T1R3 analyzed by molecular docking technology. Molecular descriptor feature data of the physicochemical properties of umami peptides are obtained based on molecular descriptors. Then, a machine learning algorithm is used to establish umami peptide screening sub-models for each of the obtained different data. The umami peptide screening sub-models are integrated to establish an umami peptide screening model. The umami peptide screening model is used to screen and identify umami peptides. Umami peptides can be quickly and accurately identified, the screening accuracy of umami peptides is significantly improved, and the screening time is significantly shortened, thereby reducing the economic cost of the screening process.

[0025] 2. The present invention discloses a umami peptide screening model, comprising a webpage display system and a back-end computing and analysis system. The back-end computing and analysis system employs the umami peptide screening method when performing data calculation and analysis. Specifically, the back-end computing and analysis system employs a specific umami peptide screening method when performing data calculation and analysis. This model integrates all processes into a streamlined platform operation interface, enabling efficient screening with minimal or no reliance on manual experience. Based on the type of activity required, specialized screening can be performed by uploading corresponding data, and the model has strong portability. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematic diagram of the structure of the umami peptide screening model in the present invention.

[0027] Figure 2 This is a flow chart of the operating steps of the umami peptide screening model of the present invention.

[0028] Figure 3 This is an interface for preparing data for the umami peptide screening model of the present invention.

[0029] Figure 4 This is the interface for displaying the raw data output results of the umami peptide screening model in the present invention.

[0030] Figure 5 This is the data input interface for the umami peptide screening model of the present invention.

[0031] Figure 6 This is the data result display interface of the umami peptide screening model in the present invention.

[0032] Figure 7 This is a comparison chart of the simulation experiment integrated model results in the present invention and the existing model results. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0034] Example 1:

[0035] A method for screening umami peptides comprises the following steps:

[0036] S1: Organize existing umami peptide data and establish a database;

[0037] Specifically, the present invention collected 244 reported umami peptide data, sorted out the peptide taste, sequence, publication source and other information, and established a database. Some umami peptide data are shown in Table 1 below. In Table 1, the PepName column shows the peptide data in FASTA format.

[0038] Table 1 Data of some umami peptides

[0039] index PepName TasteinUmami_Peptides Taste in Bitter_Peptides 0 AE 1 1 0 1 RL 1 1 0 2 NY 1 1 0 3 DK 1 1 0 4 DD 1 1 0 5 DE 1 1 0 6 CE 1 1 0 7 EA 1 1 0 8 EN 1 1 0 9 ED 1 1 0 10 EE 1 1 0 11 EK 1 1 0 12 EP 1 1 0 13 ES 1 1 0 14 ET 1 1 0 15 GD 1 1 0 16 LD 0 0 0 17 QP 1 1 0 18 LEL 0 0 0 19 KG 1 1 0 20 PE 1 1 0 21 TE 1 1 0 22 HS 1 1 0 23 LM 0 0 0 24 VT 1 1 0 25 RFPHADF 0 0 0 26 AEA 1 1 0 27 AED 1 1 0 28 ADA 1 1 0 29 ADE 1 1 0 30 DAG 1 1 0

[0040] S2: Constructing molecular fingerprint feature data of umami peptides based on the existing structural fragments of umami peptides themselves;

[0041] Specifically, the molecular fingerprint feature data includes Morgan148, 322, 428, 509, 598, 650, 805, 952, 1150, 1409, 1573, 1687, 1706, 1907 and 2017.

[0042] S3: Based on the interaction mode between umami peptides and umami receptors T1R1 / T1R3 analyzed by molecular docking technology, the characteristic data of intermolecular interaction residues were constructed;

[0043] Specifically, the intermolecular interaction residue feature data include: HI_A_2_LEU, HI_A_3_LEU, HI_A_108_ASP, HI_A_154_THR, HI_A_157_ALA, HI_A_158_LEU, HI_A_161_PRO, HI_A_163_LEU, HI_A_179_LYS, HI_A_181_GLN, HI_A_182_TYR, HI_A_183_PRO, HI_A_218_ASP, HI_A_246_PRO, HI_A_419_TRP, HI_B_19_THR, HI_B_56_ARG, HI_B_57_P RO, HI_B_106_PRO, HI_B_107_VAL, HI_B_152_VAL, HI_B_155_LYS, HI_B_156_PHE, HI_B_179_THR, HI_B_245_LEU, HdB_A_48_SER, HdB_A_50_CYS, Hd B_A_52_GLN, HdB_A_107_SER, HdB_A_109_SER, HdB_A_148_SER, HdB_A_150_ASN, HdB_A_151_ARG, HdB_A_161_PRO, HdB_A_217_SER, HdB_A_218_ASP, HdB_A_219_ASP, HdB_A_222_GLN, HdB_A_247_PHE, HdB_A_248_SER, HdB_A_249_ALA, HdB_A_276_SER, HdB_A_278_GLN, HdB_B_15_LEU, HdB_B_17_PR O, HdB_B_56_ARG, HdB_B_57_PRO, HdB_B_58_SER, HdB_B_146_SER, HdB_B_148_GLU, HdB_B_155_LYS, HdB_B_178_GLU, HdB_B_179_THR, HdB_B_215_A SP, HdB_B_217_GLU, HdB_B_221_GLN, pSp_A_247_PHE, SB_A_151_ARG, SB_B_155_LYS, SB_B_220_ARG, SB_B_247_ARG, SB_B_252_ARG; if no interaction occurs between the umami peptide and the umami receptor T1R1 / T1R3, the corresponding intermolecular interaction residue feature data is 0 (if there is an interaction, the output result is the number of interaction types shown in the above feature data); where A is T1R1 protein, B is T1R3 protein, HdB is hydrogen bond interaction, HI is hydrophobic interaction, SB is salt bridge, and pSp is Π-stacking.

[0044] S4: Based on the molecular descriptors, obtain the molecular descriptor feature data of the physicochemical properties of umami peptides;

[0045] Specifically, the molecular descriptor feature data includes BCUT2D_MWLOW, BCUT2D_LOGPHI, SMR_VSA1, MinEStateIndex, VSA_EState5, VSA_EState6, VSA_EState7, MolLogP, the number of times D appears in the peptide sequence, the number of times E appears in the peptide sequence, the sum of the number of times D and E appear in the peptide sequence, the position of the first appearance of D in the peptide sequence, and the position of the first appearance of E in the peptide sequence; if the peptide sequence includes both D and E, the corresponding molecular descriptor feature data is 1.

[0046] S5: Using a machine learning algorithm to establish umami peptide screening sub-models for the data obtained in steps S2-S4;

[0047] Specifically, the umami peptide screening sub-model established for the molecular fingerprint feature data obtained in step S2 using a machine learning algorithm includes a stochastic gradient descent discriminant model, a molecular fingerprint feature data logistic regression model, a molecular fingerprint feature data gradient boosting tree model, and a Gaussian distribution naive Bayes discriminant model.

[0048] Among them, the parameters of the logistic regression model of molecular fingerprint feature data are as follows: the hyperparameters of the logistic regression model are as follows after grid search: the penalty function is l2 regularization (penalty = 'l2'), the tolerance of the stopping criterion is 0.0001 (tol = 0.0001), the intercept is added to the decision function (fit_intercept = True), the solver selects the Newton joint CG conjugate gradient method (solver = 'newton-cg'), the weight of the penalty function is assigned according to the number of categories in the sample (class_weight = 'balanced'), the maximum number of passes of the training data (also called epochs) is 10,000 (max_iter = 5,000), and the previous solution is not called as the initialization parameter during training (warm_start = False).

[0049] The parameters of the Gaussian distribution naive Bayesian discriminant model of molecular fingerprint feature data are as follows: the size of the prior part is not adjusted according to the positive and negative sample distributions of the data, and the maximum allowed variance is 0.000000001 (var_smoothing=0.000000001).

[0050] The parameters of the stochastic gradient descent discriminant model for molecular fingerprint feature data are as follows: the loss function is a smooth curve that includes tolerance for outliers and probability estimation (loss = 'modified_huber'), the regularization term uses elasticnet regularization (penalty = 'elasticnet'), the constant multiplied by the regularization term is 0.0001 (a = 0.0001), the intercept needs to be estimated in advance during fitting (fit_intercept: True), the maximum number of passes through the training data (also called epochs) is 10,000 (max_iter = 10,000), the stopping criterion for training is set to 0.001 (tol = 0.001), the training data is shuffled after each training batch (shuffle = True), early stopping is disabled (early_stopping = False), the previous solution is used as the initialization parameter during training (warm_start = True), and the penalty function weights are allocated according to the number of classes in the sample (class_weight = 'balanced').

[0051] The parameters of the gradient boosting tree model for molecular fingerprint feature data are as follows: an exponential form of loss function is used (loss = 'exponential'), a learning rate of 0.1 (learning_rate = 0.1), the number of boosting stages to be performed by the estimator is 70 (n_estimators = 70), the quality of the data segmentation of the system is measured using the 'friedman_mse' (default = 'friedman_mse'), the minimum number of samples required to split an internal node is 2 (min_samples_split = 2), the minimum number of samples required for a leaf node is 1 (min_samples_leaf = 1), the minimum weighted fraction of the sum of the weights (of all input samples) required for a leaf node is 0 (min_weight_fraction_leaf = 0), the maximum depth of a single regression estimator is 3 (max_depth = 3), and the number of features to consider when finding the best split is the square root of the number of features (max_features = 'sqrt').

[0052] A random forest model was used to screen umami peptides using machine learning algorithms for intermolecular interaction residue feature data. The specific parameters of the random forest model were as follows: the number of estimators in the forest was 126 (n_estimators = 126), the Gini function was selected as the function for measuring segmentation quality (criterion = "gini"), the maximum tree depth was 7 layers (max_depth = 7), the minimum number of samples required to split a node was 2 (min_samples_split = 2), the minimum number of samples required for a leaf node was 1 (min_samples_leaf = 1), the minimum weighted fraction of the sum of weights (of all input samples) required for a leaf node was 0 (min_weight_fraction_leaf = 0), the number of features considered when finding the optimal split was the square root of the number of features (max_features = 'sqrt'), bootstrap samples were used when building the tree (bootstrap = True), out-of-bag samples were not used to estimate the generalization score (oob_score = False), and the previous solution was not used as an initialization parameter during training (warm_start = False).

[0053] The umami peptide screening sub-models established using machine learning algorithms for molecular descriptor feature data include a molecular descriptor feature data logistic regression model and a molecular descriptor feature data gradient boosting tree model;

[0054] Among them, the parameters of the logistic regression model of molecular descriptor feature data are as follows: The hyperparameters of the logistic regression model are as follows after grid search: the penalty function is l2 regularization (penalty = 'l2'), the tolerance of the stopping criterion is 0.0001 (tol = 0.0001), the intercept is added to the decision function (fit_intercept = True), the solver selects the Newton joint CG conjugate gradient method (solver = 'newton-cg'), the weight of the penalty function is assigned according to the number of categories in the sample (class_weight = 'balanced'), the maximum number of passes through the training data (also called epochs) is 10,000 (max_iter = 5,000), and the previous solution is not called as an initialization parameter during training (warm_start = False).

[0055] The parameters of the gradient boosting tree model for molecular descriptor feature data are as follows: 'criterion':'friedman_mse', 'loss':'exponential', 'max_features':'sqrt', 'n_estimators':50. The loss function is constructed using an exponential form (loss='exponential'), the learning rate is 0.1 (learning_rate=0.1), the number of boosting epochs to be performed by the estimator is 70 (n_estimators=70), the quality of the data segmentation is measured using the 'friedman_mse' (default='friedman_mse'), the minimum number of samples required to split an internal node is 2 (min_samples_split=2), the minimum number of samples required to be a leaf node is 1 (min_samples_leaf=1), the minimum weighted fraction of the sum of the weights (of all input samples) required for a leaf node is 0 (min_weight_fraction_leaf=0), the maximum depth of a single regression estimator is 3 (max_depth=3), and the number of features to consider when finding the best split is the square root of the number of features (max_features='sqrt'). All of the above submodels output a 1 for a positive sample and a 0 otherwise.

[0056] S6: Use the support vector machine algorithm to integrate the umami peptide screening sub-models to establish an umami peptide screening model;

[0057] Specifically, the outputs of all sub-models in step S5 are used as input, and a support vector machine algorithm (kernel = rbf, gamma = scale) is used to coordinate the outputs of all sub-models to establish an umami peptide screening model. The seven sub-models established in step S4 will output a result (positive result is 1, negative result is 0). These results form a 7-bit matrix that is input into the umami peptide screening model (SVM model) to obtain the final result, where umami is 1 and bitter is 0.

[0058] S7: Screening umami peptides using the umami peptide screening model established in step S6.

[0059] Specifically, the FASTA format sequence of the peptide to be tested is input into the model, for example, for the peptide Gly-Glu, input GE. The model will calculate and give the category judgment result (the value is 1 for umami, otherwise 0) and the confidence probability (size is 0-100%).

[0060] Example 2:

[0061] In the second embodiment, a umami peptide screening model is provided, which includes the umami peptide screening method described in the first embodiment.

[0062] Specifically, as attached Figure 1 As shown, the screening model includes a web page display system and a back-end calculation and analysis system, and the back-end calculation and analysis system adopts the umami peptide screening method when performing data calculation and analysis.

[0063] The web page display system consists of three modules: user input data integration, input TPDM information transmission, and TPDM model judgment result display. These three modules are released with Django 3.2 and Streamlit.

[0064] The back-end computing and analysis system mainly consists of four parts:

[0065] ① Receive the input of the peptide to be tested, integrate the information and output the peptide sequence output module that can be directly input into the TPDM model for processing.

[0066] ② Molecular characterization module, which is specifically divided into three parts, corresponding to the molecular fingerprint feature data, intermolecular interaction residue feature data and molecular descriptor feature data and their corresponding sub-models in Example 1;

[0067] 1) Molecular fingerprints were generated by calling the RDKit.AllChem.GetMorganFingerprintAsBitVect module to generate a 2048-bit extended concatenated fingerprint. Based on parameter differences, ECFP4 was selected as the fingerprint format, with useFeatures = True and radius = 2. Morgans 148, 322, 428, 509, 598, 650, 805, 952, 1150, 1409, 1573, 1687, 1706, 1907, and 2017 were selected as feature data. If the potential umami peptide small molecule contained the above structure, the byte in the matrix was set to 1; otherwise, it was set to 0. Thus, each small molecule was represented by a 15-byte matrix based on its substructure.

[0068] 2) Molecular interaction residues were constructed based on the six interaction force patterns and the number of umami peptide residue sites. Molecular docking was first performed. Potential umami peptides were used as docking ligands. Their structures were generated by calling Chem.MolFromFASTA to read the peptide names (short letters, in FASTA format) to create a peptide list. The Chem.AllChe module was then called using the EmbedMolecule function. The function uses the Experimental-Torsion Basic Knowledge Distance Geometry (ETKDG) algorithm to generate 3D conformations based on the modified distance geometry algorithm (Friedrich et al., 2017), and the sum of the atomic overlap distances between conformations is set to 1. Finally, the MMFFOptimizeMolecule module is used to call the MMFF94 force field to optimize the small molecule structure and energy (Halgren & Nachbar, 1996). The protein receptor in the umami judgment model is T1R1 / T1R3-VFT, and its structure is generated by trRosetta based on the protein sequence (Yang, Anishchenko, Park, Peng, Ovchinnikov, & Baker, 2020). The docking center position was selected as follows: center-x = 87.77, center-y = 45.93, center-z = 96.48, and the docking box size was as follows: size-x = 120, size-y = 120, size-z = 120. Twenty conformations were generated for each docking. The exhaustiveness of the docking was set to 80, and 20 conformations were generated for each docking. Smina was used as the docking software (Masters, Eagon, & Heying, 2020). The protein-ligand interaction profiler (plip) was used for analysis, and the interaction forces between proteins and peptides were divided into six categories: hydrophobic interactions (HI), hydrogen bonds (HdB), pi-stacking (pSp), electron-rich pi-neighboring cation interactions (pCp), halogen bonds (HB) and salt bridges (SB), which were then cross-generated with the amino acid sequence to generate permutation groups.After model feature screening, the features of the TPDM-Umami model in molecular docking parameters were finally determined as follows: HI_A_2_LEU, HI_A_3_LEU, HI_A_108_ASP, HI_A_154_THR, HI_A_157_ALA, HI_A_158_LEU, HI_A_161_PRO, HI_A_163_LEU, HI_A_179_LYS, HI_A_181_GLN, HI_A_182_TYR, HI_A_183_PRO, HI_A_218_ASP, HI_A_246_PRO , HI_A_419_TRP, HI_B_19_THR, HI_B_56_ARG, HI_B_57_PRO, HI_B_106_PRO, HI_B_107_VAL, HI_B_152_VAL, HI_B_155_LYS, HI_B_1 56_PHE, HI_B_179_THR, HI_B_245_LEU, HdB_A_48_SER, HdB_A_50_CYS, HdB_A_52_GLN, HdB_A_107_SER, HdB_A_109_SER, HdB_A_148 _SER, HdB_A_150_ASN, HdB_A_151_ARG, HdB_A_161_PRO, HdB_A_217_SER, HdB_A_218_ASP, HdB_A_219_ASP, HdB_A_222_GLN, HdB_A _247_PHE, HdB_A_248_SER, HdB_A_249_ALA, HdB_A_276_SER, HdB_A_278_GLN, HdB_B_15_LEU, HdB_B_17_PRO, HdB_B_56_ARG, HdB_B _57_PRO, HdB_B_58_SER, HdB_B_146_SER, HdB_B_148_GLU, HdB_B_155_LYS, HdB_B_178_GLU, HdB_B_179_THR, HdB_B_215_ASP, HdB _B_217_GLU, HdB_B_221_GLN, pSp_A_247_PHE, SB_A_151_ARG, SB_B_155_LYS, SB_B_220_ARG, SB_B_247_ARG, SB_B_252_ARG, docking score.The molecular docking parameters of the TPDM-Bitter model are as follows: HI_A_79_GLU, HI_A_82_PHE, HI_A_85_LEU, HI_A_89_TRP, HI_A_152_ILE, HI_A_156_ILE, HI_A_159_TYR, HI_A_172_PHE, HI_A_175_PHE, HI_A_266_GLN, HdB_A_69_SER, HdB_A_79 _GLU, HdB_A_85_LEU, HdB_A_86_THR, HdB_A_159_TYR, HdB_A_176_SER, HdB_A_180_VAL, HdB_A_254_SER, HdB_A_266_GLN, pSp_A_76_PHE, pSp_A_89_TRP, pSp_A_159_TYR, pSp_A_172_PHE, pSp_A_247_PHE, docking score.

[0069] The sequence of the protein is as follows:

[0070] VFT region sequence of T1R1

[0071] >sp|Q7RTX1|TS1R1_HUMAN Taste receptor type 1member 1OS=Homo sapiensOX=9606GN=TAS1R1 PE=2SV=1MLLCTARLVGLQLLISCCWAFACHSTESSPDFTLPGDYLLAGLFP

[0072] LHSGCLQVRHRPEVT

[0073] LCDRSCSFNEHGYHLFQAMRLGVEEINNSTALLPNITLGYQLYD

[0074] VCSDSANVYATLRVLS

[0075] LPGQHHIELQGDLLHYSPTVLAVIGPDSTNRAATTAALLSPFLVP

[0076] MISYAASSETLSVKR

[0077] QYPSFLRTIPNDKYQVETMVLLLQKFGWTWISLVGSSDDYGQL

[0078] GVQALENQATGQGICIA

[0079] FKDIMPFSAQVGDERMQCLMRHLAQAGATVVVVFSSRQLARVF

[0080] FESVVLTNLTGKVWVAS

[0081] EAWALSRHITGVPGIQRIGMVLGVAIQKRAVPGLKAFEEAYARA

[0082] DKKAPRPCHKGSWCSS

[0083] NQLCRECQAFMAHTMPKLKAFSMSSAYNAYRAVYAVAHGLHQ

[0084] LLGCASGACSRGRVYPWQ

[0085] LLEQIHKVHFLLHKDTVAFNDNRDPLSSYNIIAWDWNGPKWTFT

[0086] VLGSSTWSPVQLNINE

[0087] TKIQWHGKDNQVPKSVCSSDCLEGHQRVVTGFHHCCFECVPCG

[0088] AGTFLNKSDLYRCQPCG

[0089] KEEWAPEGSQTCFPRTVVFLALREHTSWVLLAANTLLLLLLLGT

[0090] AGLFAWHLDTPVVRSA

[0091] GGRLCFLMLGSLAAGSGSLYGFFGEPTRPACLLRQALFALGFTIF

[0092] LSCLTVRSFQLIIIF

[0093] KFSTKVPTFYHAWVQNHGAGLFVMISSAAQLLICLTWLVVWTP

[0094] LPAREYQRFPHLVMLEC

[0095] TETNSLGFILAFLYNGLLSISAFACSYLGKDLPENYNEAKCVTFS

[0096] LLFNFVSWIAFFTTA

[0097] VFT region sequence of TT1R3 of SVYDGKYLPAANMMAGLSSLSSGFGGYFLPKCYVILCRPDLNSTEHFQASIQDYTRRCGS

[0098] >sp|Q7RTX0|TS1R3_HUMAN Taste receptor type 1 member 3 OS=Homo sapiens OX=9606 GN=TAS1R3 PE=1 SV=2 MLGPAVLGLSLWALLHPGTGAPLCLSQQLRMKGDYVLGGLFPL

[0099] GEAEEAGLRSRTRPSSP

[0100] VCTRFSSNGLLWALAMKMAVEEINNKSDLLPGLRLGYDLFDTC

[0101] SEPVVAMKPSLMFLAKA

[0102] GSRDIAAYCNYTQYQPRVLAVIGPHSSELAMVTGKFFSFFLMPQ

[0103] VSYGASMELLSARETF

[0104] PSFFRTVPSDRVQLTAAAELLQEFGWNWVAALGSDDEYGRQGL [[ID=​​​​​​​​​​​​​​ERLKIRWHTSDNQKPVSRCSRQCQEGQVRRVKGFHSCCYDCVDCEAGSYRQNPDDIACTF

[0111] CGQDEWSPERSTRCFRRRSRFLAWGEPAVLLLLLLLSLALGLVLAALGLFVHHRDSPLVQ

[0112] ASGGPLACFGLVCLGLVCLSVLLFPGQPSPARCLAQQPLSHLPLTGCLSTLFLQAAEIFV

[0113] ESELPLSWADRLSGCLRGPWAWLVVLLAMLVEVALCTWYLVAFPPEVVTDWHMLPTEALV

[0114] HCRTRSWVSFGLAHATNATLAFLCFLGTFLVRSQPGCYNRARGLTFAMLAYFITWVSFVP

[0115] LLANVQVVLRPAVQMGALLLCVLGILAAFHLPRCYLLMRQPGLNTPEFFLGGGPGDAQGQNDGNTGNQGKHE

[0116] 3) Molecular descriptors: Eight metrics generated by the RDkit and Molecular.DescriptorCalculator modules include BCUT2D_MWLOW, BCUT2D_LOGPHI, SMR_VSA1, MinEStateIndex, VSA_EState5, VSA_EState6, VSA_EState7, and MolLogP, abbreviated as BM, PV14, SV, ME, VS5, VS6, VS7, and ML, respectively. The D / E feature descriptors primarily focus on factors such as the distribution and quantity of amino acids. The definitions of these descriptors are shown in Table 2.

[0117] Table 2 Eigenvalue definitions and calculation modules

[0118]

[0119] The references in Table 2 are as follows:

[0120] 1.Beno, BR; Mason, JS, The design of combinatorial libraries using properties and 3D pharmacophore fingerprints. Drug Discovery Today 2001, 6(5), 251-258.

[0121] 2.Hall,LH;Mohney,B.;Kier,LB,The Electrotopological State:An AtomIndex for QSAR.Quantitative Structure-Activity Relationships 1991,10(1),43-51.

[0122] 3.Labute,P.,A widely applicable set of descriptors.J Mol GraphModel2000,18(4-5),464-77.

[0123] 4. Wildman, SA; Crippen, GM, Prediction of Physicochemical Parameters by Atomic Contributions. Journal of Chemical Information and Computer Sciences1999, 39(5), 868-873.

[0124] ③Analysis model judgment submodule: This module is responsible for sending the statistical data obtained by the molecular characterization module into the corresponding model for modeling. The relationship between the data matrix and the corresponding submodel construction is shown in the attached figure. Figure 1 As indicated by the green arrow.

[0125] ④SVM ensemble learning model judgment module: All sub-model output results (judgment category and prediction confidence possibility) are used as input data and substituted into the SVM model for secondary modeling, and finally a unified model output result is obtained.

[0126] Furthermore, the research steps of this model are as follows: Figure 2As shown. First, the dataset was randomly divided into training and test sets. The data were then converted into corresponding digital forms using various molecular characterization schemes. Triple cross-validation of the feature training data was used to find the optimal hyperparameters for each classifier. Multiple classification algorithms including gradient boosting (GTB), LR, RF, GNB, stochastic gradient descent (SGD) were applied, combined with four molecular characterization schemes. Ultimately, seven optimal sub-classifiers were selected. The hyperparameters of each classifier are unique and independent to ensure their specificity and predictive ability. The model proposed by TPDM is constructed using an ensemble learning method based on the SVM algorithm.

[0127] Furthermore, the parameters of all sub-models in the umami peptide screening model of the present invention are as follows:

[0128] 1) Molecular fingerprint feature data

[0129] TPDM uses a logistic regression algorithm for molecular fingerprinting. The logistic regression model was grid-searched with the following hyperparameters: the penalty function was l2 regularization (penalty = 'l2'), the stopping criterion tolerance was 0.0001 (tol = 0.0001), an intercept was added to the decision function (fit_intercept = True), the solver was Newton-CG conjugate gradient method (solver = 'newton-cg'), the penalty function weights were assigned based on the number of classes in the sample (class_weight = 'balanced'), the maximum number of passes through the training data (also known as epochs) was 10,000 (max_iter = 5000), and the previous solution was not used as an initialization parameter during training (warm_start = False). A gradient boosting tree model was grid-searched with the following hyperparameters: 'criterion': 'friedman_mse', 'loss': 'exponential', 'max_features': 'sqrt', 'n_estimators': 50. The exponential form of the loss function is used (loss='exponential'), the learning rate is 0.1 (learning_rate=0.1), the number of boosting stages to be performed by the estimator is 70 (n_estimators=70), the 'friedman_mse' is used to measure the quality of the data segmentation of the system (default='friedman_mse'), the minimum number of samples required to split an internal node is 2 (min_samples_split=2), the minimum number of samples required to be a leaf node is 1 (min_samples_leaf=1), the minimum weighted fraction of the sum of the weights (of all input samples) required for a leaf node is 0 (min_weight_fraction_leaf=0), the maximum depth of a single regression estimator is 3 (max_depth=3), and the number of features to consider when finding the best split is the square root of the number of features (max_features='sqrt').

[0130] 2) Molecular descriptor feature data

[0131] TPDM uses two sub-models, logistic regression and gradient boosting tree discriminant, in terms of descriptors. The hyperparameters of the logistic regression model are grid searched as follows: the penalty function is l2 regularization (penalty = 'l2'), the tolerance of the stopping criterion is 0.0001 (tol = 0.0001), the intercept is added to the decision function (fit_intercept = True), the solver selects the Newton-CG conjugate gradient method (solver = 'newton-cg'), the weight of the penalty function is assigned according to the number of categories in the sample (class_weight = 'balanced'), the maximum number of passes through the training data (also called epochs) is 10,000 (max_iter = 5,000), and the previous solution is not called as an initialization parameter during training (warm_start = False).

[0132] The gradient boosting tree model was grid searched with the following hyperparameters: 'criterion':'friedman_mse', 'loss':'exponential', 'max_features':'sqrt', 'n_estimators':50. The exponential form of the loss function is used (loss='exponential'), the learning rate is 0.1 (learning_rate=0.1), the number of boosting stages to be performed by the estimator is 70 (n_estimators=70), the 'friedman_mse' is used to measure the quality of the data segmentation of the system (default='friedman_mse'), the minimum number of samples required to split an internal node is 2 (min_samples_split=2), the minimum number of samples required to be a leaf node is 1 (min_samples_leaf=1), the minimum weighted fraction of the sum of the weights (of all input samples) required for a leaf node is 0 (min_weight_fraction_leaf=0), the maximum depth of a single regression estimator is 3 (max_depth=3), and the number of features to consider when finding the best split is the square root of the number of features (max_features='sqrt').

[0133] 3) Intermolecular interaction residue feature data

[0134] The number of estimators in the TPDM forest is 126 (n_estimators=126), the Gini function is used as the function to measure the quality of the segmentation (criterion="gini"), the maximum depth of the tree is 7 layers (max_depth=7), the minimum number of samples required to split a node is 2 (min_samples_split=2), the minimum number of samples required for a leaf node is 1 (min_samples_leaf=1), and the minimum weighted fraction of the sum of weights (of all input samples) required for a leaf node is 0 (min_weight_fraction_leaf=0). The number of features to consider when finding the best split is the square root of the number of features (max_features='sqrt'), bootstrap samples are used when building the tree (bootstrap=True), out-of-bag samples are not used to estimate the generalization score (oob_score=False), and the previous solution is not used as an initialization parameter during training (warm_start=False).

[0135] The code used in the molecular docking part of the present invention is as follows:

[0136] #Start docking

[0137] pbb_name='statics / TasteppetidesDM / t1r1t1r3_vfd.pdb'config_txt='statics / TasteppetidesDM / T1R1.txt'

[0138] out_sdf_name=workfile_location+" / "+str(sub_fasta)+"_docked.sdf" out_log_name=workfile_location+" / "+str(sub_fasta)+".log"

[0139] os.system("smina--seed 0--cpu 1--config"+config_txt+"-l"+ligand_input_name+"-r"+pbb_name+"-o"+out_sdf_name+"--log"+out_log_name)

[0140] Furthermore, the method for using the umami peptide screening model of the present invention comprises the following steps:

[0141] (1) Data preparation

[0142] The data preparation page of the umami peptide screening model in the present invention is as shown in the attached Figure 3 As shown in the figure, refer to the example in the "Example csv" file in the "1.YYDS_input_make" section, upload the corresponding csv file in the Upload File interface. If the column name is not "PepName", you need to specify it in "Specify fasta columns". After submitting the csv file in "2.Name INPUT", click "submit" to submit the file, and you will get the following result. Figure 4 The raw data output results are shown.

[0143] (2) Data input

[0144] The data input interface of the umami peptide screening model of the present invention is shown in the attached figure. Figure 5 As shown in the figure, paste the data prepared in (1) into the "Input Here" on the left, and then click "Submit" to submit.

[0145] (3) Data result display interface

[0146] The progress bar on the left allows you to monitor the run progress. When both progress bars reach 100%, the run is complete. The calculation results are displayed on the right, including "PepName," "Descriptor," "Taste_Umami," "Pro_Umami," "Taste_bitter," and "Pro_Bitter," which represent the peptide sequence, peptide description, TPDM-Umami predicted flavor, TPDM-Umami predicted confidence probability, TPDM-Bitter predicted flavor, and TPDM-Bitter predicted confidence probability, respectively. Click "Press to download" to download the entire CSV file.

[0147] Simulation experiment:

[0148] (1) Accuracy of model judgments of reported umami peptides

[0149] The umami peptide screening model (TPDM model) in the present invention is a bitterness / umami prediction and judgment model obtained by integrating the DA, MD, and FP sub-models through the support vector machine algorithm. Tables 3 and 4 below summarize the performance of the TPDM-Umami and TPDM-Bitter total models and sub-models. It can be seen from Tables 3 and 4 that the model training set accuracy in this application is 97.95%, and the AUC value is 0.98; the test set accuracy is 92.86%, and the AUC is 0.93. Among them, the bitterness judgment model is TPDM-Bitter, and the training set accuracy of this model is 96.93%, and the AUC value is 0.99; the test set accuracy is 91.84%, and the AUC is 0.92. Overall, TPDM has achieved excellent results in indicator methods such as ACC, AUC, REC, PRE, and MCC. The confusion matrix results of TPDM show that the model has comparable judgment capabilities for umami and bitterness. In addition, the ACC and AUC values ​​of the training set and the test set are similar, both greater than 0.9. By combining many diverse and excellent models, ensemble models can explore patterns in high-dimensional data spaces, thereby improving performance. The performance of TPDM-Umami and TPDM-Bitter in various metrics is nearly equivalent to the best sub-model.

[0150] Table 3 Comparison of accuracy and AUC performance among umami judgment models

[0151]

[0152] Table 4 Comparison of accuracy and AUC performance among bitterness judgment models

[0153]

[0154] In order to better compare the reliability of the flavor prediction models, several representative flavor prediction models were selected for comparison. The results are shown in the attached figure. Figure 7 As shown in the attached Figure 7It can be seen that in terms of umami prediction, TPDM has an advantage in the comprehensive comparison with Umami_YYDS and UMPred-FRL (Charoenkwan, P.; Nantasenamat, C.; Hasan, MM; Moni, MA; Manavalan, B.; Shoombuatong, W., UMPred-FRL: A New Approach for Accurate Prediction of UmamiPeptides Using Feature Representation Learning. International Journal of Molecular Sciences 2021, 22(23), 13124.). Umami_YYDS is a GTB model based on two-dimensional molecular descriptors of physicochemical properties, while the data of TPDM includes MD data, DA and FP data. Therefore, the performance of Umami_YYDS is not as good as TPDM. The REC value of Umami_YYDS is similar to that of TPDM, but its higher ACC and MCC also show the advantages of the integrated model over the individual sub-models. Although UMPred-FRL1 is also an integrated model, its performance is still lower than TPDM. It is speculated that this is because the model has too many sub-models and the sub-model performance is average (the accuracy of 42 sub-models is less than 0.8), and the quality of its training data (140:304) is not as good as that of Umami_peptidesDB (244:257). In bitterness prediction, TPDM is still better than Q-Value and Umami_YYDS. Q-value judgment is a method based on amino acid scoring matrix and is widely used in bitterness prediction. However, the Q-value method is not as good as TPDM and Umami_YYDS, as shown in the attached figure. Figure 7 As shown in (d). TPDM collects data from multiple sources and comprehensively describes protein-ligand data to achieve high performance, which is consistent with the multi-layer perceptual descriptor method summarized by Xiong et al. (Xiong, G.; Shen, C.; Yang, Z.; Jiang, D.; Liu, S.; Lu, A.; Chen, X.; Hou, T.; Cao, D., Featurization strategies for protein–ligand interactions and their applications in scoring function development. WIREs Computational Molecular Science 2022, 12(2), e1567.).

[0155] (2) Accurate judgment of umami / bitter flavor peptides

[0156] There are 31 umami / bitter dual-flavor peptides in the TPDB (sequences: DR, AD, AH, DA, DL, EG, EL, EY, EV, GE, LE, LV, VD, VE, VG, VV, EGF, GEG, GGY, SEEK, RPLGNC, DEESLA, TYLPVH, PAATIPE, AGLQFPVGR, EQLEATVQKLDESR, GENEEEDSGAIVTVF, HV, KTGLSPDQF, KTDLNFENL, RLGSSEVEQVQ). The ability of the model to simultaneously identify the flavor characteristics of these dual-flavor peptides is a challenge. Table 5 below summarizes the accuracy of different models for these 31 peptides. TPDM was the most accurate model for both umami and bitterness predictions, with an accuracy of 0.90 for the umami model TPDM-Umami and 0.94 for the bitter model TPDM-Bitter. The accuracy rates for Umami_YYDS, UMPred-FRL, and Q were only 0.77, 0.77, and 0.74, respectively. Further evaluation of the dual-flavor peptide sources revealed that TPDM-Umami achieved an accuracy of 0.92 on the training set and 0.83 on the test set, a relatively small difference and demonstrating that the model was not overfitting. The other models, however, had lower judgment capabilities and their accuracy still did not exceed 80%.

[0157] Table 5 Summary of the accuracy of umami / bitter peptides in different detection models

[0158]

[0159] Table 5

[0160]

[0161] (3) Select unreported umami peptides and verify the accuracy of the model through artificial sensory testing

[0162] The present invention randomly synthesized 6 peptides, performed taste prediction and verification, and compared them with other machine learning modeling results. The results are shown in Table 6. As can be seen from Table 6, TPDM-Umami and TPDM-Bitter have the best model judgment accuracy.

[0163] Table 6 Synthetic peptide taste prediction comparison table

[0164]

[0165] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for screening umami peptides, characterized in that: The following steps are included: S1: Organize existing umami peptide data and establish a database; S2: Constructing molecular fingerprint feature data of umami peptides based on the existing structural fragments of umami peptides themselves; S3: Based on the interaction mode between umami peptides and umami receptors T1R1 / T1R3 analyzed by molecular docking technology, the characteristic data of intermolecular interaction residues were constructed; S4: Based on the molecular descriptors, obtain the molecular descriptor feature data of the physicochemical properties of umami peptides; S5: Using a machine learning algorithm to establish umami peptide screening sub-models for the data obtained in steps S2-S4; S6: Use the support vector machine algorithm to integrate the umami peptide screening sub-models to establish an umami peptide screening model; S7: Screening umami peptides using the umami peptide screening model established in step S6.

2. The method for screening umami peptides according to claim 1, wherein: The molecular fingerprint feature data described in step S2 include Morgan148, 322, 428, 509, 598, 650, 805, 952, 1150, 1409, 1573, 1687, 1706, 1907 and 2017.

3. The method for screening umami peptides according to claim 2, wherein: In step S5, the umami peptide screening sub-model established for the molecular fingerprint feature data obtained in step S2 using a machine learning algorithm includes a stochastic gradient descent discriminant model, a molecular fingerprint feature data logistic regression model, a molecular fingerprint feature data gradient boosting tree model and a Gaussian distribution naive Bayes discriminant model.

4. The method for screening umami peptides according to claim 1, wherein: The intermolecular interaction residue feature data described in step S3 include: HI_A_2_LEU, HI_A_3_LEU, HI_A_108_ASP, HI_A_154_THR, HI_A_157_ALA, HI_A_158_LEU, HI_A_161_PRO, HI_A_163_LEU, HI_A_179_LYS, HI_A_181_GLN, HI_A_182_TYR, HI_A_183_PRO, HI_A_218_ASP, HI_A_246_PRO, HI_A_419_TRP, HI_B_19_THR, HI_B_56_ARG, HI_B_57_PRO, HI_B_106_PRO, HI_B_107_VAL, HI_B_152_VAL, HI_B_155_LYS, HI_B_156_PHE, HI_B_179_THR, HI_B_245_LEU, HdB_A_48_SER, Hd B_A_50_CYS, HdB_A_52_GLN, HdB_A_107_SER, HdB_A_109_SER, HdB_A_148_SER, HdB_A_150_ASN, HdB_A_151_ARG, HdB_A_161_PRO, HdB_A_217_S ER, HdB_A_218_ASP, HdB_A_219_ASP, HdB_A_222_GLN, HdB_A_247_PHE, HdB_A_248_SER, HdB_A_249_ALA, HdB_A_276_SER, HdB_A_278_GLN, HdB _B_15_LEU, HdB_B_17_PRO, HdB_B_56_ARG, HdB_B_57_PRO, HdB_B_58_SER, HdB_B_146_SER, HdB_B_148_GLU, HdB_B_155_LYS, HdB_B_178_GLU, H The number of times dB_B_179_THR, HdB_B_215_ASP, HdB_B_217_GLU, HdB_B_221_GLN, pSp_A_247_PHE, SB_A_151_ARG, SB_B_155_LYS, SB_B_220_ARG, SB_B_247_ARG, and SB_B_252_ARG occurred. If no interaction between the umami peptide and the umami receptor T1R1 / T1R3 occurs, the corresponding intermolecular interaction residue feature data is 0; where A is T1R1 protein, B is T1R3 protein, HdB is hydrogen bond interaction, HI is hydrophobic interaction, SB is salt bridge, and pSp is Π-stacking.

5. The method for screening umami peptides according to claim 4, wherein: In step S5, the umami peptide screening sub-model established using a machine learning algorithm for the intermolecular interaction residue feature data obtained in step S3 is a random forest model.

6. The method for screening umami peptides according to claim 1, wherein: The molecular descriptor feature data described in step S4 include BCUT2D_MWLOW, BCUT2D_LOGPHI, SMR_VSA1, MinEStateIndex, VSA_EState5, VSA_EState6, VSA_EState7, MolLogP, the number of times D appears in the peptide sequence, the number of times E appears in the peptide sequence, the sum of the number of times D and E appear in the peptide sequence, the position of the first appearance of D in the peptide sequence, and the position of the first appearance of E in the peptide sequence; if the peptide sequence includes both D and E, the corresponding molecular descriptor feature data is 1.

7. The method for screening umami peptides according to claim 1, wherein: In step S5, the umami peptide screening sub-model established for the molecular descriptor feature data obtained in step S4 using a machine learning algorithm includes a molecular descriptor feature data logistic regression model and a molecular descriptor feature data gradient boosting tree model.

8. A umami peptide screening model, comprising the umami peptide screening method according to any one of claims 1 to 7.

9. A umami peptide screening model, characterized in that: The screening model includes a web page display system and a back-end calculation and analysis system, and the back-end calculation and analysis system adopts the umami peptide screening method when performing data calculation and analysis.