Umami peptide prediction method and system based on deep learning
By constructing the Umami-Transformer deep learning framework and combining physical and chemical characteristics with molecular docking technology, the problems of low efficiency and insufficient prediction accuracy in existing umami peptide screening were solved, and efficient and accurate umami peptide prediction and verification were achieved.
Patent Information
- Application Number
- CN202510817727.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing umami peptide screening technologies are costly, time-consuming, and have low throughput. In addition, existing pre-trained models lack generalization capabilities in umami peptide prediction and are difficult to support actual industrial applications.
A Umami-Transformer deep learning framework was constructed, integrating the Transformer architecture with 8 key physicochemical features. An exhaustive method was used to screen dipeptides to pentapeptides, and molecular docking technology was used to analyze the interaction mechanism between candidate peptides and the umami receptors T1R1/T1R3.
Significantly improve the prediction accuracy of umami peptides, realize the full process research from theoretical prediction to biological experimental verification, and provide an innovative technical path for the efficient mining and functional application of umami peptides.
Smart Images

Figure CN120766752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for predicting umami peptides based on deep learning, and belongs to the technical field of intersection of food science and artificial intelligence. Background Art
[0002] Umami peptides are a class of bioactive peptides that can significantly enhance the umami taste of food, usually derived from protein hydrolysates. They can not only enhance the umami perception of food, but also effectively reduce bitterness and improve the overall flavor palatability. For example, known umami peptides such as NKF and LGLR can enhance the salty perception of low-salt solutions by interacting with taste receptors, providing a potential solution for the development of salt-reduced foods. At present, umami peptides are widely present in natural organisms such as edible fungi and oysters. Their research is not only limited to food flavor optimization, but also involves reducing sodium intake, developing healthy functional foods and other fields, and is of great value to food science and public health.
[0003] Traditional umami peptide screening technology mainly relies on sensory-guided separation combined with mass spectrometry analysis. Although this method can effectively identify active peptides, it has problems such as high cost, long time consumption, and low throughput. In addition, the complex separation steps during the experiment can easily lead to the loss of active peptides. In recent years, with the development of artificial intelligence technology, machine learning models have gradually been applied to peptide flavor classification. For example, models such as multi-layer perceptron (MLP) and long short-term memory network (LSTM) have been used for umami peptide prediction, but RNN-type models are difficult to effectively process long sequence data due to the gradient disappearance problem. In addition, although pre-trained language models (such as BERT) perform well in natural language processing, due to the limited size of the umami peptide database, it is difficult to directly migrate to the field of peptide sequence prediction, and the sequence correlation between different peptides is significantly different from natural language, which further limits the application effect of pre-trained models.
[0004] The Transformer neural network has shown significant advantages in the emerging field of bioinformatics due to its unique multi-head attention mechanism and parallel computing architecture. Its multi-head attention mechanism can simultaneously capture the local and long-range dependencies of amino acids in peptide sequences and accurately analyze complex sequence features; the parallel computing design greatly improves training efficiency, and there is no risk of gradient vanishing or explosion, which is particularly suitable for processing variable-length peptide sequence prediction tasks. However, existing studies mostly focus on single sequence information, and rarely integrate the physicochemical characteristics of peptide segments (such as lipophilicity, charge distribution, etc.) to optimize model performance. In addition, due to the scarcity of umami peptide data and research costs, existing prediction models generally have problems such as insufficient generalization ability and limited prediction accuracy, making it difficult to support actual industrial applications. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the present invention provides a deep learning-based method and system for predicting umami peptides. This system proposes the Umami-Transformer deep learning framework, which, by integrating the Transformer architecture with eight key physicochemical characteristics (including molecular lipophilicity, charge distribution, and polarity parameters), can efficiently process peptide sequence information and significantly improve the accuracy of umami peptide prediction. Furthermore, an exhaustive approach is used to screen all dipeptides to pentapeptides, and molecular docking technology is used to analyze the interaction mechanism between candidate peptides and the umami receptors T1R1 / T1R3. Ultimately, this system achieves a full-scale research process from theoretical prediction to biological experimental verification, providing an innovative technical path for the efficient discovery and functional application of umami peptides.
[0006] Specifically, the following technical solutions are included:
[0007] In a first aspect, the present invention provides a method for predicting umami peptides based on deep learning, comprising the following steps:
[0008] S1. Dataset Construction: Collect experimentally verified umami peptide and non-umami peptide sequences to construct the UMP499 dataset containing umami peptides and non-umami peptides;
[0009] S2, feature selection and processing: standardize the peptide sequence truncation and padding, and calculate eight physicochemical features;
[0010] S3. Constructing an Umami-Transformer model: Constructing an Umami-Transformer model comprising a feature processing and generation module, a sequence information processing module, a feature information processing module, and a result output module; wherein the sequence information processing module uses a Transformer encoder; the feature information processing module uses a multi-layer fully connected network; and the result output module uses an MLP network with a sigmoid activation function to output umami peptide probabilities;
[0011] S4. Model training: The Umami-Transformer model is trained using a 5-fold cross-validation method and the BCELoss function as the loss function until the classification accuracy of the Umami-Transformer model on the test set exceeds 95%;
[0012] S5. Umami peptide prediction: The peptide sequence to be tested is input into the trained Umami-Transformer model to obtain the probability prediction result of it being an umami peptide.
[0013] In one embodiment of the present invention, constructing the data set in step S1 specifically includes:
[0014] S11. Data collection: We collected experimentally verified umami and non-umami peptide sequences from public literature and the BIOPEP-UWM database to construct the UMP499 dataset, which contains 212 umami peptides and 287 non-umami peptides.
[0015] S12. Data preprocessing, including:
[0016] Remove outliers: remove peptide sequences containing non-standard amino acid symbols to ensure data accuracy and consistency;
[0017] Remove redundancy: Use sequence alignment tools to remove duplicate sequences, eliminate redundancy, and avoid the impact of duplicate data on model training;
[0018] Stratification: The dataset was stratified into a UMP-CV training set and a UMP-TS test set in an 80:20 ratio to ensure the distribution consistency of the training set and the test set.
[0019] In one embodiment of the present invention, the feature selection and processing in step S2 specifically includes:
[0020] S21. Feature calculation: 8 physicochemical features of each peptide sequence were calculated using the RDKit cheminformatics toolkit, including: molecular weight correlation feature: BCUT2D_MWLOW; charge distribution index: PEOE_VSA14, VSA_EState5, VSA_EState6, VSA_EState7, MinEStateIndex; polarity parameter: SMR_VSA1; molecular lipophilicity: MolLogP;
[0021] S22. Sequence normalization:
[0022] Standardized truncation and padding of variable-length peptide sequences were performed to ensure that all peptide sequences had the same length. The specific operations were as follows: all peptide sequences were unified to 50 in length, and 20 standard amino acids were learnably encoded to generate a 16-dimensional feature vector;
[0023] S23, Positional encoding: A sine-cosine positional encoding strategy is used to generate a 16-dimensional encoding vector for each amino acid position to compensate for the missing positional information of the Transformer architecture and ensure that the model can effectively capture the positional features in the amino acid sequence.
[0024] In one embodiment of the present invention, in step S3:
[0025] The feature processing and generation module is used to process and generate relevant features, providing rich feature data for the model; the input layer of the feature processing and generation module receives the 16-dimensional encoding vector of the peptide sequence and 8 physicochemical features, and the physicochemical features are spliced with the sequence code after normalization;
[0026] The sequence information processing module uses a 12-layer Transformer encoder, with each layer containing 8 attention heads and compensating for positional information through sine-cosine positional encoding. This multi-head attention mechanism deeply analyzes and efficiently processes peptide sequence information, extracting correlations between amino acids. The Transformer encoder can simultaneously focus on different parts of the sequence, capturing the complex interrelationships between amino acids, including local key structural features and long-range correlations.
[0027] The feature information processing module uses a multi-layer fully connected network to further optimize and integrate the eight input physical and chemical features to improve the expressiveness and effectiveness of the features. The Dropout mechanism is introduced during the training process to prevent overfitting and enhance the generalization ability of the model.
[0028] The result output module adopts an MLP network with a Sigmoid activation function, and the Sigmoid activation function outputs a probability value in the range of [0, 1].
[0029] In one embodiment of the present invention, in the result output module, a probability threshold of 0.5 is set to classify the peptides. Peptides with a probability greater than 0.5 are determined to be umami peptides, and vice versa.
[0030] In one embodiment of the present invention, the model training in step S4 specifically includes:
[0031] S41. Training Method: The Umami-Transformer model was trained using a 5-fold cross-validation method, using the BCELoss function as the loss function to measure the difference between the model's predicted value and the true value. During training, the convergence of the loss curve was monitored and the model performance was optimized by adjusting the hyperparameters until the Umami-Transformer model achieved a classification accuracy of over 95% on the test set.
[0032] S42. Model optimization, including:
[0033] Learning rate adjustment: dynamically adjust the learning rate according to the change in loss during training to ensure that the model can converge quickly and achieve optimal performance;
[0034] Batch size setting balances training speed and model stability to improve training efficiency;
[0035] Regularization techniques are applied to prevent the model from overfitting and enhance its ability to generalize to new data.
[0036] In one embodiment of the present invention, it further comprises:
[0037] S51. Performance Evaluation Indicators: A comprehensive quantitative evaluation of the trained model is performed using a variety of evaluation indicators, including accuracy, precision, recall, F1 score, and Matthews correlation coefficient. These indicators reflect the classification performance of the model from different perspectives, ensuring its reliability and effectiveness in practical applications.
[0038] S52, Ablation experiment: Verify the impact of feature selection on model performance through ablation experiments;
[0039] S53, all dipeptides to pentapeptides were input into the model to determine the umami peptides, and finally DD, DDE, DDED, and DDEDD were selected for synthesis verification;
[0040] S54. The taste profile of the synthetic peptides was analyzed by electronic tongue and sensory evaluation, and the interaction mechanism between the candidate peptides and the umami receptors T1R1 / T1R3 was analyzed by molecular docking technology.
[0041] In a second aspect, the present invention provides an umami peptide prediction system based on deep learning, comprising:
[0042] Data construction module: Collect experimentally verified umami peptide and non-umami peptide sequences to construct the UMP499 dataset containing multiple umami peptides and multiple non-umami peptides;
[0043] Feature selection and processing module: used to perform standardized truncation and padding of peptide sequences and calculate various physical and chemical characteristics;
[0044] Model construction module: used to construct an Umami-Transformer model including a feature processing and generation module, a sequence information processing module, a feature information processing module, and a result output module; wherein the sequence information processing module uses a Transformer encoder; the feature information processing module uses a multi-layer fully connected network; the result output module uses an MLP network with a sigmoid activation function to output the probability of umami peptides;
[0045] Model training module: The Umami-Transformer model is trained using a 5-fold cross-validation method and the BCELoss function as the loss function until the classification accuracy of the Umami-Transformer model on the test set exceeds 95%;
[0046] Umami peptide prediction module: The peptide sequence to be tested is input into the trained Umami-Transformer model to obtain the probability prediction result of it being an umami peptide.
[0047] In a third aspect, the present application provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the deep learning-based umami peptide prediction method when executing the computer program.
[0048] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the deep learning-based umami peptide prediction method when executed by a processor.
[0049] The present application has the following beneficial effects:
[0050] The present application proposes a deep learning framework of Umami-Transformer, which can efficiently process peptide sequence information and significantly improve the prediction accuracy of umami peptides by fusing the Transformer architecture with eight key physicochemical characteristics (including molecular lipophilicity, charge distribution, and polarity parameters, etc.). Further combined with the exhaustive method to screen all dipeptides to pentapeptides, and through the molecular docking technology to analyze the interaction mechanism of candidate peptides and umami receptor T1R1 / T1R3, the whole process research from theoretical prediction to biological experimental verification is finally realized, which provides an innovative technical path for efficient mining and functional application of umami peptides. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 Fig. 1 is a schematic diagram of the architecture of the Umami-Transformer model.
[0052] Figure 2 Fig. 4 is a schematic diagram of the loss curve and classification accuracy change during the model training process.
[0053] Figure 3 Fig. 6 is a schematic diagram of the ablation experiment results.
[0054] Figure 4 Fig. 7 is a schematic diagram of the performance comparison of different umami peptide prediction models.
[0055] Figure 5 Fig. 8 is a schematic diagram of the molecular docking interaction between umami peptides and umami receptor T1R1 / T1R3.
[0056] Figure 6 Fig. 9 is a schematic diagram of the sensory evaluation and electronic tongue analysis results of the synthesized umami peptides.
[0057] Figure 7 Fig. 10 is a schematic diagram of the network server interface of the Umami-Transformer model. DETAILED DESCRIPTION
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] like Figures 1 to 7 As shown, the present invention provides a method and system for predicting umami peptides based on deep learning. Figure 1 This is a schematic diagram of the Umami-Transformer model architecture, showing the four core components of the model and their functions. Figure 2 The figure shows the loss curve and classification accuracy changes during model training, indicating that the model converges quickly during training and achieves a classification accuracy of over 95% on the test set. Figure 3 This is a schematic diagram of the ablation experiment results, showing the impact of different features on model performance and verifying the importance of the selected features. Figure 4 Schematic diagram of the performance comparison of different umami peptide prediction models, highlighting the superiority of the Umami-Transformer model in various evaluation indicators. Figure 5 Schematic diagram of the molecular docking interaction between umami peptides and umami receptors T1R1 / T1R3, revealing the binding mechanism between umami peptides and receptors. Figure 6 Schematic diagram of the sensory evaluation and electronic tongue analysis results of the synthetic umami peptides, verifying the accuracy of the model predictions. Figure 7 A diagram of the web server interface for the Umami-Transformer model, demonstrating the model's online prediction capabilities.
[0060] Example 1
[0061] This embodiment provides a method for predicting umami peptides based on deep learning, comprising the following steps:
[0062] S1. Dataset Construction: Collect experimentally verified umami peptide and non-umami peptide sequences to construct the UMP499 dataset containing multiple umami peptides and non-umami peptides;
[0063] S2, feature selection and processing: standardize truncation and padding of peptide sequences, and calculate various physicochemical features;
[0064] S3. Constructing an Umami-Transformer model: Constructing an Umami-Transformer model comprising a feature processing and generation module, a sequence information processing module, a feature information processing module, and a result output module; wherein the sequence information processing module uses a Transformer encoder; the feature information processing module uses a multi-layer fully connected network; and the result output module uses an MLP network with a sigmoid activation function to output umami peptide probabilities;
[0065] S4. Model training: The Umami-Transformer model is trained using a 5-fold cross-validation method and the BCELoss function as the loss function until the classification accuracy of the Umami-Transformer model on the test set exceeds 95%;
[0066] S5. Umami peptide prediction: The peptide sequence to be tested is input into the trained Umami-Transformer model to obtain the probability prediction result of it being an umami peptide.
[0067] Optionally, the step S1 of constructing the data set specifically includes:
[0068] S11. Data collection: We collected experimentally verified umami and non-umami peptide sequences from public literature and the BIOPEP-UWM database to construct the UMP499 dataset, which contains 212 umami peptides and 287 non-umami peptides.
[0069] S12. Data preprocessing, including:
[0070] Remove outliers: remove peptide sequences containing non-standard amino acid symbols (such as "B", "U", "X", "Z") to ensure data accuracy and consistency;
[0071] Remove redundancy: Use sequence alignment tools (such as CD-HIT) to remove duplicate sequences, eliminate redundancy, and avoid the impact of duplicate data on model training;
[0072] Hierarchical partitioning: The dataset is hierarchically partitioned into a UMPCV training set and a UMPTS test set in an 80:20 ratio to ensure the distribution consistency of the training set and the test set to improve the generalization ability of the model.
[0073] Optionally, the feature selection and processing in step S2 specifically includes:
[0074] S21. Feature calculation: RDKit cheminformatics toolkit was used to calculate eight physicochemical features of each peptide sequence, including: molecular weight correlation feature: BCUT2D_MWLOW; charge distribution index: PEOE_VSA14, VSA_EState5, VSA_EState6, VSA_EState7, MinEStateIndex; polarity parameter: SMR_VSA1; molecular lipophilicity: MolLogP; the above features can comprehensively describe the physicochemical properties of the peptide from different angles and provide rich information for the model.
[0075] S22. Sequence normalization:
[0076] Standardized truncation and padding of variable-length peptide sequences were performed to ensure that all peptide sequences had the same length for ease of processing by the machine learning model. The specific operations were as follows: all peptide sequences were unified to 50 in length, and 20 standard amino acids were learnably encoded to generate a 16-dimensional feature vector;
[0077] S23, Positional encoding: A sine-cosine positional encoding strategy is used to generate a 16-dimensional encoding vector for each amino acid position to compensate for the missing positional information of the Transformer architecture and ensure that the model can effectively capture the positional features in the amino acid sequence.
[0078] Optionally, in step S3:
[0079] The feature processing and generation module is used to process and generate relevant features, providing rich feature data for the model; the input layer of the feature processing and generation module receives the 16-dimensional encoding vector and 8 physicochemical features of the peptide sequence, and the physicochemical features are normalized (Z-Score normalization) and then spliced with the sequence code;
[0080] The sequence information processing module uses a 12-layer Transformer encoder, with each layer containing 8 attention heads and compensating for positional information through sine-cosine positional encoding. This multi-head attention mechanism deeply analyzes and efficiently processes peptide sequence information, accurately extracting correlations between amino acids. The Transformer encoder can simultaneously focus on different parts of the sequence, capturing the complex interrelationships between amino acids, whether local key structural features or long-range correlations.
[0081] The feature information processing module uses a multi-layer fully connected network to further optimize and integrate the eight input physical and chemical features to improve the expressiveness and effectiveness of the features. The Dropout mechanism is introduced during the training process to prevent overfitting and enhance the generalization ability of the model.
[0082] The result output module adopts an MLP network with a Sigmoid activation function. The Sigmoid activation function outputs a probability value in the range of [0, 1]. The peptides are classified by setting a probability threshold of 0.5. Peptides with a probability greater than 0.5 are judged to be umami peptides, otherwise they are non-umami peptides.
[0083] Optionally, the model training in step S4 specifically includes:
[0084] S41. Training method: The Umami-Transformer model is trained using a 5-fold cross-validation method. The BCELoss function is used as the loss function to measure the difference between the model's predicted value and the true value. During the training process, the convergence of the loss curve is monitored and the model performance is optimized by adjusting the hyperparameters until the classification accuracy of the Umami-Transformer model on the test set exceeds 95% (e.g. Figure 2 shown);
[0085] S42. Model optimization, including:
[0086] Learning rate adjustment: dynamically adjust the learning rate according to the change in loss during training to ensure that the model can converge quickly and achieve optimal performance;
[0087] Batch size setting: select an appropriate batch size to balance training speed and model stability to improve training efficiency;
[0088] Regularization technology, such as L2 regularization, is used to prevent the model from overfitting and enhance its generalization ability on new data.
[0089] Optionally, the umami peptide prediction method further comprises:
[0090] S51. Performance Evaluation Indicators: A comprehensive quantitative evaluation of the trained model is performed using a variety of evaluation indicators, including accuracy, precision, recall, F1 score, and Matthews correlation coefficient. These indicators reflect the classification performance of the model from different perspectives, ensuring its reliability and effectiveness in practical applications.
[0091] S52. Ablation experiment: The influence of feature selection on model performance was verified by ablation experiment. The results showed that the eight selected physical and chemical features made significant contributions to the various performance indicators of the model, among which VSA_EState6 and BCUT2D_MWLOW features had the most critical impact on model performance (such as Figure 3 shown);
[0092] S53, all dipeptides to pentapeptides were input into the model to determine the umami peptides, and finally DD, DDE, DDED, and DDEDD were selected for synthesis verification;
[0093] S54, analyze the taste profile of the synthetic peptides by electronic tongue and sensory evaluation, and analyze the interaction mechanism between the candidate peptides and the umami taste receptors T1R1 / T1R3 by molecular docking technology;
[0094] S55. Comparison with other models: Compared with other existing umami peptide prediction models (such as iUmami-SCM, Umami_YYDS, and Umami-BERT), Umami-Transformer performed well in terms of accuracy, precision, recall, Matthews correlation coefficient, and F1 score, reaching 0.965, 1.0, 0.825, 0.889, and 0.903, respectively, significantly outperforming other models (such as Figure 4 shown).
[0095] Example 2
[0096] This embodiment provides a method for predicting and screening umami peptides based on deep learning according to the umami peptide prediction method described in Example 1, comprising:
[0097] 2.1 Large-scale peptide prediction:
[0098] The Umami-Transformer model was used to comprehensively predict all dipeptides to pentapeptides, totaling 3,368,400 peptides (400 dipeptides, 8,000 tripeptides, 160,000 tetrapeptides, and 3,200,000 pentapeptides). Through prediction, peptides with potential umami activity were screened out, providing candidates for subsequent experimental verification and application development.
[0099] 2.2 Molecular docking analysis:
[0100] Molecular docking analysis of the top-ranked peptides showed that these peptides could stably bind to the umami taste receptors T1R1 / T1R3 with low binding energies. For example, the binding energies -CDOCKERENERGY of dipeptide DD, tripeptide DDE, tetrapeptide DDED, and pentapeptide DDEDD were 59.3507 kcal / mol, 80.4485 kcal / mol, 95.5644 kcal / mol, and 97.9330 kcal / mol, respectively, and the binding energies -CDOCKERENERGY of dipeptide DD, tripeptide DDE, tetrapeptide DDED, and pentapeptide DDEDD were 38.0929 kcal / mol, 55.1156 kcal / mol, 75.3756 kcal / mol, and 65.6673 kcal / mol, respectively (e.g., Figure 5 shown).
[0101] 2.3 Peptide synthesis and taste verification:
[0102] 2.3.1 Peptide Synthesis
[0103] The purity of the four highest-ranked peptides (DD, DDE, DDED, and DDEDD) was no less than 95%.
[0104] 2.3.2 Sensory evaluation
[0105] Sensory evaluation experiments were conducted to verify the taste properties of the synthesized peptides. The results showed that these peptides have significant umami and salty flavors. The umami intensity of DDE and DDED exceeded that of 3% MSG, and their taste thresholds were low, at 0.29 mg / mL and 0.19 mg / mL, respectively. The sensory evaluation of the peptides is shown in Table 1 below:
[0106] Table 1. Sensory evaluation of peptides
[0107]
[0108] 2.3.3 Electronic tongue analysis
[0109] The taste characteristics of the synthetic peptides were quantitatively analyzed using an electronic tongue device, and the results were consistent with the sensory evaluation, further verifying the accuracy of the model prediction. For example, the umami intensity score of DDE was 12.06, which was higher than the 11.56 of 3mg / mL MSG and close to the umami value of 5mg / mL, which was 12.60. Figure 6 shown.
[0110] Example 3
[0111] This example provides an umami peptide prediction system based on deep learning. The system implementation and deployment includes web server development: A web server for the Umami-Transformer model is developed, providing a user-friendly interface and comprehensive documentation to facilitate researchers to efficiently analyze umami peptides. The server supports the submission of single or batch peptide sequences. Users can download standardized input templates and obtain analysis results in the "Result" section after submitting the sequence. Figure 7 shown.
[0112] A deep learning-based umami peptide prediction system, comprising:
[0113] Data construction module: Collect experimentally verified umami peptide and non-umami peptide sequences to construct the UMP499 dataset containing multiple umami peptides and multiple non-umami peptides;
[0114] Feature selection and processing module: used to perform standardized truncation and padding of peptide sequences and calculate various physical and chemical characteristics;
[0115] Model construction module: used to construct an Umami-Transformer model including a feature processing and generation module, a sequence information processing module, a feature information processing module, and a result output module; wherein the sequence information processing module uses a Transformer encoder; the feature information processing module uses a multi-layer fully connected network; the result output module uses an MLP network with a sigmoid activation function to output the probability of umami peptides;
[0116] Model training module: The Umami-Transformer model is trained using a 5-fold cross-validation method and the BCELoss function as the loss function until the classification accuracy of the Umami-Transformer model on the test set exceeds 95%;
[0117] Umami peptide prediction module: The peptide sequence to be tested is input into the trained Umami-Transformer model to obtain the probability prediction result of it being an umami peptide.
[0118] Example 4
[0119] This example provides the results of docking between DD, DDE, DDED, and DDEDD and umami taste receptors.
[0120] 4.1 Docking Method
[0121] Molecular docking experiments were conducted using Discovery Studio 2019 software, targeting the homology model of the T1R1 / T1R3 umami taste receptor. Four umami peptides (DD, DDE, DDED, and DDEDD) were docked. Before docking, energy minimization optimization was performed on the receptor and ligand to ensure structural stability. During the docking process, the CDOCKER program was used to find the optimal binding conformation between the ligand and receptor.
[0122] 4.2. Docking results with umami receptors
[0123] The binding of glutamate to umami taste receptors on the taste cell membrane initiates an intracellular signaling cascade that mediates the activation of ion channels, particularly transient receptor potential (TRPM) channels. This signaling process subsequently induces a calcium-dependent depolarization of taste receptor cells, ultimately leading to neurotransmission to the taste cortex area via afferent fibers of the facial and glossopharyngeal nerves. The perception of taste is significantly influenced by T1R1 / T1R3 receptors, which are primarily located on the anterior portion of the tongue.
[0124] Homology modeling and computational molecular docking are advanced methods used to identify possible binding sites between umami peptides and taste receptors. In binding to umami receptors T1R1 / T1R3, DD uses four different forms of hydrogen bonds: salt bridges, van der Waals, traditional hydrogen bonds, and carbon-hydrogen bonds ( Figure 5 The carbon-hydrogen bond is a covalent bond with high bond energy, approximately 413 kJ / mol, while the strength of a traditional hydrogen bond is typically between 5 and 30 kJ / mol. The strength of a salt bridge is approximately 5 to 6 kJ / mol. The strength of a van der Waals force is typically in the range of 0.4 kJ / mol to 4 kJ / mol. The oxygen atom in the R group on the Asp residue forms a strong carbon-hydrogen bond with ARG 712, GLY 695, and GLN 853.
[0125] The establishment of a salt bridge between the oxygen atom and ARG 712 is of significant importance. According to the research results, ARG 712 constitutes an important binding site. Meanwhile, ARG 712 of T1R1 / T1R3 also binds with the umami peptides PQFAPEED, EEHPVLLTEA, and DQAIPNKPEE. As shown in FIG. 1, the oxygen atom of the R group on the Asp residue forms a strong carbon-hydrogen bond with ARG 712, GLY 695, and GLN 853. Figure 5 As shown in FIG. 1, DDE binds with the umami receptor T1R1 / T1R3 at several key sites, including ARG 712, TYR 694, HIS 915, SER 693, ALA 856, VAL 714, GLY 695, HIS 914, GLY 911, GLY 855, LEU 912, ARG 854, PRO 715, LEU 852, GLN 853, LEU 967, ARG 591, and PHE 851.
[0126] The main interactions formed include van der Waals forces, carbon-hydrogen bonds, attractive charge-charge interactions, unfavorable positive-positive interactions, traditional hydrogen bonds, and pi-anion interactions. Attractive charge-charge interactions are stronger than carbon-hydrogen bonds, while pi-anion interactions are stronger than salt bridges but weaker than traditional hydrogen bonds. The research shows that hydrogen bonds play an important role in the perception process caused by the binding of umami peptides with umami receptors. Notably, ARG 591 and ARG 712 form attractive charge-charge interactions, TYR 694 and HIS 915 form pi-anion interactions with DDE. In addition, GLY 695 forms a strong traditional hydrogen bond with the oxygen atom on the R group of Asp. ARG 712 forms an unfavorable positive-positive interaction with the N atom of DDE, which refers to the repulsive force between two positive charges, requiring additional energy to maintain this interaction.
[0127] The binding force of the tetrapeptide DDED to T1R1 / T1R3 includes all favorable interactions of the dipeptide DD and the tripeptide DDE, such as van der Waals forces, traditional hydrogen bonds, salt bridges, carbon-hydrogen bonds, attractive charge-charge interactions, and pi-anion interactions, as shown in FIG. 1. Figure 5As shown in EF in . ARG591, ARG854, HIS915 and ARG712 form salt bridges, attractive charge interactions and Pi anion interactions with DDED. The presence of D / E residues enhances electrostatic interactions in umami peptides, reduces the HOMO-LUMO energy gap, and regulates receptor recognition pathways through orbital redistribution effects. Compared with DDED, the pentapeptide DDEDD forms fewer salt bridge interactions with umami receptors. Overall, VAL714, TYR694, SER693, LEU912, LEU852, HIS915, HIS914, GLY911, GLY855, GLY695, GLN853, ARG854 and ARG712 form binding interactions with all four peptides, which may be key sites for umami perception ( Figure 5 In addition, DDE, DDED, and DDEDD also bind to PRO715, PHE851, and LEU967 on the umami taste receptor, which may contribute to their higher umami taste intensity compared to DD. These findings suggest that there is no correlation between peptide chain length and umami taste receptor activation efficiency, with shorter peptides (2-5 residues) exhibiting superior binding affinity, likely due to their optimal steric compatibility with the receptor's hydrophobic pocket.
[0128] Example 5
[0129] This embodiment provides an application case of the deep learning-based umami peptide prediction method according to Example 1, including:
[0130] 1. Development of low-sodium foods: By predicting and screening umami peptides with salty taste enhancement effects, such as DDE, DDED, and DDEDD, they can be used as functional salt substitutes to partially replace sodium chloride in foods and achieve the development of low-sodium foods.
[0131] 2. Taste regulation research: Using the umami peptides predicted by the Umami-Transformer model, we will conduct in-depth research on the molecular mechanism of umami perception and provide theoretical support for the development of food flavorings.
[0132] In summary, to address the low efficiency and high cost of traditional umami peptide screening methods, the present application combines the Transformer architecture with eight physicochemical descriptors (including molecular lipophilicity, charge distribution, and molecular weight-related features) to construct a high-precision umami peptide prediction model. The model is trained using five-fold cross-validation, with a classification accuracy of 0.965, an F1 score of 0.903, and a Matthews correlation coefficient of 0.889, significantly better than existing methods. Further based on the model, four candidate umami peptides (DD, DDE, DDED, and DDEDD) are screened from all 2-5 peptides. After solid-phase synthesis and sensory evaluation verification, their umami intensity exceeds 3% MSG, among which DDE and DDED have umami thresholds of 0.29 mg / mL and 0.19 mg / mL, respectively, and exhibit a synergistic salt-enhancing effect, which can be applied to low-sodium food development. Molecular docking shows that the Asp / Glu residues at the ends of the peptide chains bind stably to the umami receptors T1R1 / T1R3 through hydrogen bonds, salt bridges, and charge complementation, revealing the taste perception mechanism. The present application provides an innovative technical solution for efficient prediction, flavor optimization, and salt reduction strategies of umami peptides, with significant application value in the food industry.
[0133] The present application also has the following beneficial effects:
[0134] 1. Expanding peptide diversity: Further expanding the training data set to cover more types of peptide sequences, improving the generalization ability and prediction accuracy of the model.
[0135] 2. In-depth study of taste mechanism: Combining computational simulation and experimental techniques, in-depth exploration of the interaction mechanism between umami peptides and taste receptors, providing a basis for model optimization and development of new taste modulators.
[0136] 3. Interdisciplinary cooperation: Strengthening interdisciplinary cooperation in food science, bioinformatics, chemical informatics, and other disciplines, promoting innovation and development in the field of taste research, and bringing more innovative products and solutions to the food industry.
[0137] Through the above specific implementation schemes, the present application successfully establishes a machine learning-based umami peptide prediction method and system, providing an efficient and accurate tool for the research and application of umami peptides, and promoting the progress of taste regulation and low-sodium food development in the field of food science.
[0138] Further, the present application also provides a computer device, which can include a processor, a memory, a network interface and a database connected through a system bus. Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, causes the processor to perform the steps of the deep learning-based umami peptide prediction method according to any one of the above embodiments.
[0139] The working process, working details and technical effects of the computer device provided by the present embodiment can be referred to the above embodiments of the deep learning-based umami peptide prediction method, which will not be repeated here.
[0140] Further, the present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program, when executed by a processor, implements the steps of the deep learning-based umami peptide prediction method according to any one of the above embodiments. Wherein, the computer readable storage medium refers to a carrier for storing data, which can include, but is not limited to, floppy disks, optical discs, hard disks, flash memories, USB flash disks and / or memory sticks, etc. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.
[0141] The working process, working details and technical effects of the computer readable storage medium provided by the present embodiment can be referred to the above embodiments of the deep learning-based umami peptide prediction method, which will not be repeated here.
[0142] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0143] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for predicting umami peptides based on deep learning, characterized in that: The steps include: S1. Dataset Construction: Collect experimentally verified umami peptide and non-umami peptide sequences to construct the UMP499 dataset containing umami peptides and non-umami peptides; S2, feature selection and processing: standardize the peptide sequence truncation and padding, and calculate 8 physicochemical features; S3. Constructing an Umami-Transformer model: Constructing an Umami-Transformer model comprising a feature processing and generation module, a sequence information processing module, a feature information processing module, and a result output module; wherein the sequence information processing module uses a Transformer encoder; the feature information processing module uses a multi-layer fully connected network; and the result output module uses an MLP network with a sigmoid activation function to output umami peptide probabilities; S4. Model training: The Umami-Transformer model is trained using a 5-fold cross-validation method and the BCELoss function as the loss function until the classification accuracy of the Umami-Transformer model on the test set exceeds 95%; S5. Umami peptide prediction: The peptide sequence to be tested is input into the trained Umami-Transformer model to obtain the probability prediction result of it being an umami peptide.
2. The method for predicting umami peptides based on deep learning according to claim 1, characterized in that: The step S1 of constructing the data set specifically includes: S11. Data collection: We collected experimentally verified umami and non-umami peptide sequences from public literature and the BIOPEP-UWM database to construct the UMP499 dataset, which contains 212 umami peptides and 287 non-umami peptides. S12. Data preprocessing, including: Remove outliers: remove peptide sequences containing non-standard amino acid symbols to ensure data accuracy and consistency; Remove redundancy: Use sequence alignment tools to remove duplicate sequences, eliminate redundancy, and avoid the impact of duplicate data on model training; Stratification: The dataset was stratified into a UMP-CV training set and a UMP-TS test set in an 80:20 ratio to ensure the distribution consistency of the training set and the test set.
3. The method for predicting umami peptides based on deep learning according to claim 2, characterized in that: The feature selection and processing in step S2 specifically includes: S21. Feature calculation: 8 physicochemical features of each peptide sequence were calculated using the RDKit cheminformatics toolkit, including: molecular weight correlation feature: BCUT2D_MWLOW; charge distribution index: PEOE_VSA14, VSA_EState5, VSA_EState6, VSA_EState7, MinEStateIndex; polarity parameter: SMR_VSA1; molecular lipophilicity: MolLogP; S22. Sequence normalization: Standardized truncation and padding of variable-length peptide sequences were performed to ensure that all peptide sequences had the same length. The specific operations were as follows: all peptide sequences were unified to 50 in length, and 20 standard amino acids were learnably encoded to generate a 16-dimensional feature vector; S23, Positional encoding: A sine-cosine positional encoding strategy is used to generate a 16-dimensional encoding vector for each amino acid position to compensate for the missing positional information of the Transformer architecture and ensure that the model can effectively capture the positional features in the amino acid sequence.
4. The method for predicting umami peptides based on deep learning according to claim 3, characterized in that: In the step S3: The feature processing and generation module is used to process and generate relevant features, providing rich feature data for the model; the input layer of the feature processing and generation module receives the 16-dimensional encoding vector of the peptide sequence and 8 physicochemical features, and the physicochemical features are spliced with the sequence code after normalization; The sequence information processing module uses a 12-layer Transformer encoder, with each layer containing 8 attention heads, and compensates for position information through sine-cosine position encoding. The multi-head attention mechanism deeply analyzes and efficiently processes peptide sequence information to extract the correlation between amino acids. The Transformer encoder can simultaneously focus on different parts of the sequence and capture the complex relationships between amino acids, including local key structural features and long-range correlations; The feature information processing module uses a multi-layer fully connected network to further optimize and integrate the eight input physical and chemical features to improve the expressiveness and effectiveness of the features. The Dropout mechanism is introduced during the training process to prevent overfitting and enhance the generalization ability of the model. The result output module adopts an MLP network with a Sigmoid activation function, and the Sigmoid activation function outputs a probability value in the range of [0, 1].
5. The method for predicting umami peptides based on deep learning according to claim 4, characterized in that: In the result output module, a probability threshold of 0.5 is set to classify the peptides. Peptides with a probability greater than 0.5 are classified as umami peptides, and those with a probability less than 0.5 are classified as non-umami peptides.
6. The method for predicting umami peptides based on deep learning according to claim 5, characterized in that: The model training in step S4 specifically includes: S41. Training Method: The Umami-Transformer model was trained using a 5-fold cross-validation method, using the BCELoss function as the loss function to measure the difference between the model's predicted value and the true value. During training, the convergence of the loss curve was monitored and the model performance was optimized by adjusting the hyperparameters until the Umami-Transformer model achieved a classification accuracy of over 95% on the test set. S42. Model optimization, including: Learning rate adjustment: dynamically adjust the learning rate according to the change in loss during training to ensure that the model can converge quickly and achieve optimal performance; Batch size setting balances training speed and model stability to improve training efficiency; Regularization techniques are applied to prevent the model from overfitting and enhance its ability to generalize to new data.
7. The method for predicting umami peptides based on deep learning according to claim 6, characterized in that: Also includes: S51. Performance Evaluation Indicators: A comprehensive quantitative evaluation of the trained model is performed using a variety of evaluation indicators, including accuracy, precision, recall, F1 score, and Matthews correlation coefficient. These indicators reflect the classification performance of the model from different perspectives, ensuring its reliability and effectiveness in practical applications. S52, Ablation experiment: Verify the impact of feature selection on model performance through ablation experiments; S53, all dipeptides to pentapeptides were input into the model to determine the umami peptides, and finally DD, DDE, DDED, and DDEDD were selected for synthesis verification; S54. The taste profile of the synthetic peptides was analyzed by electronic tongue and sensory evaluation, and the interaction mechanism between the candidate peptides and the umami receptors T1R1 / T1R3 was analyzed by molecular docking technology.
8. A deep learning-based umami peptide prediction system, characterized in that: include: Data construction module: Collect experimentally verified umami peptide and non-umami peptide sequences to construct the UMP499 dataset containing multiple umami peptides and multiple non-umami peptides; Feature selection and processing module: used to perform standardized truncation and padding of peptide sequences and calculate various physical and chemical characteristics; Model construction module: used to construct an Umami-Transformer model including a feature processing and generation module, a sequence information processing module, a feature information processing module, and a result output module; wherein the sequence information processing module uses a Transformer encoder; the feature information processing module uses a multi-layer fully connected network; the result output module uses an MLP network with a sigmoid activation function to output the probability of umami peptides; Model training module: The Umami-Transformer model is trained using a 5-fold cross-validation method and the BCELoss function as the loss function until the classification accuracy of the Umami-Transformer model on the test set exceeds 95%; Umami peptide prediction module: The peptide sequence to be tested is input into the trained Umami-Transformer model to obtain the probability prediction result of it being an umami peptide.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the umami peptide prediction method based on deep learning are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the umami peptide prediction method based on deep learning are implemented.
Citation Information
Patent Citations
Polypeptide hemolytic prediction method fusing multi-view feature deep ensemble learning
CN118136112A
Taste peptide design method based on deep learning
CN119943158A
Deep learning-based method for predicting binding affinity between human leukocyte antigens and peptides
US20220028487A1
Cited By
Oyster flavor peptide screening method based on multi-mode umami intensity characteristic recognition
CN121601028A
A method for screening oyster taste presenting peptide based on multi-modal umami intensity feature recognition
CN121601028B