Taste peptide design method based on deep learning

By establishing a systematic taste peptide database and feature analysis system, and proposing a loss-supervised adaptive variant autoencoder (LA-VAE) model, the time-consuming and limited problems of traditional taste peptide identification methods are solved, and the precise regulation of taste characteristics and the design and verification of new functional peptides are achieved.

CN119943158APending Publication Date: 2025-05-06HUNAN NORMAL UNIVERSITY

Patent Information

Application Number
CN202510157337.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional taste peptide identification methods are time-consuming and resource-intensive, with limited peptide samples produced, and existing calculation methods are limited to identifying two tastes, which cannot meet actual needs, and cannot accelerate the discovery of potential taste peptides.

Method used

A systematic taste peptide database and feature analysis system was established, including a multi-source data integration system, sequence standardization processing and taste characteristic analysis method, and a loss supervision adaptive variable autoencoder (LA-VAE) model was proposed, using a dual-mode training strategy, a dynamic loss supervision mechanism and a comparative learning taste control.

Benefits of technology

73 new functional peptides of sweet, salty and fresh were successfully designed and verified, achieving accurate control of taste characteristics, reflecting the advantages of this method in precise control of taste characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to a multi-taste peptide design method based on deep learning, and belongs to the field of bioinformatics and food science and technology. Aiming at the technical problems that a traditional taste peptide development method is long in time consumption and low in efficiency and an existing calculation method is difficult to predict and generate multiple taste characteristic peptide fragments at the same time, the invention provides a solution based on a loss supervision adaptive variational auto-encoder (LA-VAE). According to the scheme, the standardized database is constructed by integrating multi-source data, model training is carried out by adopting a dual-mode training strategy and a dynamic loss supervision mechanism, and accurate control over various taste characteristics is achieved. 73 kinds of novel functional peptide fragments are successfully designed and verified by applying the method, and the peptide fragments show sweet taste, fresh taste and salty taste characteristics under different concentrations, have good biological safety and can be applied to the field of salt reduction and sugar reduction in the food industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of food biotechnology, and in particular to a taste peptide design method based on deep learning. Background Art

[0002] Taste peptides, as a new type of natural flavoring, have attracted much attention due to their unique sensory properties, good safety and potential health benefits. These bioactive peptides are usually composed of 2-20 amino acid residues and can trigger a variety of taste perceptions including sweetness, umami and saltiness, without the shortcomings of traditional flavorings. Compared with traditional flavorings (such as sodium chloride, sucrose and monosodium glutamate), taste peptides have significant advantages: first, due to their natural amino acid composition, they are easy to metabolize and absorb; second, they can present multiple taste modalities at the same time, reducing the need to use multiple flavorings; in addition, they have additional health-promoting functions such as antioxidant and anti-inflammatory properties, which provides a promising solution to the taste-health paradox in the modern food industry.

[0003] Taste peptides have shown important application value in the food industry. In salt reduction applications, specific peptide sequences can reduce the sodium chloride content of meat products while maintaining sensory quality. Studies have shown that a decapeptide (0.4 mg / mL) in fermented tofu can increase the saltiness perception of 50 mM sodium chloride to the equivalent of 63 mM. In terms of sugar reduction, sweet peptides such as aspartame are widely used, and sweet peptides in mulberry seed protein show six times higher sweetness than 0.1 g / mL sucrose.

[0004] However, traditional methods for identifying taste peptides face significant challenges. Conventional pretreatment, extraction, purification, synthesis, and sensory evaluation workflows are time-consuming and resource-intensive, and the output of peptide samples is limited. In addition, the complexity of biological samples and experimental variation may lead to application limitations such as toxicity risks, stability issues, or insufficient taste properties of the obtained peptides, which significantly increases development costs and complexity.

[0005] Existing computational methods have obvious limitations: first, the existing prediction models are of limited complexity and can only recognize two tastes at most, which cannot meet actual needs; second, many existing studies construct a binary classification framework using bitterness as the opposite of umami, ignoring the fact that multiple tastes may coexist; finally, relying solely on prediction models may not effectively accelerate the discovery of potential taste peptides because they can only classify existing sequences and do not have the ability to generate new sequences. Summary of the invention

[0006] The present invention establishes a systematic taste peptide database and feature analysis system, which specifically includes: Multi-source data integration system: integrating professional databases, prediction model data sets and literature mining data; Sequence standardization: Establish a mechanism for sequence screening and comprehensive annotation of taste characteristics; Taste characteristics analysis method: Construct an amino acid composition distribution analysis system to reveal the regularity of characteristic amino acid residues.

[0007] This paper proposes a Loss-Supervised Adaptive Variational Autoencoder (LA-VAE) model with the following innovative features: Dual-mode training strategy: Single mode: precise training for specific taste combinations; Multimodal: Adopt inclusive strategies to expand sequence diversity; Dynamic loss supervision mechanism: to accurately capture the model state; Contrastive learning of taste control: Towards a framework for structured contrastive representations.

[0008] The method of the present invention achieves remarkable results in practical applications: The method of the present invention successfully designed and verified 73 new functional peptides of sweet, salty and umami, which showed sweet and umami characteristics at a concentration of 0.1 mg / mL and salty characteristics at a concentration of 1 mg / mL. In particular, some peptides showed strong sweetness and umami intensity at low concentrations, and showed differentiated salty intensity at high concentrations, reflecting the advantages of this method in the precise regulation of taste characteristics. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below.

[0010] Figure 1 A flow chart for the construction and analysis of the dataset of the present invention, showing the sequence characteristics and taste property distribution of taste peptides, including: sequence length distribution (88.54% of the sequences are less than 10 amino acids in length), quantitative distribution of the five basic tastes (575 for umami, 541 for bitter, 201 for sour, 162 for sweet, and 141 for salty), amino acid composition feature analysis, and multiple taste classification statistics of 1131 taste peptides; Figure 2 This is a schematic diagram of the structure of the LA-VAE model of the present invention, showing the core components and workflow of the model, including: the architecture design of the encoder, latent space sampler and decoder, and the implementation mechanism of the training strategy based on single mode and multi-mode; Figure 3This is a performance evaluation diagram of the LA-VAE model of the present invention, showing the optimization process of the model under different training rounds, including: the loss function convergence curve (positive sample loss dropped from 2.6708 to 0.1817, negative sample loss dropped from 1.2410 to 0.1720) and the principal component analysis results of the feature space distribution; Figure 4 The taste property verification diagram of the peptides generated for the present invention shows the electronic tongue analysis results of 73 peptides at two concentrations (0.1 mg / mL and 1 mg / mL). The quartile distribution of taste intensity is represented by the size of the colored dots, verifying the differentiated taste properties of the peptides at different concentrations. DETAILED DESCRIPTION

[0011] In order to allow those skilled in the art to understand the present invention more clearly and intuitively, the present invention will be further described below in conjunction with the accompanying drawings.

[0012] Construction of the dataset: like Figure 1 As shown, the present invention first extracts basic data from professional databases such as BIOPEP-UWM and TastePeptidesDB, and integrates data sets of prediction models such as Umami-MRNN, VirtuousUmami and IUP-BERT. Literature mining is performed through PubMed and Google Scholar, and the combination of "Tastes", "Sour", "Sweet", "Bitter", "Salty", "Umami" and "Peptides" is used to retrieve supplementary data. The collected sequences are standardized, sequences containing non-standard amino acid residues are eliminated, and the sequence length is limited to no more than 15 amino acids (accounting for more than 97%). Figure 2 As shown, the final dataset contained umami (575), bitter (541), sour (201), sweet (162), and salty (141) peptides.

[0013] Taste Property Annotation System: A five-digit coding system (>abcde) was designed to correspond to the five basic tastes of sour, sweet, bitter, salty, and umami. The annotation was performed using a ternary labeling method of "1" (present), "0" (absent), and "x" (undetermined). For example, peptides with salty and umami tastes were annotated as ">xxx11", and peptides with sweet tastes were annotated as ">x1xxx". A classification system for multi-taste combination patterns was established, including single taste (770 types), dual taste (250 types), triple taste (94 types), and quadruple taste (17 types) sequences.

[0014] LA-VAE model construction: Training modes include: Single Pattern Mode: Precise training for specific taste combinations, such as using only dual-flavor peptide data when developing sweet and sour dual-flavor peptides; Improve the prediction accuracy of target features by limiting the range of training data; Suitable for targeted development needs with clear target taste characteristics; Multiple Pattern Mode: Adopt inclusive training strategies, such as integrating single sour peptides, single sweet peptides and their combination data when developing sweet and sour dual-flavor peptides; Expand training data resources and enhance the model's ability to explore sequence space; Suitable for development scenarios where greater sequence diversity is required.

[0015] The core components of the model include: Encoder: One-dimensional convolutional layer (Conv1D, filters=32, kernel_size=3); Fully connected layers; Map the amino acid sequence to a 2000-dimensional Gaussian latent space; Decoder: Adopt mirror structure; Reconstruct the sequence probability distribution through inverse mapping; Loss monitoring mechanism: Divide the training process into exploration phases; Convergence optimization phase; Keep tracking the global optimal loss value; Contrastive learning mechanism: By constructing positive and negative sample pairs; Towards a framework for structured contrastive representation in latent space.

[0016] Sequence optimization and screening: The hierarchical clustering system based on sequence similarity uses an improved Needleman-Wunsch global alignment algorithm to optimize sequences through a scoring matrix (matching score 2.0, mismatch penalty -1.0) and an affine gap penalty (open penalty -0.5, extension penalty -0.1). A dual distance evaluation system is established to calculate the average Euclidean distance between the candidate sequence and the positive and negative sample sets based on k nearest neighbors (k=5) for sequence screening.

[0017] Experimental verification: like Figure 4As shown, the taste properties of the generated peptides were verified by electronic tongue analysis, and the test was carried out at two concentrations of 0.1 mg / mL and 1 mg / mL. The taste property analysis showed that at a concentration of 0.1 mg / mL, it exhibited sweet and umami properties, and at a concentration of 1 mg / mL, it exhibited salty properties. It was proved that the method of the present invention successfully designed and verified 73 new functional peptides of sweet, salty and umami, reflecting the advantages of this method in the precise regulation of taste properties.

Claims

1. A method for designing and predicting taste peptides based on deep learning, characterized in that: The method comprises the following steps: Build a multi-source data integration system to obtain training data from the BIOPEP-UWM database, TastePeptidesDB database, Umami-MRNN model dataset, VirtuousUmami model dataset, and PubMed and Google Scholar literature libraries; Construct a five-dimensional vector representation system for taste characteristics, and encode them in the order of the five basic tastes: sour, sweet, bitter, salty, and umami, where "1" indicates that the taste characteristic exists, "0" indicates that the taste characteristic does not exist, and "x" indicates that the taste characteristic is undetermined; One-hot encoding method was used to numerically preprocess the amino acid sequence; Sequence design and optimization using deep generative models.

2. The method according to claim 1, characterized in that There are two training modes: Single feature mode: the training set only contains sequences that fully match the target taste feature vector; Feature combination mode: the training set contains sequences corresponding to all valid subsets of the target taste feature vector; Feature contrast learning mechanism: When the target feature vector contains "0", positive and negative sample pairs are automatically constructed for contrast learning.

3. The method according to claim 1, characterized in that Adopts the variational autoencoder (LA-VAE) architecture, including: Encoding network: One-dimensional convolution layer, number of feature maps 32, convolution kernel size 3; Fully connected mapping layer; 2000-dimensional Gaussian distribution parameter layer, output mean vector and logarithmic variance vector; Decoding network: Symmetric reverse mapping structure; Sequence probability distribution reconstruction layer; Network Regularization: Convolutional layer dropout regularization; The fully connected layer is L1 regularized with a coefficient of λ=0.

01.

4. The method according to claim 3, characterized in that Model training strategies include: Parameter optimization: Adam adaptive optimizer; Initial learning rate η=0.001; Objective function: Sequence reconstruction loss term: binary cross entropy; Latent space regularization term: KL divergence relative to the standard normal distribution; Sequence length adaptive normalization term; Training Monitoring: Status parameter records; Model checkpoint saving; Sequence generation monitoring; Distribution status visualization.

5. The method according to claim 1, characterized in that Hierarchical clustering systems include: Latent Space Sequence Screening: Standard mode: retain the top 25% sequences of the training manifold neighbor degree; Avoidance mode: Building a dual distance measurement framework; Sequence homology analysis: Improved Needleman-Wunsch global alignment algorithm; Scoring matrix: match score 2.0, mismatch penalty -1.0; Affine gap penalty: open penalty -0.5, extension penalty -0.1; Clustering optimization strategy: Similarity threshold 0.7; Cluster representative selection mechanism based on node centrality.

6. The method according to claim 5, characterized in that Taste signature avoidance patterns include: Double distance calculation: Average Euclidean distance of positive samples; Average Euclidean distance of negative samples; Nearest Neighbor Analysis: The number of nearest neighbors k=5; k nearest neighbor distance calculation of positive and negative sample sets; Sequence evaluation criteria: Distance difference metric = average Euclidean distance of positive samples - average Euclidean distance of negative samples; A sequence sorting mechanism based on a difference measure.

7. A system for implementing the method according to any one of claims 1 to 6, characterized in that: include: Data integration module: realize unified management of multi-source data; Feature encoding module: realizes the vectorized representation of taste features; Mode selection module: realizes the switching between single feature and feature combination mode; Deep learning module: realize sequence generation and optimization; Cluster analysis module: realize sequence screening and redundancy removal.

8. The system according to claim 7, characterized in that It has the following functions: The precise definition and combination of taste feature vectors; Parallel processing of multi-modal sequence generation tasks; Feature avoidance mechanism based on contrastive learning; Hierarchical analysis of sequence similarity.

9. The system according to claim 7, characterized in that Data processing capabilities include: Batch parallel processing of sequence data; generate visual analysis of the results; Automated evaluation of clustering results; Real-time monitoring of the training process.

10. The system according to claim 7, characterized in that System outputs include: A set of candidate peptide sequences for target taste characteristics; Sequence clustering analysis report; Sequence similarity network topology diagram; A set of metrics for evaluating model performance.

Citation Information

Patent Citations

  • Umami peptide as well as preparation method and application thereof

    CN113651869A

  • Method for constructing polypeptide molecule and electronic equipment

    CN114155909A

  • Intelligent screening and identifying method and system for microorganism source umami peptide

    CN117877588A

  • New bitter peptide screening method based on peptiomics and machine learning

    CN118841089A

  • Polypeptide immunocompetence prediction and generation method fusing polypeptide physicochemical properties, sequence characteristics and word vector embedding

    CN119068997A

Cited By

  • Umami peptide prediction method and system based on deep learning

    CN120766752A

  • Small-molecular sweet peptide DSTT and application thereof

    CN121135817A

  • A small molecule sweet peptide DSTT and its application

    CN121135817B