A bitterness peptide screening method based on peptidomics and machine learning

By constructing a bitter peptide dataset and combining it with the LightGBM algorithm model, we achieved efficient screening of bitter peptides, solving the problem of time-consuming and expensive methods in existing technologies, and discovering novel bitter peptides with application value.

CN118841089BActive Publication Date: 2026-06-19DALIAN INSTITUTE OF CHEMICAL PHYSICS CHINESE ACADEMY OF SCIENCES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN INSTITUTE OF CHEMICAL PHYSICS CHINESE ACADEMY OF SCIENCES
Filing Date
2023-04-24
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing methods for identifying bitter peptides are time-consuming and expensive, and there is limited room for improvement in the performance of sequence-based machine learning models, resulting in a lack of efficient methods for screening bitter peptides.

Method used

A dataset containing both bitter and non-bitter peptides was constructed. Bitter peptides were quantified using characteristic factors and combined with the Light Gradient Boosting Machine (LightGBM) algorithm model to develop a high-throughput screening method based on peptidomics and machine learning, thereby identifying novel bitter peptides.

Benefits of technology

Rapid and accurate screening of bitter peptides was achieved, and three novel bitter peptides were discovered, which have potential application value in the food and pharmaceutical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118841089B_ABST
    Figure CN118841089B_ABST
Patent Text Reader

Abstract

This invention relates to a novel method for screening bitter peptides based on peptidomics and machine learning. In this invention, we established a more comprehensive benchmark dataset of bitter and non-bitter peptides, created a new set of bitter peptide characteristic factors, and constructed a classification and prediction model based on a novel data processing method and the Light Gradient Boosting Machine (LightGBM) algorithm for screening bitter peptides. Simultaneously, by combining machine learning and peptidomics technologies, high-throughput identification and prediction of potential bitter peptides are achieved. This method has potential applications in the food and pharmaceutical fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a novel method for the identification, prediction, and verification of bitter peptides based on peptidomics technology and machine learning, as well as the discovery and application of novel bitter peptides. Background Technology

[0002] Bitter peptides are derived from protein breakdown, and their bitterness is primarily determined by their amino acid composition and properties. Besides affecting taste, some bitter peptides also possess various biological activities, such as angiotensin-converting enzyme inhibition, hypoglycemic activity, antioxidant activity, and gastrointestinal digestive regulation activity. Therefore, bitter peptides play an important role in drug development and nutritional research, and their identification, prediction, and validation methods have high research value.

[0003] While traditional laboratory methods are considered reliable for identifying bitter peptides, they are often both time-consuming and expensive. Therefore, it is essential to develop rapid and accurate identification methods for predicting potential bitter peptides. Several computational methods based on quantitative structure-activity relationship (QSAR) models and machine learning have been developed to predict bitter peptides, such as iBitter-SCM, iBitter-Fuse, BERT4Bitter, and iBitter-DRLF. Despite significant progress in this field, there is still considerable room for improvement in the performance of sequence-based machine learning models for bitter peptide identification. Summary of the Invention

[0004] The purpose of this invention is to provide a method for screening bitter peptides based on peptidomics and machine learning, and to discover three novel bitter peptides using this method. In this invention, we constructed a more comprehensive benchmark dataset to develop a classification and prediction method for bitter peptides based on novel data processing methods and the Light Gradient Boosting Machine (LightGBM) algorithm. Simultaneously, by combining the classification and prediction model with peptidomics technology, a new high-throughput screening method for potential bitter peptides was developed.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for screening bitter peptides based on peptidomics and machine learning includes the following steps:

[0007] A dataset containing bitter peptides and non-bitter peptides was constructed, and the bitter peptide and non-bitter peptide datasets were quantified using bitter peptide feature factors;

[0008] The bitter peptides and non-bitter peptides, quantified by different combinations of bitter peptide feature factors, are input into the machine learning model for training. The model outputs the prediction results of bitter peptides, and the parameters of the machine learning model are adjusted based on the training results to obtain the optimal combination of bitter peptide feature factors.

[0009] The bitter peptides and non-bitter peptides in the test set are quantified according to the optimal combination of bitter peptide feature factors, and then input into the trained machine learning model to output the prediction results of bitter peptides.

[0010] The peptide to be tested is quantified using bitter peptide characteristic factors and input into a trained machine learning model to obtain the prediction results of bitter peptides.

[0011] The bitter peptides and non-bitter peptides, quantified from the different combinations of bitter peptide characteristic factors, are obtained through the following steps:

[0012] A separability experiment was conducted on the characteristic factors of bitter peptides to screen for combinations of bitter peptide characteristic factors that can more significantly distinguish between bitter peptides and non-bitter peptides. The datasets of bitter peptides and non-bitter peptides were then quantified according to the selected combinations of bitter peptide characteristic factors.

[0013] The quantification of bitter peptide and non-bitter peptide datasets based on a combination of bitter peptide characteristic factors includes the following steps:

[0014] The bitter peptide characteristic factors include:

[0015] Q is used to represent the average hydrophobicity of a polypeptide;

[0016] Q1 is used to represent the percentage of amino acids in a polypeptide with a Q value < 0;

[0017] Q2 is used to represent the percentage of amino acids in a polypeptide with a Q value between 0 and 1000;

[0018] Q3 is used to represent the percentage of amino acids in a polypeptide with a Q value between 1000 and 2000;

[0019] Q4 is used to represent the percentage of amino acids in a polypeptide with a Q value between 2000 and 3000;

[0020] AH is used to represent the average hydrophobicity of the peptide obtained from the amino acid descriptor COWR900101 in the reference database AAindex.

[0021] N is used to represent the hydrophobicity of the N-terminal amino acid of a polypeptide;

[0022] C is used to represent the hydrophobicity of the C-terminal amino acid of a polypeptide;

[0023] Percentage-BAA is used to express the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in a polypeptide.

[0024] N-basic AA is used to indicate whether the N-terminal amino acid of a polypeptide is a basic amino acid;

[0025] LFIYWV-C is used to indicate whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine.

[0026] Percentage-FWY is used to express the percentage of phenylalanine, tryptophan, and tyrosine in a polypeptide;

[0027] PXC is used to indicate whether proline is located at the second position of the C-terminus of a polypeptide;

[0028] RP is used to indicate whether there are adjacent RPs in a polypeptide;

[0029] Any combination of multiple different bitter peptide characteristic factors can be used as input for a machine learning model.

[0030] The different combinations of bitter peptide characteristic factors quantify the bitter peptides and non-bitter peptides, and the combination of bitter peptide characteristic factors is one of the following:

[0031] 1) Q is used to represent the average hydrophobicity of a peptide; Q1 represents the percentage of amino acids with a Q value < 0; Q2 represents the percentage of amino acids with a Q value between 0 and 1000; Q3 represents the percentage of amino acids with a Q value between 1000 and 2000; Q4 represents the percentage of amino acids with a Q value between 2000 and 3000; AH represents the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA represents the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; N-basic AA indicates whether the N-terminal amino acid of a peptide is a basic amino acid; LFIYWV-C indicates whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY indicates the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC indicates whether proline is located at the second position at the C-terminus of the peptide.

[0032] 2) Q represents the average hydrophobicity of the peptide; Q2 represents the percentage of amino acids with Q values ​​between 0 and 1000; Q3 represents the percentage of amino acids with Q values ​​between 1000 and 2000; Q4 represents the percentage of amino acids with Q values ​​between 2000 and 3000; AH represents the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA represents the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; N-basic AA indicates whether the N-terminal amino acid of a peptide is a basic amino acid; LFIYWV-C indicates whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY indicates the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC indicates whether proline is located at the second position at the C-terminus of the peptide.

[0033] 3) Q represents the average hydrophobicity of the peptide; Q2 represents the percentage of amino acids with Q values ​​between 0 and 1000; Q4 represents the percentage of amino acids with Q values ​​between 2000 and 3000; AH represents the average hydrophobicity of the peptide; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide obtained from the amino acid descriptor COWR900101 in the AAindex database; Percentage-BAA represents the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; N-basic AA indicates whether the N-terminal amino acid of a peptide is a basic amino acid; LFIYWV-C indicates whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY indicates the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC indicates whether proline is located at the second position at the C-terminus of the peptide.

[0034] 4) Q is used to represent the average hydrophobicity of the peptide; Q2 is used to represent the percentage of amino acids in the peptide with a Q value between 0 and 1000; Q4 is used to represent the percentage of amino acids in the peptide with a Q value between 2000 and 3000; AH is used to represent the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N is used to represent the hydrophobicity of the N-terminal amino acid of the peptide; C is used to represent the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA is used to represent the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; LFIYWV-C is used to indicate whether the amino acid located at the C-terminus of the peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY is used to represent the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC is used to indicate whether proline is located at the second position at the C-terminus of the peptide.

[0035] 5) Q is used to represent the average hydrophobicity of the peptide; Q2 is used to represent the percentage of amino acids in the peptide with a Q value between 0 and 1000; Q4 is used to represent the percentage of amino acids in the peptide with a Q value between 2000 and 3000; AH is used to represent the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N is used to represent the hydrophobicity of the N-terminal amino acid of the peptide; C is used to represent the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA is used to represent the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; LFIYWV-C is used to indicate whether the amino acid located at the C-terminus of the peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY is used to represent the percentage of phenylalanine, tryptophan, and tyrosine in the peptide.

[0036] The prediction results for the bitter peptides include:

[0037] 1 indicates bitter peptide;

[0038] 0 indicates a non-bitter peptide.

[0039] The process involves quantifying bitter peptides and non-bitter peptides in the test set according to the optimal combination of bitter peptide feature factors, inputting this quantification into the trained machine learning model, outputting the prediction results of bitter peptides, and then obtaining the evaluation index of the bitter peptide classification prediction model based on the prediction results to validate the machine learning model.

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046] Wherein, ACC, PRE, SN, SP, F1, and MCC represent the accuracy, precision, sensitivity, specificity, F1 score, and Matthews correlation coefficient of the machine learning model, respectively; TP and TN represent the number of correctly predicted true bitter peptides and true non-bitter peptides, respectively; FP represents the number of non-bitter peptides predicted as bitter peptides; and FN represents the number of bitter peptides predicted as non-bitter peptides.

[0047] A bitter peptide screening system based on peptidomics and machine learning, comprising:

[0048] The dataset construction module is used to construct datasets containing bitter peptides and non-bitter peptides, and to quantify the bitter peptide and non-bitter peptide datasets using bitter peptide feature factors.

[0049] The model training module is used to input bitter peptides and non-bitter peptides, which are quantified according to different combinations of bitter peptide feature factors, into the machine learning model for training, output the prediction results of bitter peptides, and adjust the parameters of the machine learning model based on the training results to obtain the optimal combination of bitter peptide feature factors; the module quantifies bitter peptides and non-bitter peptides in the test set according to the optimal combination of bitter peptide feature factors, inputs them into the trained machine learning model, and outputs the prediction results of bitter peptides.

[0050] The bitter peptide screening module is used to quantify the target peptides using bitter peptide characteristic factors, input them into the trained machine learning model, and obtain the prediction results of bitter peptides.

[0051] A bitter peptide, wherein the bitter peptide is FFVAPFPEVFGKE, FALPQYL, or EMPFPKYP.

[0052] An application of the aforementioned bitter peptide, wherein one or more of the bitter peptides are used as excipients in food and / or pharmaceuticals.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] This invention constructs a new benchmark dataset for bitter peptides and non-bitter peptides, a new set of characteristic factors for bitter peptides, and develops a classification prediction model for predicting bitter peptides. Combining the prediction model with peptidomics enables a novel high-throughput screening method for bitter peptides. The three newly discovered bitter peptides have potential applications in the food and pharmaceutical industries. Attached Figure Description

[0055] Appendix Figure 1 A roadmap for building classification prediction models based on novel data processing methods and the LightGBM algorithm.

[0056] Appendix Figure 2 The figure shows the results of the Euclidean distance-based verification of the separability of bitter peptide characteristic factors.

[0057] Appendix Figure 3 The figure shows the effect of four potential bitter peptides on calcium ion mobilization in HEK293T cells expressing the human bitter taste receptor T2R4.

[0058] Appendix 1 is the benchmark dataset for bitter peptides;

[0059] Appendix 2 is the benchmark dataset for non-bitter peptides;

[0060] Appendix 3 shows the amino acid values ​​used to calculate the Q value of the polypeptide;

[0061] Appendix Table 4 shows the amino acid values ​​used in descriptor COWR900101 to calculate the average hydrophobicity of peptides. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0063] This invention provides a high-throughput method for predicting and screening potential bitter peptides by combining peptidomics technology with a classification prediction model based on the LightGBM algorithm. The construction of this classification prediction model includes establishing a benchmark dataset, selecting characteristic factors of bitter peptides, and using these characteristic factors to construct a LightGBM algorithm model for classification prediction. The method construction process is detailed in the appendix. Figure 1 .

[0064] Example 1: Construction of a classification prediction model based on a novel data processing method and the LightGBM algorithm

[0065] 1.1 Construction of the benchmark dataset

[0066] The benchmark dataset used in this study is based on BTP640 and expanded by collecting data from Biopep, BitterDB, the flavor database developed by the Center for Innovative Food Taste Perception at Shanghai Jiao Tong University, and published literature. This dataset contains 720 peptide sequences, including 360 bitter peptides (Appendix 1) and 360 non-bitter peptides (Appendix 2), randomly divided into training and independent test datasets at an 8:2 ratio. The training dataset includes 288 bitter peptides and 288 non-bitter peptides, while the independent test dataset includes 72 bitter peptides and 72 non-bitter peptides.

[0067] 1.2 Selection of Characteristic Factors for Bitter Peptides

[0068] Bitter peptides exhibit certain special tendencies in their amino acid properties, particularly the hydrophilicity and hydrophobicity of amino acids, as well as the composition or position of hydrophobic amino acids, which can be used to distinguish bitter peptides from non-bitter peptides. For example, the average hydrophilicity of a peptide, usually defined as Q, is an important indicator for assessing the bitterness of a peptide. Therefore, this study used the Q value and a series of factors derived from the Q value. Based on their corresponding values ​​(Appendix Table 3), amino acids were divided into four fractions (<0, 0~1000, 1000~2000, and 2000~3000). Q1, Q2, Q3, and Q4 represent the percentage of amino acids in each fraction of the peptide, respectively. The values ​​of amino acids in the descriptor COWR900101 in the AAindex database (Appendix Table 4) were also used to assess the average hydrophilicity of amino acids. Furthermore, alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan were defined as bitter amino acids. Based on published literature and relevant databases, we innovatively proposed a set of potential bitterness characteristic factors (Table 1). These feature factors are used to quantify bitter and non-bitter peptides in the benchmark dataset.

[0069] Table 1. Set of characteristic factors of bitter peptides.

[0070]

[0071] These characteristic factors are manually removed in subsequent steps.

[0072] This feature factor was automatically eliminated by the subsequent LightGBM algorithm.

[0073] Separability verification based on Euclidean distance is a common method for verifying data separability. Two data points... and , respectively representing the characteristic factors of two different bitter peptides; , Then the Euclidean distance Defined as:

[0074]

[0075] According to the definition of Euclidean distance, the shorter the distance between two data points, the higher their similarity. Therefore, Euclidean distance can be used to determine whether the characteristic factors of bitter peptides can be accurately classified (see appendix). Figure 2 Based on the separability verification results, some feature factors are similar between bitter peptides and non-bitter peptides in the dataset, meaning that these factors cannot distinguish between bitter peptides and non-bitter peptides very well. Therefore, we progressively eliminated the five most similar factors to find the optimal combination of feature factors.

[0076] 1.3 LightGBM Model Construction

[0077] In this study, the LightGBM algorithm was used for model construction. Based on LightGBM feature selection, training datasets quantized with different feature factor subsets were input into the LightGBM classifier, and the optimal subset for predicting bitter peptides was found by continuously optimizing hyperparameters. The optimal parameters of the model were selected through 10-fold cross-validation and grid search. The five different feature factor sets used for model training are shown in Table 2.

[0078] Table 2 Five different sets of feature factors

[0079]

[0080] The outputs of both the model training and screening phases are predictions of bitter peptides, including: 1 for bitter peptides and 0 for non-bitter peptides.

[0081] Based on the predicted values ​​of bitter peptides (1 represents bitter peptides, 0 represents non-bitter peptides) and the actual values ​​of bitter peptides, the following parameters can be calculated for performance evaluation:

[0082] The correct predicted number of true bitter peptides (TP), the number of true non-bitter peptides (TN), the number of non-bitter peptides predicted as bitter peptides (FP), and the number of bitter peptides predicted as non-bitter peptides (FN).

[0083] 1.4 Performance Evaluation

[0084] To evaluate the predictive power of the model, we used the following six widely used metrics to solve the binary classification prediction problem (Equations 1-6):

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] In this context, ACC, PRE, SN, SP, F1, and MCC represent accuracy, precision, sensitivity, specificity, F1 score, and Matthews correlation coefficient, respectively. Precision and sensitivity are conflicting measures, and the F1 score is used to comprehensively evaluate these two metrics. MCC is commonly used as a measure of binary and multi-class classification quality and is generally considered a balanced measure. TP and TN represent the number of correctly predicted true bitter peptides and true non-bitter peptides, respectively. Meanwhile, FP represents the number of non-bitter peptides predicted as bitter peptides, and FN represents the number of bitter peptides predicted as non-bitter peptides. The area under the receiver operating characteristic (ROC) curve (AUC) is used to evaluate predictive performance, where AUC values ​​of 0.5 and 1 represent a randomized model and a perfect model, respectively.

[0092] Table 3 shows the results of 10-fold cross-validation and independent testing, with the highest values ​​for each indicator indicated by bold and underline. Based on the cross-validation results, the highest ACC, PRE, SN, SP, F1, MCC, and AUC values ​​were observed to be 84.9%, 85.4%, 84.9%, 84.9%, 84.8%, 69.8%, and 86.8%, respectively. After manual and automatic removal by the algorithm, the optimal LightGBM model was finally constructed using 10 feature factors (Q, Q2, Q4, AH, N, C, Percentage-BAA, N-basic AA, Percentage-FWY, PXC) to predict bitter peptides. For the independent test set, Table 2 shows that the overall prediction performance was relatively better than the 10-fold cross-validation results, with the highest ACC, PRE, SN, SP, F1, MCC, and AUC values ​​being 90.3%, 98.3%, 81.9%, 98.6%, 89.4%, 81.6%, and 90.5%, respectively.

[0093] Table 3 shows the cross-validation and independent testing results after progressively removing the five lowest-contribution eigenfactors.

[0094]

[0095] Example 2: Identification and Screening of Potential Bitter Peptides in Spoiled Milk

[0096] Peptidomics based on RPLC-MS / MS was used to analyze the peptide profiles of fresh UHT milk (FM) and spoiled UHT milk (RM) samples after centrifugation, ultrafiltration, and RPLC-MS / MS analysis. Differential peptides in RM compared to FM were screened, as these peptides may be the cause of the bitter taste in RM.

[0097] 2.1 Sample Preparation

[0098] FM and RM samples were stored at 4°C for 1 hour, and then at 15000× g Centrifuge at 4°C for 20 minutes to remove fat and precipitate. The supernatant containing peptides is then centrifuged in 10-kDa ultrafiltration centrifuge tubes at 4000× g The peptides were obtained by ultrafiltration at 20°C for 40 minutes. The extracted peptides were lyophilized and stored at -80°C. Each sample was tested five times.

[0099] A commercially available C18 solid-phase extraction column (10 mg) was activated and equilibrated using methanol and 0.1% TFA (v / v), respectively. 1 mg of lyophilized peptide was redissolved in 0.1% TFA (v / v) and loaded onto the extraction column. Desalting was performed by washing twice with 200 μL of 0.1% TFA (v / v). The peptide was eluted with 1.5 mL of 80% acetonitrile (ACN, v / v) containing 0.1% TFA (v / v). The purified peptide was collected, lyophilized, and stored at -80°C.

[0100] 2.2 RPLC-MS / MS Analysis

[0101] Frozen samples were reconstituted in 0.1% (v / v) FA-H2O solution and analyzed using an UltiMate 3000 RSLCnano liquid chromatography system in tandem with an LTQ-OrbitrapElite mass spectrometer. Samples were loaded at 1 μg increments. Mobile phase A in the liquid chromatography was an aqueous solution (v / v) containing 0.1% FA, and mobile phase B was 80% ACN (v / v) containing 0.1% FA. The gradient elution program was as follows: 0–2 min, 2%–8% B phase (v / v); 2–102 min, 8%–45% B phase; 102–105 min, 45%–95% B phase.

[0102] Mass spectrometry was performed in positive ion data-dependent (DDA) mode. The full MS scan range was set to m / z 350–2000. CID fragmentation was performed on the 20 most abundant precursors with the following parameters: minimum intensity 500, isolation width 2, normalized collision energy 35, and dynamic exclusion enabled (repetition count 1, repetition duration 30 s, exclusion duration 40 s).

[0103] 2.3 Data Retrieval

[0104] RAW files use MaxQuant TM The (v.1.5.3.30) software performed a search in the wheat database (wheat, 379 proteins). Search parameters were as follows: no standard quantification, no restriction enzyme digestion, no maximum missed digestion count, and no fixed modifications; the variable modification was set to methionine oxidation (+15.9949 Da). The mass tolerance for precursor ions was 20 ppm, and for fragment ions, it was 0.5 Da. Peptides with a PSM false positive rate (FDR) <1% were considered valid data for analysis.

[0105] 2.4 Data Analysis

[0106] 1280 and 1072 peptides were identified in FM and RM samples, respectively. Differential peptides in RM compared with FM were screened according to the following criteria: (1) peptides identified only in RM; (2) peptides identified in both FM and RM, but the peptide intensity in RM was at least 5 times higher than that in FM. A total of 724 differential peptides were screened according to the above criteria.

[0107] 1280 and 1072 peptides were identified in FM and RM samples, respectively. Differential peptides in RM compared with FM were screened according to the following criteria: (1) peptides identified only in RM; (2) peptides identified in both FM and RM, but the peptide intensity in RM was at least 5 times higher than that in FM. A total of 724 differential peptides were screened according to the above criteria.

[0108] The 724 differentially expressed peptides were quantified according to the optimal combination of bitter peptide characteristic factors (Q, Q2, Q4, AH, N, C, Percentage-BAA, N-basic AA, Percentage-FWY, PXC) selected in Example 1. The quantified differentially expressed peptides were then input into the machine learning model constructed in Example 1, which output bitter peptide prediction results: 1 indicates a bitter peptide; 0 indicates a non-bitter peptide. Statistical analysis of the model output showed that 180 of the 724 differentially expressed peptides were predicted to be bitter peptides. To further verify the accuracy of the model prediction, we selected one reported bitter peptide (FALPQYLK) and three new, unreported potential bitter peptides (FFVAPFPEVFGKE, FALPQYL, and EMPFPKYP) for a cellular calcium ion mobilization experiment, with FALPQYLK serving as a positive control. Four test peptides (FALPQYLK, FFVAPFPEVFGKE, FALPQYL and EMPFPKYP) were synthesized by Nanjing Jietai Biotechnology Co., Ltd. using a solid-phase method, with purities of 95.86%, 97.29%, 99.85% and 96.62%, respectively.

[0109] The specific experimental steps for the cell calcium ion assay are as follows: 1. Cell culture and transfection:

[0110] Human embryonic kidney cells (HEK293T cells) were used at a concentration of 1.0 × 10⁻⁶. 5 Cells were seeded at a density of 10 cells / well in poly-L-lysine-precoated 96-well plates and incubated in DMEM medium containing 10% FBS at 37°C and 5% CO2 for 24 hours. FLAG-T2R4 and Gα16 / 44-FLAG cells were transiently transfected using Lipofectamine 2000 (experimental group). HEK293T cells transfected only with Gα16 / 44-FLAG served as the negative control group.

[0111] 2. Calcium ion mobilization experiment

[0112] Cells in both the experimental and control groups were incubated at 37°C for 30 min with Fluo-4 AM dye containing probenecid (2.5 mM) and Pluronic F-127 (0.05%, w / v), followed by incubation at room temperature for 30 min. Cells were then treated with different concentrations (0.1 mM, 1 mM, and 5 mM) of potential bitter peptides. Calcium ion mobilization levels were measured at 525 nm (excited at 494 nm) using a microplate reader (Biotek Synergy H1, Winoosky, VT, USA). The response of the negative control group was used as a blank response, and the results were calculated by subtracting the blank response from the maximum response in the experimental group. Fluorescence intensity. Data from at least three independent experiments.

[0113] 3. Experimental Results

[0114] During a 300-second monitoring period, calcium ion mobilization was measured in the experimental groups (FFVAPFPEVFGKE, FALPQYL, and EMPFPKYP) and the control group (FALPQYLK), which were respectively incubated with four peptides. (See attached image.) Figure 3 The study showed the calcium mobilization of the four peptides at different concentrations (0.1 mM, 1.0 mM, and 5.0 mM) in the experimental and control groups. The positive control FALPQYLK showed the most significant change in calcium signaling, followed by FFVAPFPEVFGKE, FALPQYL, and EMPFPKYP. Peptides FALPQYLK, FALPQYL, and FFVAPFPEVFGKE all showed dose-dependent effects at all three different concentrations (0.1 mM, 1 mM, and 5 mM). Meanwhile, peptide EMPFPKYP... Fluorescence intensity also showed a dose-dependent effect at different concentrations (0.1 mM and 1.0 mM). The results indicated that FALPQYLK, FALPQYL, and FFVAPFPEVFGKE have a distinct bitter taste. Meanwhile, EMPFPKYP also exhibited a certain degree of bitterness at relatively low concentrations (0.1 and 1.0 mM). These results demonstrate the effectiveness of the classification and prediction model and further prove the feasibility of the entire method for identifying, predicting, and validating bitter peptides.

[0115] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

[0116] Appendix 1: Bitter Peptide Benchmark Dataset

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127] Appendix 2 Non-bitter peptide benchmark dataset

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138] Appendix 3: Amino acid values ​​used to calculate peptide Q value

[0139]

[0140] Appendix Table 4: Amino acid values ​​used in descriptor COWR900101 for calculating the average hydrophobicity of peptides

[0141]

Claims

1. A bitterness peptide screening method based on peptidomics and machine learning, characterized by, Includes the following steps: A dataset containing bitter peptides and non-bitter peptides was constructed, and the bitter peptide and non-bitter peptide datasets were quantified using bitter peptide feature factors; The bitter peptides and non-bitter peptides, quantified by different combinations of bitter peptide feature factors, are input into the machine learning model for training. The model outputs the prediction results of bitter peptides, and the parameters of the machine learning model are adjusted based on the training results to obtain the optimal combination of bitter peptide feature factors. The bitter peptides and non-bitter peptides in the test set are quantified according to the optimal combination of bitter peptide feature factors, and then input into the trained machine learning model to output the prediction results of bitter peptides. The peptide to be tested is quantified using bitter peptide characteristic factors and input into the trained machine learning model to obtain the prediction results of bitter peptides. The datasets of bitter peptides and non-bitter peptides are quantified based on a combination of bitter peptide characteristic factors, including the following steps: The bitter peptide characteristic factors include: Q is used to represent the average hydrophobicity of a polypeptide; Q1 is used to represent the percentage of amino acids in a polypeptide with a Q value < 0; Q2 is used to represent the percentage of amino acids in a polypeptide with a Q value between 0 and 1000; Q3 is used to represent the percentage of amino acids in a polypeptide with a Q value between 1000 and 2000; Q4 is used to represent the percentage of amino acids in a polypeptide with a Q value between 2000 and 3000; AH is used to represent the average hydrophobicity of the peptide obtained from the amino acid descriptor COWR900101 in the reference database AAindex. N is used to represent the hydrophobicity of the N-terminal amino acid of a polypeptide; C is used to represent the hydrophobicity of the C-terminal amino acid of a polypeptide; Percentage-BAA is used to express the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in a polypeptide. N-basic AA is used to indicate whether the N-terminal amino acid of a polypeptide is a basic amino acid; LFIYWV-C is used to indicate whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine. Percentage-FWY is used to express the percentage of phenylalanine, tryptophan, and tyrosine in a polypeptide; PXC is used to indicate whether proline is located at the second position of the C-terminus of a polypeptide; RP is used to indicate whether there are adjacent RPs in a polypeptide; Any combination of multiple different bitter peptide characteristic factors can be used as input for a machine learning model.

2. The method of claim 1, wherein the method is characterized by, The bitter peptides and non-bitter peptides, quantified from the different combinations of bitter peptide characteristic factors, are obtained through the following steps: A separability experiment was conducted on the characteristic factors of bitter peptides to screen for combinations of bitter peptide characteristic factors that can more significantly distinguish between bitter peptides and non-bitter peptides. The datasets of bitter peptides and non-bitter peptides were then quantified according to the selected combinations of bitter peptide characteristic factors.

3. The method according to claim 1 or 2, wherein the method is characterized by, The different combinations of bitter peptide characteristic factors quantify the bitter peptides and non-bitter peptides, and the combination of bitter peptide characteristic factors is one of the following: 1) Q is used to represent the average hydrophobicity of a peptide; Q1 represents the percentage of amino acids with a Q value < 0; Q2 represents the percentage of amino acids with a Q value between 0 and 1000; Q3 represents the percentage of amino acids with a Q value between 1000 and 2000; Q4 represents the percentage of amino acids with a Q value between 2000 and 3000; AH represents the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA represents the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; N-basic AA indicates whether the N-terminal amino acid of a peptide is a basic amino acid; LFIYWV-C indicates whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY indicates the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC indicates whether proline is located at the second position at the C-terminus of the peptide. 2) Q represents the average hydrophobicity of the peptide; Q2 represents the percentage of amino acids with Q values ​​between 0 and 1000; Q3 represents the percentage of amino acids with Q values ​​between 1000 and 2000; Q4 represents the percentage of amino acids with Q values ​​between 2000 and 3000; AH represents the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA represents the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; N-basic AA indicates whether the N-terminal amino acid of a peptide is a basic amino acid; LFIYWV-C indicates whether the amino acid at the C-terminus of a peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY indicates the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC indicates whether proline is located at the second position at the C-terminus of the peptide. 3) Q represents the average hydrophobicity of the peptide; Q2 represents the percentage of amino acids in the peptide with Q values ​​between 0 and 1000; Q4 represents the percentage of amino acids in the peptide with Q values ​​between 2000 and 3000; AH represents the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA represents the hydrophobicity of alanine and phenylalanine in the peptide. Percentages of amino acids, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan; N-basicAA indicates whether the N-terminal amino acid of the peptide is a basic amino acid; LFIYWV-C indicates whether the amino acid at the C-terminus of the peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY indicates the percentages of phenylalanine, tryptophan, and tyrosine in the peptide; PXC indicates whether proline is located at the second position at the C-terminus of the peptide. 4) Q represents the average hydrophobicity of the peptide; Q2 represents the percentage of amino acids with Q values ​​between 0 and 1000; Q4 represents the percentage of amino acids with Q values ​​between 2000 and 3000; AH represents the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N represents the hydrophobicity of the N-terminal amino acid of the peptide; C represents the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA represents the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; LFIYWV-C represents whether the amino acid at the C-terminus of the peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY represents the percentage of phenylalanine, tryptophan, and tyrosine in the peptide; PXC represents whether proline is located at the second position at the C-terminus of the peptide. 5) Q is used to represent the average hydrophobicity of the peptide; Q2 is used to represent the percentage of amino acids in the peptide with a Q value between 0 and 1000; Q4 is used to represent the percentage of amino acids in the peptide with a Q value between 2000 and 3000; AH is used to represent the average hydrophobicity of the peptide obtained by referring to the amino acid descriptor COWR900101 in the AAindex database; N is used to represent the hydrophobicity of the N-terminal amino acid of the peptide; C is used to represent the hydrophobicity of the C-terminal amino acid of the peptide; Percentage-BAA is used to represent the percentage of alanine, phenylalanine, glycine, isoleucine, leucine, methionine, proline, valine, tyrosine, and tryptophan in the peptide; LFIYWV-C is used to indicate whether the amino acid located at the C-terminus of the peptide is leucine, phenylalanine, isoleucine, tyrosine, tryptophan, or valine; Percentage-FWY is used to represent the percentage of phenylalanine, tryptophan, and tyrosine in the peptide.

4. The method of claim 1, wherein the method is characterized by, The prediction results for the bitter peptides include: 1 represents a bitter peptide; 0 represents a non-bitter peptide.

5. The method of claim 1, wherein the method is characterized by, The process involves quantifying bitter peptides and non-bitter peptides in the test set according to the optimal combination of bitter peptide feature factors, inputting this quantification into the trained machine learning model, outputting the prediction results of bitter peptides, and then obtaining the evaluation index of the bitter peptide classification prediction model based on the prediction results to validate the machine learning model. ; ; ; ; ; ; Wherein, ACC, PRE, SN, SP, F1, and MCC represent the accuracy, precision, sensitivity, specificity, F1 score, and Matthews correlation coefficient of the machine learning model, respectively; TP and TN represent the number of correctly predicted true bitter peptides and true non-bitter peptides, respectively; FP represents the number of non-bitter peptides predicted as bitter peptides; and FN represents the number of bitter peptides predicted as non-bitter peptides.

6. A bitterness peptide screening system based on peptidomics and machine learning, the system being used to implement a bitterness peptide screening method based on peptidomics and machine learning according to any one of claims 1-5, characterized in that, include: The dataset construction module is used to construct datasets containing bitter peptides and non-bitter peptides, and to quantify the bitter peptide and non-bitter peptide datasets using bitter peptide feature factors. The model training module is used to input the quantified bitter peptides and non-bitter peptides with different combinations of bitter peptide feature factors into the machine learning model for training, output the prediction results of bitter peptides, and adjust the parameters of the machine learning model through the feedback of the training results to obtain the optimal combination of bitter peptide feature factors. The bitter peptides and non-bitter peptides in the test set are quantified according to the optimal combination of bitter peptide feature factors, and then input into the trained machine learning model to output the prediction results of bitter peptides. The bitter peptide screening module is used to quantify the target peptides using bitter peptide characteristic factors, input them into the trained machine learning model, and obtain the prediction results of bitter peptides.

7. A bitter tasting peptide, characterized in that The bitter peptides are FFVAPFPEVFGKE, FALPQYL, or EMPFPKYP.

8. Use of a bitter peptide according to claim 7, characterized in that: One or more of the bitter peptides are used as excipients in food and / or pharmaceuticals.