Tea processing parameter regulation method based on metabolomics and wgcna algorithm
By using metabolomics and the WGCNA algorithm, electronic sensory and non-volatile metabolite data of tea samples were obtained. A weighted gene co-expression network was constructed, metabolites were screened and clustered, and processing parameters were determined. This solved the problem of dynamic balance of tea flavor and enabled precise control of tea processing and optimal flavor effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUIZHOU UNIV
- Filing Date
- 2026-03-19
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes. Traditional methods rely heavily on experience and sensory judgment, lacking quantitative basis and making it difficult to accurately control tea flavor.
Using metabolomics and the WGCNA algorithm, electronic sensory data and non-volatile metabolite data of tea samples were obtained. Differential metabolites were screened through multidimensional statistical differential analysis, a weighted gene co-expression network was constructed, metabolites were clustered and module feature vectors were extracted, the dynamic trend of key metabolite content was analyzed, and target processing parameters were determined.
It achieves a dynamic balance between bitterness and sweetness in tea, improves the precise control of tea processing parameters, and ensures that tea achieves the best flavor at different stages.
Smart Images

Figure CN122490133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of process optimization technology, and in particular to a method for regulating tea processing parameters based on metabolomics and the WGCNA algorithm. Background Technology
[0002] As a globally important beverage, tea's flavor and quality directly impact consumers' sensory experience and economic value. The bitterness, sweetness, fullness, and aroma characteristics of tea are primarily determined by non-volatile metabolites such as catechins, amino acids, tea polyphenols, and soluble sugars, as well as volatile aroma compounds. Different processing steps, such as fixation, shaping, and aroma enhancement, cause dynamic changes in the metabolites within the tea leaves, thus affecting the taste and aroma of the tea infusion. Tea harvested in summer and autumn, in particular, tends to have relatively low amino acid content but high catechin and caffeine content, often resulting in a bitter and astringent tea with poor flavor harmony. Therefore, scientific and technological methods are urgently needed to optimize and control the flavor and quality of tea.
[0003] Currently, research on tea flavor formation relies heavily on sensory evaluation and chemical analysis methods. Sensory measurement techniques such as electronic tongue and electronic nose can quantitatively assess the taste attributes of tea, including acidity, sweetness, bitterness, astringency, and umami, as well as its aroma. High-performance liquid chromatography (HPLC) and gas chromatography-mass spectrometry (GC-MS) are widely used for the quantitative analysis of major chemical components in tea.
[0004] However, existing practices still have significant shortcomings. First, most studies focus only on a single processing stage or a single type of metabolite, making it difficult to reflect the overall dynamic evolution of metabolites during tea processing. Second, traditional statistical methods cannot fully analyze the complex network of relationships between metabolites and their systematic relationship with sensory flavor. Third, the optimization of processing parameters relies heavily on experience and sensory judgment, lacking quantitative evidence and making it difficult to accurately control the balance between bitterness and sweetness in tea. Therefore, how to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this application is to provide a method for regulating tea processing parameters based on metabolomics and the WGCNA algorithm, aiming to solve the technical problem of how to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes.
[0006] To achieve the above objectives, this application proposes a method for regulating tea processing parameters based on metabolomics and the WGCNA algorithm, the method comprising: Electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times were obtained; Multidimensional statistical difference analysis was performed on the non-volatile metabolite data to obtain a set of differentially expressed metabolites; The power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set are calculated using the WGCNA algorithm, and the power-weighted correlation coefficients are used as edge weights to construct a weighted gene co-expression network. The differentially expressed metabolites in the weighted gene co-expression network are clustered into different co-expression modules, and the feature vectors of the modules are extracted. Calculate the first Pearson correlation coefficient between the feature vector of the module and the preset flavor trait quantitative score, and screen out target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than the preset correlation threshold; Analyze the dynamic trend curves of the content of key metabolites within the target metabolic module, and determine the target processing parameters based on the extreme points of the dynamic trend curves.
[0007] In one embodiment, the step of calculating the power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set using the WGCNA algorithm, and constructing a weighted gene co-expression network using the power-weighted correlation coefficients as edge weights, includes: Calculate the second Pearson correlation coefficient between each pair of differential metabolites in the differential metabolite set to obtain the Pearson correlation coefficient matrix; The Pearson correlation coefficient matrix is subjected to power exponentiation based on a preset soft threshold to obtain the power-weighted correlation coefficient. The adjacency degree between each of the differential metabolite nodes is calculated based on the power-weighted correlation coefficient to obtain the adjacency matrix; The topological overlap index between each differential metabolite is calculated based on the adjacency matrix to obtain the topological overlap matrix; Using the topological overlap values in the aforementioned topological overlap matrix as edge weights, the corresponding metabolite nodes are connected to obtain a weighted gene co-expression network.
[0008] In one embodiment, the step of clustering the differentially expressed metabolites in the weighted gene co-expression network into different co-expression modules and extracting module feature vectors includes: The topological dissimilarity between each pair of the differential metabolites is calculated based on the topological overlap matrix to obtain the topological dissimilarity matrix; The topological dissimilarity matrix is used as a distance metric to perform hierarchical clustering on the differential metabolites, resulting in a hierarchical clustering dendrogram. The hierarchical clustering tree diagram is divided into multiple initial metabolic modules by performing topological identification and path cutting on the branches using a dynamic pruning tree algorithm. Calculate the eigenvalue correlation coefficient between the first principal components of each initial metabolic module, and merge the initial metabolic modules whose eigenvalue correlation coefficient is greater than a preset similarity threshold to obtain multiple co-expression modules; The first principal component of the metabolite expression matrix within each co-expression module is extracted using the principal component analysis algorithm, and the first principal component is used as the module feature vector.
[0009] In one embodiment, the step of performing topological identification and path cutting on the branches of the hierarchical clustering dendrogram using a dynamic pruning tree algorithm to divide it into multiple initial metabolic modules includes: The branch paths of the hierarchical clustering tree diagram are scanned to identify potential branch clusters that meet the preset minimum module size; Dynamic path cutting is performed on the potential branch clusters according to a preset shearing height threshold to obtain multiple independent topological branches; Module identifiers are assigned to the differential metabolites within each of the aforementioned topological branches, resulting in multiple initial metabolic modules.
[0010] In one embodiment, the step of performing multidimensional statistical difference analysis on the non-volatile metabolite data to screen and obtain a set of differentially expressed metabolites includes: Principal component analysis was performed on the non-volatile metabolite data to calculate the metabolic characteristic distribution map of tea samples at different processing stages; Based on the inter-group clustering information in the metabolic feature distribution map of the samples, an orthogonal partial least squares discriminant analysis model is constructed. The variable projection importance value, inter-group difference multiple, and significance check value of each metabolite were calculated using the orthogonal partial least squares discriminant analysis model. The non-volatile metabolite data are filtered based on the variable projection importance value, the inter-group difference multiple, and the significance check value to obtain a set of differential metabolites.
[0011] In one embodiment, the step of analyzing the dynamic trend curve of the content of key metabolites within the target metabolic module and determining the target processing parameters based on the extreme points of the dynamic trend curve includes: Calculate the third Pearson correlation coefficient between the content distribution of each differential metabolite within the target metabolic module and the sweetness quantification score; Flavor-enhancing amino acids or soluble sugars with a third Pearson correlation coefficient greater than or equal to a preset contribution threshold are designated as key metabolites. The content change trajectory of the key metabolite under different processing times was fitted to generate a dynamic trend curve of the content. Identify the x-axis positions where the expression level reaches a maximum or minimum value in the dynamic trend curve of the content, and determine the extreme points; The processing time corresponding to the extreme point is combined with the preset temperature to perform parameter mapping, and the mapping result is obtained; Based on the mapping results, target processing parameters are determined, including target blanching time, target stripping time, and target aroma enhancement time.
[0012] In one embodiment, the step of obtaining electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times includes: Send processing control commands to the tea processing equipment to obtain tea samples for a preset processing time; The tea sample was scanned for aroma and taste using an electronic tongue and electronic nose system to generate raw electronic sensory signals. The original electronic sensory signals are subjected to analog-to-digital conversion and feature dimensionality reduction to generate electronic sensory data; The tea samples were subjected to non-targeted metabolomics detection using liquid chromatography-mass spectrometry to extract the mass spectrometric characteristic peaks of non-volatile components; Peak extraction and normalization alignment were performed on the mass spectrometry characteristic peaks to obtain non-volatile metabolite data.
[0013] Furthermore, to achieve the above objectives, this application also proposes a tea processing parameter regulation device based on metabolomics and the WGCNA algorithm, the device comprising: The data acquisition module is used to acquire electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times; The difference analysis module is used to perform multidimensional statistical difference analysis on the non-volatile metabolite data and filter to obtain a set of differential metabolites; The network construction module is used to calculate the power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set using a weighted gene co-expression network analysis algorithm, and to construct a weighted gene co-expression network using the power-weighted correlation coefficients as edge weights. The clustering extraction module is used to cluster the differentially expressed metabolites in the weighted gene co-expression network into different co-expression modules and extract the module feature vectors. The correlation screening module is used to calculate the first Pearson correlation coefficient between the feature vector of the module and the preset flavor trait quantitative score, and to screen out target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than the preset correlation threshold. The parameter determination module is used to analyze the dynamic trend curve of the content of key metabolites in the target metabolic module, and determine the target processing parameters based on the extreme points of the dynamic trend curve.
[0014] Furthermore, to achieve the above objectives, this application also proposes a tea processing parameter regulation device based on metabolomics and the WGCNA algorithm. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm as described above.
[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the tea processing parameter regulation method based on metabolomics and WGCNA algorithm as described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm as described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: First, by acquiring electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times, the changes in taste, aroma, and composition of tea at each processing stage can be comprehensively recorded, providing a complete raw data foundation for subsequent analysis. Second, multidimensional statistical difference analysis is performed on the non-volatile metabolite data to screen and obtain a set of differentially expressed metabolites. This allows for the identification of key chemical components that change significantly under different processing conditions, thereby focusing on metabolites that have the most significant impact on tea flavor and improving analytical efficiency. Then, the power-weighted correlation coefficients between pairs of metabolites in the differentially expressed metabolite set are calculated using the WGCNA algorithm. These power-weighted correlation coefficients are then used as edge weights to construct a weighted gene co-expression network, revealing synergistic relationships between metabolites and helping to identify groups of metabolites that may jointly regulate flavor. Subsequently, the differentially expressed metabolites in the weighted gene co-expression network are clustered into different co-expression modules, and module feature vectors are extracted. This quantifies the overall expression trend within a module into a single vector, facilitating subsequent correlation analysis and improving the simplicity and reliability of data processing. Next, the first Pearson correlation coefficient between the feature vector of the calculation module and the preset flavor trait quantitative score is used to screen out target metabolic modules with an absolute value greater than the preset correlation threshold. This identifies metabolite modules highly correlated with flavor characteristics, improving the resolution of flavor formation mechanisms. Finally, the dynamic trend curves of key metabolite content within the target metabolic modules are analyzed, and target processing parameters are determined based on the extreme points of these curves. This allows for precise control of the processing technology, enabling tea to achieve optimal flavor at different processing stages. By quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes, a dynamic balance between bitterness and sweetness in tea is achieved. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating Example 1 of the tea processing parameter regulation method based on metabolomics and WGCNA algorithm of this application. Figure 2 This is a flowchart illustrating Example 2 of the tea processing parameter regulation method based on metabolomics and WGCNA algorithm of this application. Figure 3This is a schematic diagram of the module structure of the tea processing parameter regulation device based on metabolomics and WGCNA algorithm according to an embodiment of this application. Figure 4 This is a schematic diagram of the hardware operating environment of the tea processing parameter regulation method based on metabolomics and WGCNA algorithm in the embodiments of this application.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] It should be noted that the executing entity of this application embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or parameter control system capable of realizing the above functions. The following description uses a parameter control system as an example to illustrate this embodiment and the subsequent embodiments.
[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0026] Based on this, embodiments of this application provide a method for regulating tea processing parameters based on metabolomics and the WGCNA algorithm, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm of this application.
[0027] In this embodiment, the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm includes steps S10 to S60: Step S10: Obtain electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times; Step S20: Perform multidimensional statistical difference analysis on the non-volatile metabolite data to screen and obtain a set of differential metabolites; Step S30: Calculate the power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set using the WGCNA algorithm, and construct a weighted gene co-expression network using the power-weighted correlation coefficients as edge weights. Step S40: Cluster the differentially expressed metabolites in the weighted gene co-expression network into different co-expression modules and extract the module feature vectors; Step S50: Calculate the first Pearson correlation coefficient between the feature vector of the module and the preset flavor trait quantitative score, and screen out target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than the preset correlation threshold. Step S60: Analyze the dynamic trend curve of the content of key metabolites in the target metabolic module, and determine the target processing parameters based on the extreme points of the dynamic trend curve.
[0028] It should be noted that the preset processing time refers to different processing times determined based on experimental design or experience during tea processing, used to control the duration of key processes such as fixation, shaping, and aroma enhancement. Electronic sensory data refers to quantitative signal data on the taste and aroma of tea samples acquired through instruments such as E-tongue and E-nose, used to simulate human sensory perception of tea infusion. Non-volatile metabolite data refers to the detection data of non-volatile chemical components in tea, including catechins, amino acids, tea polyphenols, and soluble sugars. These components mainly affect the bitterness, astringency, sweetness, and body of tea. The differential metabolite set refers to the set of non-volatile metabolites with significant differences in content identified through statistical analysis among different processing conditions or treatment groups.
[0029] WGCNA (Weighted Gene Co-Expression Network Analysis) is an algorithm used to construct correlation networks between metabolites or genes and identify functional modules. The power-weighted correlation coefficient, obtained by exponentiation of the Pearson correlation coefficients between metabolites in WGCNA, is used as the edge weights between metabolite nodes to reflect connection strength. A weighted gene co-expression network is a network where metabolites are nodes, and the weights of edges between nodes are determined by the power-weighted correlation coefficient, representing the overall co-expression relationship between metabolites. A co-expression module refers to a group of metabolites with similar expression patterns identified through cluster analysis in the weighted gene co-expression network, reflecting their potential synergistic effects or functional correlations. A module feature vector is a numerical vector representing the overall expression level of a co-expression module; typically, the first principal component of the metabolite expression matrix within the module is selected as the feature vector. Preset flavor trait quantification scoring refers to the quantitative scoring of tea in flavor traits such as sweetness, bitterness, and astringency based on sensory evaluation, serving as reference data for flavor attributes.
[0030] The first Pearson correlation coefficient refers to the degree of linear correlation between the module feature vector and the preset flavor trait quantitative score, calculated using the Pearson correlation coefficient. The preset correlation threshold is the absolute correlation coefficient threshold set when screening metabolic modules, used to determine whether the association between the module feature vector and the flavor score is significant. The target metabolic module refers to a co-expressed module whose absolute correlation with the flavor trait quantitative score is greater than the preset correlation threshold, and is considered to have a significant impact on tea flavor formation. Key metabolites refer to metabolites in the target metabolic module that contribute significantly to flavor traits, typically exhibiting significant correlation or biological function. The content dynamic trend curve is a curve recording the content changes of key metabolites under different processing times, used to show their dynamic evolution during processing. Target processing parameters refer to the optimized processing conditions determined based on the dynamic trend and extreme points of key metabolite content, including, for example, fixation time, shaping time, and aroma enhancement time, to achieve an ideal balance of tea flavor.
[0031] Understandably, the process involves several steps. First, the parameter control system processes tea leaves according to a preset processing time, collecting samples after each processing period. Each sample is then scanned for taste and aroma using an electronic sensory evaluation device, while quantitative data on non-volatile metabolites are obtained using a metabolomics analyzer. This systematically reflects changes in tea flavor and composition under different processing conditions. Second, the system performs multidimensional statistical difference analysis (e.g., principal component analysis or analysis of variance) on the collected non-volatile metabolite data to identify metabolites that show significant changes under different processing conditions. This allows for the selection of a set of differentially expressed metabolites, focusing on the chemical components that most significantly affect flavor. Finally, the system uses the WGCNA algorithm to calculate the power-weighted correlation coefficient between every two metabolites in the set of differentially expressed metabolites. This coefficient is then used as edge weights to construct a weighted gene co-expression network, capturing the overall synergistic relationships between metabolites and facilitating the discovery of metabolic modules that may collectively influence flavor.
[0032] Subsequently, the system clusters the constructed weighted gene co-expression network, dividing metabolites with similar expression patterns into different co-expression modules. A feature vector is extracted from each module to represent the overall trend of change within that module using a single vector, simplifying subsequent analysis. Next, the parameter control system calculates the first Pearson correlation coefficient between each module's feature vector and the tea sample's quantitative scores for preset flavor traits such as sweetness, bitterness, and astringency. By comparing the correlation coefficient with a preset correlation threshold, target metabolic modules are selected to identify the most critical metabolite groups for flavor formation. Finally, the system analyzes the content changes of key metabolites within the target metabolic modules over processing time, plots dynamic trend curves, identifies extreme points, and determines target processing parameters (e.g., optimal times for fixation or shaping) based on the processing time corresponding to these extreme points to optimize and balance tea flavor.
[0033] As an example, the step of performing multidimensional statistical difference analysis on the non-volatile metabolite data to screen and obtain a set of differentially expressed metabolites includes: performing principal component analysis on the non-volatile metabolite data to calculate the metabolic characteristic distribution map of tea samples at different processing stages; constructing an orthogonal partial least squares discriminant analysis model based on the inter-group clustering information in the sample metabolic characteristic distribution map; calculating the variable projection importance value, inter-group difference multiple, and significance check value of each metabolite using the orthogonal partial least squares discriminant analysis model; and screening the non-volatile metabolite data based on the variable projection importance value, the inter-group difference multiple, and the significance check value to obtain a set of differentially expressed metabolites.
[0034] It should be noted that the metabolic characteristic distribution map is a graph generated by projecting non-volatile metabolite data of tea samples at different processing stages into a low-dimensional space through principal component analysis. It is used to visualize the overall metabolic differences and trends among samples. Inter-group clustering information refers to identifying the similarities and differences between different processing conditions or treatment groups based on the relative positions or clustering results of tea samples in the metabolic characteristic distribution map. This is used to determine which samples belong to similar or dissimilar categories at the metabolic level. The orthogonal partial least squares discriminant analysis model is a statistical analysis model established using the orthogonal partial least squares method. It is used to distinguish different groups of tea samples in the metabolite space and quantify the contribution of each metabolite to the inter-group differences. The variable projection importance value is the contribution index of each metabolite in the model calculated by the orthogonal partial least squares discriminant analysis model, used to measure the importance of the metabolite in distinguishing different sample groups. The inter-group difference multiple refers to the ratio of the average content of the same metabolite in different processing groups, used to reflect the increase or decrease of the metabolite under different treatment conditions. The significance test value refers to the p-value or similar index obtained by statistically testing the differences between groups, and is used to determine whether the difference in the content of the metabolite is statistically significant.
[0035] Understandably, the parameter control system first performs principal component analysis on the collected non-volatile metabolite data. Specifically, it centers and standardizes the metabolite content matrix of each tea sample at different processing stages, then calculates the covariance matrix and extracts its eigenvalues and eigenvectors. The eigenvectors are then used to project the high-dimensional metabolite data into a low-dimensional space, generating a metabolic feature distribution map to visually display the overall metabolic differences and distribution trends among samples (e.g., observing whether samples with different processing times show significant clustering). Secondly, the system determines the inter-group differences based on the clustering shown in the metabolic feature distribution map and constructs an orthogonal partial least squares discriminant analysis model. Specifically, it uses the metabolite data of each processing group as the independent variable and the flavor group label as the dependent variable. The data matrix is decomposed using orthogonal partial least squares regression to extract variables related to inter-group differences while removing noise irrelevant to group classification. This forms a model that can effectively distinguish different processing groups, quantifying the contribution of each metabolite in different processing conditions and reducing interference.
[0036] Then, the parameter control system uses the constructed orthogonal partial least squares discriminant analysis model to calculate the variable projection importance value of each metabolite, measuring its overall explanatory power for the model. Simultaneously, it calculates the inter-group difference fold to reflect the magnitude of metabolite content changes among different processing groups and performs significance testing to obtain a significance check value, ensuring that the screened metabolites have statistically reliable differences. Finally, the system filters out metabolites with variable projection importance values greater than a preset importance threshold, inter-group difference folds exceeding preset increase or decrease thresholds, and significance check values less than a preset significance threshold. These eligible metabolites are then integrated to obtain a differential metabolite set. This ensures that subsequent analysis focuses on metabolites with the greatest impact on flavor, significant changes, and statistical reliability, thus providing a scientific basis for network analysis and processing parameter optimization.
[0037] As an example, the step of analyzing the dynamic trend curve of the content of key metabolites within the target metabolic module and determining the target processing parameters based on the extreme points of the content dynamic trend curve includes: calculating the third Pearson correlation coefficient between the content distribution of each differential metabolite within the target metabolic module and the sweetness quantification score; identifying flavor amino acids or soluble sugars with the third Pearson correlation coefficient greater than or equal to a preset contribution threshold as key metabolites; fitting the content change trajectory of the key metabolites under different processing times to generate a content dynamic trend curve; identifying the abscissa position of the expression level at a maximum or minimum value in the content dynamic trend curve to determine the extreme points; mapping the processing time corresponding to the extreme points with a preset temperature to obtain the mapping result; and determining the target processing parameters, including the target blanching time, the target stripping time, and the target aroma enhancement time, based on the mapping result.
[0038] It should be noted that content distribution refers to the arrangement of the quantitative content and its range of variation of each differential metabolite within the target metabolic module at different samples or processing time points during tea processing, reflecting the concentration fluctuations of metabolites during processing. Sweetness quantification score refers to the numerical value obtained by quantifying the sweetness intensity of tea samples through sensory evaluation or an electronic tongue system, used to measure the relative level of sweetness in tea under different processing conditions. The third Pearson correlation coefficient is the linear correlation coefficient calculated between the content distribution of each differential metabolite within the target metabolic module and the sweetness quantification score, used to assess the contribution of each metabolite to sweetness. The preset contribution threshold is the minimum correlation standard used to screen key metabolites; when the third Pearson correlation coefficient between a metabolite and sweetness is greater than or equal to this threshold, the metabolite is considered to contribute significantly to sweetness. Key metabolites refer to flavor amino acids or soluble sugars within the target metabolic module that are highly correlated with the sweetness quantification score and play an important role in tea flavor formation.
[0039] The content change trajectory refers to the curve plotted showing the quantitative content of key metabolites over time at different processing durations, used to display their dynamic evolution. The preset temperature refers to the temperature conditions pre-set for the fixation, shaping, or aroma enhancement processes during tea processing, used to map processing parameters to extreme points. The mapping result refers to the parameter information obtained by mapping the extreme point x-axis position of the content dynamic trend curve to the actual processing time and combining it with the preset temperature, used to guide specific processing operations. The target fixation time refers to the optimal fixation processing time determined by analyzing the dynamic trend curve and extreme points of key metabolites, used to achieve the ideal tea flavor. The target shaping time refers to the optimal processing time for the shaping process determined based on the extreme points of key metabolite content, used to regulate tea shape and flavor. The target aroma enhancement time refers to the optimal processing time for the aroma enhancement process determined by combining the extreme points of key metabolite content and the preset temperature, used to optimize tea aroma and overall sensory quality.
[0040] Understandably, the system first collects quantitative content data of each differential metabolite within the target metabolic module at different processing times. This data is then paired with the corresponding tea sample's sweetness quantitative score. A linear correlation analysis is performed between the content of each metabolite and the sweetness score, quantifying its contribution to sweetness by calculating the third Pearson correlation coefficient. This allows for the screening of metabolites with the most significant impact on sweetness. Secondly, the system compares all third Pearson correlation coefficients with a preset contribution threshold. Flavor amino acids or soluble sugars greater than or equal to this threshold are identified as key metabolites, and these metabolites are used for subsequent analysis. This ensures a focus on the components that truly affect sweetness. Finally, the system performs curve fitting on the content data of key metabolites at different processing times, generating a dynamic trend curve to visually display the pattern and rate of change in content with processing time (e.g., by smoothing curves or using polynomial fitting to process noisy data, making the trend clearer).
[0041] Subsequently, the system scans the expression values of key metabolites on the generated dynamic trend curve of content, identifies the x-axis locations of maximum or minimum values, and marks these locations as extreme points. This allows the system to determine the critical time points when metabolite content reaches its peak or trough during processing. Next, the system maps the processing time corresponding to these extreme points to a preset temperature, obtaining a mapping result that directly converts changes in metabolite content into operable processing parameters. Finally, the system comprehensively determines the target processing parameters based on the mapping results, including the target fixation time, target shaping time, and target aroma enhancement time, to ensure that appropriate treatments are applied at critical time points during processing, thereby optimizing tea flavor, achieving a balance between bitterness and sweetness, and improving the final sensory quality of the tea.
[0042] As an example, the step of obtaining electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times includes: sending processing control commands to tea processing equipment to obtain tea samples under preset processing times; using an electronic tongue and electronic nose system to scan the tea samples for aroma and taste to generate raw electronic sensory signals; performing analog-to-digital conversion and feature dimensionality reduction on the raw electronic sensory signals to generate electronic sensory data; using liquid chromatography-mass spectrometry to perform non-targeted metabolomics detection on the tea samples to extract mass spectrometry characteristic peaks of non-volatile components; and performing peak extraction and standardization alignment processing on the mass spectrometry characteristic peaks to obtain non-volatile metabolite data.
[0043] It should be noted that processing control commands refer to the operating signals or instructions sent by the parameter control system to the tea processing equipment. These commands control the start of processing steps, processing time, temperature, or other processing parameters to ensure that tea samples are obtained according to preset conditions. Tea processing equipment refers to the mechanical or automated devices used to perform tea processing steps, including withering machines, leaf-forming machines, and aroma-enhancing equipment, used to control the temperature, time, and mechanical processing of tea. Raw electronic sensory signals refer to the initial electrical signals or sensor response signals of the aroma and taste of tea samples collected by electronic tongue and electronic nose systems. These signals are unprocessed and used to reflect the sensory characteristics of tea. Liquid chromatography-mass spectrometry (LC-MS / MS) refers to an analytical instrument combining high-performance liquid chromatography and mass spectrometry (HPLC-MS / MS) for separating and qualitatively and quantitatively analyzing non-volatile chemical components in tea samples. Non-volatile components refer to chemical substances in tea that are not easily volatilized under processing conditions, such as catechins, amino acids, tea polyphenols, and soluble sugars. These components mainly affect the bitterness, astringency, sweetness, and body of the tea. Mass spectrometry characteristic peaks refer to the mass spectrometry signal peaks corresponding to each non-volatile component in liquid chromatography-mass spectrometry detection, reflecting its presence and relative abundance in the sample, and can be used for further quantitative analysis.
[0044] Understandably, firstly, the parameter control system sends processing control commands to the tea processing equipment. The equipment then executes tea processing according to preset processing times and conditions, such as initiating fixation, shaping, or aroma enhancement processes. Tea samples are acquired at each processing point to ensure consistent and repeatable processing conditions across different stages. Secondly, the system uses an electronic tongue and electronic nose system to scan the collected tea samples for aroma and taste. This involves preparing the tea samples into a tea infusion and placing it in a sensor measurement chamber. A sensor array detects taste and aroma response signals, generating raw electronic sensory signals. This quantifies the sensory characteristics of tea under different processing conditions.
[0045] Then, the system performs analog-to-digital conversion on the raw electronic sensory signals, converting analog voltage or current signals into digital signals. It then extracts key information features using feature reduction methods (such as principal component analysis or linear dimensionality reduction) to generate electronic sensory data suitable for subsequent analysis, thereby reducing redundant information and noise. Next, the system uses liquid chromatography-mass spectrometry (LC-MS / MS) to perform non-targeted metabolomics detection on tea samples. Non-volatile components in the tea samples are separated, and mass spectrometry is used to obtain the characteristic peaks of each component, reflecting their presence and relative abundance in the sample. This allows for a comprehensive capture of the dynamic changes in metabolites. Finally, the system performs peak extraction processing on the mass spectrometry characteristic peaks, identifying the peak position and peak area of each metabolite peak. Peaks from different samples are standardized and aligned to ensure accurate comparison of the same metabolite across different samples, thus obtaining non-volatile metabolite data suitable for subsequent statistical analysis and network construction.
[0046] This embodiment provides a method for regulating tea processing parameters based on metabolomics and the WGCNA algorithm. First, by acquiring electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times, the changes in taste, aroma, and composition of tea at each processing stage can be comprehensively recorded, providing a complete raw data foundation for subsequent analysis. Second, multidimensional statistical difference analysis is performed on the non-volatile metabolite data to screen and obtain a set of differentially expressed metabolites. This identifies key chemical components that change significantly under different processing conditions, thus focusing on metabolites that have the most significant impact on tea flavor and improving analysis efficiency. Then, the WGCNA algorithm is used to calculate the power-weighted correlation coefficients between pairs of metabolites in the set of differentially expressed metabolites. These power-weighted correlation coefficients are then used as edge weights to construct a weighted gene co-expression network, revealing synergistic relationships between metabolites and helping to identify groups of metabolites that may jointly regulate flavor. Subsequently, the differentially expressed metabolites in the weighted gene co-expression network are clustered into different co-expression modules, and module feature vectors are extracted. This quantifies the overall expression trend within a module into a single vector, facilitating subsequent correlation analysis and improving the simplicity and reliability of data processing. Next, the first Pearson correlation coefficient between the feature vector of the calculation module and the preset flavor trait quantitative score is used to screen out target metabolic modules with an absolute value greater than the preset correlation threshold. This identifies metabolite modules highly correlated with flavor characteristics, improving the resolution of flavor formation mechanisms. Finally, the dynamic trend curves of key metabolite content within the target metabolic modules are analyzed, and target processing parameters are determined based on the extreme points of these curves. This allows for precise control of the processing technology, enabling tea to achieve optimal flavor at different processing stages. By quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes, a dynamic balance between bitterness and sweetness in tea is achieved.
[0047] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm of this application. Step S30 of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm includes steps S31 to S35: Step S31: Calculate the second Pearson correlation coefficient between each pair of differential metabolites in the differential metabolite set to obtain the Pearson correlation coefficient matrix; Step S32: Perform power exponentiation on the Pearson correlation coefficient matrix according to a preset soft threshold to obtain the power-weighted correlation coefficient; Step S33: Calculate the adjacency degree between each differential metabolite node based on the power-weighted correlation coefficient to obtain the adjacency matrix; Step S34: Calculate the topological overlap index between each differential metabolite based on the adjacency matrix to obtain the topological overlap matrix; Step S35: Using the topological overlap values in the topological overlap matrix as edge weights, connect the corresponding metabolite nodes to obtain a weighted gene co-expression network.
[0048] It should be noted that the second Pearson correlation coefficient refers to the linear correlation coefficient calculated between the levels of each pair of metabolites in a differential metabolite set under different sample or processing conditions. It is used to quantify the strength and direction of their expression relationship. The Pearson correlation coefficient matrix is a matrix organized into pairs of second Pearson correlation coefficients for all metabolites in the differential metabolite set. Rows and columns correspond to metabolites, and matrix elements represent the degree of correlation between two metabolites, reflecting the overall synergistic relationship between metabolites. The preset soft threshold is an exponential parameter set when calculating the power-weighted correlation coefficient. Its value satisfies the scale-free topological distribution characteristics, smoothing the Pearson correlation coefficient into network edge weights, making the network more consistent with the complexity of biological systems.
[0049] Adjacency degree refers to the quantified value of the connection strength between each metabolite node and other metabolite nodes in a weighted network, reflecting the degree of direct association between nodes. The adjacency matrix is a matrix recording the adjacency degrees between differential metabolite nodes, where rows and columns correspond to metabolite nodes, and matrix elements represent the connection strength between nodes; it is the basic data for constructing a weighted network. The topological overlap index measures the degree to which two metabolite nodes share neighboring nodes in the network, reflecting their similarity and modularity, and is used to enhance the stability and robustness of the network. The topological overlap matrix is a matrix arranged by the topological overlap indices of all metabolite nodes pairwise, where rows and columns correspond to metabolite nodes, and matrix elements represent the degree of topological overlap between nodes; it is used for subsequent module identification and network analysis. A metabolite node is the basic unit in a weighted gene co-expression network; each node corresponds to a differential metabolite and is used to represent the metabolite and its relationship with other metabolites in the network.
[0050] Understandably, the system first calculates the content of each pair of differentially related metabolites in all samples or under different processing conditions, obtaining a second Pearson correlation coefficient. By statistically analyzing the linear correlation between each metabolite and other metabolites, a Pearson correlation coefficient matrix is generated. This matrix is then organized by row and column corresponding to the differentially related metabolites to clearly reflect the correlation strength and direction between each pair of metabolites (e.g., positive correlation indicates synchronous content changes, negative correlation indicates opposite content changes). Secondly, the system applies a power-law processing to each correlation coefficient in the Pearson correlation coefficient matrix based on a preset soft threshold, converting the values into power-weighted correlation coefficients. This strengthens highly correlated metabolite connections while weakening low-correlation noise relationships, enabling the constructed network to approximate a scale-free topology, which better reflects the actual characteristics of node connection distribution in biological systems.
[0051] Then, the adjacency degree between each differential metabolite node and other nodes is calculated based on the power-weighted correlation coefficient, quantifying the strength of direct connections between metabolite nodes. These adjacency degrees are then organized into an adjacency matrix to provide complete node connection information required for subsequent network analysis (e.g., higher element values in the adjacency matrix indicate stronger collaboration between nodes). Subsequently, the topological overlap index between each pair of metabolite nodes is calculated using the adjacency matrix to quantify the degree to which they share neighbor nodes, and the results are formed into a topological overlap matrix. This enhances the ability to identify modular features in the network structure, making it easier to cluster closely related metabolites into the same module. Finally, using the topological overlap values in the topological overlap matrix as edge weights, metabolite nodes are connected according to the node correspondence in the matrix to construct a weighted gene co-expression network, thus obtaining a complete network structure. This network not only reflects the direct connections between metabolites but also reveals their overall collaborative patterns, providing a reliable foundation for subsequent co-expression module partitioning and module feature extraction.
[0052] As an example, the step of clustering the differentially expressed metabolites in the weighted gene co-expression network into different co-expression modules and extracting module feature vectors includes: calculating the topological dissimilarity between each pair of differentially expressed metabolites based on the topological overlap matrix to obtain a topological dissimilarity matrix; using the topological dissimilarity matrix as a distance metric to perform hierarchical clustering on the differentially expressed metabolites to obtain a hierarchical clustering dendrogram; performing topological identification and path cutting on the branches of the hierarchical clustering dendrogram using a dynamic pruning tree algorithm to divide it into multiple initial metabolic modules; calculating the eigenvalue correlation coefficient between the first principal components of each initial metabolic module, and merging the initial metabolic modules whose eigenvalue correlation coefficient is greater than a preset similarity threshold to obtain multiple co-expression modules; extracting the first principal component of the metabolite expression matrix within each co-expression module using a principal component analysis algorithm, and using the first principal component as the module feature vector.
[0053] It's important to note that topological dissimilarity, in a weighted gene co-expression network, measures the degree of structural dissimilarity between two differentially expressed metabolite nodes. A higher value indicates more different connectivity relationships between the two metabolites. The topological dissimilarity matrix is a matrix formed by calculating the topological dissimilarity between all pairs of differentially expressed metabolite nodes. Rows and columns correspond to metabolite nodes, and matrix elements represent the topological differences between nodes, used for subsequent clustering analysis. The hierarchical clustering dendrogram is a tree-like structure generated after hierarchically clustering differentially expressed metabolites based on the topological dissimilarity matrix. It displays the topological association distances and grouping hierarchy between metabolites, intuitively reflecting the aggregation relationships of similar metabolites. The dynamic pruning tree algorithm scans and identifies branch structures in the hierarchical clustering dendrogram, using an adaptive pruning method to divide the dendrogram into independent modules to obtain initial metabolic modules, ensuring similar metabolite expression patterns within each module.
[0054] The initial metabolic module refers to the set of metabolites obtained by partitioning the hierarchical clustering dendrogram branches using a dynamic pruning tree algorithm. Each module contains differentially expressed metabolites with similar topological structures in the network, used for subsequent module merging and feature extraction. The first principal component (PPC) of the initial metabolic module refers to the first PPC of the expression matrix of all metabolites within the initial metabolic module, calculated through principal component analysis. This quantifies the overall expression pattern and trend of the module. The eigenvalue correlation coefficient refers to the degree of linear correlation between the first PPCs of different initial metabolic modules, used to determine the similarity of expression patterns between modules, facilitating the merging of similar modules. The preset similarity threshold is the minimum eigenvalue correlation coefficient threshold used to determine whether initial metabolic modules should be merged. When the correlation coefficient of two modules exceeds this threshold, they are merged into a single co-expression module. The first PPC of the metabolite expression matrix refers to the first PPC obtained after performing principal component analysis on the expression matrix of metabolites within each final co-expression module. This serves as the module feature vector, representing the overall expression level and trend of metabolites within the module.
[0055] Understandably, firstly, based on the topological overlap matrix, the topological dissimilarity between nodes of each pair of metabolites in the differential metabolite set is calculated. Specifically, this is achieved by comparing the number of shared neighbor nodes and network structure differences between each pair of metabolite nodes, converting these into numerical indicators, and then organizing all the calculation results into a topological dissimilarity matrix. This quantifies the differences in the connection patterns of metabolites in the network, providing a reliable distance metric for subsequent clustering, thus more accurately reflecting the functional associations between metabolites. Secondly, using the topological dissimilarity matrix as a distance metric, hierarchical clustering is performed on the differential metabolites. By calculating the distances between metabolites and progressively merging the most similar nodes, a hierarchical clustering dendrogram is generated. The dendrogram displays the topological association distances and hierarchical distribution relationships between metabolites, providing a basis for subsequent module partitioning (e.g., observing which metabolites are highly associated in the network).
[0056] Then, a dynamic pruning tree algorithm was used to scan and identify the branches of the hierarchical clustering dendrogram. An adaptive path cutting method was employed to divide highly similar branches in the dendrogram into multiple initial metabolic modules, while removing noisy nodes and excessively small branches to ensure that metabolites within each initial metabolic module exhibit a relatively consistent expression pattern, facilitating subsequent merging and feature extraction. Subsequently, the eigenvalue correlation coefficients between the first principal components of each initial metabolic module were calculated. By comparing the similarity of the first principal components between modules, initial metabolic modules with eigenvalue correlation coefficients exceeding a preset similarity threshold were merged, resulting in multiple final co-expression modules. This ensures uniform metabolite expression patterns within modules and significant differences between modules, improving the biological interpretability of the modules. Finally, principal component analysis was performed on the metabolite expression matrix within each co-expression module, extracting the first principal component as the module's feature vector to quantify the overall expression level and trend of the module, providing simplified and reliable representative data for subsequent correlation analysis between modules and flavor traits.
[0057] As an example, the step of performing topological identification and path cutting on the branches of the hierarchical clustering tree diagram using a dynamic pruning tree algorithm to divide it into multiple initial metabolic modules includes: scanning the branch paths of the hierarchical clustering tree diagram to identify potential branch clusters that meet a preset minimum module size; performing dynamic path cutting on the potential branch clusters according to a preset pruning height threshold to obtain multiple independent topological branches; and assigning module identifiers to the differential metabolites within each topological branch to divide it into multiple initial metabolic modules.
[0058] It should be noted that the preset minimum module size refers to the lower limit of the number of metabolites set in the dynamic pruning tree algorithm to ensure that the partitioned modules are sufficiently representative. A branch is considered a potential module candidate only when the number of metabolites it contains is greater than or equal to this value. A potential branch cluster refers to the set of metabolites that meet the minimum module size requirement identified after scanning the branch paths in the hierarchical clustering tree diagram. These branch clusters may become candidate regions for the initial metabolic modules. The preset pruning height threshold refers to the branch cutting criterion set in the dynamic pruning tree algorithm, used to determine at what height in the tree diagram to cut the path of the potential branch cluster to separate mutually independent topological branches. A topological branch refers to the independent subset of metabolites obtained after cutting the potential branch cluster by applying the pruning height threshold. The expression patterns of metabolites within each topological branch are relatively consistent and serve as the basic unit for partitioning the initial metabolic modules.
[0059] Understandably, the process involves several steps. First, each branch path in the hierarchical clustering dendrogram is scanned sequentially to calculate the number of differentially expressed metabolites on each branch. Potential branch clusters with a size greater than or equal to a preset minimum module size are identified. This ensures that each candidate module contains a sufficient number of metabolites to represent the overall expression characteristics of the module, while eliminating excessively small, isolated branches to reduce noise interference. Second, a preset shearing height threshold is applied to the identified potential branch clusters for dynamic path cutting. Specifically, the cutting point is adaptively determined based on the hierarchical height of each branch in the dendrogram, dividing the branch along its height into several independent sub-branches. This results in multiple topological branches, ensuring that the expression patterns of metabolites within each branch are as similar as possible, while minimizing confounding relationships between different modules (e.g., shearing points are chosen where branch height changes significantly or where the distance from the upstream node reaches a threshold). Finally, a unique module identifier is assigned to each differentially expressed metabolite within each topological branch, ensuring that metabolites within the same branch belong to the same initial metabolic module, facilitating subsequent analysis. Finally, after organizing all the topological branches, a complete set of initial metabolic modules is formed. Each module contains differentially expressed metabolites with similar expression patterns, thus providing a reliable foundation for subsequent feature vector extraction and correlation analysis between modules.
[0060] This embodiment first calculates the second Pearson correlation coefficient between the contents of each pair of differentially expressed metabolites under different sample or processing conditions in the differentially expressed metabolite set, obtaining a Pearson correlation coefficient matrix. This matrix quantifies the linear correlation between metabolites, enabling the identification of metabolite pairs with similar expression patterns and providing a foundation for subsequent network construction. Secondly, the Pearson correlation coefficient matrix is subjected to power-law processing based on a preset soft threshold, resulting in power-weighted correlation coefficients. By strengthening the weights of highly correlated metabolites and weakening the influence of low-correlation metabolites, the network structure better conforms to scale-free topology characteristics, improving the reliability of the network in reflecting metabolic synergies in biology. Then, the adjacency degree between nodes of each differentially expressed metabolite is calculated based on the power-weighted correlation coefficients, and these adjacency degrees are organized into an adjacency matrix to quantify the direct connection strength between nodes. This helps to accurately describe the relationships between metabolites in the network. Next, the topological overlap index between each differentially expressed metabolite is calculated based on the adjacency matrix, forming a topological overlap matrix. By measuring the degree of shared neighbors between nodes, the identification capability of the network's modular features can be enhanced, thereby more accurately discovering groups of metabolites with synergistic changes. Finally, the topological overlap values in the topological overlap matrix were used as edge weights to connect the corresponding metabolite nodes and construct a weighted gene co-expression network. This network not only reflects the direct connections between metabolites but also reveals the overall synergistic pattern among metabolites, providing a stable and reliable structural foundation for subsequent module partitioning and functional analysis.
[0061] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the tea processing parameter regulation method based on metabolomics and WGCNA algorithm in this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0062] This application also provides a tea processing parameter regulation device based on metabolomics and the WGCNA algorithm. Please refer to [reference needed]. Figure 3 The tea processing parameter regulation device based on metabolomics and the WGCNA algorithm includes: Data acquisition module 10 is used to acquire electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times; The difference analysis module 20 is used to perform multidimensional statistical difference analysis on the non-volatile metabolite data and filter to obtain a set of differential metabolites; Network construction module 30 is used to calculate the power-weighted correlation coefficient between each pair of metabolites in the differential metabolite set using a weighted gene co-expression network analysis algorithm, and to construct a weighted gene co-expression network using the power-weighted correlation coefficient as the edge weight. Clustering extraction module 40 is used to cluster the differential metabolites in the weighted gene co-expression network into different co-expression modules and extract the module feature vectors; The association screening module 50 is used to calculate the first Pearson correlation coefficient between the feature vector of the module and the preset flavor trait quantitative score, and to screen out target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than the preset correlation threshold. The parameter determination module 60 is used to analyze the dynamic trend curve of the content of key metabolites in the target metabolic module, and determine the target processing parameters based on the extreme points of the dynamic trend curve.
[0063] The tea processing parameter regulation device based on metabolomics and the WGCNA algorithm provided in this application, employing the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm in the above embodiments, can solve the technical problem of how to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes. Compared with the prior art, the beneficial effects of the tea processing parameter regulation device based on metabolomics and the WGCNA algorithm provided in this application are the same as those of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm provided in the above embodiments, and other technical features in the tea processing parameter regulation device based on metabolomics and the WGCNA algorithm are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0064] This application provides a tea processing parameter regulation device based on metabolomics and the WGCNA algorithm. The tea processing parameter regulation device based on metabolomics and the WGCNA algorithm includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm in the above embodiment 1.
[0065] The following is for reference. Figure 4 This document illustrates a structural schematic diagram of a tea processing parameter regulation device based on metabolomics and the WGCNA algorithm, suitable for implementing embodiments of this application. The tea processing parameter regulation device based on metabolomics and the WGCNA algorithm in this application embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Android Devices), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The tea processing parameter regulation device based on metabolomics and WGCNA algorithm shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0066] like Figure 4As shown, the tea processing parameter control device based on metabolomics and the WGCNA algorithm may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in ROM (Read Only Memory) 1002 or the program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the tea processing parameter control device based on metabolomics and the WGCNA algorithm. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the tea processing parameter control device based on metabolomics and WGCNA algorithms to exchange data with other devices wirelessly or via wired communication. Although the figure shows a tea processing parameter control device based on metabolomics and WGCNA algorithms with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0067] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0068] The tea processing parameter regulation device based on metabolomics and the WGCNA algorithm provided in this application, employing the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm in the above embodiments, can solve the technical problem of how to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes. Compared with the prior art, the beneficial effects of the tea processing parameter regulation device based on metabolomics and the WGCNA algorithm provided in this application are the same as those of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm provided in the above embodiments, and other technical features in this tea processing parameter regulation device based on metabolomics and the WGCNA algorithm are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0069] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0070] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0071] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the tea processing parameter regulation method based on metabolomics and WGCNA algorithm in the above embodiments.
[0072] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0073] The aforementioned computer-readable storage medium may be included in a tea processing parameter regulation device based on metabolomics and WGCNA algorithm; or it may exist independently and not assembled into a tea processing parameter regulation device based on metabolomics and WGCNA algorithm.
[0074] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a tea processing parameter control device based on metabolomics and the WGCNA algorithm, the device performs the following: acquires electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times; performs multidimensional statistical difference analysis on the non-volatile metabolite data to obtain a set of differential metabolites; calculates the power-weighted correlation coefficient between each pair of metabolites in the set of differential metabolites using the WGCNA algorithm, and constructs a weighted gene co-expression network using the power-weighted correlation coefficient as the edge weight; clusters the differential metabolites in the weighted gene co-expression network into different co-expression modules and extracts module feature vectors; calculates the first Pearson correlation coefficient between the module feature vector and a preset flavor trait quantitative score, and selects target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than a preset correlation threshold; analyzes the dynamic trend curve of the content of key metabolites in the target metabolic module, and determines the target processing parameters based on the extreme points of the dynamic trend curve.
[0075] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0077] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0078] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for regulating tea processing parameters based on metabolomics and the WGCNA algorithm. This method can solve the technical problem of how to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm provided in the above embodiments, and will not be repeated here.
[0079] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm as described above.
[0080] The computer program product provided in this application can solve the technical problem of how to achieve a dynamic balance between bitterness and sweetness in tea by quantifying the correlation between the dynamic evolution of metabolites and sensory flavor attributes. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the tea processing parameter regulation method based on metabolomics and WGCNA algorithm provided in the above embodiments, and will not be repeated here.
[0081] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A tea processing parameter regulation method based on metabolomics and WGCNA algorithm, characterized in that, The method includes: Electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times were obtained; Multidimensional statistical difference analysis was performed on the non-volatile metabolite data to obtain a set of differentially expressed metabolites; The power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set are calculated using the WGCNA algorithm, and the power-weighted correlation coefficients are used as edge weights to construct a weighted gene co-expression network. The differentially expressed metabolites in the weighted gene co-expression network are clustered into different co-expression modules, and the feature vectors of the modules are extracted. Calculate the first Pearson correlation coefficient between the feature vector of the module and the preset flavor trait quantitative score, and screen out target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than the preset correlation threshold; Analyze the dynamic trend curves of the content of key metabolites within the target metabolic module, and determine the target processing parameters based on the extreme points of the dynamic trend curves.
2. The method of claim 1, wherein, The step of calculating the power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set using the WGCNA algorithm, and constructing a weighted gene co-expression network using the power-weighted correlation coefficients as edge weights, includes: Calculate the second Pearson correlation coefficient between each pair of differential metabolites in the differential metabolite set to obtain the Pearson correlation coefficient matrix; The Pearson correlation coefficient matrix is subjected to power exponentiation based on a preset soft threshold to obtain the power-weighted correlation coefficient. The adjacency degree between each of the differential metabolite nodes is calculated based on the power-weighted correlation coefficient to obtain the adjacency matrix; The topological overlap index between each differential metabolite is calculated based on the adjacency matrix to obtain the topological overlap matrix; Using the topological overlap values in the aforementioned topological overlap matrix as edge weights, the corresponding metabolite nodes are connected to obtain a weighted gene co-expression network.
3. The method of claim 2, wherein, The step of clustering the differentially expressed metabolites in the weighted gene co-expression network into different co-expression modules and extracting the module feature vectors includes: The topological dissimilarity between each pair of the differential metabolites is calculated based on the topological overlap matrix to obtain the topological dissimilarity matrix; The topological dissimilarity matrix is used as a distance metric to perform hierarchical clustering on the differential metabolites, resulting in a hierarchical clustering dendrogram. The hierarchical clustering tree diagram is divided into multiple initial metabolic modules by performing topological identification and path cutting on the branches using a dynamic pruning tree algorithm. Calculate the eigenvalue correlation coefficient between the first principal components of each initial metabolic module, and merge the initial metabolic modules whose eigenvalue correlation coefficient is greater than a preset similarity threshold to obtain multiple co-expression modules; The first principal component of the metabolite expression matrix within each co-expression module is extracted using the principal component analysis algorithm, and the first principal component is used as the module feature vector.
4. The method of claim 3, wherein, The step of performing topological identification and path cutting on the branches of the hierarchical clustering dendrogram using a dynamic pruning tree algorithm to divide it into multiple initial metabolic modules includes: The branch paths of the hierarchical clustering tree diagram are scanned to identify potential branch clusters that meet the preset minimum module size; Dynamic path cutting is performed on the potential branch clusters according to a preset shearing height threshold to obtain multiple independent topological branches; Module identifiers are assigned to the differential metabolites within each of the aforementioned topological branches, resulting in multiple initial metabolic modules.
5. The method of claim 1, wherein, The step of performing multidimensional statistical difference analysis on the non-volatile metabolite data to screen and obtain a set of differentially expressed metabolites includes: Principal component analysis was performed on the non-volatile metabolite data to calculate the metabolic characteristic distribution map of tea samples at different processing stages; Based on the inter-group clustering information in the metabolic feature distribution map of the samples, an orthogonal partial least squares discriminant analysis model is constructed. The variable projection importance value, inter-group difference multiple, and significance check value of each metabolite were calculated using the orthogonal partial least squares discriminant analysis model. The non-volatile metabolite data are filtered based on the variable projection importance value, the inter-group difference multiple, and the significance check value to obtain a set of differential metabolites.
6. The method of claim 1, wherein, The step of analyzing the dynamic trend curve of the content of key metabolites within the target metabolic module and determining the target processing parameters based on the extreme points of the dynamic trend curve includes: Calculate the third Pearson correlation coefficient between the content distribution of each differential metabolite within the target metabolic module and the sweetness quantification score; Flavor-enhancing amino acids or soluble sugars with a third Pearson correlation coefficient greater than or equal to a preset contribution threshold are designated as key metabolites. The content change trajectory of the key metabolite under different processing times was fitted to generate a dynamic trend curve of the content. Identify the x-axis positions where the expression level reaches a maximum or minimum value in the dynamic trend curve of the content, and determine the extreme points; The processing time corresponding to the extreme point is combined with the preset temperature to perform parameter mapping, and the mapping result is obtained; Based on the mapping results, target processing parameters are determined, including target blanching time, target stripping time, and target aroma enhancement time.
7. The method according to any one of claims 1 to 6, characterized in that, The steps for obtaining electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times include: Send processing control commands to the tea processing equipment to obtain tea samples for a preset processing time; The tea sample was scanned for aroma and taste using an electronic tongue and electronic nose system to generate raw electronic sensory signals. The original electronic sensory signals are subjected to analog-to-digital conversion and feature dimensionality reduction to generate electronic sensory data; The tea samples were subjected to non-targeted metabolomics detection using liquid chromatography-mass spectrometry to extract the mass spectrometric characteristic peaks of non-volatile components; Peak extraction and normalization alignment were performed on the mass spectrometry characteristic peaks to obtain non-volatile metabolite data.
8. A tea processing parameter control device, characterized in that, The device includes: The data acquisition module is used to acquire electronic sensory data and non-volatile metabolite data of tea samples under different preset processing times; The difference analysis module is used to perform multidimensional statistical difference analysis on the non-volatile metabolite data and filter to obtain a set of differential metabolites; The network construction module is used to calculate the power-weighted correlation coefficients between pairs of metabolites in the differential metabolite set using a weighted gene co-expression network analysis algorithm, and to construct a weighted gene co-expression network using the power-weighted correlation coefficients as edge weights. The clustering extraction module is used to cluster the differentially expressed metabolites in the weighted gene co-expression network into different co-expression modules and extract the module feature vectors. The correlation screening module is used to calculate the first Pearson correlation coefficient between the feature vector of the module and the preset flavor trait quantitative score, and to screen out target metabolic modules whose absolute value of the first Pearson correlation coefficient is greater than the preset correlation threshold. The parameter determination module is used to analyze the dynamic trend curve of the content of key metabolites in the target metabolic module, and determine the target processing parameters based on the extreme points of the dynamic trend curve.
9. A device for regulating tea processing parameters based on metabolomics and the WGCNA algorithm, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the tea processing parameter regulation method based on metabolomics and the WGCNA algorithm as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the tea processing parameter regulation method based on metabolomics and WGCNA algorithm as described in any one of claims 1 to 7.