Tea polyphenol and gallic acid content virtual measurement method based on machine learning model

Through the virtual measurement method of tea polyphenols and gallic acid content based on machine learning models, the prediction model is constructed and parameters are optimized, and the existing prediction model is solved, and more efficient and accurate prediction of tea polyphenols and gallic acid content is achieved.

CN120148701APending Publication Date: 2025-06-13TEA RES INST OF FUJIAN ACADEMY OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510301538.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing tea polyphenol and gallic acid content prediction models have problems such as low prediction accuracy, high application cost, long time consumption and high destructiveness to samples.

Method used

Using the virtual measurement method of tea polyphenols and gallic acid content based on machine learning models, a training sample set is constructed and a prediction model is constructed using machine learning models (such as linear regression, decision tree regression, random forest regression, support vector regression, gradient enhancement regression, etc.) is optimized to improve prediction accuracy.

Benefits of technology

It improves the prediction accuracy of tea polyphenols and gallic acid content, reduces application cost and time-consuming, and reduces the destructiveness of the sample.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148701A_ABST
    Figure CN120148701A_ABST
Patent Text Reader

Abstract

The invention relates to a tea polyphenol and gallic acid content virtual measurement method based on a machine learning model, and belongs to the technical field of tea component analysis. The method comprises the following steps: S1, obtaining each component value in a tea sample, and combining data of each component value of a plurality of tea samples to construct a sample set; s2, selecting a part of data in the sample set as a training sample set, and constructing a prediction model by adopting a machine learning model to carry out fitting analysis and prediction on the content of tea polyphenol and gallic acid so as to obtain a tea polyphenol and gallic acid prediction model; and S3, evaluating the prediction accuracy of the tea polyphenol and gallic acid of the tea polyphenol and gallic acid prediction model by adopting the evaluation sample set, the mean square error and the decision coefficient so as to obtain the tea polyphenol and gallic acid prediction model with the optimal parameter combination. Compared with an existing prediction model of tea polyphenol and gallic acid, the prediction model of tea polyphenol and gallic acid in the tea sample obtained by the method has the advantages that the prediction accuracy is improved, and the method has good application and popularization prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tea component analysis, and particularly relates to a virtual metering method for the contents of tea polyphenols and gallic acid based on a machine learning model. Background Art

[0002] Component analysis and detection mainly analyze and detect the internal nutrient substances of tea with the help of relevant detection instruments, which has the advantages of objectivity, accuracy, etc., but has the disadvantages of high cost, long time consumption, and great damage to samples. Machine learning, as an effective and economical alternative method for the traditional data processing process, has the advantages of self-learning, strong fault tolerance, and high prediction accuracy. Therefore, analyzing and summarizing a large amount of data (color difference data) in tea production through machine learning to predict its internal nutrient substances is one of the directions to solve this problem. The present invention specifically introduces the application of the current machine learning model in predicting the contents of finished tea polyphenols and gallic acid, and makes a prospect for future research directions, in order to provide corresponding reference for the in-depth research of common machine learning methods in quality prediction applications.

[0003] In the fields of tea production and research, using machine learning to predict the content of quality components is an important application. Tea samples of different varieties, origins, growth environments, and processing technologies are collected, and the contents of tea polyphenols and gallic acid are measured as target values. At the same time, characteristic data that may be related to the contents of tea polyphenols and gallic acid are collected, such as tea chemical components, growth environment parameters, and parameters during the processing process.

[0004] According to the category information and prediction methods of the data set, machine learning is divided into traditional machine learning and deep machine learning. Traditional machine learning (TML) mainly extracts features manually on a small sample data set, analyzes the data to balance the effectiveness of the learning result and the interpretability of the learning model, and provides a framework for solving learning problems when limited samples are available. Deep learning (DL) is to extract the pattern of features from the existing data according to the pre-designed feature extraction rules, obtain the corresponding deep features from it, and achieve the purpose of dimensionality reduction. According to whether the training method and training data are labeled, it can usually be divided into supervised learning and unsupervised learning. Quality prediction belongs to the supervised learning in TML, and common models include the nearest neighbor model, decision tree model, random forest model, support vector machine model, and neural convolutional network model, etc. Summary of the Invention

[0005] The object of the present invention is to provide a virtual metering method for the contents of tea polyphenols and gallic acid based on a machine learning model. This method obtains a prediction model for tea polyphenols and gallic acid in a tea sample, which improves the prediction accuracy compared with the existing prediction models for tea polyphenols and gallic acid, and has good application and promotion prospects.

[0006] To achieve the above object, the technical solution of the present invention is: a virtual metering method for the contents of tea polyphenols and gallic acid based on a machine learning model, comprising:

[0007] S1. Obtain the component values in the tea sample, and combine the component value data of multiple tea samples to construct a sample set;

[0008] S2. Select part of the data in the sample set as a training sample set, and use a machine learning model to construct a prediction model for fitting analysis to predict the contents of tea polyphenols and gallic acid, so as to obtain a prediction model for tea polyphenols and gallic acid.

[0009] In an embodiment of the present invention, the method further comprises:

[0010] S3. Based on the prediction model for tea polyphenols and gallic acid obtained in step S2, use part of the data in the sample set as an evaluation sample set, and use the mean square error and the coefficient of determination to evaluate the prediction accuracy of the tea polyphenols and gallic acid of the prediction model for tea polyphenols and gallic acid, so as to obtain a prediction model for tea polyphenols and gallic acid with the best parameter combination.

[0011] In an embodiment of the present invention, in step S1, the component values in the tea sample include L, a, b, p, hue chroma Cab, color saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L 2, a2, b2, p 2, hue chroma Cab 2, color saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p 3, color chroma Sab 3, color saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, tea polyphenols, theaflavone, amino acids, theabrownin, thearubigin, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechin, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol values in the tea sample.

[0012] In an embodiment of the present invention, in step S2, taking the values of L, a, b, p, chroma of hue Cab, saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L 2, a2, b2, p 2, chroma of hue Cab 2, saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L 3, a3, b3, p3, chroma of color Sab 3, saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, tea flavonoids, amino acids, theabrownin, thearubigin, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechins, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol in the tea samples in the training sample set as the input, and the value of tea polyphenols as the output, a machine learning model is used to construct a tea polyphenol prediction model; taking the values of L, a, b, p, chroma of hue Cab, saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L 2, a2, b2, p 2, chroma of hue Cab 2, saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L 3, a3, b3, p 3, chroma of color Sab 3, saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, tea polyphenols, tea flavonoids, amino acids, theabrownin, thearubigin, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechins, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol in the tea samples in the training sample set as the input, and the value of gallic acid as the output, a machine learning model is used to construct a gallic acid prediction model.

[0013] In an embodiment of the present invention, the machine learning model includes a linear regression model, a decision tree regression model, a random forest regression model, a support vector regression model, a gradient boosting regression model, and a neural network regression model.

[0014] In an embodiment of the present invention, in step S3, for the tea polyphenol and gallic acid prediction models constructed in step S2, part of the data in the sample set is used as an evaluation sample set, and the mean square error and the coefficient of determination are used to evaluate the prediction accuracy of tea polyphenols and gallic acid of the tea polyphenol and gallic acid prediction models, so as to obtain the optimal machine learning model for constructing the tea polyphenol and gallic acid prediction models, and based on the optimal machine learning model, the best parameter combination of the tea polyphenol prediction model and the gallic acid prediction model is obtained.

[0015] In one embodiment of the present invention, the gradient boosting regression model, i.e., the gradient boosting decision tree regression model GBDT, uses the negative gradient of the loss function to approximate the residual, so as to fit a new CART regression tree. The expression formula of the negative gradient is as follows:

[0016]

[0017] where r t,i represents the negative gradient of the loss function of the i-th sample in the t-th round, L(y i , f(x i )) is the loss function, which is used to measure the difference between the true value y i and the predicted value f(x i ), and f t-1 (x) is the predicted value of the model in the previous round of iteration; in each round of iteration, first use (x i , r t,i ) to fit a CART regression tree. Each leaf node of the regression tree will contain a predetermined range of input data, which is called the leaf node region R t,j , j = 1, 2, ……, J, where J is the number of leaf nodes and j represents the number of the leaf node; each leaf node outputs a constant value c t,j , which is obtained by minimizing the loss function. Specifically, for all samples in the leaf node region R t,j , the goal is to find a c t,j such that the loss function of all samples in the corresponding node is minimized. The formula is as follows:

[0018]

[0019] c is the output value of the current leaf node;

[0020] Next, h t (x) is expressed as the weighted sum of the output values c t,j of each leaf node, and the decision tree fitting function of this round is obtained as follows:

[0021]

[0022] where I(x ∈ R t,j ) is an indicator function, which indicates whether the sample x belongs to the corresponding leaf node region R t,j ;

[0023] In each round, the strong learner is the update of the base learner, and it is gradually optimized by adding the output of the decision tree in the current round to the previous model. The expression of the strong learner finally obtained in this round is as follows:

[0024]

[0025] In an embodiment of the present invention, the loss function adopts the mean squared error loss or the logarithmic loss.

[0026] In an embodiment of the present invention, the best parameter combination is a machine learning model with a maximum depth of 7 and a minimum sample splitting number of 2 as the prediction model for tea polyphenols and gallic acid.

[0027] The embodiment of the present invention also provides a computer system, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps as described in any one of the above can be implemented.

[0028] The embodiment of the present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method steps as described in any one of the above can be implemented.

[0029] The embodiment of the present invention also provides an electronic device, which includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor executes the method steps as described in any one of the above.

[0030] The embodiment of the present invention also provides a computer program product, including a computer program, and the computer program is stored in a computer-readable storage medium; when the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the method steps as described in any one of the above.

[0031] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention obtains a prediction model for tea polyphenols in a tea sample, which improves the prediction accuracy compared with the existing tea polyphenol prediction model and has good application and promotion prospects. Description of the Drawings

[0032] Figure 1 It is a flowchart of the method of the present invention.

[0033] Figure 2 It is a heat map of the correlation coefficient matrix of the relevant index values of each component in the tea sample related to tea polyphenols. Detailed Embodiments

[0034] Next, in combination with the drawings, the technical solutions of the present invention will be specifically described.

[0035] The present invention provides a virtual metering method for the contents of tea polyphenols and gallic acid based on a machine learning model, including:

[0036] S1. Obtain the component values in the tea sample, including: L, a, b, p, hue chroma Cab, color saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L2, a2, b2, p 2, hue chroma Cab2, color saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p 3, color chroma Sab3, color saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, tea polyphenols, tea flavonoids, amino acids, theabrownin, thearubigin, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechin, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p - coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol values, and combine the component value data of multiple tea samples to construct a sample set;

[0037] S2. Select part of the data in the sample set as the training sample set, and use a machine learning model to construct a prediction model for fitting analysis to predict the contents of tea polyphenols and gallic acid, and obtain the prediction models for tea polyphenols and gallic acid; Specifically,

[0038] Using the values of L, a, b, p, chroma Cab, saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L2, a2, b2, p2, chroma Cab 2, saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p3, chroma Sab 3, saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, theaflavone, amino acid, theabrownin, thearubigin, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechin, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol in the tea samples in the training sample set as the input and the value of tea polyphenols as the output, a machine learning model is used to construct a tea polyphenol prediction model; using the values of L, a, b, p, chroma Cab, saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L2, a2, b2, p2, chroma Cab 2, saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p3, chroma Sab 3, saturation Sab 3, hue b / a3, hue angle Hab3, F3, theabrownin 3, thearubigin 3, tea polyphenols, theaflavone, amino acid, theabrownin, thearubigin, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechin, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol in the tea samples in the training sample set as the input and the value of gallic acid as the output, a machine learning model is used to construct a gallic acid prediction model;

[0039] S3. Based on the tea polyphenol and gallic acid prediction models obtained in step S2, part of the data in the sample set is used as the evaluation sample set, and the mean square error and determination coefficient are used to evaluate the prediction accuracy of tea polyphenols and gallic acid in the tea polyphenol and gallic acid prediction models, so as to obtain the tea polyphenol and gallic acid prediction models with the best parameter combinations.

[0040] In step S1, the method for obtaining the component values in the tea samples is as follows:

[0041] Scald the evaluation tea cups and bowls with boiling water. Weigh 5.0 g of representative tea samples and place them in an 110-ml inverted bell-shaped evaluation tea cup. Quickly fill it with boiling water and cover it with a cup lid. According to the tea evaluation method, filter the tea sample soup of the first, second, and third brews (2 minutes, 3 minutes, 5 minutes) while it is hot with filter paper. Cool the filtrate to room temperature and measure the brightness (L) value, red-green chromaticity (a) value, yellow-blue chromaticity (b) value, and color depth index (p) value of the tea soup sample with a spectrophotometric color difference meter. Repeat 3 times and calculate F1 (the first brew), F2 (the second brew), and F3 value (the third brew). The specific calculation formula is as follows:

[0042] F i = b i / a i , i = 1, 2, 3.

[0043] i = 1, 2, 3 respectively refer to the content of the corresponding values in the tea soup of the first, second, and third brews (2 minutes, 3 minutes, 5 minutes) during brewing.

[0044] In addition, the following values are also calculated:

[0045] Sab i (Color saturation) = Cab i / L i , b i / a i (Hue), Hab i (Hue angle) = tan(b i / a i )

[0046] At the same time, measure and obtain theabscurochrome 1 (theabscurochrome in the first brew), thearubigins 1 (thearubigins in the first brew), theabscurochrome 2 (theabscurochrome in the second brew), thearubigins 2 (thearubigins in the second brew), theabscurochrome 3 (theabscurochrome in the third brew), thearubigins 3 (thearubigins in the third brew), as well as the tea polyphenols, theaflavins, amino acids, theabscurochrome, thearubigins, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechins, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapic acid, ECG, rutin, myricetin, theaflavin, quercetin, kaempferol values of the tea sample before brewing.

[0047] In step S3, the mean square error and the coefficient of determination are used to evaluate the prediction accuracy of the tea polyphenols and gallic acid prediction models for tea polyphenols and gallic acid, so as to obtain the best parameter combination of the tea polyphenols and gallic acid prediction models. The specific implementation is as follows:

[0048] Example 1:

[0049] In this example, the performance of the tea polyphenol prediction model is analyzed. The performance of the gradient boosting regression model, that is, the gradient boosting decision tree regression model GBDT, under different parameter combinations is shown in Table 1 as follows:

[0050] Table 1

[0051]

[0052] Mean Squared Error (MSE): The mean squared error measures the average of the squares of the deviations between the predicted values and the true values. The smaller the value, the higher the accuracy of the model prediction. From the results in Table 1, when the maximum depth is 7 and the minimum number of samples for splitting is 2, the mean squared error is the smallest, which is 0.6921039643694028. This indicates that under this parameter combination, the deviation between the model's predicted values and the true values is relatively small.

[0053] Coefficient of determination (R 2 ) : The coefficient of determination represents the degree of fit of the model to the data. The closer its value is to 1, the better the fitting effect. From the results in Table 1, when the maximum depth is 7 and the minimum number of samples for splitting is 2, the coefficient of determination is the largest, reaching 0.8953163507681694, indicating that the model has the best fitting effect on the data.

[0054] Considering the two indicators of mean squared error and coefficient of determination comprehensively, the machine learning model with a maximum depth of 7 and a minimum number of samples for splitting of 2 performs best in predicting tea polyphenols and gallic acid. Therefore, based on the mean squared error and coefficient of determination predicted by the above gradient boosting tree regression model, it can be known that when using a machine learning model to construct a tea polyphenol and gallic acid prediction model, this parameter combination can be preferentially selected for the prediction of tea polyphenols and gallic acid.

[0055] Based on the parameter combination of a maximum depth of 7 and a minimum number of samples for splitting of 2, this application also predicts tea polyphenols based on decision tree regression models, random forest regression models, gradient boosting tree regression models, etc. Through comparison, the gradient boosting regression model performs more excellently in both indicators of mean squared error and coefficient of determination, indicating that this model has better performance in predicting tea polyphenols.

[0056] The gradient boosting decision tree regression model GBDT (Gradient Boosting Decision Tree) is an ensemble learning algorithm that uses multiple decision trees to solve classification and regression problems. The core idea is to construct a new decision tree through the residuals of the previous round of the model. To improve the fitting effect, the gradient boosting tree regression model uses the negative gradient of the loss function to approximate the residuals, thereby fitting a new CART regression tree. The formula for representing the negative gradient is as follows:

[0057]

[0058] where r t,i represents the negative gradient of the loss function of the i-th sample in the t-th round, and L(y i , f(x i )) is the loss function used to measure the difference between the true value y i and the predicted value f(x i ), and f t-1 (x) is the predicted value of the model in the previous round of iteration; in each round of iteration, first use (x i , r t,i ) to fit a CART regression tree. Each leaf node of the regression tree will contain a predetermined range of input data, called the leaf node region R t,j , j = 1, 2, ……, J, where J is the number of leaf nodes and j represents the number of the leaf node; each leaf node outputs a constant value c t,j , which is obtained by minimizing the loss function. Specifically, for all samples in the leaf node region R t,j , the goal is to find a c t,j such that the loss function of all samples in the corresponding node is minimized, and the formula is as follows:

[0059]

[0060] c is the output value of the current leaf node;

[0061] Next, h t (x) is expressed as the weighted sum of the output values c t,j of each leaf node, and the decision tree fitting function for this round is obtained as follows:

[0062]

[0063] where I(x ∈ R t,j ) is an indicator function indicating whether the sample x belongs to the corresponding leaf node region R t,j ;

[0064] In each round, the strong learner is an update of the base learner, and it is gradually optimized by adding the output of the decision tree in the current round to the previous model. The expression of the strong learner finally obtained in this round is as follows:

[0065]

[0066] Such as Figure 2As shown, the component values in the tea samples also affect the accuracy of the theaflavin prediction model constructed using the gradient boosting tree regression model. Based on the influence of the component values in the tea samples on theaflavin prediction, the correlation coefficient matrix of the component values in the tea samples is calculated, and then the index pairs with the absolute value of the correlation coefficient greater than 0.8 (which can be adjusted according to the actual situation) are selected to determine the color difference indicators with high correlation.

[0067] Example 2

[0068] In this example, the performance of the gallic acid prediction model was analyzed. Based on the L, a, b, p, chroma Cab, saturation Sab, hue angle Hab, F1, a2, b2, p 2, chroma Cab 2, saturation Sab 2, hue angle Hab 2, F2, L 3, a3, b3, p 3, chroma Sab3, saturation Sab3, hue angle Hab3, F3, theabrownin 2, theabrownin 3 indicators (correlation coefficient > 0.5), linear regression models, decision tree regression models, and random forest regression models were constructed to build the gallic acid prediction model. The specific performance of each model is as follows:

[0069] (1) Linear regression model

[0070] Mean squared error (MSE): 791582.954663841, approximately 791582.95 when rounded to two decimal places. The mean squared error measures the average squared error between the predicted value and the true value. The larger the value, the greater the prediction error.

[0071] Coefficient of determination (R 2 ) : 0.4946510414642674, approximately 0.49 when rounded to two decimal places. The coefficient of determination represents the goodness of fit of the model to the data. The closer it is to 1, the better the fitting effect. This model can explain approximately 49% of the changes in gallic acid content.

[0072] (2) Decision tree regression model

[0073] Mean squared error (MSE): 1089985.0054897433, approximately 1089985.01 when rounded to two decimal places. Compared with the linear regression model, its mean squared error is larger, indicating a greater prediction error.

[0074] Coefficient of determination (R 2 ) : 0.3041502673870451, approximately 0.30 when rounded to two decimal places. This shows that the model has a weak ability to explain the changes in gallic acid content and can only explain approximately 30% of the changes.

[0075] (3) Random forest regression model

[0076] Mean Squared Error (MSE): 752658.2558020351, approximately 752658.26 when rounded to two decimal places. Among these three models, it has the smallest MSE, indicating that the prediction error is relatively the smallest.

[0077] Coefficient of determination (R 2 ): 0.5195006872471077, approximately 0.52 when rounded to two decimal places. It can explain approximately 52% of the changes in gallic acid content and has the best fitting effect among these three models.

[0078] Generally speaking, in this prediction task, the performance of the random forest regression model is the best, followed by the linear regression model, and the decision tree regression model is relatively poor. However, the coefficients of determination of each model are not very high, and there may be room for further optimization, such as performing feature engineering, adjusting model hyperparameters, etc., to improve the prediction accuracy of the model.

[0079] Example 3

[0080] In this example, the performance of the gallic acid prediction model was analyzed. Using L, a, b, p, chroma Cab, saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L2, a2, b2, p 2, chroma Cab 2, saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L 3, a3, b3, p 3, chroma Sab 3, saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3 as indicators, gallic acid prediction models were constructed based on the support vector regression model, gradient boosting regression model, and neural network regression model respectively. The specific performance of each model is as follows:

[0081] (1) Support vector regression model

[0082] Mean Squared Error (MSE): 406335.13752981165, approximately 406335.14 when rounded to two decimal places. The MSE reflects the average squared error between the predicted value and the true value. The smaller this value, the higher the accuracy of the model prediction.

[0083] Coefficient of determination (R 2 ) : 0.6248017806818604, approximately 0.62 when rounded to two decimal places. The coefficient of determination measures the goodness of fit of the model to the data. Its value ranges from 0 to 1, and the closer it is to 1, the stronger the model's ability to explain the data. Here, it shows that the model can explain approximately 62% of the changes in gallic acid content.

[0084] (2) Gradient boosting regression model

[0085] Mean Squared Error (MSE): 381021.71746703197, approximately 381021.72 when rounded to two decimal places. Among the three models, the MSE of this model is the smallest, indicating that its prediction error is relatively the smallest.

[0086] Coefficient of determination (R 2 ): 0.6481754672159475, approximately 0.65 when rounded to two decimal places. It shows that this model can explain approximately 65% of the changes in gallic acid content and has the best data fitting effect among the three models.

[0087] (3) Neural network regression model

[0088] Mean Squared Error (MSE): 470146.5670678161, approximately 470146.57 when rounded to two decimal places. Among the three models, its MSE is the largest, indicating relatively lower prediction accuracy.

[0089] Coefficient of determination (R 2 ): 0.5658801356566436, approximately 0.57 when rounded to two decimal places. It shows that this model can explain approximately 57% of the changes in gallic acid content and has the worst data fitting effect among the three models.

[0090] Overall, in the task of predicting gallic acid content this time, the gradient boosting regression model has the best performance, the support vector regression model ranks second, and the neural network regression model is relatively poor. However, there is still room for improvement in the coefficient of determination of each model. We can consider further optimizing the hyperparameters of the model, increasing feature engineering, or collecting more data to improve the prediction performance of the model.

[0091] The embodiment of the present invention also provides a computer system, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, it can implement the method steps as described in any one of the above.

[0092] The embodiment of the present invention also provides a computer-readable storage medium, on which computer program instructions executable by the processor are stored. When the processor runs the computer program instructions, it can implement the method steps as described in any one of the above.

[0093] The embodiment of the present invention also provides an electronic device, which includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the method steps as described in any one of the above.

[0094] An embodiment of the present invention further provides a computer program product, including a computer program, where the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the method steps as described in any one of the above.

[0095] The above are the preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.

Claims

1. A virtual measurement method for tea polyphenols and gallic acid content based on a machine learning model, characterized in that: include: S1. Obtain the values ​​of each component in the tea sample, and combine the data of each component value of multiple tea samples to construct a sample set; S2. Select part of the data in the sample set as the training sample set, use the machine learning model to build a prediction model for fitting analysis to predict the content of tea polyphenols and gallic acid, and obtain the prediction model of tea polyphenols and gallic acid.

2. According to claim 1, a virtual measurement method for tea polyphenols and gallic acid content based on a machine learning model is characterized in that: Also includes: S3. Based on the tea polyphenols and gallic acid prediction model obtained in step S2, part of the data in the sample set is used as an evaluation sample set, and the tea polyphenols and gallic acid prediction accuracy of the tea polyphenols and gallic acid prediction model is evaluated using mean square error and determination coefficient to obtain the tea polyphenols and gallic acid prediction model with the best parameter combination.

3. A virtual measurement method for tea polyphenols and gallic acid content based on a machine learning model according to claim 1 or 2, characterized in that: In step S1, the values ​​of each component in the tea sample include L, a, b, p, hue chroma Cab, color saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L2, a2, b2, p2, hue chroma Cab 2, color saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p3, color chroma Sab 3, color saturation Sab 3. Hue angle b / a3, hue angle Hab3, F3, theabrownin 3, thearubigins 3, tea polyphenols, theaflavones, amino acids, theabrownins, thearubigins, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechins, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapinic acid, ECG, rutin, myricetin, theaflavins, quercetin, and kaempferol values.

4. The virtual measurement method of tea polyphenols and gallic acid content based on a machine learning model according to claim 2, characterized in that: In step S2, L, a, b, p, hue chroma Cab, color saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigin 1, L2, a2, b2, p2, hue chroma Cab 2, color saturation Sab 2, hue b / a 2, hue angle Hab 2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p3, color chroma Sab 3, color saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, tea flavonoids, amino acids, theabrownins, thearubigins, gallic acid, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechins, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapinic acid, ECG, rutin, myricetin, theaflavins, quercetin, and kaempferol values ​​are input, and tea polyphenols values ​​are output. A machine learning model is used to construct a tea polyphenols prediction model; the tea samples in the training sample set are L, a, b, p, hue chroma Cab, color saturation Sab, hue b / a, hue angle Hab, F1, theabrownin 1, thearubigins 1, L2, a2, b2, p2, hue chroma Cab 2, color saturation Sab 2, hue b / a 2. Hue angle Hab2, F2, theabrownin 2, thearubigin 2, L3, a3, b3, p3, color chroma Sab 3, color saturation Sab 3, hue b / a 3, hue angle Hab 3, F3, theabrownin 3, thearubigin 3, tea polyphenols, tea flavonoids, amino acids, theabrownins, thearubigins, theobromine, GC, protocatechuic acid, theophylline, EGC, chlorogenic acid, catechin, caffeine, vanillic acid, caffeic acid, syringic acid, epicatechin, EGCG, GCG, p-coumaric acid, ferulic acid, sinapinic acid, ECG, rutin, myricetin, theaflavins, quercetin, and kaempferol values ​​are input, and gallic acid value is output. A machine learning model is used to construct a gallic acid prediction model.

5. A virtual measurement method for tea polyphenols and gallic acid content based on a machine learning model according to claim 1, 2 or 4, characterized in that: Machine learning models include linear regression models, decision tree regression models, random forest regression models, support vector regression models, gradient boosting regression models, and neural network regression models.

6. A virtual measurement method for tea polyphenols and gallic acid content based on a machine learning model according to claim 2 or 4, characterized in that: In step S3, for the tea polyphenols and gallic acid prediction model constructed in step S2, part of the data in the sample set is used as an evaluation sample set, and the mean square error and determination coefficient are used to evaluate the tea polyphenols and gallic acid prediction accuracy of the tea polyphenols and gallic acid prediction model to obtain the optimal machine learning model for constructing the tea polyphenols and gallic acid prediction model, and based on the optimal machine learning model, the tea polyphenols prediction model and gallic acid prediction model with the best parameter combination are obtained.

7. A computer system, characterized in that: The method comprises a memory, a processor and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the method steps as claimed in any one of claims 1 to 6 can be implemented.

8. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 6 can be implemented.

9. An electronic device comprising a processor and a memory, wherein: The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method steps according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, wherein the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the method steps as described in any one of claims 1-6.