Method for modulating and controlling online parameters of tea production line customized for functional components

By analyzing the various characteristics of fresh tea leaves on the tea production line and using machine learning models to predict the chemical composition of tea, the problems of unstable tea quality and low production efficiency are solved, and the rapid, accurate prediction and personalized production of chemical composition of tea are achieved.

CN119937473AInactive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202411866686.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to achieve rapid and accurate control of the functional components in tea on the tea production line, resulting in unstable tea quality and low production efficiency.

Method used

By analyzing the tenderness, moisture content, processing parameters and image information of fresh tea leaves, machine learning models (such as PLSR and recursive feature elimination methods) are used to predict the content of specific chemical components in tea leaves, such as ECG, GA, and CAF, and then optimize the process parameters to achieve personalized tea production.

Benefits of technology

It achieves rapid and accurate prediction of tea chemical composition, improves the stability and production efficiency of tea production quality, and reduces labor demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937473A_ABST
    Figure CN119937473A_ABST
Patent Text Reader

Abstract

The invention discloses an online parameter modulation and control method for a tea production line customized for functional components, and the method specifically comprises the following steps: firstly, collecting the tenderness grade, moisture content and production process parameters of fresh tea leaves and the image information of the fresh tea leaves, and providing comprehensive data support for model construction; and then, the collected data is processed, and abnormal values are eliminated to ensure the data quality, so that the model performance is improved. Through multiple feature importance analysis methods, the relationship between the input features and the target chemical components is deeply analyzed, and a scientific basis is provided for feature screening. On the basis, the data set is screened and optimized, and the prediction precision of the machine learning model is further improved. Meanwhile, a brand-new feature screening method is provided and used for screening optimal features required by modeling, so that the predictive ability of the model is enhanced. And optimizing the process parameters by using the constructed optimization optical component prediction model and combining with a multi-target particle swarm optimization algorithm. Finally, the optimization result is applied to control software of a tea production line, the production target of high-quality tea is achieved, and technical support is provided for personalized customization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of machine learning and primary processing of tea, and in particular relates to an online parameter modulation and control method for a tea production line customized for functional components. Background Art

[0002] China has a thousands-year-old custom of drinking tea. Tea is not only a refreshing beverage, but also a rich source of bioactive compounds, which have many health benefits, antioxidant and anti-inflammatory effects, and potential cancer prevention and cardiovascular health support. Therefore, tea occupies an important position in the modern food and beverage industry. This places special requirements on the quality of tea, such as the homogeneity of processed tea leaves and the ability to customize the production of tea product raw materials with special flavors according to demand. Exploring how to obtain the more personalized and stable quality tea required is of great significance for high-value processing.

[0003] At present, the traditional chemical composition analysis methods of tea, such as ultra-high performance liquid chromatography (UPLC) and gas chromatography-mass spectrometry (GC-MS), are accurate but time-consuming and costly, and are not suitable for production lines. Therefore, the present invention proposes an innovative method to predict the content of epicatechin gallate (ECG), gallic acid (GA) and caffeic acid (CAF) in tea by analyzing the tenderness of fresh tea leaves, the moisture content of fresh tea leaves, key parameters in the processing process, the color and texture characteristics of fresh tea leaves, so as to achieve the purpose of controlling the production quality of tea. Summary of the invention

[0004] In order to solve the shortcomings of the prior art, achieve the purpose of improving the quality of tea production and reducing labor, the present invention adopts the following technical solutions:

[0005] The online parameter modulation and control method of a tea production line customized for functional ingredients comprises the following steps:

[0006] Step S1: Acquire multiple unit data sets of the tea primary processing production line, including fresh tea leaf grade, fresh tea leaf moisture content, process parameters of each tea processing step, fresh leaf image texture information and chromaticity information; measure the contents of multiple chemical components of tea leaves of different grades through UPLC experiments;

[0007] Step S2: Preprocess the multiple unit data sets to remove erroneous data and reduce the impact of data noise on the model building results. According to the changes in the unit data during the production process, remove the data that does not change and abnormal values;

[0008] Step S3: The importance of each feature for each chemical component is calculated by the Pearson correlation coefficient method and the SHAP method, and a preliminary analysis of the contribution of the input variables to the chemical component content is performed;

[0009] Step S4: By applying the variable importance projection method (VIP), the importance values ​​of the features are calculated, and the features with VIP values ​​greater than 1 are screened for modeling. Then, chemical composition prediction models were established based on the complete data set, the data set screened by the VIP method, and the data set screened by the recursive feature elimination method combining the Pearson correlation coefficient, SHAP importance and VIP value. The experimental results show that the model established by the data set screened by the proposed method has the best prediction effect;

[0010] Step S5: According to the prediction model, the process parameters are optimized with the chemical composition of raw tea as the target, and the optimization results are applied to the software parameter control of the primary tea processing production line to realize personalized tea production.

[0011] Furthermore, the fresh tea leaves in step S1 are divided into six categories, with the quality ranging from level 1 to level 6 from high to low. The tenderness grade of the tea raw materials is judged according to the leaf shape of the tea leaves; the process parameters of the withering unit and the rolling unit include hot air temperature, drum speed, feeding speed, and dehumidification speed, wherein the process parameters of the rolling unit also include hot air temperature and hot air speed, and the process parameters of the primary drying unit and the secondary drying unit include hot air temperature, oil inlet temperature, oil outlet temperature, main conveying speed, hot air speed, and leaf leveling speed;

[0012] Fresh leaf image information includes color features and texture features of fresh tea leaf images, where color features include HSV, LAB and RGB, and texture features include mean, standard deviation, smoothness, third-order moment, consistency, and entropy of fresh tea leaf images;

[0013] The chemical components of raw tea include gallic acid GA, catechin C, epicatechin EC, epigallocatechin EGC, epicatechin gallate ECG, epigallocatechin gallate EGCG, theanine L-The and caffeine CAF. The above chemical component standards are prepared into standard solutions, and the contents of chemical components in tea are calculated according to the standard curve and peak area based on liquid chromatography of the standard solutions.

[0014] Furthermore, the preprocessing in step S2 includes missing value processing and outlier processing; missing values ​​refer to the missing of part of the data due to equipment failure or other reasons during the data collection process. In order to ensure the accuracy of the model, the corresponding batch of tea production data containing missing values ​​is removed, and the average value of the remaining data is used to represent the process parameters of the batch; outliers refer to those values ​​that significantly deviate from the normal range of the data set, which usually come from atypical fluctuations in machine operation or abnormal performance caused by sensor failure.

[0015] Furthermore, the SHAP value in step S3 is calculated based on the PLSR model.

[0016] Furthermore, the variable importance projection method in step S4 is a variable screening method based on the partial least squares algorithm, which is mainly used to evaluate the importance of each variable in the data set for explaining the target variable. The calculation formula is as follows:

[0017]

[0018] VIP here j corresponds to the VIP value of the jth feature; p corresponds to the total number of predictor variables, which is 41; A is the number of PLS ​​components. After cross-validation, the optimal principal component of ECG is 7, the optimal principal component of GA is 9, and the optimal principal component of CAF is 5; q a is the residual;

[0019] Furthermore, the machine learning model in step S4 is based on the least squares regression algorithm PLSR. For k chemical components {y1, y2, ..., y k}, m unit process parameters {x1, x2,…, x m} data, the matrix composed of n samples of data is Y = {y1, y2, ..., y k} n×k and X={x1,x2,…,x m} n×m In order to study the relationship between the independent variable and the dependent variable, the first principal component is extracted from X, which should contain as much information as possible about the original data, and then a linear regression equation is established between the principal component and Y. If the accuracy of the regression equation meets the requirements, the algorithm terminates; otherwise, the second principal component is extracted from the remaining independent variable X, and an equation is established until the accuracy meets the requirements. The specific calculation of PLSR includes principal component decomposition and correlation matrix B calculation. Principal component decomposition is to decompose the X and Y matrices into characteristic factors. The specific formula is as follows:

[0020]

[0021] Where T and U represent the principal component score matrices of matrices X and Y, respectively; P and Q represent the loading matrices of X and Y, respectively; E and F represent the error matrices of X and Y during the fitting process, respectively;

[0022] U=TB

[0023] B=T T U(T T T) -1

[0024] The linear regression model is built using matrices T and U, and B represents the association matrix.

[0025] Furthermore, the recursive feature elimination method combining Pearson correlation coefficient, SHAP importance, and VIP value in step S4 has the following main calculation process:

[0026] A. Calculate the importance weights of all features in three cases and normalize them.

[0027] B. Add these three values with the same weight.

[0028] C. Sort according to the weights, and sort from high to low importance. Then, perform model modeling and optimization under different numbers of features according to this sorting to obtain the optimal number of features and corresponding hyperparameters.

[0029] For the online parameter modulation and control method of the tea production line customized for functional components, the established machine learning model is evaluated for performance according to three index evaluation methods, including root mean square error RMSE, correlation coefficient r, and ratio of performance to deviation RPD; the smaller the RMSE, the closer the r value is to 1, and the larger the RPD, the better the model performance. RPD < 1.4 indicates poor model prediction; 2.0 < RPD < 2.5 indicates very good model quantification; RPD > 2.5 indicates excellent model performance. Their calculations are as follows:

[0030]

[0031]

[0032] where N is the number of samples in the dataset, y i and are the true value and predicted value of the i-th sample point, and are the average values of the true vector y and the predicted vector respectively. The control module applies the optimization result to the software parameter control of the primary tea production line to achieve high-quality tea production.

[0033] For the online parameter modulation and control method of the tea production line customized for functional components, the chemical components are measured by the ultra-high performance liquid chromatography instrument WATER ACQUITY UPLC, which is used to obtain the ultra-high performance liquid chromatography of the standard solution of the chemical components of the raw tea. The chemical component content in the tea is calculated according to the standard curve and peak area. The chemical components of the raw tea include gallic acid GA, catechin C, epicatechin EC, epigallocatechin EGC, epicatechin gallate ECG, epigallocatechin gallate EGCG, theanine L-The, and caffeine CAF. The specific elution procedure is as follows:

[0034] The eluent consisted of mobile phase A (formic acid / water, 1 / 999, mL / mL) and mobile phase B (formic acid / acetonitrile, 1 / 999, mL / mL). The column temperature was set at 25 °C. The mobile phase gradient was: 0–3 min, 100 / 0 (A / B, mL / mL); 3–5 min, from 100 / 0 to 92 / 8; 5–9 min, from 92 / 8 to 88 / 12; 9–11 min, from 88 / 12 to 83 / 17; 11–14 min, from 83 / 17 to 80 / 20; 14–16 min, from 80 / 20 to 10 / 90; keep the ratio of 10 / 90 for 2 min; 18–18.5 min, from 10 / 90 to 100 / 0; then keep the ratio of 100 / 0 for 2 min. GA, C, EC, EGC, ECG, EGCG and CAF substances in green tea were detected using UPLC chromatogram at 280 nm, while L-The was detected using UPLC chromatogram at 210 nm.

[0035] The advantages and beneficial effects of the present invention are:

[0036] The present invention proposes an online parameter modulation and control method for tea production line customized for functional components. First, a series of tea samples were collected, and their ECG, GA, EGC, EC, EGCG, L-THE and CAF contents were determined using UPLC technology. Subsequently, key parameters in the tea processing process were recorded, including the tenderness of fresh tea leaves, the main uniform leaf change speed, the main conveying speed and other main processing parameters in the production process of tea leaf withering, rolling, primary baking and re-baking, as well as multimodal processing information such as image chromaticity characteristics and texture characteristics of fresh tea leaves; then, the main influencing factors affecting the chemical composition of tea were judged by analyzing the Pearson correlation coefficient and SHAP importance value between the input variables and all chemical components; finally, the PLSR model was trained and the input parameters and chemical components were modeled using different data sets to achieve successful prediction of ECG, GA and CAF.

[0037] The innovation of the present invention lies in: 1. By using easily accessible data such as tea tenderness, moisture content, processing parameters and image information, a rapid prediction of ECG, GA and CAF contents is achieved. By taking chemical composition as the key indicator of quality analysis, a novel and efficient method is provided for the rapid quality assessment of tea. 2. The image information is captured by an ordinary industrial camera. Compared with the existing component analysis method that relies on a hyperspectral camera, the cost is significantly reduced, making this solution more economical and practical, and has a high potential for integrated application in production lines. 3. The tea processing parameters are directly associated with the chemical component content, and the most important processing parameters for predicting the chemical component content are identified, providing a scientific basis for the optimization of the tea processing process. The method proposed in this study can effectively improve the production quality of tea, and can autonomously adjust the chemical composition in tea according to specific needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 A flow chart of the online parameter modulation and control method for a tea production line customized for functional ingredients of the present invention;

[0040] Figure 2 It is a structural schematic diagram of the tea leaf image acquisition system of the present invention;

[0041] Figure 3 It is a Pearson correlation coefficient heat map of the input variables and chemical components of the present invention;

[0042] Figure 4a-4c They are respectively the shap importance bee swarm diagrams of the input variables and chemical components of the present invention;

[0043] Figure 5 It is a bar chart of the VIP values ​​of ECG, GA and CAF obtained based on the variable importance projection method of the present invention. DETAILED DESCRIPTION

[0044] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the inspiration of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

[0045] like Figure 1 As shown, the online parameter modulation and control method of a tea production line customized for functional ingredients includes the following steps:

[0046] Step S1: Acquire multiple unit data sets of the tea primary processing production line, including fresh leaf tenderness grade, fresh leaf moisture content, process parameters of each unit of the tea production line, fresh leaf image texture information and color information, and multiple chemical component contents of the final raw tea.

[0047] The tenderness level of fresh leaves is determined by workers according to the morphological characteristics of the fresh leaves, specifically the buds and leaves of the fresh leaves. Tea with one bud and one leaf is rated as tender level one, usually used to make high-grade green tea; tea with one bud and two leaves is rated as tender level two; tea with one bud and three leaves is rated as tender level three; double leaves: two leaves growing opposite each other, with an undeveloped bud in the middle, are rated as tender level four; tea with only mature leaves but no buds is called pure leaves, with a tenderness level of five; very mature leaves, with a harder texture and darker color, are called old leaves, with a tenderness level of six.

[0048] The moisture content of fresh tea leaves is calculated by weighing method, and the specific steps are as follows: First, place the picked fresh tea leaves in a dry and clean envelope, remove the weight of the envelope and record the initial total weight of the tea leaves (recorded as m1). Next, put the envelope and tea leaves into an oven set at 180℃ for drying. The drying time is two hours to ensure that the moisture in the tea leaves is completely evaporated. After drying, remove the envelope and tea leaves from the oven, cool them to room temperature, and weigh them again, recording the total weight after drying (recorded as m2).

[0049] The calculation formula of fresh leaf moisture content is as follows:

[0050]

[0051] Here M is the moisture content of tea leaves, m1 is the weight of tea leaves before drying, and m2 is the weight of tea leaves after drying.

[0052] The process parameters of the tea production line are established by using the tea production line control software developed based on the C# language to establish communication with the equipment of the tea production line. The frequency of collecting processing parameters of each unit is set to 1min, and the data is stored in the MicrosoftSQL database. In terms of data interaction, Huawei Datacom Intelligent Selection Switch is used to connect to the master control switch on the production line. The specific model is S1730S-L8P2S-A1. The power supply and communication functions are realized with multiple devices in the production workshop through the POE protocol. In terms of hardware, Ethernet and RS485 are used for communication: Ethernet is responsible for transmitting the parameters collected by the equipment to the computer through the network cable, and RS485 is used for data collection between multiple sensors inside each device. In terms of communication protocol, MODBUS TCP / IP protocol and serial communication protocol are used for communication with the equipment respectively.

[0053] The tea production process mainly includes spreading, killing, moisture retention, rolling, rolling, moisture retention, primary baking, moisture retention, re-baking, and moisture retention. In the quality control of tea production, the killing unit determines the inhibitory effect of enzyme activity contained in tea and the retention of aroma precursor substances, which directly affects the aroma and color of tea. The rolling unit is a key link in the morphological shaping of tea leaves, and its process parameters have a significant impact on the appearance of tea leaves and the degree of internal cell damage. The primary baking unit regulates the moisture content of tea leaves through preliminary drying, laying the foundation for subsequent quality stability. The re-baking unit is the last step to improve the internal quality of tea leaves and stabilize the quality. Its temperature and time control directly affect the aroma, taste and storage performance of tea leaves. Therefore, the present invention selects the process parameters of these four units to model the quality of tea leaves, and specifically analyzes the specific impact of each process unit on the quality, providing a scientific basis for optimizing process parameters and improving tea quality. Table 1 shows the names of process parameters and their range of variation.

[0054] Table 1 Process parameters of each unit

[0055]

[0056] The image information of fresh tea leaves was obtained using a self-built experimental platform, which consists of an industrial camera (MV-CS060-10GM / GC, Hikvision, Gigabit Ethernet), a ring-shaped high-brightness diffuse reflection light source (XS-WAR08036-90), a connecting plate, and an aluminum profile bracket. The power supply of the ring-shaped high-brightness diffuse reflection light source is realized by connecting to the 24V power supply of each equipment control cabinet in the factory, and the industrial camera is powered by POE, which has the advantage of realizing power supply and data interaction at the same time. After the image acquisition is completed, six color features and six texture features are extracted. The gray level co-occurrence matrix is ​​used to extract the texture features of the image. The texture features are mean, standard deviation, smoothness, third-order moment, consistency, and entropy; the color features include the information of the three channels of R, G and B and the information of the three channels of H, S, and V.

[0057] The chromaticity information of tea leaves is tested using NR10QC. The chromaticity values ​​of five different positions of the fresh tea leaf sample are measured, and the average value is taken to represent the chromaticity of the tea leaves. The three chromaticity indicators are L*, a*, and b*. L* represents the brightness of the object being measured, a* represents the red and green degree of the object being measured, and b* represents the blue and yellow degree of the object being measured.

[0058] The present invention adopts UPLC method to determine the chemical composition of finished raw tea, and the instrument model is water ACQUITYUPLC TMThe chemical components of raw tea determined included gallic acid (GA), catechin (C), epicatechin (EC), epigallocatechin (EGC), epicatechin gallate (ECG), epigallocatechin gallate (EGCG), theanine (L-The), and caffeine (CAF).

[0059] First, you need to prepare the sample. Weigh 0.2g of raw tea and put it into a centrifuge tube of about 10ml. Add ultrapure water. Wrap the test tube cap with sealing tape and seal it. Then place the centrifuge tube in a test tube rack. The test tube rack is placed in an ultrasonic water bath heating pot for water bath heating and extraction for 60 minutes. The water temperature is set to 70℃. Then prepare the tea supernatant by centrifugation. The centrifugal speed is set to 12000rpm and the time is 10min. During centrifugation, pay attention to the weight difference of the symmetrically placed test tubes not exceeding 0.1g. After centrifugation, use a syringe to absorb 1.5ml of supernatant, and then use a water filter membrane to filter the supernatant and store it in a 2ml brown reagent bottle. The prepared samples need to be stored at 4℃ for use.

[0060] The above chemical components need to be prepared with standard solutions to obtain standard curves of different concentrations. A series of standard solutions with standard concentrations of 0.05-0.5 mg / mL, 0.005-0.1 mg / mL, 0.05-0.5 mg / mL, 0.15-1.2 mg / mL, 0.025-0.25 mg / mL, 0.3-1.5 mg / mL, 0.05-0.5 mg / mL and 0.2-1.0 mg / mL are prepared and placed in ultra-high performance liquid chromatography. The standard curve is calculated based on the actual content of each component and its corresponding peak area. Then, the corresponding concentration of the chemical component of each tea sample is calculated based on the peak area corresponding to the standard curve.

[0061] Step S2: Preprocess the multiple unit data sets to remove erroneous data and reduce the impact of data noise on the model building results. According to the changes in the unit data during the production process, remove the data that does not change substantially and the abnormal values.

[0062] Specifically, preprocessing includes missing value processing and outlier processing. Missing values ​​refer to the missing data due to equipment failure and other reasons during the data collection process. In order to ensure the accuracy of the model, the tea production data of the batch containing missing values ​​is removed, and the average value of the remaining data represents the process parameters of the batch. Outliers refer to values ​​that are significantly deviated from the real data. Unchanged data refers to the situation where the process parameters of some units in the actual production process do not change. This is because they have little impact on the process, and changing them has little effect on the quality of the tea. Moreover, data that do not change in the actual modeling process is meaningless, so they need to be eliminated.

[0063] Step S3: The importance of each feature for each chemical component was calculated by the Pearson correlation coefficient method and the SHAP method, and a preliminary analysis of the contribution of the input variables to the prediction results was performed.

[0064] To model the quality of tea production, we first need to determine the contribution of each feature to the chemical composition. The Pearson correlation coefficient method is a model for measuring the linear correlation between variables, which helps to screen the features. The Pearson correlation coefficient is a statistical tool for quantifying the strength of the linear relationship between two variables. The correlation coefficient ranges from 0 (no correlation) to 1 (perfect positive correlation) and is expressed through a color gradient map (such as Figure 3 The vertical axis variables in the figure are composed of tea parameters (water content, image and other information) and process parameters of the four links of withering, rolling, primary baking and re-baking. The horizontal axis variables represent the 8 chemical components of the raw tea measured. The value in each cell is the correlation coefficient between the input variable and the chemical component, and * represents the degree of significance. In the figure, an asterisk (*) indicates a statistically significant correlation (p<0.05), two asterisks (**) indicate a highly statistically significant correlation (p<0.01), and three asterisks (***) indicate an extremely high statistically significant correlation (p<0.001). The characteristic importance of each chemical component was calculated using the SHAP method. Figure 4a-4c These are the bee swarm diagrams of the feature importance of the three chemical components ECG, GA, and CAF. The distribution of the results in the diagram is arranged according to the feature importance. The wider the distribution range of the feature's scatter points on the x-axis, the more important the feature is. And the color of the scatter points along the way is determined according to the size of the sample. Specifically, small sample values ​​are represented by blue, and large sample values ​​are represented by red. In this way, the specific contribution of the feature size to the feature change can be judged based on the distribution scatter diagram of different samples of each feature. From the figure, we can see that the tenderness of tea leaves is the most important influencing factor for the change of each chemical component.

[0065] Step S4: By applying the variable importance projection method (VIP), the importance values ​​of the features are calculated, and the features with VIP values ​​greater than 1 are selected for modeling. Then, the chemical composition prediction model was established based on the complete data set, the data set selected by the VIP method, and the data set selected by the recursive feature elimination method combining the Pearson correlation coefficient, SHAP importance and VIP value.

[0066] (1) Partial Least Squares Regression (PLSR)

[0067] For k chemical compositions and moisture contents {y1,y2,…,y k}, m unit process parameters {x1, x2,…, x m} data, the matrix composed of n samples of data is Y = {y1, y2, ..., y k} n×k and X={x1,x2,…,x m} n×m In order to study the relationship between the independent variable and the dependent variable, the first principal component is extracted from X, which should contain as much information as possible about the original data, and then a linear regression equation is established between the principal component and Y. If the accuracy of the regression equation meets the requirements, the algorithm terminates; otherwise, the second principal component is extracted from the remaining independent variables and the equation is established until the accuracy meets the requirements. The specific calculation steps of PLSR are divided into principal component decomposition and correlation matrix B calculation. The specific formula is as follows: Decompose the X and Y matrices into characteristic factors:

[0068]

[0069] Where T and U represent the principal component score matrices of matrices X and Y, respectively; P and Q represent the loading matrices of X and Y, respectively; E and F represent the error matrices of X and Y in the fitting process, respectively; the linear regression model is established using matrices T and U; and B represents the correlation matrix.

[0070] U=TB (2)

[0071] B=T T U(T T T) -1 (3)

[0072] (2) Recursive feature elimination based on Pearson correlation coefficient, SHAP importance, and VIP value

[0073] Method

[0074] Recursive Feature Elimination (RFE) is a model-based feature selection method that selects the optimal feature subset by repeatedly training the model and eliminating the least important features. The RFE algorithm can avoid overfitting problems and improve the generalization ability of the model. At the same time, since it can select the most important features from all features, it can improve the efficiency and accuracy of the model. The steps are as follows:

[0075] Initialize the model: First, train an initial model using the training data and calculate the importance score of each feature (e.g., the coefficients of the model, the impact of the feature on the target variable, etc.).

[0076] Feature sorting: Then, the features are sorted according to their importance scores, and several features with the lowest scores are selected as the features to be eliminated.

[0077] Feature removal: Retrain the model on the remaining features and calculate new feature importance scores.

[0078] Iteration process: If the number of features reaches the preset target or all features have been eliminated, stop the algorithm; otherwise, return to the second step.

[0079] Result selection: Finally, the remaining features are selected as the final feature subset

[0080] The feature importance calculated by the RFE method depends on the type of model used. When the model used is a linear model, the importance of the feature is evaluated by the absolute value of the model coefficient. The larger the absolute value of the coefficient, the greater the influence of the feature on the model; when the model used is a tree model, such as a decision tree or a random forest, the importance of the feature is evaluated by the number of times the feature is used as a split node in the tree, or by the degree to which the feature reduces impurity; when the model used is a support vector machine (SVM), the importance of the feature is evaluated by the Lagrange multiplier (corresponding to the coefficient of the feature). The recursive feature elimination (RFE) method calculates the importance of features and screens them by recursively removing the features that contribute the least to the model. Although its idea is simple and intuitive, it has some disadvantages. First, RFE depends on the performance of the selected model and is highly sensitive to the model type and hyperparameters, which may lead to unstable results; second, since the model needs to be trained multiple times to gradually remove features, the computational cost is high, especially when the amount of data is large or the number of features is large; in addition, RFE only considers the contribution of features to the current model and may ignore the interaction between features.

[0081] In contrast, feature screening methods based on Pearson correlation coefficient, SHAP (Shapley Additive Explanations) importance, and VIP (Variable Importance in Projection) value have the following advantages:

[0082] Pearson correlation coefficient: It can directly reflect the linear correlation between the feature and the target variable. It is simple and efficient to calculate and is suitable for quickly screening features with strong linear correlation.

[0083] SHAP Importance: By assigning the contribution of features to model predictions, it can explain the feature importance of complex models (such as integrated models), taking into account the interaction between features, and the results are more interpretable.

[0084] The VIP value is designed for partial least squares regression (PLSR) and measures the importance of a feature by combining its weight and variance contribution in the model. It not only considers the correlation between the feature and the target variable, but also reflects its global contribution in model prediction.

[0085] In summary, Pearson importance, SHAP importance, and VIP value have more significant advantages in computational efficiency, model adaptability, and feature interpretability, and can provide a more comprehensive and reliable reference for feature screening.

[0086] The results of the established models are shown in Tables 2, 3, and 4, which are respectively based on the dataset established with all features, the dataset established with features with VIP values ​​greater than one, and the dataset screened by the recursive feature elimination method based on Pearson correlation coefficient, SHAP importance, and VIP value. The results show that the model constructed using the dataset established by the proposed method has the best effect.

[0087] Table 2. Performance of the raw tea chemical composition prediction model based on the complete dataset

[0088]

[0089] Note: LV stands for latent variable

[0090] Table 3. Performance of the raw tea chemical composition prediction model based on the dataset screened by the VIP method

[0091]

[0092] Note: LV stands for latent variable

[0093] Table 4. Performance of the raw tea chemical composition prediction model based on the dataset screened by the RFE-PSV method

[0094]

[0095] Note: LV stands for latent variable

[0096] Step S5: According to the prediction model, with the chemical composition of raw tea as the target, a heuristic algorithm is used to optimize the process parameters, and the optimization results are applied to the software parameter control of the tea primary processing production line to realize personalized tea production; through a multi-objective particle swarm algorithm, according to the fresh tea leaves and the set production goals (three chemical components of raw tea), the process parameters of each unit are optimized with the actual application range of the process parameters of each unit as the constraint condition, and the production of each lower machine is controlled by the production line control software.

[0097] For the machine learning parameter optimization model, the specific implementation method is:

[0098] According to step S4, an optimal chemical composition prediction model is established, and the machine learning model is saved as a .pkl file through the joblib library. The model is embedded in the multi-objective particle swarm optimization algorithm to determine the independent variables and dependent variables of the optimization algorithm. After the variables are determined, the algorithm starts to optimize. First, the speed and position of the particle are initialized, and the individual optimal position and fitness value are set for each particle. At the same time, the algorithm also initializes the global optimal solution and the Pareto solution set as the starting point of the optimization. Next, in each iteration, the algorithm calculates the objective function value according to the current position of the particle, and continuously updates the individual optimal solution of the particle by comparing the current fitness value with the individual historical optimal value. In addition, the algorithm also checks whether the current solution is dominated by the existing Pareto solution set. If the current solution is not dominated, it is added to the Pareto solution set, and the old solution dominated by the new solution is removed to ensure that the Pareto solution set always retains a valid solution. In order to further optimize the candidate solutions in the Pareto solution set, the algorithm calculates the crowdedness of the solution and combines the roulette selection method to determine the global optimal solution. Using the global optimal solution and the individual optimal solution, the algorithm updates the particle's velocity and position, while ensuring that the particle's position always remains within the specified range to avoid out-of-bounds problems. Finally, after multiple iterations, the algorithm outputs the Pareto solution set obtained during the optimization process and the corresponding processing parameters.

[0099] After the calculation is completed, the appropriate unit process parameters can be selected according to the fresh tea leaves and production goals, thus guiding tea production and achieving the optimal configuration for personalized tea production according to requirements.

[0100] Of course, the above description is only a specific embodiment of the present invention and is not intended to limit the scope of implementation of the present invention. All equivalent changes or modifications made according to the structure, characteristics and principles described in the patent application scope of the present invention should be included in the patent application scope of the present invention.

[0101] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. Online parameter modulation and control method for tea production line customized for functional ingredients, characterized by The steps include: Step S1: Acquire multiple unit data sets of a tea production line, including fresh tea leaf grade, fresh tea leaf moisture content, process parameters of each tea processing step, fresh leaf image texture information and chromaticity information; measure the contents of multiple chemical components of tea leaves of different grades through UPLC experiments; Step S2: preprocessing multiple unit data sets to remove outliers; Step S3: The importance of each feature for each chemical component is calculated by the Pearson correlation coefficient method and the SHAP method, and a preliminary analysis of the contribution of the input variables to the prediction results is performed; Step S4: Calculate the importance value of the feature by applying the variable importance projection method VIP, and select the features with VIP value greater than 1 for modeling; Chemical composition prediction models were established based on the complete dataset, the dataset screened by the VIP method, and the dataset screened by the recursive feature elimination method based on the Pearson correlation coefficient, SHAP importance, and VIP value. Step S5: According to the chemical composition prediction model, taking the chemical composition of raw tea as the target, optimizing the process parameters, and applying the optimization results to the software parameter control of the primary tea processing production line to realize personalized tea production.

2. The method for online parameter modulation and control of a tea production line customized for functional ingredients according to claim 1 is characterized in that: The fresh tea leaves in step S1 are divided into six categories, with the quality ranging from level 1 to level 6 from high to low, and are judged according to the leaf shape of the tea leaves; the process parameters of the withering unit and the rolling and drying unit include hot air temperature, drum speed, feeding speed, and dehumidification speed, wherein the process parameters of the rolling and drying unit also include hot air speed, and the process parameters of the primary drying unit and the secondary drying unit include hot air temperature, oil inlet temperature, oil outlet temperature, main conveying speed, hot air speed, and leaf leveling speed; The texture information and chromaticity information of fresh leaf images include the color features and texture features of fresh tea leaf images, wherein the color features include HSV, LAB and RGB, and the texture features include the mean, standard deviation, smoothness, third-order moment, consistency and entropy of fresh tea leaf images; The chemical components of raw tea include gallic acid GA, catechin C, epicatechin EC, epigallocatechin EGC, epicatechin gallate ECG, epigallocatechin gallate EGCG, theanine L-The and caffeine CAF. The above chemical component standards are prepared into standard solutions, and the contents of chemical components in the raw tea are calculated according to the standard curve and peak area based on liquid chromatography of the standard solutions.

3. The method for online parameter modulation and control of a tea production line for functional ingredient customization according to claim 1 is characterized in that: The preprocessing in step S2 is missing value processing and outlier processing. The missing value processing is to remove the corresponding batch of tea production data containing missing values, and use the average value of the remaining data to represent the process parameters of the batch; Outlier processing is to remove values ​​that are significantly deviated from the true data.

4. The method for online parameter modulation and control of a tea production line for functional ingredient customization according to claim 1 is characterized in that: The different data sets in step S4 include all data sets, data sets screened by using the VIP method, and data sets screened by using a recursive feature elimination method based on Pearson values, SHAP values, and VIP values.

5. The method for online parameter modulation and control of a tea production line for functional ingredient customization according to claim 1, characterized in that: The step S4 adopts a machine learning model including partial least squares regression PLSR, for k chemical components and moisture contents {y1, y2, ..., y k }, m unit process parameters {x1, x2,…, x m } data, the matrix composed of n samples of data is Y = {y1, y2, ..., y k } n×k and X={x1,x2,…,x m } n×m , extract the first principal component from X, and then establish a linear regression equation between the principal component and Y. If the accuracy of the regression equation meets the requirements, the algorithm terminates; otherwise, extract the second principal component from the remaining X and establish an equation until the accuracy meets the requirements. The specific calculation of PLSR includes principal component decomposition and correlation matrix B calculation. Principal component decomposition is to decompose the X and Y matrices into characteristic factors. The specific formula is as follows: Where T and U represent the principal component score matrices of matrices X and Y, respectively; P and Q represent the loading matrices of X and Y, respectively; and E and F represent the error matrices of X and Y during the fitting process, respectively; U=TB B=T T U(T T T) -1 The linear regression model is built using matrices T and U, and B represents the association matrix.

6. The method for online parameter modulation and control of a tea production line for functional ingredient customization according to claim 2, characterized in that: The fresh leaf image acquisition hardware includes an aluminum profile, a camera body, a support frame and an annular light source. The support frame is fixed on the aluminum profile, and the camera body and the annular light source are fixed on the support frame to obtain fresh leaf image information.

7. The method for online parameter modulation and control of a tea production line for functional ingredient customization according to claim 4 is characterized in that: The recursive feature elimination method based on Pearson correlation coefficient, SHAP importance and VIP value in step S4 comprises the following steps: A. Calculate the importance weights of all features in three cases and normalize them; B. Add these three values ​​with equal weight; C. Sort by weight, sort by importance from high to low, and then build and optimize the model according to this sorting under different numbers of features to obtain the optimal number of features and corresponding hyperparameters.

8. The method for online parameter modulation and control of a tea production line for functional ingredient customization according to claim 1, characterized in that: The step S5 adopts a multi-objective particle swarm optimization algorithm. After multiple iterations, the algorithm outputs the Pareto solution set and corresponding processing parameters obtained during the optimization process, and produces the optimal configuration of tea leaves according to the requirements.

Citation Information

Patent Citations

  • Tea primary processing production line parameter cooperative control method and system based on data driving

    CN118444649A

  • Tea production line multi-target process parameter optimization method based on improved particle swarm

    CN118627851A

  • Component design method and preparation method of titanium alloy with high comprehensive performance

    CN118711730A

Cited By

  • Device and method for detecting, regulating and controlling tea material-water ratio

    CN121411346A

  • Layered analysis and interpretable modeling method, system and equipment for multi-source driving factors of power system and medium

    CN121980275A