Data-driven parameter coordination control method and system for tea leaf primary production line

A data-driven method using machine learning models optimizes tea production parameters across multiple units, addressing labor shortages and enhancing tea quality and efficiency.

JP2026016284AActive Publication Date: 2026-02-03ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024229754
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2024-12-26
Publication Date
2026-02-03
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing tea production relies heavily on manual labor and traditional methods, lacking efficient and intelligent processing equipment to optimize process parameters across multiple units in the production line, leading to labor shortages and suboptimal tea quality.

Method used

A data-driven parameter coordination control method using machine learning models to predict moisture content and chemical components, optimizing process parameters through a multi-objective optimization algorithm, integrating data collection, preprocessing, feature selection, and model training to enhance tea quality and reduce labor.

Benefits of technology

The method improves tea production quality and efficiency by predicting moisture content and chemical components, reducing labor requirements and enhancing the intelligence and cost-effectiveness of the tea production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016284000001_ABST
    Figure 2026016284000001_ABST
Patent Text Reader

Abstract

To provide a parameter cooperative control method and system for a tea leaf primary production line based on data drive.SOLUTION: Collecting engineering parameters, image information and water content information in the production process of tea leaves through experimental equipment and production line software, pre-processing the collected data, removing abnormal data to improve the performance of the model, selecting important features and chemical components through analyzing the importance and correlation of eigenvalues, applying the tea leaf process data set to train the machine learning model, and obtaining optimal hyperparameters of a plurality of models using a grid optimization method; The modeling results of different models in different units are compared, the optimal model of the roll-drying unit is PLSR, the water content of each unit and the chemical composition of the tea are taken as targets, the process parameters are optimized by using a multi-purpose algorithm, and the optimization result is applied to the tea leaf production line control software to achieve high-quality tea leaf production.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of machine learning and tea first manufacturing processing, and more particularly to a method and system for data-driven parameter coordination control of a tea first manufacturing production line. [Background technology]

[0002] The main components of tea leaves include flavonoids, alkaloids, phenols, and theanine, which have antioxidant, anti-inflammatory, anti-cancer, cardioprotective, and antibacterial properties. The functions of these parts of the tea leaf are closely related to the processing process. Currently, tea processing relies mainly on the traditional "see, touch, and smell" method, which requires a high level of experience. However, there is currently a labor shortage in tea production, and the lack of experienced workers is a bottleneck in the high-quality and rapid development of the tea industry. Therefore, it is particularly important to adopt efficient, clean, and intelligent processing equipment to replace manual labor and realize automated, digital, and intelligent tea processing lines.

[0003] Existing technologies mainly detect the tea ingredient content and moisture content through spectral collection of tea leaves at the discharge stage and neural networks, but there has been no precedent for modeling all units in the entire production line and building neural network models to use optimization algorithms to optimize process parameters for multiple units. Summary of the Invention [Problem to be solved by the invention]

[0004] In order to solve the deficiencies of existing technology, improve tea production quality and reduce labor, the present invention adopts the following technical solutions: [Means for solving the problem]

[0005] A data-driven parameter coordination control method for a tea leaf initial production line, comprising the following steps: Step S1: Obtain a data set of multiple units of a tea leaf initial production line, including process parameters of each unit, moisture content before and after processing, raw leaf image information, tea quality, and multiple chemical component contents of the final tea leaf; Step S2: Preprocessing the data sets of the multiple units to remove erroneous data and reduce the impact of equipment shortages and noise on the model building results; Step S3: sorting the data, and based on the data changes of each unit in the production process, removing basically unchanged data and / or data with the same change trend, and using machine learning to predict the importance of the process parameters and raw leaf image information on the corresponding moisture content and chemical components, and removing data with low importance, and selecting chemical components related to the quality, and these chemical components have no value for prediction or optimization, Step S4: Construct a machine learning model to relate the sorted process parameters and raw leaf image information to moisture content, and relate the sorted process parameters and raw leaf image information to chemical components related to quality after sorting; Step S5: Based on the machine learning model, construct a moisture content prediction model for each unit and a plurality of chemical component prediction models for the final tea leaves, evaluate the performance of the prediction models based on RMSE and R2, select the machine learning model with the best performance, and display the relationship between the process parameters and the moisture content and chemical components; Step S6: Based on the prediction model, the moisture content of each unit and the chemical components of the tea are targeted, and the process parameters are optimized based on the raw leaf image information. The optimization results are applied to the software parameter control of the tea leaf initial production line to achieve high-quality tea leaf production.

[0006] Furthermore, the multiple units in step S1 include a killing unit, a drum roasting unit, a primary roasting unit, and a second roasting unit, and these four units have large changes in moisture content and are highly related to quality formation. The process parameters of the killing unit and the drum roasting unit both include temperature, drum speed, input speed, and moisture removal speed, where the process parameters of the drum roasting unit further include hot air temperature and hot air speed, and the process parameters of the primary roasting unit and the second roasting unit both include hot air temperature, oil input temperature, oil output temperature, main conveying speed, hot air speed, and leaf homogenization speed. The raw leaf image information includes color features and texture features of the raw tea leaf image, where the color features include HSV, LAB, and RGB, and the texture features include mean value, standard deviation, smoothness, third moment, consistency, and entropy of the raw tea leaf image; The quality of Maocha is graded based on the appearance and internal quality of the tea leaves. The appearance includes the stripes, color, luster, degree of crushing, and cleanliness of the tea leaves. The internal quality includes the color, aroma, taste, and base of the tea leaves. The chemical components of Mao Tea include gallic acid (GA), catechin (C), epicatechin (EC), epigallocatechin (EGC), epicatechin gallate (ECG), epigallocatechin gallate (EGCG), theanine (L-The), and caffeine (CAF). Standards of the above chemical components are prepared as standard solutions, and the chemical component contents in the tea leaves are calculated from the standard curve and peak area based on liquid phase chromatography of the standard solutions.

[0007] Furthermore, the pre-processing in step S2 includes missing value processing and outlier processing. Missing values ​​are data that is missing during the data collection process due to reasons such as equipment failure. In order to ensure the accuracy of the model, the tea leaf production data for the relevant lot containing missing values ​​is deleted, and the average value of the remaining data is used to represent the process parameters for that lot. Outliers refer to values ​​that are clearly significantly different from the actual data. The process parameters need to be adjusted multiple times to obtain ideal parameters at the start of processing, and data before the ideal parameters are obtained is considered abnormal data. Errors in moisture content are caused by the accuracy of the equipment. The actual production situation and the range of change in moisture content must be taken into consideration comprehensively.

[0008] Furthermore, the machine learning model in step S4 includes a support vector machine (SVM), and the SVR regression function of the support vector machine SVM for a set of data sets is as follows:

[0009]

number

[0010]

number

[0011] where ω represents the feature weight vector, C, εi and εi * are the penalty parameter and two slack variables, respectively, δ represents the insensitive interval of the slack variable, Xi represents the feature vector of the process parameters and raw leaf image information in the dataset, yi represents the moisture content feature vector of the corresponding tea leaves, and n represents the number of feature vectors, which are calculated using the Lagrangian function as follows:

[0012]

number

[0013] In the formula, αi and αi *represents two Lagrange multipliers, b represents deviation, and K(Xi,Xi ' ) represents the kernel function and is calculated as follows:

[0014]

number

[0015] where Xi and Xi' are two feature vectors in the input space, and γ represents the coefficient of the kernel function. Common kernel functions include linear, sigmoid, and radial basis functions. When constructing and analyzing an SVR model, it is necessary to optimize the hyperparameters C, δ, and γ.

[0016] Furthermore, the machine learning model in step S4 includes partial least squares regression (PLSR), and k chemical components and moisture contents {y1, y2, ..., y k}, m unit process parameters {x1, x2, ..., x m}, the matrix consisting of n sample data is Y={y1, y2, ..., y k} n×k and X={x1, x2, ..., x m} n×m To study the relationship between the independent and dependent variables, extract the first principal component from X, which requires information that contains as much of the original data as possible. Then, construct a linear regression equation between the principal component and Y. If the accuracy of the regression equation meets the requirements, the algorithm is terminated. If the accuracy of the regression equation does not meet the requirements, extract the second principal component from the remaining independent variables X until the accuracy meets the requirements. The specific calculation of PLSR involves principal component decomposition and the calculation of the related matrix B. Principal component decomposition is to perform feature factorization on the X and Y matrices. The specific formula is as follows:

[0017]

number

[0018] where T and U represent the principal component score matrices of X and Y respectively, P and Q represent the weight matrices of X and Y respectively, and E and F represent the error matrices during the fitting process in X and Y respectively.

[0019]

number

[0020] A linear regression model is constructed using matrices T and U, and B represents the association matrix.

[0021] Furthermore, the machine learning model in step S4 includes multiple linear regression (MLR), which is used to obtain fitting relationships between variables and can explain how a single dependent variable linearly depends on multiple predictor variables, and the formula for multiple linear regression with n predictor variables is as follows:

[0022]

number

[0023] In the formula, x represents the process parameters and raw leaf image information of each unit, y represents the moisture content after processing, and β0, β1, β2, ..., β n represents the polynomial coefficients, and ε represents the residual term of the regression model.

[0024] The parameter coordinated control of the tea leaf initial production line based on data driving includes a collection module, a pre-processing module, a data selection module, a machine learning module, a prediction module, and a control module. The collecting module acquires a plurality of unit data sets of the primary tea production line, including process parameters of each unit of the plurality of units, moisture content before and after processing, raw leaf image information, quality of the tea leaves, and the contents of a plurality of chemical components of the final tea leaves; The pre-processing module pre-processes the plurality of unit data sets, removes erroneous data, and reduces the influence of equipment shortages and noise on the model construction results; The data selection module, based on the data changes of each unit in the production process, removes data that does not change fundamentally and / or data that has the same change trend, and uses machine learning to predict the importance of the process parameters and raw leaf image information on the corresponding moisture content and chemical components, removes data with low importance, and selects chemical components related to the quality as chemical components, and these chemical components have no value for prediction or optimization. the machine learning module is for relating the sorted process parameters and the fresh leaf image information to moisture content, and for relating the sorted process parameters and the fresh leaf image information to chemical components related to quality after sorting; The prediction module builds a moisture content prediction model for each unit and a multiple chemical component prediction model for the final tea based on a machine learning model, and calculates RMSE and R 2 The performance of the prediction model is evaluated based on the results, and the best-performing machine learning model is selected. The relationship between the process parameters and the moisture content and chemical components is displayed. The moisture content and chemical components of each unit are targeted, and the process parameters are optimized based on the raw leaf image information. The optimization results are obtained. The control module applies the optimization results to software parameter control of the tea leaf initial production line to achieve high quality tea leaf production.

[0025] Further, the collection module includes a camera body, a camera lens, and a ring light, wherein the camera body is connected to the camera lens and the ring light is installed around the camera lens, and is used to acquire raw leaf image information including color features and texture features of the raw tea leaf image, wherein the color features include HSV, LAB, and RGB, and the texture features include the mean value, standard deviation, smoothness, third-order moment, consistency, and entropy of the raw leaf image.

[0026] Furthermore, the multiple units of the primary tea production line include a green kill unit, a drum roast unit, a primary roast unit, and a second roast unit, and these four units have large changes in moisture content and are highly related to the formation of quality. The process parameters of the green kill unit and the drum roast unit all include temperature, drum speed, input speed, and moisture removal speed, where the process parameters of the drum roast unit further include hot air temperature and hot air speed, and the process parameters of the primary roast unit and the second roast unit all include hot air temperature, oil input temperature, oil output temperature, main conveying speed, hot air speed, and leaf homogenization speed.

[0027] Furthermore, the collection module includes an ultra-high efficiency liquid phase chromatograph device, which is used to obtain ultra-high efficiency liquid chromatographs of the standard solution of chemical components of Mao Tea, and calculate the chemical component contents in the tea leaves based on the standard curve and peak area, where the chemical components of Mao Tea include gallic acid GA, catechin C, epicatechin EC, epigallocatechin EGC, epicatechin gallate ECG, epigallocatechin gallate EGCG, theanine L-The, and caffeine CAF. [Effects of the Invention]

[0028] The advantages and beneficial effects of the present invention are as follows: This invention proposes a data-driven parameter coordination control method and system for a tea production line. First, data is collected during the production process, and the data is preprocessed and feature-selected. Then, different machine learning models are used to establish the relationship between the process parameters and the production targets, thereby realizing the prediction of the moisture content of each unit and multiple chemical components of the tea. Finally, a multi-objective optimization algorithm is combined to optimize the process parameters, and the optimization results are sent to multiple lower-level computer devices through a computer program for actual production, thereby improving the quality of tea production and reducing labor. The method is characterized by low cost, high efficiency, and intelligence. [Brief explanation of the drawings]

[0029] In order to more clearly describe the specific embodiments of the present invention or the technical solutions of the prior art, the following briefly introduces drawings used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and a person skilled in the art can obtain other drawings based on these drawings without any creative efforts.

[0030] [Figure 1] FIG. 2 is a flowchart of the data-driven collaborative control method for primary tea production line parameters according to the present invention; [Figure 2] 1 is a structural schematic diagram of the tea leaf image collection system of the present invention; [Figure 3] [Fig. 3A] A diagram showing the results of analyzing abnormal moisture content data for each tea leaf production unit in the present invention using a box plot. [Fig. 3B] A diagram showing the results of analyzing abnormal moisture content changes for each tea leaf production unit in the present invention. [Figure 4] FIG. 1 is a distribution diagram of importance of chemical components based on SHAP of the present invention. [Figure 5] FIG. 1 is a diagram of the correlation between process parameters and chemical composition of the present invention. [Figure 6] FIG. 1 is a block diagram of the software architecture of the data-driven primary tea production line parameter coordination control system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031] Having described the present invention through several embodiments, those skilled in the art will understand that various modifications or equivalent substitutions can be made to these features and examples without departing from the spirit and scope of the present invention. Furthermore, the features and examples can be modified to adapt a particular situation and material to the teachings of the present invention without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific examples disclosed herein, and all examples falling within the scope of the claims of this application are within the scope of protection of the present invention.

[0032] As shown in Figure 1, the parameter coordination control method for the tea leaf initial production line based on data driving includes the following steps: Step S1: Obtain multiple unit data sets of the tea leaf initial production line, including process parameters of each unit, moisture content before and after processing, raw leaf image information, tea quality, and multiple chemical component contents of the final tea leaf.

[0033] The tea production line consists of multiple units: sterilization, drum roasting, primary roasting, and re-roasting. These four units experience significant changes in moisture content and are highly correlated with quality. The process parameters for each unit in the initial tea processing are determined by experienced operators based on information such as the tea leaf's freshness. In step S6, the tea production line control software communicates with each piece of tea production equipment, collecting process parameter data for each unit every minute and storing it in the computerized system. The process parameters for each unit are listed in Table 1. The moisture content of each unit before and after tea processing is measured using an automatic moisture analyzer (MA150). The analyzer's operating temperature is set to 130°C. 3g tea leaf samples are uniformly selected from the production line and spread evenly on a sample tray. The moisture evaporation channel is used to rapidly roast the tea leaf samples, and the moisture content of the sample is displayed in real time on the device's interface.

[0034] [Table 1]

[0035] Images of fresh tea leaves were captured using a homemade experimental platform. This platform consisted of an industrial camera body (MV-CS060-10GM / GC), a camera lens (MVL-HF1228M-6MPE), a high-brightness ring-shaped diffuse light source (XS-WAR08036-90), a connecting plate (3), and an aluminum frame stand (2). The lens was connected to the camera with screws, and the camera was fixed to the connecting plate (3). The light source (5) was fixed below the connecting plate (3), and the connecting plate (3) was further fixed to the aluminum frame stand (2) with a trapezoidal nut. The specific connection structure is shown in Figure 2. Ten grams of fresh tea leaves from each batch were uniformly selected and images were collected. Nine color features and six texture features were then extracted from the images using a Python program. The color features included HSV, LAB, and RGB, and the texture features included mean, standard deviation, smoothness, third-order moment, consistency, and entropy.

[0036] The grade of Mao Tea is determined by five professional tea judges, who evaluate the appearance (stripe, color, degree of crushing, cleanliness) and internal quality (color of the tea brew, aroma, taste, bottom of the leaf) according to the sensory evaluation requirements. Mao Tea is divided into six grades, including 14 special A grade, 22 special grade, 14 special first grade, 27 first grade, 7 second grade, and 16 third grade.

[0037] The chemical components of Mao Tea to be measured included gallic acid (Ga), catechin (C), epicatechin (EC), epigallocatechin (EGC), epicatechin gallate (ECG), epigallocatechin gallate (EGCG), theanine (L-The), and caffeine (CAF). Standard solutions were prepared using the above chemical component standards. A series of standard solutions with standard concentrations of each component of 0.05-0.5 mg / L, 0.005-0.1 mg / L, 0.05-0.5 mg / L, 0.15-1.2 mg / L, 0.025-0.25 mg / L, 0.3-1.5 mg / L, 0.05-0.5 mg / L, and 0.2-1.0 mg / L, respectively, were prepared and applied to an ultra-high-performance liquid chromatography system. A standard curve was calculated based on the actual content of each component and its corresponding peak area. 0.2g of tea leaves was placed in a centrifuge tube, ultrapure water was added, and the tea was extracted in a water bath for 60 minutes. After that, the sample was prepared by centrifugation and extraction of the supernatant. The sample was analyzed using an ultra-high-efficiency liquid phase chromatography system, Water ACQUITY UPLC™, and the contents of eight chemical components in the tea leaves were calculated based on the standard curve and the peak areas of the sample.

[0038] Step S2: Preprocess the acquired process parameter dataset of the tea leaf production line and remove erroneous data to reduce the impact of equipment defects and noise on the model construction results.

[0039] Specifically, preprocessing includes missing value processing and outlier processing. Missing values ​​are data that is missing during the data collection process due to factors such as equipment failure. To ensure the accuracy of the model, the tea production data for the corresponding batch containing missing values ​​is deleted and the average value of the remaining data is used to represent the process parameters for that batch. Outliers refer to values ​​that are clearly significantly different from the actual data. Process parameters need to be adjusted multiple times at the start of processing to obtain ideal parameters, and data before the ideal parameters are obtained is considered abnormal data. Errors in moisture content are caused by equipment precision, and the actual production situation and the range of change in moisture content must be comprehensively considered. The analysis results are shown in Figures 3a and 3b. In the box plots, some moisture content points for the drum roast, first roast, and second roast units are outside the range of the box plots. This data is considered to be outliers, and the data for the relevant lots were excluded. In the moisture content change diagrams for each unit, the moisture content of the tea leaves in some lots increased from the killing process to the drum roasting process. This did not match the actual production situation, so the data for the relevant lots were excluded as they were considered to be outliers.

[0040] Step S3: Combine the actual production process and feature importance to select features related to moisture content and chemical components to build a model, and select important chemical components based on the relationship between chemical components and quality.

[0041] Feature selection is required for machine learning models. First, selection is based on the change in characteristics during actual production. Features that do not change during actual production or that have the same change trend are removed. For example, the five temperatures during the greening process tend to be 10°C higher in the previous area than in the next area. Next, feature selection is performed on tea leaf production data. Each machine learning prediction model is then used to calculate the importance of each feature using the Feature Importances function in the Sklearn library. Based on the feature importance of each model, features are reduced in order, removing the least important feature and using the remaining features as model inputs to train the model. This process is repeated until the model achieves the highest accuracy, thereby ensuring model performance. When building chemical component models, chemical components not related to quality must be removed. These chemical components have no value in prediction or optimization. The feature importance of each chemical component is calculated using the SHAP method. Figure 4 shows the distribution of feature importance for chemical components. The results show that ECG, CAF, EGCG, and GA are chemical components highly related to quality. Figure 5 shows the correlation between process parameters and chemical components. The first 26 variables in the figure consist of tea leaf parameters (information such as moisture content and images) and process parameters for the four processes of sterilization, roasting, primary roasting, and secondary roasting. The last four variables represent the chemical components of Mao-cha tea: GA, CAF, EGCG, and ECG. Of these, GA, CAF, and ECG have strong correlations with process parameters, making them suitable for building multi-task regression prediction models.

[0042] Step S4: The above dataset is used to train a machine learning model, including a multiple linear regression (MLR), a partial least squares regression (PLSR), a support vector machine (SVM), a random forest (RF), and a multi-task random forest (MRF) model, and the hyperparameters of the above model are optimized.

[0043] (1) Support Vector Machine (SVM) SVM is a type of supervised machine learning algorithm that uses a dataset D={(X1,y1), (X2,y2), ..., (X n ,y n )}, where X i ∈R m represents the input parameter vector, such as process parameters and raw leaf image information, m represents the number of input parameters, and y i ∈R represents the moisture content of tea leaves, which is the output parameter, and its SVR regression function is expressed as follows:

[0044]

number

[0045] where ω represents the feature weight vector and φ(X i ) represents a function that maps the data to a high-dimensional feature space through a nonlinear transformation, and b represents the bias. To find a suitable SVR regression function f(xi), the problem can be transformed into the following equation:

[0046]

number

[0047] In the formula, C, ε i and ε i * where represent the penalty parameter and the two slack variables, respectively, and δ represents the insensitive interval of the slack variables, which are solved using the Lagrangian function, and the final result is the solution of the following function:

[0048]

number

[0049] In the formula, α i and α i * represents the Lagrange multiplier, b is a constant value, and K(Xi , X i ' ) represents the kernel function.

[0050]

number

[0051] In the formula, X i and X i ' are two feature vectors in the input space, and γ represents the coefficients of the kernel function. Common kernel functions include linear, sigmoid, and radial basis functions. When building and analyzing an SVR model, it is necessary to optimize the hyperparameters (C, δ, and γ).

[0052] (2) Partial Least Squares Regression (PLSR) k chemical components and moisture content {y1, y2, ..., y k}, m unit process parameters {x1, x2, ..., x m}, the matrix consisting of n sample data is Y={y1, y2, ..., y k} n×k and X={x1, x2, ..., x m} n×m To study the relationship between the independent and dependent variables, the first principal component is extracted from X, and information containing as much of the original data as possible is required. Next, a linear regression equation is constructed between the principal component and Y. If the accuracy of the regression equation meets the requirements, the algorithm terminates. If the accuracy of the regression equation does not meet the requirements, the second principal component is extracted from the remaining independent variables and the equation is established until the accuracy meets the requirements. The specific calculation steps of PLSR include principal component decomposition and the calculation of the related matrix B. The specific formula is as follows. Feature factorization is performed on the X and Y matrices as follows:

[0053]

number

[0054] In the formula, T and U represent the principal component score matrices of X and Y, respectively; P and Q represent the weight matrices of X and Y, respectively; E and F represent the error matrices during the fitting process in X and Y, respectively; matrices T and U are used to construct a linear regression model; and B represents the association matrix.

[0055]

number

[0056] (3) Multiple linear regression The multiple linear regression method is used to obtain the fitting relationship between variables, and can explain how a single dependent variable is linearly dependent on multiple predictor variables, and the equation for multiple linear regression with n predictor variables can be described as follows:

[0057]

number

[0058] In the formula, x represents the process parameters and raw leaf image information for each unit, y represents the moisture content after processing, and β0, β1, β2, ..., β n represents the polynomial coefficients, and ε represents the residual term of the regression model.

[0059] Using the above machine learning model and the preprocessed process parameter dataset, the data was divided into a training set and a test set in a ratio of 8:2. These models were trained and the hyperparameters of the models were optimized using the grid search method. The optimal hyperparameters for the moisture content prediction model for each unit are shown in Table 2, and the optimized hyperparameters for the chemical component prediction model are shown in Table 3. The optimized hyperparameters for the random forest are the number of decision trees (n_estimators), the maximum depth of the decision tree (max_depth), and the minimum number of samples required for node splitting (min_samples_split). In support vector machine regression, the optimized hyperparameters are the regularization parameter C, the kernel function coefficient γ, and the kernel function kernel. In partial least squares regression, the optimized hyperparameter is the number of principal components n.

[0060] [Table 2]

[0061] [Table 3]

[0062] Step S5: Based on the data set, features, and hyperparameter optimization results, a moisture content prediction model for each unit and a prediction model for multiple chemical components of the final tea are constructed, and the RMSE and R 2 The performance of the predictive models is evaluated based on the results, and the best-performing machine learning model is selected to show the relationship between process parameters and moisture content and chemical components. The results are shown in Tables 4 and 5. In terms of the accuracy of each unit model, the drum roast unit performed best, with PLSR being the optimal model, while SVM was the optimal model for kill-greening, primary roasting, and re-roasting. Comparing the two models, RF and SVM, the RF model outperformed the SVM model on test set R2 for Task 1 and Task 3, while the performance of the models was reversed for Task 2. Each model has its own advantages and disadvantages, and when compared with the multi-task MRF model, the single-task model's predictive performance for each task was inferior to that of the multi-task model.

[0063] [Table 4]

[0064] [Table 5]

[0065] Step S6: Based on the process parameters and target relationships (moisture content, chemical components) in step S5, the moisture content of each unit and the chemical components of the tea are targeted, and the process parameters are optimized using a heuristic algorithm. The optimization results are applied to the control software of the tea leaf production line to achieve high-quality tea leaf production.

[0066] Using a multi-objective particle swarm algorithm, the process parameters of each unit are optimized based on the condition of the raw tea leaves and the set production targets (moisture content of each unit and multiple chemical components of the tea leaves), with the range of process parameters actually applicable to each unit as the limiting condition, and production at each sub-machine is controlled through the production line control software.

[0067] As shown in Figure 6, the parameter coordination control of the tea production line based on data driving is achieved through the control software of the tea production line. A specific control software architecture block diagram is shown, and the system includes a collection module, a pre-processing module, a data sorting module, a machine learning module, a prediction module, and a control module.

[0068] The collection module is for acquiring multiple unit datasets of the initial tea production line, including process parameters of each unit, moisture content before and after processing, raw leaf image information, tea quality, and multiple chemical component contents of the final tea. The collection module includes a camera body 1, a camera lens 4, and a ring light 5, wherein the camera body 1 is connected to the camera lens 4, and the ring light 5 is installed around the camera lens 4, and is configured to acquire raw leaf image information including color features and texture features of the raw tea leaf images, where the color features include HSV, LAB, and RGB, and the texture features include the mean value, standard deviation, smoothness, third-order moment, consistency, and entropy of the raw leaf images.

[0069] The multiple units in the primary tea production line include a green kill unit, a drum roast unit, a primary roast unit, and a second roast unit. These four units experience significant changes in moisture content and are highly related to quality. The process parameters of the green kill unit and the drum roast unit include temperature, drum speed, input speed, and moisture removal speed. The process parameters of the drum roast unit also include hot air temperature and hot air speed, while the process parameters of the primary roast unit and the second roast unit include hot air temperature, oil input temperature, oil output temperature, main conveying speed, hot air speed, and leaf homogenization speed.

[0070] The collection module further includes an ultra-high efficiency liquid phase chromatograph device, which is used to obtain ultra-high efficiency liquid chromatographs of the standard solution of chemical components of Mao Tea, and calculates the chemical component contents in the tea leaves based on the standard curve and peak area, where the chemical components of Mao Tea include gallic acid GA, catechin C, epicatechin EC, epigallocatechin EGC, epicatechin gallate ECG, epigallocatechin gallate EGCG, theanine L-The, and caffeine CAF.

[0071] The pre-processing module pre-processes the plurality of unit data sets to remove erroneous data and reduce the influence of equipment shortages and noise on the model building results.

[0072] The data selection module removes data that basically does not change and / or data that has the same change trend based on the change status of each unit data during the production process, and uses machine learning to predict the importance of the process parameters and raw leaf image information to the corresponding moisture content and chemical components, deletes data with low importance, and selects chemical components related to the quality, which are not worth predicting or optimizing.

[0073] The machine learning module relates post-sorting process parameters and raw leaf image information to moisture content, and relates post-sorting process parameters and raw leaf image information to chemical components related to post-sorting quality.

[0074] The prediction module builds a moisture content prediction model for each unit and a multiple chemical component prediction model for the final tea based on machine learning models, and calculates the RMSE and R 2 The performance of the predictive model is evaluated based on the above, the machine learning model with the best performance is selected, the relationship between the process parameters and the moisture content and chemical components is displayed, the moisture content of each unit and the chemical components of the Mao tea are targeted, and the process parameters are optimized based on the raw leaf image information to obtain the optimization results.

[0075] The control module applies the optimization results to the software parameter control of the tea leaf initial production line to achieve high quality tea leaf production.

[0076] The system creates a computer control program for the tea leaf production line based on the C# language, and when the computer program is executed on the processor, it has the following functions:

[0077] Data interaction function: Exchanges data with multiple sensor devices and touchscreen devices in the production site. Ethernet and RS485 are used for the hardware, and MOBUS TCP / IP and serial communication protocols can be selected for communication with the devices.

[0078] Data collection function: During the operation of the tea leaf production line, each production device generates a large amount of process parameter data. In order to manage production process parameters and realize data shareability, the data is stored in a MySql database. According to production needs, a MySql database containing databases such as multiple unit process parameters and user information is designed.

[0079] Parameter optimization module: For the above machine learning model, the machine learning model is saved as a file using the joblib library, and then imported into a computer program for application. At the same time, a multi-objective optimization algorithm is incorporated into the program. In actual application, the program can optimize the process parameters of multiple units based on the status of raw tea leaves and production targets, and send them to downstream machines to produce tea leaves.

[0080] Of course, the above is merely a specific example of the present invention, and does not limit the scope of the present invention. All equivalent changes and modifications based on the structure, features and principles described in the claims of the present invention should be included in the claims of the present invention.

[0081] Finally, it should be noted that the above examples are merely specific implementations of the present invention and are intended to illustrate the technical solutions of the present invention, not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above examples, those skilled in the art may modify or easily devise the technical solutions described in the above examples within the technical scope disclosed in the present invention, and may substitute some technical features therein with equivalents. These modifications, variations, or substitutions shall all be included in the scope of protection of the present invention, provided that the essence of the corresponding technical solutions does not depart from the spirit and scope of the technical solutions of the embodiments of the present invention. Therefore, the scope of protection of the present invention shall be determined based on the scope of protection of the claims. [Explanation of symbols]

[0082] 1 camera body 2 stands 3 Connection plate 4 camera lenses 5 light sources

Claims

1. A data-driven parameter coordination control method for a tea leaf initial production line, comprising: Step S1: Acquiring a data set of multiple units of a tea leaf initial production line, including process parameters of each unit, moisture content before and after processing, raw leaf image information, quality of the tea leaves, and the content of multiple chemical components of the final tea leaves; a step S2 of pre-processing the data set of the plurality of units; Step S3: sorting the data, removing data that does not change fundamentally and / or data that has the same change tendency based on the change status of the data of each unit in the production process, and deleting data with low importance based on the importance of the process parameters and the raw leaf image information on the corresponding moisture content and chemical components, and selecting chemical components related to the quality as the chemical components; Step S4: constructing a machine learning model to relate the sorted process parameters and the raw leaf image information to moisture content and to relate the sorted process parameters and the raw leaf image information to chemical components related to quality after sorting; Step S5: building a moisture content prediction model for each unit and a plurality of chemical component prediction models for the final tea leaves based on the machine learning model; Step S6: based on the prediction model, optimize the process parameters based on the moisture content of each unit and the chemical components of the tea leaves, and apply the optimization results to the software parameter control of the tea leaf initial production line; A parameter collaborative control method for a tea leaf initial production line based on data driving, comprising:

2. The multiple units in step S1 include a blue killing unit, a drum roasting unit, a primary roasting unit, and a second roasting unit. The process parameters of the blue killing unit and the drum roasting unit all include temperature, drum speed, input speed, and moisture removal speed. The process parameters of the drum roasting unit further include hot air temperature and hot air speed. The process parameters of the primary roasting unit and the second roasting unit all include hot air temperature, oil input temperature, oil output temperature, main conveying speed, hot air speed, and leaf homogenization speed. The fresh leaf image information includes color features and texture features of the fresh tea leaf image, where the color features include HSV, LAB and RGB, and the texture features include mean value, standard deviation, smoothness, third moment, consistency and entropy of the fresh tea leaf image; The quality of Maocha is graded based on the appearance and internal quality of the tea leaves. The appearance includes the stripes, color, luster, degree of crushing, and cleanliness of the tea leaves. The internal quality includes the color, aroma, taste, and base of the tea leaves. The chemical components of Mao Tea include gallic acid (GA), catechin (C), epicatechin (EC), epigallocatechin (EGC), epicatechin gallate (ECG), epigallocatechin gallate (EGCG), theanine (L-The), and caffeine (CAF), and standards of the above chemical components are prepared as standard solutions, and the contents of the chemical components in the tea leaves are calculated from the standard curve and peak area based on liquid phase chromatography of the standard solutions.

3. The parameter coordination control method for a data-driven primary tea production line according to claim 1, characterized in that the pre-processing in step S2 includes missing value processing and outlier processing, where missing values ​​are data that is missing during the data collection process, the tea production data for the corresponding lot including the missing values ​​is deleted, and the average value of the remaining data is used to represent the process parameters of the corresponding lot, and the outlier value is a value that is clearly significantly different from the actual data.

4. The machine learning model in step S4 includes a support vector machine SVM, and the SVR regression function of the support vector machine SVM for a set of data sets is as follows: [Equation 1] [Equation 2] where ω represents the feature weight vector, C, ε i and ε i * represent the penalty parameter and two slack variables, respectively, δ represents the insensitive interval of the slack variables, and X i represents the feature vector of the process parameters and raw leaf image information in the dataset, and y i represents the moisture content feature vector of the corresponding tea leaves, and n represents the number of feature vectors. It is calculated using the Lagrangian function as follows: [Equation 3] In the formula, α i and α i * represents two Lagrange multipliers, b represents the deviation, and K(X i , X i ' ) represents the kernel function and is calculated as follows: [Equation 4] In the formula, X i and X i ' 2. The method for parameter coordination control of a tea leaf initial production line based on data driving as claimed in claim 1, characterized in that: γ are two feature vectors in the input space, and γ represents the coefficient of the kernel function.

5. The machine learning model in step S4 includes partial least squares regression (PLSR), and is used to calculate the k chemical components and the moisture content {y 1 , y 2 ,...,y k }, m unit process parameters {x 1 , x 2 , ..., x m }, the matrix consisting of n sample data is Y = {y 1 , y 2 ,...,y k } n×k and X = {x 1 , x 2 , ..., x m } n×m Extract the first principal component from X and construct a linear regression equation between the principal component and Y. If the accuracy of the regression equation meets the requirement, terminate the algorithm. If the accuracy of the regression equation does not meet the requirement, extract the second principal component from the remaining X and construct the equation until the accuracy meets the requirement. The specific calculation of PLSR includes principal component decomposition and calculation of the related matrix B. The principal component decomposition is to perform feature factor decomposition on the X and Y matrices. The specific formula is as follows: [Equation 5] where T and U represent the principal component score matrices of X and Y respectively, P and Q represent the weight matrices of X and Y respectively, and E and F represent the error matrices during the fitting process in X and Y respectively. [Equation 6] 2. The method for parameter coordination control of a tea production line based on data driving as claimed in claim 1, wherein the linear regression model is constructed using matrices T and U, and B represents the related matrix.

6. The machine learning model in step S4 includes multiple linear regression (MLR), and the formula of multiple linear regression with n predictor variables is as follows: [Equation 7] In the formula, x represents the process parameters and raw leaf image information of each unit, y represents the moisture content after processing, and β 0 , β 1 , β 2 , ..., β n 2. The method for parameter coordination control of a tea leaf initial production line based on data driving as claimed in claim 1, wherein ε represents a polynomial coefficient and ε represents a residual term of the regression model.

7. It includes a collection module, a pre-processing module, a data selection module, a machine learning module, a prediction module, and a control module. The collecting module acquires a plurality of unit data sets of the primary tea production line, including process parameters of each unit of the plurality of units, moisture content before and after processing, raw leaf image information, quality of the tea leaves, and the contents of a plurality of chemical components of the final tea leaves; the preprocessing module preprocesses the plurality of unit data sets; The data selection module removes data that does not change essentially and / or data that has the same change trend based on the change situation of data of each unit in the production process, and removes data with low importance based on the importance of process parameters and raw leaf image information on the corresponding moisture content and chemical components, and selects chemical components related to the quality as chemical components; the machine learning module is for relating the sorted process parameters and the fresh leaf image information to moisture content, and for relating the sorted process parameters and the fresh leaf image information to chemical components related to quality after sorting; The prediction module builds a moisture content prediction model for each unit and a plurality of chemical component prediction models for the final tea leaves based on the machine learning model, and optimizes the process parameters based on the image information of the raw leaves to obtain the optimization results. The control module applies the optimization results to software parameter control of the tea leaf initial production line.

8. 8. The data-driven parameter coordination control system for a primary tea production line as claimed in claim 7, characterized in that the collection module includes a camera body (1), a camera lens (4) and a ring light (5), the camera body (1) is connected to the camera lens (4) and the ring light (5) is installed around the camera lens (4), and is configured to acquire raw leaf image information including color features and texture features of the raw tea leaf image, wherein the color features include HSV, LAB and RGB, and the texture features include the mean value, standard deviation, smoothness, third-order moment, consistency and entropy of the raw leaf image.

9. 8. The data-driven parameter coordination control system for a primary tea production line as claimed in claim 7, wherein the multiple units of the primary tea production line include a green killing unit, a drum roasting unit, a primary roasting unit and a second roasting unit, the process parameters of the green killing unit and the drum roasting unit all include temperature, drum speed, input speed and moisture removal speed, wherein the process parameters of the drum roasting unit further include hot air temperature and hot air speed, and the process parameters of the primary roasting unit and the second roasting unit all include hot air temperature, oil input temperature, oil output temperature, main conveying speed, hot air speed and leaf homogenization speed.

10. The data-driven parameter coordination control system for a tea leaf primary production line as described in claim 7, characterized in that the collection module further includes an ultra-high efficiency liquid phase chromatograph device, which is used to obtain an ultra-high efficiency liquid chromatograph of a standard solution of the chemical components of Mao-cha tea, and calculates the chemical component contents in the tea leaves based on the standard curve and peak area, wherein the chemical components of Mao-cha tea include gallic acid GA, catechin C, epicatechin EC, epigallocatechin EGC, epicatechin gallate ECG, epigallocatechin gallate EGCG, theanine L-Then, and caffeine CAF.