A method and system for detecting moisture content of tea kernels based on microwave sweep technology and stacking integrated model

Through microwave sweeping technology and Stacking integrated model, the problem of insufficient non-destructive detection accuracy in tea seed kernel moisture content detection is solved, and fast and accurate tea seed kernel moisture content detection is achieved, reducing costs and improving detection accuracy.

CN115541621BActive Publication Date: 2025-08-19ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211328547.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2025-08-19
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

The existing non-destructive testing technology has problems of insufficient accuracy and high cost in the moisture detection of agricultural materials. Especially in the moisture content detection of tea seed kernels, microwave method has not been used, and electrical measurement, machine vision method and spectroscopy method have their own shortcomings.

Method used

Using microwave sweep technology combined with Stacking integrated model, we design a basic learner for selecting mixed evaluation indexes to establish a relational model of moisture content of tea seeds to improve detection accuracy by detecting the amplitude and phase changes of microwave signals before and after penetrating tea seeds.

Benefits of technology

It realizes rapid and accurate non-destructive testing of the moisture content of tea seed kernels, fills the gap in the field of tea seed quality testing, reduces instrument costs and improves detection accuracy and objectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115541621B_ABST
    Figure CN115541621B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting the moisture content of tea seed kernels based on microwave frequency sweeping technology and a Stacking integration model, and belongs to the field of tea seed quality detection. Specifically, the method and system include: 1. obtaining a microwave frequency sweeping data set; 2. normalizing processing and performing feature selection based on the Ridge‑RFE algorithm to generate a training subset and a test subset; 3. constructing a Stacking integration model with a two-layer structure based on the training subset, wherein a hybrid evaluation index IN3 is calculated to select the algorithm of the base learner of the first layer of the training model; 4. using the trained Stacking integration model to detect the moisture content of the tea seed kernel sample to be tested. The present invention fills the technical gap in the field of non-destructive detection of tea seed kernel quality using microwave methods, designs a hybrid evaluation index to evaluate the overall performance of the model during cross-validation and testing, and improves the objectivity of Stacking integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of tea seed quality detection, and in particular to a tea seed kernel moisture content detection method and system based on microwave sweep technology and Stacking integrated model. Technical Background

[0002] Kernel moisture content is a key monitoring indicator in tea seed oil production and processing. During the shelling process, excessively high kernel moisture content can make shelling difficult, while too low a content can lead to excessive crushing. Therefore, the moisture content must be controlled at around 7%. During the flavoring and aromatizing process, the kernel moisture content must also be kept between 5% and 6% to ensure it is optimal for oil extraction. Rapid, accurate, and non-destructive moisture content testing of tea seed kernels is crucial for promoting high-quality and efficient tea seed oil production.

[0003] Current research on nondestructive moisture testing of agricultural materials primarily relies on electrical measurement, machine vision, and spectroscopy. However, each of these methods has its own limitations in practical application. Electrical measurement, which can be divided into capacitance and resistance, is simple, fast, and convenient, but is susceptible to multiple factors, such as material type, firmness, temperature, and humidity, resulting in poor measurement stability. Machine vision determines the true moisture content by extracting surface texture or color features associated with the material's moisture content. This, however, loses some of the deeper features that reflect the material's internal moisture information, limiting detection accuracy. Spectroscopic methods primarily include near-infrared and hyperspectral methods. Near-infrared light has a micron-level penetration depth, making it difficult to obtain measurement information representative of the material's overall moisture content. Hyperspectral methods, which combine spectroscopy and imaging techniques, can obtain richer moisture-related information. However, the high cost of purchasing and maintaining hyperspectral equipment, coupled with the difficulty of transmitting and processing massive amounts of measurement data, also hinder its practical application.

[0004] Microwaves are electromagnetic waves with frequencies between 300MHz and 300GHz, capable of penetrating depths up to centimeters. When microwaves penetrate water-containing materials, most of their energy is absorbed by highly polar water molecules, causing changes in parameters such as the amplitude and phase of the microwave signal, allowing the moisture content of the material to be inferred. Microwaves offer the advantages of being non-contact, highly penetrating, and responsive, while also having lower instrument costs than hyperspectral technology. This makes them an ideal method for moisture detection in agricultural materials. However, existing research has not explored their application to tea seed moisture detection. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of existing nondestructive testing technologies for agricultural material moisture detection and to provide a method and system for detecting the moisture content of tea kernels based on microwave sweeping technology and a Stacking integrated model. The present invention detects the amplitude and phase changes of a set of microwave sweeping signals before and after they penetrate the sample to be tested, thereby obtaining rich correlation information that can reflect the overall moisture distribution within the sample. The Stacking integrated model is used to establish a relationship model between microwave signal characteristics and the moisture content of tea kernel samples, enabling a more accurate estimation of the moisture content of tea kernels than a single machine learning algorithm. When selecting the base learner for the Stacking integrated model, a hybrid evaluation metric based on the model's performance on a cross-validation set and a test set is designed to reduce the empirical and subjective nature of the general Stacking method in selecting the base learner.

[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In one aspect, the present invention provides a method for detecting the moisture content of tea kernels based on microwave sweep technology and a stacking integrated model, which specifically includes:

[0008] (1) Using a microwave detection platform, we measured a set of parameter change data before and after a frequency-sweep microwave signal penetrated tea seed kernel samples, including amplitude and phase, and marked the moisture content attribute values of all samples to obtain a frequency-sweep microwave dataset S0∈R m*2n , where m represents the number of samples and n represents the number of frequency points in the swept microwave signal;

[0009] (2) The swept frequency microwave data set S0 is randomly divided into a training set S1 and a test set S2 according to a certain ratio. The training set S1 is first normalized. After the coefficient of the normalization formula is determined based on the training set S1, the same formula is applied to the normalization of the test set S2 to obtain the preprocessed training set S′1 and test set S′2.

[0010] (3) Use the ridge regression-recursive feature elimination algorithm to perform packaged feature selection on the preprocessed training set S′1 to obtain the training subset FS′1. Record the serial numbers of all features in FS′1 in S′1, and then filter out the features with the same serial numbers in the test set S′2 to obtain the test subset FS′2.

[0011] (4) Using Q candidate regression algorithms, train Q candidate models based on the K-fold cross-validation method on the training subset FS′1 after feature selection, and calculate the evaluation index IN1 of the model training effect; then use the candidate models trained on the training subset FS′1 after feature selection to fit the data in the test subset FS′2 after feature selection, and calculate the evaluation index IN2 of the model test effect;

[0012] (5) Based on IN1 and IN2, the hybrid evaluation index IN3 of the model performance is calculated. According to the value of IN3, all candidate models are sorted in reverse order, and the candidate models ranked in the top L are selected;

[0013] (6) Using the feature-selected training subset FS′1 obtained in step (3) and the regression algorithm corresponding to the L candidate models selected in step (5) to train the base learner of the first layer of the Stacking ensemble model, and using the multivariate linear regression algorithm to train the meta-learner of the second layer, a Stacking ensemble model with a two-layer structure is obtained by training;

[0014] (7) Use the Stacking ensemble model trained in step (6) to detect the moisture content of the tea seed kernel sample to be tested.

[0015] Furthermore, the swept-frequency microwave data set S0 obtained in step (1) is expressed as:

[0016] S0={A1,A2,…,A i ,…,A n ,P1,P2,…,P j ,…,P n}

[0017] Among them, A i ={a i1 ,a i2 ,…,a im} represents the attenuation eigenvector corresponding to the i-th frequency signal, a im represents the microwave attenuation value of the i-th frequency signal corresponding to the m-th sample; P j ={p j1 ,p j2 ,…,p jm} represents the phase shift eigenvector corresponding to the jth frequency signal, p jm Indicates the microwave phase shift value of the i-th frequency signal corresponding to the m-th sample.

[0018] Furthermore, the normalization formula in step (2) is specifically:

[0019]

[0020] Among them, X i represents the i-th eigenvector, x ik represents the i-th eigenvalue of the k-th sample, x′ ik Represents the i-th eigenvalue of the k-th sample after normalization; max(X i ) and min(X i) two coefficients, and then directly apply the above formula to the normalization processing of the test set S2.

[0021] Furthermore, the packaged feature selection method in step (3) is specifically as follows:

[0022] (3.1) Use the ridge regression algorithm to train the regression prediction model based on the preprocessed training set S′1 and generate the corresponding weights σ for each feature in S′1 i , remove σ i The smallest feature to update S′1;

[0023] (3.2) Set the number of selected features to N f , repeat step (3.1) until the number of features in S′1 is reduced to N f , and obtain the training subset FS′1 after feature selection.

[0024] Furthermore, the calculation formulas of the model evaluation indicators IN1 and IN2 in step (4) are respectively:

[0025]

[0026]

[0027] in, and Respectively represent the mean absolute error, root mean square error and coefficient of determination of the model on the training subset after feature selection; MAE p , RMSE p and They represent the mean absolute error, root mean square error, and coefficient of determination of the model on the test subset after feature selection.

[0028] Furthermore, the calculation formula of the hybrid evaluation index IN3 in step (5) is:

[0029] IN3=IN1+IN2.

[0030] Furthermore, the training steps of the Stacking integration model in step (6) are specifically as follows:

[0031] (6.1) After feature selection, the training subset FS′1 is randomly divided into K subsets of equal size, and one of them is randomly selected as the validation subset FS′ 12 , the remaining K-1 combinations are the training subset FS′ 11 , there are K kinds of FS′ 11 and FS′ 12 The division method;

[0032] (6.2) Use the P regression algorithms selected in step (5) and the K-fold cross-validation method in step (6.1) to train the base learners of the first layer of the Stacking ensemble model. For any selected algorithm k, first train the training subset FS′ 11 After training the corresponding model, the validation subset FS′ 12 The data is used for prediction and the verification result set is obtained in, Indicates the model trained based on algorithm k in the i-th cross-validation for the i-th validation subset FS′ 12 The prediction results;

[0033] (6.3) Vertically splice the validation result sets of L models in the first layer of the Stacking ensemble model to generate the training set P′={P1; P2; ...; P L}, use the multivariate linear regression algorithm to train the meta-learner of the second layer of the Stacking ensemble model, and train the corresponding model on the training set P′.

[0034] Furthermore, the prediction steps of the Stacking integration model in step (7) are specifically as follows:

[0035] (7.1) For any algorithm k trained in the first layer of the Stacking ensemble model, use the K-fold cross validation in step (6.2) to calculate the K-fold cross validation algorithm in FS′. 11 The K models trained above predict the sample data to be tested after normalization and feature selection, and obtain the test result set in, Indicates the prediction results of the model trained based on algorithm k for the sample to be tested in the i-th cross-validation;

[0036] (7.2) Longitudinal splicing T k The average value of , generates the test set The multivariate linear regression model trained in the second layer of the Stacking ensemble model is used to predict the data in the test set T′, and the prediction results of the Stacking ensemble model for tea seed kernel samples with unknown moisture content are output.

[0037] Furthermore, the candidate regression algorithms in step (4) include decision tree, random forest, support vector machine, lightweight gradient boosting machine, extreme gradient boosting machine and multi-layer perceptron.

[0038] On the other hand, the present invention provides a tea seed kernel moisture content detection system based on microwave sweeping technology and Stacking integrated model. The system can be stored in a computer-readable storage medium in the form of software or other functional units, and calls and executes lower-level program modules according to computer instructions to implement the tea seed kernel moisture content detection method based on microwave sweeping technology and Stacking integrated model, which specifically includes:

[0039] Data acquisition module: used to measure the parameter change data of a set of swept-frequency microwave signals before and after they penetrate the tea seed kernel samples, including amplitude and phase, and mark the moisture content attribute values of all samples to obtain the swept-frequency microwave data set S0;

[0040] Feature Engineering Module: This module randomly divides the swept-frequency microwave dataset S0 into a training set S1 and a test set S2 in proportion. It first normalizes the training set S1, determines the coefficients of the normalization formula based on the training set S1, and then applies the same formula to the normalization of the test set S2, obtaining the preprocessed training set S′1 and test set S′2. It also performs packaged feature selection on the preprocessed training set S′1 using the ridge regression-recursive feature elimination algorithm to obtain a training subset FS′1. The sequence numbers of all features in FS′1 in S′1 are recorded, and then features with the same sequence numbers are selected from the test set S′2 to obtain a test subset FS′2.

[0041] The candidate model selection module is used to train Q candidate models using Q candidate regression algorithms based on the K-fold cross-validation method on the training subset FS′1 after feature selection, and calculate the evaluation index IN1 of the model training effect; then use the candidate models trained on the training subset FS′1 after feature selection to fit the data in the test subset FS′2 after feature selection, and calculate the evaluation index IN2 of the model test effect; based on IN1 and IN2, calculate the hybrid evaluation index IN3 of the model performance, sort all candidate models in reverse order according to the value of IN3, and select the candidate models ranked in the top L;

[0042] Model training module: This module is used to train the base learner of the first layer of the Stacking ensemble model using the training subset FS′1 after feature selection and the regression algorithm corresponding to the selected L candidate models, and to train the meta-learner of the second layer using the multivariate linear regression algorithm, thereby obtaining a Stacking ensemble model with a two-layer structure.

[0043] Moisture content model prediction module: used to predict the moisture content of the tea seed kernel sample to be tested using the trained Stacking ensemble model.

[0044] The present invention realizes moisture content detection of tea seed kernels based on microwave sweeping technology and Stacking integration model. Its specific advantages are: filling the technical gap of microwave method in the field of non-destructive testing of tea seed quality, and obtaining richer moisture-related information through a set of swept microwave signals; targeting the empirical and ambiguous selection of the base learners in the first layer of traditional Stacking integration, a hybrid evaluation index is designed to evaluate the overall performance of the model during cross-validation and testing, and based on this index, an ideal base learner training algorithm is selected, thereby improving the objectivity of Stacking integration. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is the attenuation characteristic spectrum diagram obtained by testing tea seed kernel samples with different moisture contents using swept frequency microwave technology;

[0046] Figure 2 This is a flow chart of a method for detecting moisture content of tea kernels based on microwave sweep technology and Stacking integrated model proposed by the present invention;

[0047] Figure 3 This is a schematic diagram of a tea seed kernel moisture content detection system based on microwave sweep technology and Stacking integrated model proposed in the present invention. DETAILED DESCRIPTION

[0048] In order to enable those skilled in the art to understand and implement the present invention more clearly, the present invention will be further described below in conjunction with the specific implementation methods and related drawings. It should be understood that the specific factual examples described below are only used to illustrate and explain the present invention, and are not limited to the present invention.

[0049] like Figure 2 As shown, this embodiment provides a method for detecting the moisture content of tea kernels based on microwave sweep technology and a stacking integrated model, which specifically includes the following steps:

[0050] Step 1: Use a self-made microwave detection platform to measure the parameter changes before and after a set of swept-frequency microwave signals penetrate the sample to be tested, including amplitude and phase, and mark the moisture content attribute values of all samples.

[0051] In this example, tea seed kernels prepared through shelling and drying processes were used as test subjects. First, approximately 10 g of a sample was taken and its initial moisture content (7.74%) was determined based on the current national standard GB 5009.3-2016. This was used to represent the initial moisture content of the entire batch of samples. The sample was then humidified to obtain 11 remaining samples with moisture contents of 9.80%, 11.73%, 12.95%, 15.16%, 17.26%, 18.69%, 20.40%, 22.56%, 23.95%, 25.46%, and 26.73%, respectively. Thus, a total of 12 groups of tea seed kernel samples with different moisture contents (7.74%-26.73%) were prepared.

[0052] In this embodiment, a homemade microwave detection platform is used to measure the amplitude and phase changes of a set of swept-frequency microwave signals (2.00-10.00 GHz, a total of 801 frequencies) before and after penetrating the sample to be tested, and the corresponding attenuation characteristics A and phase shift characteristics P are calculated, as shown in FIG. Figure 1 Five parallel samples were prepared at each moisture content, and each parallel sample was measured 5 times. A total of 300 sets of swept-frequency microwave attenuation data and 300 sets of swept-frequency microwave phase shift data were obtained to generate the swept-frequency microwave data set S0∈R 300 *1602 , which can be specifically expressed as:

[0053] S0={A1,A2,…,A i ,…,A 801 ,P1,P2,…,P j ,…,P 801}

[0054] Among them, A i ={a i1 ,a i2 ,…,a i300}, represents the attenuation feature vector corresponding to the i-th frequency signal; P j ={p j1 ,p j2 ,…,p j300}, represents the phase shift eigenvector corresponding to the j-th frequency signal.

[0055] Step 2: After randomly dividing the swept frequency microwave data set S0 into a training set S1 and a test set S2 at a ratio of 4:1, the training set S1 is normalized first. After determining the coefficient of the normalization formula based on the training set S1, the same formula is applied to the normalization of the test set S2 to obtain the preprocessed training set S′1∈R 240*1602 and the test set S′2∈R 60*1602 , the specific normalization formula is as follows:

[0056]

[0057] Among them, X i represents the i-th eigenvector, x ik represents the i-th eigenvalue of the k-th sample, x′ ik represents the i-th eigenvalue of the k-th sample after normalization; wherein, max(X i ) and min(X i ) two coefficients, and then directly apply the above formula to the normalization processing of the test set S2.

[0058] Step 3: Use the Ridge Regression-Recursive Feature Elimination (Ridge-RFE) algorithm to perform packaged feature selection on the preprocessed training set S′1 to obtain the training subset FS′1. Record the serial numbers of all features in FS′1 in S′1, and then filter out the features with the same serial numbers in the preprocessed test set S′2 to obtain the test subset FS′2. The feature selection method is as follows:

[0059] 3.1, use the Ridge algorithm to train the regression prediction model based on the preprocessed training set S′1, and generate the corresponding weight σ of each feature in S′1 i , remove σ i The smallest feature to update S′1;

[0060] 3.2, set the number of selected features to N f = 10, repeat step 3.1 until the number of features in S′1 is reduced to 10, and obtain the training subset FS′1∈R after feature selection 240*10 .

[0061] Step 4: Use decision tree (DT), random forest (RF), support vector machine (SVR), lightweight gradient boosting machine (LGBM), extreme gradient boosting machine (XGBoost) and multi-layer perceptron (MLP) to train candidate models based on the 5-fold cross-validation method on the training subset FS′1 after feature selection, and calculate the evaluation index IN1 of the model training effect (see Table 1 for the results); then use the candidate models trained based on FS′1 to fit the test data of the test subset FS′2 after feature selection, and calculate the evaluation index IN2 of the model test effect (see Table 1 for the results); where the calculation formulas for IN1 and IN2 are:

[0062]

[0063]

[0064] in, and Respectively represent the mean absolute error, root mean square error, and coefficient of determination of the model on the training subset after feature selection. The average score of the 5-fold cross validation is calculated here; MAE p , RMSE p and They represent the mean absolute error, root mean square error, and coefficient of determination of the model on the test subset after feature selection.

[0065] Step 5. Based on IN1 and IN2, the hybrid evaluation index IN3 of the model performance is calculated (see Table 1 for the results). All candidate models are sorted in reverse order according to the value of IN3, and the top 3 SVR, XGBoost, and RF algorithms are selected to train the base learners of the stacking ensemble model in the next step. The calculation formula of IN3 is:

[0066] IN3=IN1+IN2

[0067] Table 1 Evaluation index scores of candidate models

[0068]

[0069]

[0070] Step 6: Use the SVR, XGBoost, and RF algorithms selected in step 5 to train the base learners of the first layer, use the multivariate linear regression (MLR) algorithm to train the meta-learner of the second layer, and train the Stacking ensemble model with a two-layer structure on the training subset FS′1 after feature selection based on the principle of the Stacking method. The specific training steps are as follows:

[0071] 6.1, randomly divide FS′1 into 5 subsets of equal size, and randomly select one of them as the validation subset FS′ 12 ∈R 48*10 , the remaining 4 combinations are the training subset FS′ 11 ∈R 192*10 , there are 5 types of FS′ 11 and FS′ 12 The division method;

[0072] 6.2, use the SVR, XGBoost and RF algorithms selected in step 5 and the 5-fold cross-validation method in step 6.1 to train the base learners of the first layer. For any algorithm k, first 11 Train the corresponding model and then FS′ 12 The data is used for prediction and the verification result set is obtained in, Indicates the model trained based on algorithm k in the i-th cross validation for the i-th FS′ 12 The prediction results;

[0073] 6.3, vertically concatenate the verification results of the three algorithms in the first layer to generate the training set P′={P1;P2;P3}∈R 240*30 , the MLR algorithm is used to train the meta-learner of the second layer, and the corresponding model is trained on the training set P′.

[0074] Step 7: Use the Stacking ensemble model trained in step 6 to detect the moisture content of unknown tea seed kernel samples.

[0075] Taking the test subset FS′2 after feature selection as an example, the output model predicts the moisture content of the unknown tea seed kernel sample in FS′2. The specific steps are as follows:

[0076] 7.1, for any algorithm k, use the 5-fold cross validation in step 6.2 to get the 5 FS′ 11 The five models trained above predict the data of FS′2 respectively and obtain the test result set in, represents the prediction result of FS′2 by the model trained based on algorithm k in the i-th cross-validation;

[0077] 7.2, vertical splicing test result set T of three algorithms k The average value of , generates the test set Use the MLR model trained on the training set P′ obtained in step 6.3 to predict the data of the test set T′, and output the prediction results of the Stacking ensemble model for tea seed kernel samples with unknown moisture content.

[0078] Step 7 is not limited to testing the moisture content of samples in the feature-selected test subset FS′2; it can also be applied to other tea kernel samples to be tested. For example, a swept-frequency microwave dataset S′0 of the tea kernel samples to be tested is first obtained and normalized. Features with the same feature numbers as those in the training subset FS′1 are then filtered from this normalized swept-frequency microwave dataset S′0 to obtain a test set. The Stacking ensemble model trained in step 6 is then used to test the moisture content of unknown tea kernel samples in the test set. The Stacking ensemble model's processing is identical to that described in steps 7.1 and 7.2 above and will not be further elaborated.

[0079] Step 8. In order to evaluate the prediction performance of the Stacking ensemble model designed and built based on the hybrid evaluation index in this embodiment, its actual performance on the training subset FS′1 and the test subset FS′2 was compared with that of a general single model and a Stacking ensemble model built based on a randomly selected base learner. The results are shown in Table 2.

[0080] Table 2 Comparison of prediction performance of Stacking ensemble model with other models

[0081]

[0082] As can be seen from Table 2, most models have achieved good training results on the training subset, but compared with the base learners that constitute the Stacking ensemble model, all Stacking ensemble models have achieved lower errors and higher determination coefficients on the test subset, which proves the effectiveness of the Stacking ensemble method in improving model accuracy. Secondly, compared with the Stacking ensemble model built by the randomly selected base learner algorithm, the hybrid evaluation index IN designed by the present invention is better than that of the original model. S The Stacking ensemble model built by the selected SVR, XGBoost and RF algorithms achieved the best prediction performance (MAE p =0.486, RMSE p =0.768, ), indicating that this method can select the base learner algorithm suitable for the Stacking ensemble model, thereby improving the prediction accuracy and generalization of the Stacking ensemble.

[0083] like Figure 3 As shown, this embodiment provides a tea seed kernel moisture content detection system based on microwave sweep technology and a stacking integrated model. This system is used to implement the above-mentioned embodiments. The terms "module," "submodule," etc. used below may refer to a combination of software and / or hardware that implements the predetermined functions. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible.

[0084] A tea seed kernel moisture content detection system based on microwave sweep technology and stacking integrated model, specifically including the following modules:

[0085] Data acquisition module S1: used to obtain microwave frequency sweep data set S0;

[0086] Feature engineering module S2: used to perform data normalization and feature selection based on the Ridge-RFE algorithm on S0 to generate the training subset FS′1 and the test subset FS′2;

[0087] Model training module S3: used to train a stacking ensemble model with a two-layer structure on FS′1. The algorithm used to train the base learner of the first layer of the model is selected by calculating the hybrid evaluation index IN3, and the algorithm used to train the meta-learner of the second layer of the model is MLR.

[0088] Model prediction module S4: used to fit the test data of FS′2 using the trained Stacking ensemble model and output the fusion prediction result of the moisture content of the unknown tea seed kernel sample.

[0089] In a specific implementation of the present invention, the model training module further includes:

[0090] The candidate model selection submodule selects base learners. It uses Q candidate regression algorithms to train Q candidate models using K-fold cross-validation on the feature-selected training subset FS′1, calculating the model training performance evaluation index IN1. The candidate models trained on the feature-selected training subset FS′1 are then used to fit the data on the feature-selected test subset FS′2, calculating the model test performance evaluation index IN2. Based on IN1 and IN2, a hybrid evaluation index IN3 of model performance is calculated. All candidate models are sorted in reverse order according to the value of IN3, and the top L candidate models are selected. For detailed implementation, refer to steps 4 and 5 of the method section.

[0091] Furthermore, the data acquisition module uses a custom-built microwave detection platform to measure the parameter changes of a set of swept-frequency microwave signals before and after they penetrate tea seed kernel samples, including amplitude and phase. The data is then labeled with the moisture content attribute values of all samples to obtain a swept-frequency microwave dataset S0. The specific implementation method is described in step 1 of the Method section.

[0092] The feature engineering module is used to select effective features from the swept-frequency microwave dataset S0. First, the swept-frequency microwave dataset S0 is randomly divided into a training set S1 and a test set S2. Training set S1 is normalized. After determining the coefficients of the normalization formula based on training set S1, the same formula is applied to the normalization of test set S2, resulting in preprocessed training set S′1 and test set S′2. Furthermore, the module uses the ridge regression-recursive feature elimination algorithm to perform packaged feature selection on the preprocessed training set S′1, obtaining a training subset FS′1. The sequence numbers of all features in FS′1 in S′1 are recorded, and then features with the same sequence numbers are selected from the test set S′2 to obtain the test subset FS′2. For detailed implementation, refer to steps 2 and 3 of the method section.

[0093] The model training module uses the feature-selected training subset FS′1 and the regression algorithms corresponding to the L selected candidate models to train the base learners in the first layer of the stacking ensemble model. It then uses a multivariate linear regression algorithm to train the meta-learners in the second layer, resulting in a two-layer stacking ensemble model. For detailed implementation, refer to step 6 of the Methods section.

[0094] Moisture Content Model Prediction Module: This module uses the trained Stacking ensemble model to predict the moisture content of the tea seed kernel sample. For detailed implementation, refer to step 7 of the Methods section.

[0095] The implementation process of the functions and effects of each module in the above-mentioned system is specifically detailed in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here. For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0096] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, it can include a processor, memory, network interface, and non-volatile memory, etc. In the embodiment, any device with data processing capabilities in which the system is located can also include other hardware according to the actual functions of the device with data processing capabilities, which will not be described in detail.

[0097] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for detecting the moisture content of tea kernels based on microwave sweep technology and Stacking integrated model, characterized in that: include: (1) Using a microwave detection platform, we measured a set of parameter change data before and after a frequency-sweep microwave signal penetrated tea seed kernel samples, including amplitude and phase, and marked the moisture content attribute values of all samples to obtain a frequency-sweep microwave dataset S0∈R m*2n , where m represents the number of samples and n represents the number of frequency points in the swept microwave signal; (2) The swept frequency microwave data set S0 is randomly divided into a training set S1 and a test set S2 according to a certain ratio. The training set S1 is first normalized. After the coefficient of the normalization formula is determined based on the training set S1, the same formula is applied to the normalization of the test set S2 to obtain the preprocessed training set S′1 and test set S′2. (3) Use the ridge regression-recursive feature elimination algorithm to perform packaged feature selection on the preprocessed training set S′1 to obtain the training subset FS′1. Record the serial numbers of all features in FS′1 in S′1, and then filter out the features with the same serial numbers in the test set S′2 to obtain the test subset FS′2. (4) Using Q candidate regression algorithms, train Q candidate models based on the K-fold cross-validation method on the training subset FS′1 after feature selection, and calculate the evaluation index IN1 of the model training effect; then use the candidate models trained on the training subset FS′1 after feature selection to fit the data in the test subset FS′2 after feature selection, and calculate the evaluation index IN2 of the model test effect; The calculation formulas for model evaluation indicators IN1 and IN2 are: in, and Respectively represent the mean absolute error, root mean square error and coefficient of determination of the model on the training subset after feature selection; MAE p , RMSE p and They represent the mean absolute error, root mean square error, and coefficient of determination of the model on the test subset after feature selection; (5) Based on IN1 and IN2, the hybrid evaluation index of model performance is calculated as IN3 = IN1 + IN2. All candidate models are sorted in reverse order according to the value of IN3, and the candidate models ranked in the top L are selected; (6) Using the feature-selected training subset FS′1 obtained in step (3) and the regression algorithm corresponding to the L candidate models selected in step (5) to train the base learner of the first layer of the Stacking ensemble model, and using the multivariate linear regression algorithm to train the meta-learner of the second layer, a Stacking ensemble model with a two-layer structure is obtained by training; The training steps of the Stacking ensemble model are as follows: (6.1) After feature selection, the training subset FS′1 is randomly divided into K subsets of equal size, and one of them is randomly selected as the validation subset FS′ 12 , the remaining K-1 combinations are the training subset FS′ 11 , there are K kinds of FS′ 11 and FS′ 12 The division method; (6.2) Use the regression algorithms corresponding to the L candidate models selected in step (5) and the K-fold cross-validation method in step (6.1) to train the base learners of the first layer of the Stacking ensemble model. For any selected algorithm k, first train the training subset FS′ 11 After training the corresponding model, the validation subset FS′ 12 The data is used for prediction and the verification result set is obtained in, Indicates the model trained based on algorithm k in the i-th cross-validation for the i-th validation subset FS′ 12 The prediction results; (6.3) Vertically splice the validation result sets of L models in the first layer of the Stacking ensemble model to generate the training set P′={P1; P2; ...; P L }, use the multivariate linear regression algorithm to train the meta-learner of the second layer of the Stacking ensemble model, and train the corresponding model on the training set P′; (7) Use the Stacking ensemble model trained in step (6) to detect the moisture content of the tea seed kernel sample to be tested.

2. The method for detecting moisture content of tea kernels based on microwave sweep technology and stacking integrated model according to claim 1, characterized in that: The swept frequency microwave data set S0 obtained in step (1) is expressed as: S0={A1,A2,…,A i ,…,A n ,P1,P2,…,P j ,…,P n } Among them, A i ={a i1 ,a i2 ,…,a im } represents the attenuation eigenvector corresponding to the i-th frequency signal, a im represents the microwave attenuation value of the i-th frequency signal corresponding to the m-th sample; P j ={p j1 ,p j2 ,…,p jm } represents the phase shift eigenvector corresponding to the jth frequency signal, p jm Indicates the microwave phase shift value of the i-th frequency signal corresponding to the m-th sample.

3. The method for detecting moisture content of tea kernels based on microwave sweep technology and stacking integrated model according to claim 1, characterized in that: The normalization formula in step (2) is specifically: Among them, X i represents the i-th eigenvector, x ik represents the i-th eigenvalue of the k-th sample, x′ ik Represents the i-th eigenvalue of the k-th sample after normalization; max(X i ) and min(X i ) two coefficients, and then directly apply the above formula to the normalization processing of the test set S2.

4. The method for detecting moisture content of tea kernels based on microwave sweep technology and stacking integrated model according to claim 1, characterized in that: The packaged feature selection method in step (3) is specifically as follows: (3.1) Use the ridge regression algorithm to train the regression prediction model based on the preprocessed training set S′1 to generate S1 ′ The corresponding weight σ of each feature in i , remove σ i The smallest feature to update S′1; (3.2) Set the number of selected features to N f , repeat step (3.1) until the number of features in S′1 is reduced to N f , and obtain the training subset FS′1 after feature selection.

5. The method for detecting moisture content of tea kernels based on microwave sweep technology and stacking integrated model according to claim 1, characterized in that: The prediction steps of the Stacking ensemble model in step (7) are specifically as follows: (7.1) For any algorithm k trained in the first layer of the Stacking ensemble model, use the K-fold cross validation in step (6.2) to calculate the K-fold cross validation algorithm in FS′. 11 The K models trained above predict the sample data to be tested after normalization and feature selection, and obtain the test result set in, Indicates the prediction results of the model trained based on algorithm k for the samples to be tested in the ioth cross-validation; (7.2) Longitudinal splicing T k The average value of , generates the test set The multivariate linear regression model trained in the second layer of the Stacking ensemble model is used to predict the data in the test set T′, and the prediction results of the Stacking ensemble model for tea seed kernel samples with unknown moisture content are output.

6. The method for detecting moisture content of tea kernels based on microwave sweep technology and stacking integrated model according to claim 1, characterized in that: The candidate regression algorithms in step (4) include decision tree, random forest, support vector machine, lightweight gradient boosting machine, extreme gradient boosting machine and multi-layer perceptron.

7. A tea seed kernel moisture content detection system based on microwave sweep technology and Stacking integrated model, characterized in that: include: Data acquisition module: used to measure the parameter change data of a set of swept-frequency microwave signals before and after they penetrate the tea seed kernel samples, including amplitude and phase, and mark the moisture content attribute values of all samples to obtain the swept-frequency microwave data set S0; Feature Engineering Module: This module randomly divides the swept-frequency microwave dataset S0 into a training set S1 and a test set S2 in proportion. It first normalizes the training set S1, determines the coefficients of the normalization formula based on the training set S1, and then applies the same formula to the normalization of the test set S2, obtaining the preprocessed training set S′1 and test set S′2. , used to perform packaged feature selection on the preprocessed training set S′1 using the ridge regression-recursive feature elimination algorithm to obtain the training subset FS′1, record the sequence numbers of all features in FS′1 in S′1, and then filter out the features with the same sequence numbers in the test set S′2 to obtain the test subset FS′2; The candidate model selection module is used to train Q candidate models using Q candidate regression algorithms based on the K-fold cross-validation method on the training subset FS′1 after feature selection, and calculate the evaluation index IN1 of the model training effect; then use the candidate models trained on the training subset FS′1 after feature selection to fit the data in the test subset FS′2 after feature selection, and calculate the evaluation index IN2 of the model test effect; based on IN1 and IN2, calculate the hybrid evaluation index IN3 of the model performance, sort all candidate models in reverse order according to the value of IN3, and select the candidate models ranked in the top L; The calculation formulas for model evaluation indicators IN1 and IN2 are: in, and Respectively represent the mean absolute error, root mean square error and coefficient of determination of the model on the training subset after feature selection; MAE p , RMSE p and They represent the mean absolute error, root mean square error, and coefficient of determination of the model on the test subset after feature selection; IN3=IN1+IN2; The model training module is used to train the base learner of the first layer of the Stacking ensemble model using the training subset FS′1 after feature selection and the regression algorithm corresponding to the selected L candidate models, and to train the meta-learner of the second layer using the multivariate linear regression algorithm, thereby obtaining a Stacking ensemble model with a two-layer structure. The working process of the model training module is as follows: The training subset FS′1 after feature selection is randomly divided into K subsets of the same size, and any one of them is taken as the validation subset FS′ 12 , the remaining K-1 combinations are the training subset FS′ 11 , there are K kinds of FS′ 11 and FS′ 12 The division method; The regression algorithms corresponding to the L candidate models and the K-fold cross-validation method are used to train the base learners of the first layer of the Stacking ensemble model. For any selected algorithm k, first 11 After training the corresponding model, the validation subset FS′ 12 The data is used for prediction and the verification result set is obtained in, Indicates the model trained based on algorithm k in the i-th cross-validation for the i-th validation subset FS′ 12 The prediction results; The validation result sets of L models in the first layer of the stacking ensemble model are vertically spliced to generate the training set P′={P1;P2;…;P L }, use the multivariate linear regression algorithm to train the meta-learner of the second layer of the Stacking ensemble model, and train the corresponding model on the training set P′; Moisture content model prediction module: used to predict the moisture content of the tea seed kernel sample to be tested using the trained Stacking ensemble model.

Citation Information

Patent Citations

  • Method for measuring grain moisture content based on microwave frequency sweep technology

    CN109632834A

  • Second-order frequency selection method and device for microwave frequency sweep data

    CN111812122A