A method for determining moisture content of honeysuckle by combining two-stage characteristic wavelength screening
Patent Information
- Application Number
- CN202610778620.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]金银花属于典型的植物类中药材,其组织结构复杂、内部成分多样,水分在不同组织中的分布状态存在差异,传统水分检测方法在实际应用中普遍存在检测周期长、样品易受破坏及难以满足快速批量检测需求等问题
[0012]因此,本发明采用上述一种结合二阶段特征波长筛选测定金银花水分含量的方法,采用误差反馈竞争自适应重加权采样与逐步回归相结合的二阶段特征波长筛选策略,对金银花近红外光谱中的特征变量进行筛选,并建立金银花水分含量检测模型。该方法能够实现对不同含水量金银花样本水分含量的快速、高效测定,提高检测效率和识别精度,为不同水分含量金银花的高精度鉴别提供了一种新的技术方案。
Smart Images

Figure CN122591605A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of near-infrared spectroscopy analysis technology, and in particular to a method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening. Background Technology
[0002] Moisture content is a crucial indicator for evaluating the quality of agricultural products. It not only directly affects the stability of their active ingredients but is also closely related to quality and safety issues such as mold and pests during storage. In recent years, with the increasing market demand for honeysuckle, the moisture content of honeysuckle varies significantly under different harvesting and storage conditions, resulting in inconsistent product quality in the market. This necessitates faster and more accurate detection of honeysuckle moisture content.
[0003] Honeysuckle is a typical plant-based traditional Chinese medicine with a complex tissue structure and diverse internal components. The distribution of moisture varies across different tissues. Traditional moisture detection methods generally suffer from long detection cycles, sample damage, and difficulty in meeting the demands of rapid, batch testing. Based on these characteristics, it is necessary to introduce a moisture detection technology suitable for honeysuckle and similar traditional Chinese medicines. Near-infrared spectroscopy, based on the absorption characteristics of vibrational overtones and combination frequencies of hydrogen-containing functional groups (OH, CH, NH, etc.) in molecules, collects near-infrared spectral information from medicinal samples and combines it with chemometric methods to achieve a comprehensive characterization of the sample's internal chemical information. In particular, the OH bond in water molecules exhibits significant and stable absorption characteristics in the near-infrared band, giving near-infrared spectroscopy advantages such as speed and non-destructiveness in the detection of moisture content in medicinal materials, providing an effective technical approach for rapid moisture content detection in traditional Chinese medicine. Summary of the Invention
[0004] The purpose of this invention is to provide a method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening. It utilizes near-infrared spectroscopy technology with an error feedback competitive adaptive reweighted sampling method for characteristic wavelength screening, suppresses the interference of variable noise, and introduces a stepwise regression algorithm to further reduce the redundancy between characteristic wavelengths and alleviate the multicollinearity problem, thereby improving the model's prediction accuracy and generalization ability. This enables accurate, rapid, and non-destructive detection of the moisture content of different honeysuckle samples, meeting the market demand for large-scale detection of honeysuckle moisture content.
[0005] To achieve the above objectives, this invention provides a method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening, comprising the following steps: S1. Collect honeysuckle samples and obtain near-infrared spectral data under different moisture content gradients. At the same time, use the standard drying method to determine the moisture content of the corresponding samples as a reference value. S2. The Kennard-Stone algorithm is used to divide all samples into training and test sets. S3. The near-infrared spectral data is screened for the first stage of feature wavelengths using the error feedback competitive adaptive reweighted sampling method EF-CARS to obtain an initial set of feature wavelengths. The error feedback competitive adaptive reweighted sampling method introduces a prediction error feedback mechanism to dynamically update the variable weights. S4. Based on the initial set of characteristic wavelengths, the stepwise regression SR algorithm is used to perform the second stage of characteristic wavelength screening in order to further select the optimal combination of characteristic wavelengths from the initial set of characteristic wavelengths. S5. Construct a multiple linear regression (MLR) model using the optimal combination of characteristic wavelengths to determine the moisture content of honeysuckle. S6. Import the test set samples and evaluate the predictive performance of the model.
[0006] Preferably, in S1, the near-infrared spectral data acquisition process is as follows: the fresh honeysuckle sample is dried, and multiple samples are taken during the drying process. After cooling to room temperature, the sample is laid flat and scanned with a near-infrared spectrometer. The spectrum of each sample is collected three times and the average value is taken to obtain near-infrared spectral data with different moisture content gradients. The total moisture content of the sample is determined based on the difference between the initial mass and the mass after drying to constant weight. Reference values for the moisture content of the sample at different drying stages are then determined by combining the weighing results at each stage.
[0007] Preferably, in S2, the samples are divided according to the gradient of moisture content, and the training set and the test set are consistent in terms of moisture content distribution, origin and physical state; the training set is used for feature wavelength screening and model building, and the test set is used to independently evaluate the model performance.
[0008] Preferably, in S3, the first stage of feature wavelength screening specifically includes the following: S31. Divide the spectral data matrix X The variable weights are initialized using the actual values of the corresponding target variables as input, and the maximum number of iterations is set. S32. In each iteration, from the spectral matrix X The training set and corresponding target variable subset are extracted from the dataset. Based on the variable subset retained in the current iteration, the training set is further filtered to obtain subset data. A partial least squares regression model is then established using the subset data and the target variable subset, and the prediction error of the model is calculated through cross-validation. ; S33. Based on the cross-validation prediction error, construct an error feedback function: ; in,k Indicates the number of iterations; This is the error feedback adjustment coefficient, used to control the intensity of the influence of prediction error on variable weights; S34. Update the variable weights based on the regression coefficients and error feedback function of the multiple linear regression model to obtain the comprehensive weights of the variables: ; in, Indicates the first k In the second modeling j The regression coefficients of the variables, Indicates the updated variable weights; S35. Employing an exponential decay function Control the proportion of variables retained in each iteration, where, Indicates the first k The proportion of variables retained in each iteration a and b A constant parameter to control the decay rate; and retaining the top weights according to the comprehensive weighting. The proportional variable proceeds to the next iteration; specifically, a and b Each by and Sure, p The initial total number of variables, N This represents the total number of CARS iterations. S36. After the iteration is completed, select the subset of variables with the smallest cross-validation prediction error as the initial feature wavelength set.
[0009] Preferably, in S4, the second stage of characteristic wavelength screening specifically includes the following steps: S41. Data Preparation and Initialization: Extract the spectral data corresponding to the initial feature wavelength set and construct a new feature matrix. X ′ and the target variable moisture content reference value y , the feature matrix X ′ and target variable y The dataset is divided into training and validation sets, with an empty variable set or a univariate set containing only the intercept term as the initial model state for stepwise regression. S42. Variable Introduction: Based on the variables included in the current model, iterate through the feature wavelengths not yet included in the model, and construct a temporary multiple regression model for each feature wavelength; based on the training set data, calculate and compare the preset evaluation indicators for each temporary model, including the root mean square error of cross-validation (RMSECV) and the coefficient of determination (R²). 2 Select the wavelength that will most significantly improve model performance and officially add it to the current model. S43. Variable Removal: For all wavelengths currently in the model, temporarily remove them one by one; build a new model after removing each wavelength; calculate and compare the evaluation index of the model after removal based on the training set data. If the evaluation index of the model does not decrease significantly or even improves after removing a certain variable, it is determined that the variable does not contribute significantly to the model and is permanently removed from the current model. S44. Repeat S42 and S43, performing variable filtering in each iteration. The iteration process stops when any of the following preset termination conditions are met: If the improvement in the model evaluation metrics is less than the preset small threshold, it indicates that the model performance has stabilized; or that the preset maximum number of iterations has been reached. S45. After the iteration is completed, the set of all variables that are finally retained in the model is used as the optimal combination of feature wavelengths obtained after the second stage of screening, and is used to establish the final multivariate linear regression MLR quantitative correction model.
[0010] Preferably, in S5, the expression for the multiple linear regression (MLR) model is: ; in, y i For the first i Reference values for moisture content of each sample. x ij For the first i The sample at the th j Spectral values at the optimal characteristic wavelengths β 0 represents the intercept term. β j For regression coefficients, ε i For random error term, m The number of variables for the optimal combination of characteristic wavelengths.
[0011] Preferably, in S6, the metrics for evaluating prediction performance include: root mean square error of cross-validation (RMSECV), root mean square error of prediction set (RMSEP), and coefficient of determination (R²). 2 Relative prediction bias (RPD).
[0012] Therefore, this invention employs the aforementioned method for determining the moisture content of honeysuckle using a two-stage characteristic wavelength screening approach. It utilizes a two-stage characteristic wavelength screening strategy combining error feedback competitive adaptive reweighted sampling and stepwise regression to screen characteristic variables in the near-infrared spectrum of honeysuckle and establish a honeysuckle moisture content detection model. This method enables rapid and efficient determination of the moisture content of honeysuckle samples with different moisture contents, improving detection efficiency and identification accuracy, and providing a new technical solution for high-precision identification of honeysuckle with different moisture contents.
[0013] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0014] Figure 1 This is a picture of a honeysuckle sample; Figure 2 This is the original near-infrared spectrum of honeysuckle according to an embodiment of the present invention; Figure 3 This invention relates to the EF-CARS-SR two-stage screening method; Figure 4 This is the characteristic wavelength screening result of an embodiment of the present invention; Figure 5 This is a scatter plot of the model prediction results in an embodiment of the present invention. Detailed Implementation
[0015] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0016] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0017] Example 1 This embodiment provides a method for determining the moisture content of honeysuckle using a two-stage characteristic wavelength screening approach. The study used honeysuckle samples from 68 different origins as the research object. (The honeysuckle samples are shown in the image.) Figure 1 As shown, the specific steps are as follows: S1. Honeysuckle sample preparation, near-infrared spectral data acquisition, and moisture reference value determination.
[0018] S11. Sample Preparation and Spectral Acquisition. Fresh honeysuckle samples were collected and numbered, and their initial mass was recorded. The samples were placed in a 40℃ constant temperature oven for drying. During the drying process, the samples were removed at set time intervals and allowed to cool naturally to room temperature.
[0019] The cooled honeysuckle samples were evenly spread on the test bench, and spectral data were acquired using a near-infrared spectrometer. The near-infrared spectrometer used halogen light as the light source, and the acquired spectral wavelength range was 360 nm to 1100 nm, with a spectral resolution of 0.5 nm. The final spectral data consisted of 1400 wavelength points. The original near-infrared spectrum of the honeysuckle is shown below. Figure 2 As shown.
[0020] During the acquisition process, the spectrometer probe is aimed at the sample surface and scanned. After each scan, the position of the light spot is moved to reduce the impact of local differences in the sample on the spectral data. Each sample is scanned three times, and the spectral data obtained from the three scans are averaged to obtain the representative near-infrared raw spectral data of the sample at this stage.
[0021] After completing a single spectral acquisition, the sample was placed back in a 40℃ constant temperature oven to continue drying. The above cycle of "drying-cooling-spectral acquisition" was repeated to obtain near-infrared spectral data of honeysuckle at different moisture content stages.
[0022] S12. Moisture content reference value determination. Simultaneously with near-infrared spectroscopy acquisition, the moisture content reference value of the honeysuckle samples at the corresponding stage is determined. Specifically: after each drying and cooling to room temperature, the sample is weighed and its mass is recorded, followed by near-infrared spectroscopy acquisition; after spectral acquisition, the sample is returned to the oven to continue drying until the sample mass reaches constant weight.
[0023] The total moisture content of the sample is calculated based on the difference between the initial mass and the final constant mass. The moisture content gradient of the sample at different drying stages is determined by combining the weighing results at each stage. The obtained moisture content data is used as the reference true value for model establishment and performance evaluation, and is used for the training and validation of the subsequent near-infrared quantitative model.
[0024] S2. The Kennard-Stone algorithm is used to divide all samples into training and test sets in a 7:3 ratio.
[0025] The Kennard–Stone (KS) algorithm was used to partition the dataset of 68 honeysuckle samples. All samples were divided into training and test sets according to the gradient of moisture content, with the training set accounting for 70% and the test set accounting for 30%. The partitioning process ensured that the training and test sets were consistent in terms of moisture content distribution, origin, and physical condition. The training set was used for feature wavelength selection and the construction of the EF-CARS-SR-MLR joint model, while the test set was used to independently evaluate the robustness and generalization ability of the model.
[0026] S3. The first stage of feature wavelength selection is performed using the EF-CARS algorithm, as follows: Figure 3 As shown.
[0027] The error feedback competitive adaptive reweighted sampling method EF-CARS is used to perform the first stage of feature wavelength coarse screening on the original near-infrared spectral data of the training set to obtain the initial feature wavelength set. The specific steps are as follows: S31. Data Input and Initialization: Input the pre-divided training set spectral data matrix. X Using the corresponding moisture content reference value as input, the variable weights are initialized to obtain the initial weight vector, and the maximum number of iterations is set.
[0028] S32, Variable Subset Modeling: In the... k During the secondary Monte Carlo sampling process, from the spectral matrix X The training set and corresponding target variable subset are extracted from the dataset. Based on the variable subset retained in the current iteration, the training set is further filtered to obtain subset data. A partial least squares regression model is built using the subset data and the target variable subset, and the prediction error of the model is calculated through cross-validation. ,in k Indicates the number of iterations.
[0029] S33. Construction of Error Feedback Function: Construct the error feedback function based on the cross-validation prediction error. ;in, λ This is the error feedback adjustment coefficient, used to control the intensity of the influence of prediction error on variable weights. Its value is set according to the numerical range of prediction error to ensure that the error feedback function changes within a reasonable range and to avoid excessive amplification or suppression during the weight update process.
[0030] S34. Variable Weight Update: Update the variable weights based on the regression coefficients and error feedback function obtained from the multiple linear regression model to obtain the comprehensive weights of the variables. ; in, Indicates the first k In the second modeling j The regression coefficients of the variables, This represents the updated variable weights.
[0031] S35. Variable Selection: Using the Exponential Decay Function Control the proportion of variables retained in each iteration, where, Indicates the first k The proportion of variables retained in each iteration a and b A constant parameter to control the decay rate; and retaining the top-ranked parameters according to the comprehensive weighting. The proportional variable proceeds to the next iteration. Specifically, a and b Each by and Sure, p The initial total number of variables, N This represents the total number of CARS iterations. In this embodiment, the initial number of variables is 1400, and the number of CARS iterations is 50, therefore a≈1.1430, b≈0.1337. The maximum number of iterations within CARS is set to 50. To improve the stability of the variable selection results, this embodiment further repeats CARS 30 times, selecting the optimal subset of variables from the multiple selection results based on the cross-validation error.
[0032] S36. Determining the optimal subset of variables: After completing all iterations, analyze the prediction errors corresponding to each iteration. The variables with the smallest cross-validation prediction error are compared and selected as the initial feature wavelength set. The feature wavelength selection results are as follows: Figure 4 As shown, CARS, CARS-SR, EF-CARS, and EF-CARS-SR are used to select characteristic wavelengths from the raw near-infrared spectrum. The four subplots in the figure show the location of the characteristic wavelengths selected by each method and the spectral response. The black curve represents the average spectral absorbance of the sample, and the colored vertical lines represent the characteristic wavelengths selected by the corresponding method.
[0033] Figure 4 (a) The characteristic wavelengths screened by the CARS method alone are relatively dispersed, covering multiple ranges of the spectrum, showing the broad characteristics of the preliminary screening. Figure 4 (b) The CARS-SR method combines the CARS algorithm with the SR algorithm to perform a second-stage feature wavelength screening, which significantly reduces the number of feature wavelengths and presents a more concentrated and representative distribution. Figure 4 (c) The EF-CARS method dynamically updates variable weights by introducing a prediction error feedback mechanism, optimizes the selection of feature wavelengths, and removes redundant information. Figure 4 (d) The EF-CARS-SR method adopts a two-stage screening strategy, and the final characteristic wavelengths are concentrated near the key absorption peaks in the entire spectral range, which reduces the interference of non-information bands and further improves the accuracy and stability of the model.
[0034] S4. Use the SR algorithm for the second stage of feature wavelength selection.
[0035] Based on the initial set of characteristic wavelengths obtained by the EF-CARS algorithm, the stepwise regression (SR) algorithm is used for the second stage of refined selection of characteristic wavelengths, further selecting the optimal combination of characteristic wavelengths from the initial set. The specific steps are as follows: S41. Data Preparation and Initialization: Extract the spectral data corresponding to the initial feature wavelength set and construct a new feature matrix. X′ and the target variable moisture content reference value y , the feature matrix X ′ and target variable y The dataset is divided into training and validation sets, with an empty variable set or a univariate set containing only the intercept term as the initial model state for stepwise regression.
[0036] S42. Variable Introduction: Based on the variables included in the current model, iterate through all candidate feature wavelengths not yet included in the model, and construct a temporary multiple regression model for each feature wavelength; based on the training set data, calculate and compare the preset evaluation indicators for each temporary model, including the root mean square error of cross-validation (RMSECV) and the coefficient of determination (R²). 2 Select the wavelength that most significantly improves model performance and formally add it to the current model.
[0037] S43. Variable Removal: For all wavelengths currently in the model, temporarily remove them one by one; build a new model after removing each wavelength; calculate and compare the evaluation index of the model after removal based on the training set data. If the evaluation index of the model does not decrease significantly or even improves after removing a certain variable, it is determined that the variable does not contribute significantly to the model and is permanently removed from the current model.
[0038] S44. Repeat steps S42 and S43, dynamically selecting and removing variables in each iteration. The iteration process stops when either of the following preset termination conditions is met: the improvement in the model evaluation index is lower than a preset small threshold, indicating that the model performance has stabilized; or the preset maximum number of iterations has been reached.
[0039] S45. After the iteration is completed, the set of all variables that are finally retained in the model is used as the optimal combination of feature wavelengths obtained after the second stage of screening, and is used to establish the final multivariate linear regression MLR quantitative correction model.
[0040] S5. Construction of a prediction model for the moisture content of honeysuckle.
[0041] Based on the optimal combination of characteristic wavelengths obtained in the second stage of screening, a multiple linear regression (MLR) quantitative correction model was established to achieve a quantitative mapping between spectral data and the moisture content of honeysuckle. The model expression is as follows: ; in, y i For the first i Reference values for moisture content of each sample. x ij For the first i The sample at the th j Spectral values at the optimal characteristic wavelengths β0 represents the intercept term. β j For regression coefficients, ε i For random error term, m The number of variables for the optimal combination of characteristic wavelengths.
[0042] The model parameters are solved by partial least squares regression to minimize the sum of squared residuals between the predicted and reference values, thus obtaining the regression coefficient vector. β The regression coefficients and model expressions are used as the final quantitative calibration model output for predicting the moisture content of honeysuckle.
[0043] S6. Import honeysuckle samples into the test set and evaluate the predictive performance of the trained model using the root mean square error of cross-validation (RMSECV), the root mean square error of the prediction set (RMSEP), and the coefficient of determination (R²). 2 The relative prediction deviation (RPD) is used as an evaluation index, and the definitions of each index are as follows: Root mean square error of cross-validation (RMSECV): ; Root mean square error of prediction set (RMSEP): ; Coefficient of determination ( ): ; Relative Prediction Deviation (RPD): ; In the formula, y i For the first i Reference moisture content value for each sample These are the model's predicted values. These are the predicted values obtained during the cross-validation process. This is the average of the reference values; n c and n p represents the number of samples in the calibration set and the prediction set, respectively, and SD is the standard deviation of the reference values of the prediction set samples.
[0044] Among them, the smaller the RMSECV and RMSEP values, the lower the model prediction error; the coefficient of determination R... 2 The closer the value is to 1, the better the model fits; the larger the RPD value, the stronger the model's predictive ability.
[0045] To verify the advantages of the two-stage feature wavelength screening method of this invention, this embodiment establishes moisture content prediction models based on the same dataset using four feature wavelength screening strategies: CARS, CARS-SR, EF-CARS, and the EF-CARS-SR of this invention. The performance evaluation indicators and the number of feature variables for each model on the test set are shown in the table below:
[0046] Experimental results show that the proposed EF-CARS-SR model exhibits significant advantages in both prediction performance and model simplicity. In the prediction set, the model's root mean square prediction error (RMSEP) is 0.0288, and the coefficient of determination (R²) is [missing value]. 2 The accuracy was 0.9739, and the relative prediction deviation (RPD) was 6.22, both better than the comparison model, indicating that the model has high prediction accuracy and good generalization ability. Furthermore, the EF-CARS-SR model ultimately retained 29 feature wavelengths, significantly reducing model complexity while maintaining prediction performance. The scatter plot of the model prediction results is shown below. Figure 5 As shown in the figure, the correlation between the predicted and measured values of the test set exhibits a clear linear relationship in the scatter plot. The scatter points are basically distributed along the diagonal, indicating that the modeling method can accurately reflect the actual moisture content of the sample. In the low moisture content range (approximately 0–0.3) and the medium moisture content range (approximately 0.3–0.6), the predicted values are basically consistent with the measured values with small deviations; in the high moisture content range (approximately 0.6–0.8), the predicted values fluctuate to some extent, but still maintain a linear trend overall. Therefore, the method described in this application can predict the moisture content of honeysuckle, has good stability, and provides a feasible solution for rapid online detection of moisture content.
[0047] Therefore, this invention employs a method for determining the moisture content of honeysuckle using a two-stage characteristic wavelength screening approach. Utilizing near-infrared spectroscopy, it employs a two-stage characteristic wavelength screening strategy. First, a coarse screening is performed using an error feedback competitive adaptive reweighted sampling method to eliminate redundant wavelengths. Then, a second-stage fine screening is performed using a stepwise regression algorithm to optimize variable combinations. This method effectively suppresses variable noise interference, reduces redundancy between characteristic wavelengths, improves the model's prediction accuracy and generalization ability, and enables rapid, non-destructive, and accurate detection of the moisture content of honeysuckle samples from different origins.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening, characterized in that, Includes the following steps: S1. Collect honeysuckle samples and obtain near-infrared spectral data under different moisture content gradients. At the same time, use the standard drying method to determine the moisture content of the corresponding samples as a reference value. S2. The Kennard-Stone algorithm is used to divide all samples into training and test sets. S3. The near-infrared spectral data is screened for the first stage of feature wavelengths using the error feedback competitive adaptive reweighted sampling method EF-CARS to obtain an initial set of feature wavelengths. The error feedback competitive adaptive reweighted sampling method introduces a prediction error feedback mechanism to dynamically update the variable weights. S4. Based on the initial set of characteristic wavelengths, the stepwise regression SR algorithm is used to perform the second stage of characteristic wavelength screening in order to further select the optimal combination of characteristic wavelengths from the initial set of characteristic wavelengths. S5. Construct a multiple linear regression (MLR) model using the optimal combination of characteristic wavelengths to determine the moisture content of honeysuckle. S6. Import the test set samples and evaluate the predictive performance of the model.
2. The method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening according to claim 1, characterized in that, In S1, the near-infrared spectral data acquisition process is as follows: the fresh honeysuckle sample is dried, and multiple samples are taken during the drying process. After cooling to room temperature, the sample is laid flat and scanned with a near-infrared spectrometer. The spectrum of each sample is collected three times and the average value is taken to obtain near-infrared spectral data with different moisture content gradients. The total moisture content of the sample is determined based on the difference between the initial mass and the mass after drying to constant weight. Reference values for the moisture content of the sample at different drying stages are then determined by combining the weighing results at each stage.
3. The method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening according to claim 1, characterized in that, In S2, the samples are divided according to the gradient of moisture content, and the training set and the test set are consistent in terms of moisture content distribution, origin and physical condition; the training set is used for feature wavelength selection and model building, and the test set is used to independently evaluate the model performance.
4. The method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening according to claim 1, characterized in that, In S3, the first stage of feature wavelength screening specifically includes the following: S31. Divide the spectral data matrix X The variable weights are initialized using the actual values of the corresponding target variables as input, and the maximum number of iterations is set. S32. In each iteration, from the spectral matrix X The training set and corresponding target variable subset are extracted from the dataset. Based on the variable subset retained in the current iteration, the training set is further filtered to obtain subset data. A partial least squares regression model is then established using the subset data and the target variable subset, and the prediction error of the model is calculated through cross-validation. ; S33. Based on the cross-validation prediction error, construct an error feedback function: ; in, k Indicates the number of iterations; This is the error feedback adjustment coefficient, used to control the intensity of the influence of prediction error on variable weights; S34. Update the variable weights based on the regression coefficients and error feedback function of the multiple linear regression model to obtain the comprehensive weights of the variables: ; in, Indicates the first k In the second modeling j The regression coefficients of the variables, Indicates the updated variable weights; S35. Employing an exponential decay function Control the proportion of variables retained in each iteration, where, Indicates the first k The proportion of variables retained in each iteration a and b A constant parameter to control the decay rate; and retaining the top weights according to the comprehensive weighting. The proportional variable proceeds to the next iteration; specifically, a and b Each by and Sure, p The initial total number of variables, N This represents the total number of CARS iterations. S36. After the iteration is completed, select the subset of variables with the smallest cross-validation prediction error as the initial feature wavelength set.
5. The method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening according to claim 1, characterized in that, In S4, the second stage of characteristic wavelength screening specifically includes the following steps: S41. Data Preparation and Initialization: Extract the spectral data corresponding to the initial feature wavelength set and construct a new feature matrix. X ′ and the target variable moisture content reference value y , the feature matrix X ′ and target variable y The dataset is divided into training and validation sets, with an empty variable set or a univariate set containing only the intercept term as the initial model state for stepwise regression. S42. Variable Introduction: Based on the variables included in the current model, iterate through the feature wavelengths not yet included in the model, and construct a temporary multiple regression model for each feature wavelength; based on the training set data, calculate and compare the preset evaluation indicators for each temporary model, including the root mean square error of cross-validation (RMSECV) and the coefficient of determination (R²). 2 Select the wavelength that will most significantly improve model performance and officially add it to the current model. S43. Variable Removal: For all wavelengths currently in the model, temporarily remove them one by one; build a new model after removing each wavelength; calculate and compare the evaluation index of the model after removal based on the training set data. If the evaluation index of the model does not decrease significantly or even improves after removing a certain variable, it is determined that the variable does not contribute significantly to the model and is permanently removed from the current model. S44. Repeat S42 and S43, performing variable filtering in each iteration. The iteration process stops when any of the following preset termination conditions are met: If the improvement in the model evaluation metrics is less than the preset small threshold, it indicates that the model performance has stabilized; or that the preset maximum number of iterations has been reached. S45. After the iteration is completed, the set of all variables that are finally retained in the model is used as the optimal combination of feature wavelengths obtained after the second stage of screening, and is used to establish the final multivariate linear regression MLR quantitative correction model.
6. The method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening according to claim 1, characterized in that, In S5, the expression for the multiple linear regression (MLR) model is: ; in, y i For the first i Reference values for moisture content of each sample. x ij For the first i The sample at the th j Spectral values at the optimal characteristic wavelengths β 0 represents the intercept term. β j For regression coefficients, ε i For random error term, m The number of variables for the optimal combination of characteristic wavelengths.
7. The method for determining the moisture content of honeysuckle by combining two-stage characteristic wavelength screening according to claim 1, characterized in that, In S6, the metrics for evaluating prediction performance include: root mean square error of cross-validation (RMSECV), root mean square error of prediction set (RMSEP), and coefficient of determination (R²). 2 Relative prediction bias (RPD).