Rice origin identification method based on dynamic feature screening and weighted integrated model

Through wavelet denoising and improved SNV correction methods combined with dynamic feature screening and multi-stage mixed classification model, the problems of fluctuations in the accuracy of the rice origin identification method and low computational efficiency of the mid-infrared spectroscopic rice origin identification method are solved, and efficient and accurate rice origin identification is achieved.

CN120408341BActive Publication Date: 2025-08-22CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510884668.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-22
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing mid-infrared spectroscopic rice production identification methods have problems such as large fluctuations in accuracy, poor model mobility, and low computing efficiency. Especially when the differences in adjacent production areas are small, the classification specificity is significantly reduced.

Method used

The wavelet denoising and improved SNV correction method are used for spectral preprocessing, combined with dynamic feature screening and multi-level mixed classification model, and the feature subset optimization and model integration are achieved through dynamic weight sets, and a rice origin identification method based on dynamic feature screening and weighted integration model is constructed.

Benefits of technology

The accuracy and stability of the mid-infrared spectroscopic rice origin identification is improved, the computing time is reduced, the model's mobility and computing efficiency are enhanced, the feature retention rate and noise reduction effect are significantly improved, and the cross-device accuracy fluctuates by less than 1.5%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408341B_ABST
    Figure CN120408341B_ABST
Patent Text Reader

Abstract

A method for identifying the origin of rice based on dynamic feature screening and weighted integrated model. It belongs to the field of agricultural product traceability technology, and specifically relates to the field of intelligent identification technology of rice origin. It solves the prominent problems of existing mid-infrared spectral identification methods, such as large fluctuations in accuracy, poor model transferability, and low computational efficiency. The method comprises the following steps: using a mid-infrared spectrometer equipped with an ATR accessory to collect spectra of rice and rice flour samples, and performing spectral standardization processing; using wavelet denoising and improved SNV correction method to reduce noise and correct the standardized spectra to obtain reduced-noise corrected spectra; performing dynamic feature screening on the reduced-noise corrected spectra to obtain an optimized feature subset with low redundancy and high independence; by optimizing the feature subset, a number of classification models are integrated into a multi-level hybrid classification model by using a dynamic weight set; and the multi-level hybrid classification model is used to identify the origin of rice and rice flour samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agricultural product traceability, and specifically relates to the technical field of intelligent identification of rice origin. Background Art

[0002] In recent years, mid-infrared spectroscopy (MIR) technology has been increasingly applied to agricultural product origin traceability due to its ability to sensitively reflect molecular vibrational information in organic matter. Existing methods for rice origin identification based on MIR spectroscopy typically follow a conventional process of "spectral acquisition - preprocessing - feature extraction - classification modeling," but practical applications still face several technical bottlenecks.

[0003] Traditional preprocessing methods (such as single standard normal transformation (SNV) or multivariate scatter correction (MSC)) have limited effectiveness in removing baseline drift and stray light interference unique to mid-infrared spectroscopy. For example, while conventional SNV processing can eliminate scattering effects, it can overcorrect weak peaks associated with origin (such as bending vibration signals of COH bonds), resulting in loss of effective chemical information. Furthermore, there is a lack of targeted algorithms for removing high-frequency noise in mid-infrared spectra (primarily arising from contact fluctuations between the ATR crystal and the sample), which directly affects the accuracy of subsequent analysis.

[0004] Existing methods often use full-spectrum analysis or simple thresholding to screen for characteristic wavelengths, failing to fully consider the synergistic effects between variables. For example, dimensionality reduction methods based on principal component analysis (PCA) tend to overlook key discriminant information in minor components, while the variable importance projection (VIP) algorithm, directly transplanted from near-infrared analysis, fails to adjust the threshold strategy to account for the high dimensionality and high collinearity of mid-infrared spectra. This makes the model susceptible to interference from redundant variables, especially when the differences in origin are small (such as between adjacent provinces), significantly reducing classification specificity.

[0005] Existing technologies mostly use a single classifier (such as support vector machine (SVM) or random forest), whose performance is limited by the inherent defects of the algorithm. SVM is sensitive to kernel function parameters and is prone to overfitting in small sample scenarios. Although random forest is noise-resistant, it lacks the ability to distinguish subtle difference features.

[0006] In addition, traditional cross-validation methods find it difficult to effectively balance model complexity and generalization ability, resulting in poor reproducibility of data from different instruments or batches in practical applications (relative standard deviation RSD often exceeds 5%).

[0007] Mid-infrared spectral data is generally high-dimensional (typically containing over 2,000 data points). Existing methods often use a serial computational model during feature extraction and model training. Experiments have shown that traditional methods can take up to 2-3 seconds to process a single sample, making them difficult to meet real-time detection requirements.

[0008] These technical deficiencies lead to significant issues with existing mid-infrared spectroscopy identification methods, including wide accuracy fluctuations (85%-92%), poor model transferability, and low computational efficiency. Therefore, a new mid-infrared spectroscopy analysis method that combines high precision and high efficiency is needed to enable rapid and accurate identification of rice origin. Summary of the Invention

[0009] In order to solve the prominent problems of existing mid-infrared spectroscopy identification methods, such as large accuracy fluctuations (85%-92%), poor model transferability, and low computational efficiency, the present invention provides a rice origin identification method based on dynamic feature screening and weighted integrated model, the method comprising the following steps:

[0010] S1. Use a mid-infrared spectrometer equipped with an ATR accessory to collect spectra of rice and rice flour samples and perform spectral standardization.

[0011] S2, using wavelet denoising and improved SNV correction method to reduce noise and correct the standardized spectrum to obtain the noise-reduced and corrected spectrum;

[0012] S3, dynamic feature screening is performed on the noise reduction and correction spectrum to obtain an optimized feature subset with low redundancy and high independence;

[0013] S4. By optimizing feature subsets and adopting dynamic weight set, several classification models are integrated into a multi-level hybrid classification model;

[0014] S5. Use a multi-class mixed classification model to identify the origin of rice and rice flour samples.

[0015] Furthermore, the spectrum normalization process is performed by: Conduct, among which represents the normalized spectral intensity vector, represents the original spectral intensity vector, Indicates taking the minimum / maximum value of the spectrum vector.

[0016] Furthermore, the improved SNV correction method is specifically as follows: introducing a baseline protection factor into the SNV correction method , ,in, Indicates the wavenumber position of the spectrum, unit: , the value range is consistent with the collected spectrum range, Indicates the permitted range, the value range is 0.1-0.3, Indicates the central wave number of the characteristic peak concentration area, Represents the standard deviation of the wavenumber distribution in the characteristic peak area, and introduces the baseline protection factor Finally, the correction formula of the improved SNV correction method is: ;in, represents the spectral intensity value after improved SNV correction, Indicates that the spectrum The absorbance value at and are the mean and standard deviation of the spectral vector respectively.

[0017] Furthermore, the dynamic feature screening is specifically as follows:

[0018] S31. Quantify each variable through variable projection importance analysis The contribution to the classification model is recorded as ;variable The spectral intensity value of a single wave number point representing the noise reduction and correction spectrum;

[0019] S32, using sliding window to calculate dynamic threshold value, through dynamic threshold function Get the threshold of the current window , Indicates the The central wave value of the window is and Indicates the scope of permission, , ;

[0020] S33. Perform dynamic feature screening by collinearity elimination and set the retained variable conditions: Condition 1: > and +Condition 2: Correlation coefficient of adjacent variables , filter the variables;

[0021] in, Indicates the The peak significance at the wave number point, represents the standard deviation of the noise area, , Indicates the The normalized spectral intensity value at the wave number point is represents the average spectral intensity in the baseline region.

[0022] Furthermore, in step S4, a training set and a validation set are constructed by optimizing the feature subsets, the training set is used to train each classification model, the validation set is used to adjust the weights of each classification model when adjusting the dynamic weight set, and the prediction results of each classification model are fused by decision making through the weights of each classification model to obtain the final rice origin identification result.

[0023] Furthermore, the classification model set includes SVM model, RF model, LightGBM model and 1D-CNN model.

[0024] Furthermore, when training the SVM model, all the features of the training set are input; when training the RF model, all the features of the training set are input, and all the features of the training set are compressed to 50 dimensions using PCA before input; when training the LightGBM model, all the features of the training set are input; when training the 1D-CNN model, the standardized spectrum is input.

[0025] Furthermore, when using the validation set to adjust the dynamic weight set, the weights of each classification model are specifically:

[0026] Adjust the weights of each classification model, where Indicates the The dynamic weight of each branch model, Indicates the The macro of the branch model on the validation set, 1.5 represents the weight adjustment index.

[0027] Furthermore, the prediction results of each classification model are fused by weights of each classification model to obtain the final rice origin identification result. accomplish, Indicates the final predicted origin category label, Indicates the rice production area category, Indicates the Rice origin classification The predicted probability of Indicates the category that maximizes the sum of weighted probabilities .

[0028] The beneficial effects of the method of the present invention are:

[0029] (1) Composite preprocessing technology: the first tandem structure of wavelet denoising and improved SNV, solving the problem of traditional SNV overcorrection through baseline protection factor;

[0030] (2) Dynamic VIP screening mechanism: Proposing a wave number adaptive VIP threshold function ,breaking the bottleneck of feature omission caused by breaking the fixed threshold;

[0031] (3) Heterogeneous feature hybrid model: Full features and compressed features are processed separately through multiple models, and complementary advantages are achieved through dynamic weights.

[0032] (4) Incorporate model training time into the Bayesian optimization objective function to balance accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Flowchart of a method for identifying rice origin based on dynamic feature screening and weighted integrated model in an embodiment of the present invention;

[0034] Figure 2 A diagram showing the construction of a multi-level hybrid classification model in an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0036] Example 1

[0037] This embodiment provides a method for identifying rice origin based on dynamic feature screening and weighted integrated model. Figure 1 As shown, the method includes the following steps:

[0038] S1. Use a mid-infrared spectrometer equipped with an ATR accessory to collect spectra of rice and rice flour samples and perform spectral standardization.

[0039] S2, using wavelet denoising and improved SNV correction method to reduce noise and correct the standardized spectrum to obtain the noise-reduced and corrected spectrum;

[0040] S3, dynamically screening the noise reduction and correction spectrum to obtain an optimized feature subset with low redundancy and high independence;

[0041] S4. By optimizing feature subsets, several classification models are integrated into a multi-level hybrid classification model using a dynamic weight set approach;

[0042] S5. Use a multi-class hybrid classification model to identify the origin of rice and rice flour samples and output the prediction results.

[0043] Example 2

[0044] This embodiment further limits the embodiment 1. A mid-infrared spectrometer equipped with an ATR accessory is used to collect spectra of rice and rice flour samples, and the parameters are specifically controlled as follows:

[0045] Wave number range: 4000-400 ;

[0046] Resolution: 0.09 ;

[0047] Number of scans: 32 (permitted range 16-64).

[0048] The spectrum normalization process is performed by: Conduct, among which represents the normalized spectral intensity vector, represents the original spectral intensity vector, Indicates taking the minimum / maximum value of the spectrum vector.

[0049] Example 3

[0050] This embodiment further limits the embodiment 1.

[0051] The decomposition parameters of wavelet denoising are specifically:

[0052] Wavelet basis: Daubechies5 (licensed range db3-db10);

[0053] Number of decomposition layers: 6 layers (permitted range 5-8 layers).

[0054] The threshold processing of wavelet denoising is specifically as follows:

[0055] The SURE (Stein Unbiased Risk Estimation) threshold method is used for threshold processing:

[0056] ;

[0057] in, Indicates the Layer wavelet coefficient threshold;

[0058] : The standard deviation of the noise in this layer (estimated by Median Absolute Deviation);

[0059] : No. The number of layer detail coefficients.

[0060] The improved SNV correction method is specifically: introducing a baseline protection factor into the SNV correction method , ,in, Indicates the wavenumber position of the spectrum, unit: , the value range is consistent with the collected spectrum range, Indicates the permitted range, the value range is 0.1-0.3, Indicates the central wave number of the characteristic peak concentration area, Represents the standard deviation of the wavenumber distribution in the characteristic peak area, and introduces the baseline protection factor Finally, the correction formula of the improved SNV correction method is: ;in, represents the spectral intensity value after improved SNV correction, Indicates that the spectrum The absorbance value at and are the mean and standard deviation of the spectral vector, respectively.

[0061] Example 4

[0062] This embodiment further limits the first embodiment, and the dynamic feature screening is specifically as follows:

[0063] S31. Quantify each variable through variable projection importance analysis The contribution to the classification model is recorded as ;variable Represents the spectral intensity value of a single wave number point of the noise reduction and correction spectrum;

[0064] For example, the spectral range 4000-600 (Step 2 ), there are 1701 variables in total ( Corresponding to 4000 , Corresponding to 600 ).

[0065] S32, using sliding window to calculate dynamic threshold value, through dynamic threshold function Get the threshold of the current window , Indicates the The central wave value of the window, in units of ,in and Indicates the scope of permission, , ;

[0066] For example, in In a window, if the window range is 1700-1650 ,but =1675 , the sliding window is applied to the spectral axis for noise reduction correction, a variable Corresponding to one , through the dynamic threshold at the center wave number of the current window , filter out all variables in the window of .

[0067] S33. Perform dynamic feature screening by collinearity elimination and set the retained variable conditions: Condition 1: > and +Condition 2: Correlation coefficient of adjacent variables , filter the variables;

[0068] in, Indicates the The peak significance at the wave number point, represents the standard deviation of the noise area, , Indicates the The normalized spectral intensity value at the wave number point is represents the average spectral intensity of the baseline region.

[0069] Example 5

[0070] This embodiment is a further limitation of Embodiment 1. In step S4, a training set and a validation set are constructed by optimizing feature subsets, the training set is used to train each classification model, and the validation set is used to adjust the weights of each classification model when adjusting the dynamic weight set. The prediction results of each classification model are fused by decision-making using the weights of each classification model to obtain the final rice origin identification result.

[0071] like Figure 2 As shown, the classification model set includes SVM model, RF model, LightGBM model and 1D-CNN model.

[0072] The optimization objective function is defined as:

[0073] ;

[0074] : Macro F1-score of the validation set (range 0-1);

[0075] : Single training time (seconds), is its coefficient;

[0076] : Model complexity penalty term, As its coefficient, the calculation formula is:

[0077] ;

[0078] : the number of support vectors of support vector machine;

[0079] : number of random forest decision trees;

[0080] N: number of training set samples;

[0081] Dataset: Rice spectral data from 12 rice-producing areas in Jilin Province (N=1200, training set: validation set: test set = 7:1.5:1.5);

[0082] Number of optimization iterations: 50.

[0083] The optimization parameter trajectory is shown in Table 1:

[0084] Table 1:

[0085]

[0086] The performance comparison of the optimal parameters is shown in Table 2:

[0087] Table 2:

[0088]

[0089] When training the SVM model, the details are as follows:

[0090] Input: all features of the training set;

[0091] Kernel function: RBF (Radial Basis Function) kernel function;

[0092] Regularization parameter: , which is used to balance the complexity of the classifier and the tolerance of training error. The kernel function is too large ( ): Strictly fit the training data, may overfit, the kernel function is too small ( ): Allow more misclassification and improve generalization;

[0093] Kernel function width parameter : , which controls the distribution of data mapped to high-dimensional space. The kernel function width parameter is too large ( ): The influence range of a single sample is small, the decision boundary is complex, and the kernel function width parameter is too small ( ): A single sample has a large influence range and a smooth decision boundary.

[0094] When training the RF model, the details are as follows:

[0095] Input: All features of the training set. Before input, all features of the training set are compressed to 50 dimensions using PCA (the permitted range is 30-70 dimensions).

[0096] Parameters: Number of trees = 100 (allowed range is 50-200), Maximum depth = 15 (allowed range is 10-20).

[0097] When training the LightGBM model, the details are as follows:

[0098] Input: all features of the training set;

[0099] Parameters: number of leaves = 128, learning rate = 0.05.

[0100] When training a 1D-CNN model, the following steps are performed:

[0101] Input: normalized spectrum;

[0102] Structure: From input to output, it is: 3 layers of convolution (kernel size is 5), maximum pooling and full connection. When using the validation set to adjust the dynamic weight set, the weights of each classification model are as follows:

[0103] Adjust the weights of each classification model, where Indicates the The dynamic weight of each branch model, Indicates the The macro F1-score of the branch model on the validation set, 1.5 represents the weight adjustment index, which is used to amplify the contribution of the high F1-score model.

[0104] The prediction results of each classification model are fused by weights of each classification model to obtain the final rice origin identification result. accomplish, Indicates the final predicted origin category label, Indicates the rice production area category, Indicates the Rice origin classification The predicted probability of Indicates the category that maximizes the sum of weighted probabilities .

[0105] Example 6

[0106] This embodiment is a further limitation of embodiment 1. In order to speed up the output of prediction results, GPU can be used to accelerate model training. Figure 1 As shown, MATLAB 2023a Bayesian Optimization Toolbox is used for parallelization. The specific parameters are:

[0107] Dataset: Rice spectral data from 12 rice-producing areas in Jilin Province (N=1200, training set: validation set: test set = 7:1.5:1.5)

[0108] Number of sample blocks: ( is the total number of samples);

[0109] Number of threads: 80% of the GPU's maximum threads (70-90% permitted range);

[0110] If the maximum output probability is <0.7, the manual review process is triggered and the low-confidence samples are recorded for model iteration.

[0111] Objective function:

[0112] ;

[0113] Based on the global optimization of Example 5, the time penalty coefficient is verified separately Regarding the impact of efficiency and precision, Example 6 further improves the energy efficiency ratio (F1 / Time) of the global optimization results of Example 5 by refining the efficiency penalty term in the objective function. The combined model further improves the ability to distinguish between adjacent production areas.

[0114] The specific implementation effects verified by the examples are shown in Table 3:

[0115] Table 3:

[0116]

[0117] In addition, the present invention has the following technical advantages: for the first time, it achieves a double breakthrough in feature retention rate (98.2%) and noise reduction effect (SNR≥35dB) in mid-infrared spectral processing; dynamic VIP screening increases the capture rate of key discriminant features (such as the amide II band at 1580 cm⁻¹) to 99.3%; and the hybrid model is significantly more robust to differences in spectrometer models (accuracy fluctuation across devices is <1.5%).

Claims

1. A rice origin identification method based on dynamic feature screening and weighted integrated model, characterized in that: The method comprises the following steps: S1. Use a mid-infrared spectrometer equipped with an ATR accessory to collect spectra of rice and rice flour samples and perform spectral standardization. S2, using wavelet denoising and improved SNV correction method to reduce noise and correct the standardized spectrum to obtain the noise-reduced and corrected spectrum; The improved SNV correction method is specifically: introducing a baseline protection factor into the SNV correction method , ,in, Indicates the wavenumber position of the spectrum, unit: , the value range is consistent with the collected spectrum range, Indicates the permitted range, the value range is 0.1-0.3, Indicates the central wave number of the characteristic peak concentration area, Represents the standard deviation of the wavenumber distribution in the characteristic peak area, and introduces the baseline protection factor Finally, the correction formula of the improved SNV correction method is: ;in, represents the spectral intensity value after improved SNV correction, Indicates that the spectrum The absorbance value at and are the mean and standard deviation of the spectral vector respectively; S3, dynamically screening the noise reduction and correction spectrum to obtain an optimized feature subset with low redundancy and high independence; S4. By optimizing feature subsets and adopting dynamic weight set, several classification models are integrated into a multi-level hybrid classification model; S5. Use a multi-class mixed classification model to identify the origin of rice and rice flour samples.

2. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 1, characterized in that: The spectrum normalization process is performed by: Conduct, among which represents the normalized spectral intensity vector, represents the original spectral intensity vector, Indicates taking the minimum / maximum value of the spectrum vector.

3. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 2, characterized in that: The dynamic feature screening is specifically as follows: S31. Quantify each variable through variable projection importance analysis The contribution to the classification model is recorded as ;variable The spectral intensity value of a single wave number point representing the noise reduction and correction spectrum; S32, using sliding window to calculate dynamic threshold value, through dynamic threshold function Get the threshold of the current window , Indicates the The central wave value of the window is and Indicates the scope of permission, , ; S33. Perform dynamic feature screening by collinearity elimination and set the retained variable conditions: Condition 1: > and +Condition 2: Correlation coefficient of adjacent variables , filter the variables; in, Indicates the The peak significance at the wave number point, represents the standard deviation of the noise area, , Indicates the The normalized spectral intensity value at the wave number point is represents the average spectral intensity in the baseline region.

4. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 3, characterized in that: In step S4, a training set and a validation set are constructed by optimizing the feature subsets, the training set is used to train each classification model, the validation set is used to adjust the weights of each classification model when adjusting the dynamic weight set, and the prediction results of each classification model are fused by decision making through the weights of each classification model to obtain the final rice origin identification result.

5. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 4, characterized in that: The classification model set includes SVM model, RF model, LightGBM model and 1D-CNN model.

6. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 5, characterized in that: When training the SVM model, all features of the training set are input; When training the RF model, all the features of the training set are input, and all the features of the training set are compressed to 50 dimensions using PCA before input; when training the LightGBM model, all the features of the training set are input; When training the 1D-CNN model, the normalized spectrum is input.

7. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 6, characterized in that: When using the validation set to adjust the dynamic weight set, the weights of each classification model are as follows: Adjust the weights of each classification model, where Indicates the The dynamic weight of each branch model, Indicates the The macro of the branch model on the validation set, 1.5 represents the weight adjustment index.

8. The method for identifying rice origin based on dynamic feature screening and weighted integrated model according to claim 7, characterized in that: The prediction results of each classification model are fused by weights of each classification model to obtain the final rice origin identification result. accomplish, Indicates the final predicted origin category label, Indicates the rice production area category, Indicates the Rice origin classification The predicted probability of Indicates the rice origin category that maximizes the weighted probability sum .

Citation Information

Patent Citations

  • Traceable method for rice origin and application thereof

    CN105021562A

  • Method and system for identifying geology of traditional Chinese medicinal materials and computer readable medium

    CN114624207A