Blueberry variety identification and quality nondestructive detection method based on hyperspectral imaging
By combining hyperspectral imaging technology with BOSVM and SOCAR-Net models, the problems of low efficiency and low accuracy in blueberry variety identification and quality detection have been solved, realizing non-destructive, rapid and high-precision multi-task collaborative detection, which is suitable for the intelligent upgrading of the blueberry industry.
Patent Information
- Application Number
- CN202510797311.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing technologies for blueberry variety identification and quality testing suffer from problems such as low efficiency, high sample loss, inability to provide real-time feedback, low classification accuracy, and weak model generalization ability. In particular, in scenarios involving mixed planting of multiple varieties, traditional methods struggle to achieve non-destructive, rapid, and high-precision multi-task collaborative detection.
A hyperspectral imaging-based approach was adopted. By sharing hyperspectral data input, the BOSVM model was used for variety classification and the SOCAR-Net model was used for quality prediction. By combining texture and color features for feature fusion, the BOSVM and SOCAR-Net models were constructed to independently complete the tasks of variety identification and quality prediction.
It enables rapid, non-destructive, and high-precision identification and quality testing of blueberry varieties, supports rapid identification in mixed multi-variety scenarios, has high accuracy in quality prediction, and possesses efficient, non-destructive, and high-precision multi-task collaborative testing capabilities.
Smart Images

Figure CN120314300B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural testing technology, specifically to a method for blueberry variety identification and non-destructive quality testing based on hyperspectral imaging. Background Technology
[0002] As an agricultural product, blueberries require careful variety identification and quality grading. Traditional testing methods rely on manual experience or destructive chemical analysis, which suffer from low efficiency, high sample loss, and lack of real-time feedback.
[0003] Especially in mixed-variety planting scenarios, manual identification is easily affected by subjective factors, resulting in low classification accuracy and difficulty in simultaneously detecting multiple quality indicators such as soluble solids, vitamin C, and anthocyanins. While near-infrared spectroscopy can achieve some non-destructive testing, it has poor adaptability to complex surface textures (such as the waxy layer and uneven skin of blueberries), and single spectral features are easily affected by varietal differences, leading to weak model generalization ability. Although hyperspectral imaging technology can provide spectral-spatial multidimensional information, existing methods do not fully integrate features, resulting in low utilization of key bands. Therefore, there is an urgent need for a non-destructive, high-precision, multi-task collaborative detection method to meet the needs of the entire blueberry industry chain, from sorting to quality inspection. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method for blueberry variety identification and multi-quality non-destructive testing based on hyperspectral imaging. By sharing hyperspectral data input, the method independently completes variety classification and quality prediction tasks using BOSVM and SOCAR-Net respectively.
[0005] To achieve the above objectives, the present invention provides a method for blueberry variety identification and non-destructive quality testing based on hyperspectral imaging, comprising the following steps:
[0006] 1. Hyperspectral data acquisition and preprocessing
[0007] Spectral data of blueberry samples were acquired using a hyperspectral imaging device (wavelength range 400-1000 nm, spectral resolution 1.9 nm), covering the visible light and part of the near-infrared band. The raw data underwent preprocessing, including using first derivatives to eliminate spectral noise and competitive adaptive reweighted sampling (CARS) to select feature bands. Texture and color features were extracted from the blueberry hyperspectral images.
[0008] 2. Training and Testing of the Variety Identification Model (BOSVM)
[0009] The optimal gamma and C are automatically searched using a Bayesian algorithm, with the optimization objective being to maximize classification accuracy. Input features include spectral data, texture features, and color features. The normalized fused data is input into the BOSVM model, which outputs the blueberry variety name. The test set classification accuracy reaches 100%, with a kappa coefficient of 1, supporting rapid identification in mixed multi-variety scenarios.
[0010] 3. Training and Testing of the Quality Prediction Model (SOCAR-Net)
[0011] The SOCAR-Net model comprises the following core modules: SEBlock (dynamically weighted key spectral bands to suppress irrelevant noise); and residual connections (downsampling the original input through 1×1 convolutions and adding it to deep features to alleviate gradient vanishing). The SOCAR-Net model also optimizes its convolutional layer design, employing three layers of convolutions and pooling, with a pooling layer added after each convolutional layer to more effectively extract multi-scale features from the input data. Hyperspectral data and measured quality values are input into the SOCAR-Net model, which outputs predicted values for soluble solids, vitamin C, and anthocyanin content, respectively. The coefficients of determination for both the training and test sets are greater than 0.94, the root mean square error is less than 0.82, and the relative analysis error is greater than 4. This achieves lossless, fast, and accurate prediction of blueberry quality.
[0012] 4. Results Integration and Output
[0013] By correlating the BOSVM classification results with the SOCAR-Net predicted values, a test report containing variety names and quality parameters is generated.
[0014] This invention provides a method for blueberry variety identification and non-destructive quality testing based on hyperspectral imaging. It has the following beneficial effects:
[0015] 1. This invention constructs BOSVM and SOCAR-Net models by sharing hyperspectral data input, which can be independently applied to blueberry variety identification and quality prediction. This achieves efficient, non-destructive, and high-precision multi-task collaborative detection, providing core technical support for the intelligent upgrading of the blueberry industry. Attached Figure Description
[0016] Figure 1 This is a structural diagram of the BOSVM model according to an embodiment of the present invention;
[0017] Figure 2 These are spectral diagrams of different varieties of blueberries according to embodiments of the present invention;
[0018] Figure 3 Average spectra of different blueberry varieties in embodiments of the present invention;
[0019] Figure 4These are texture feature images of different varieties of blueberries according to embodiments of the present invention;
[0020] Figure 5 These are color feature diagrams of different varieties of blueberries according to embodiments of the present invention;
[0021] Figure 6 These are the spectra of different blueberry varieties after processing with the first derivative according to an embodiment of the present invention.
[0022] Figure 7 This is a diagram illustrating the process of selecting CARS characteristic bands from blueberry hyperspectral data used for variety identification in an embodiment of the present invention.
[0023] Figure 8 This is a schematic diagram illustrating the principle of fusion between blueberry hyperspectral data and texture color feature data according to an embodiment of the present invention.
[0024] Figure 9 This is a confusion matrix diagram of the BOSVM model test set according to an embodiment of the present invention;
[0025] Figure 10 This is a structural diagram of the SOCAR-Net model according to an embodiment of the present invention;
[0026] Figure 11 This is a blueberry spectrum according to an embodiment of the present invention;
[0027] Figure 12 This is a diagram illustrating the process of selecting CARS characteristic bands from blueberry hyperspectral data used for quality testing in an embodiment of the present invention.
[0028] Figure 13 This is a fitted scatter plot of the soluble solids content of blueberries in an embodiment of the present invention;
[0029] Figure 14 This is a fitted scatter plot of the vitamin C content of blueberries according to an embodiment of the present invention;
[0030] Figure 15 This is a fitted scatter plot of the anthocyanin content of blueberries in an embodiment of the present invention;
[0031] Figure 16 This is a flowchart of the detection method according to an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Example:
[0034] Please see the appendix Figure 1 - Appendix Figure 16 This invention provides a method for blueberry variety identification and non-destructive quality testing based on hyperspectral imaging, including the establishment of a BOSVM model, wherein the establishment of the BOSVM model specifically involves:
[0035] 1. Bayesian optimization (BO) is a global optimization method based on a probabilistic model, particularly suitable for optimization problems where the objective function is expensive or non-differentiable. The core idea of BO is to construct a surrogate model to approximate the objective function and select the next evaluation point that is most likely to improve the objective function through a sampling function. The computational process of Bayesian optimization is as follows:
[0036] (1) Proxy model
[0037] Boolean algorithm (BO) uses a probabilistic model to approximate the objective function f(x), typically a Gaussian process (GP). A Gaussian process is a nonparametric model defined as:
[0038] (3-1)
[0039] in It is a mean function, and its value is usually 0. It is a kernel function (or covariance function) that describes the input. and The similarity between them.
[0040] Given a set of training data and corresponding observations The posterior distribution of a Gaussian process can be calculated using Bayes' theorem:
[0041] (3-2)
[0042] in It is the likelihood of the observed value y given the objective function value f, and is usually assumed to be Gaussian noise. This is the prior distribution of the objective function value. According to the properties of Gaussian processes, the posterior distribution... It is still a Gaussian process, and the predicted mean of the objective function can be calculated. and variance .
[0043] (2) Acquisition function
[0044] Boolean optimization (BO) improves efficiency by selecting the sampling point that best improves the objective function, typically using a sampling function to guide the selection of the next sampling point. Common sampling functions include expected improvement and probabilistic improvement. Expected improvement measures the anticipated improvement of a given point relative to the current optimum. Using the posterior distribution of a Gaussian process, expected improvement can be expressed as:
[0045] (3-3)
[0046] in It is the currently known optimal objective value. These are exploration parameters (often used to balance exploration and development). It is the cumulative distribution function of the standard normal distribution. It is the probability density function of the standard normal distribution. The probability improvement measures the improvement probability of a point relative to the current best point, and is usually defined as:
[0047] (3-4)
[0048] (3) Select the next sampling point
[0049] BO selects the next evaluation point by maximizing the acquisition function. :
[0050] (3-5)
[0051] (4) Update the proxy model
[0052] BO updates the surrogate model after selecting new sampling points and evaluating the objective function at each step.
[0053] (5) Stopping conditions
[0054] The optimization process will stop when the stopping condition is met. The stopping condition is usually the maximum number of iterations, the convergence of the objective function value, or the reaching of the maximum number of evaluations.
[0055] 2. Boolean optimization (BO) can adaptively adjust the hyperparameters of an SVM to find the optimal parameter values that maximize SVM performance, reducing the complexity of manual tuning and hyperparameter search. The main process of BO in optimizing an SVM is as follows: Figure 1As shown, in the Boolean Optimization (BO) process, the hyperparameters to be optimized are first defined. For SVM models, the two most important hyperparameters are usually C and gamma. C controls the tolerance for error; the larger the C value, the more the SVM focuses on fitting the training data. Gamma controls the sensitivity of the Gaussian radial basis function (RBF kernel) to the distance between data points. Next, the objective function is defined, which evaluates the model performance under the currently selected hyperparameters (C and gamma). Cross-validation is used to evaluate the model's loss function, ensuring the model's generalization ability on unseen data. The core of BO is to use a surrogate model to approximate the objective function and select the next sampling point by maximizing the sampling function, i.e., selecting a set of hyperparameters for evaluation. As optimization progresses, the surrogate model becomes increasingly accurate, thus converging to the optimal solution faster. During this process, BO gradually adjusts the hyperparameters by evaluating the objective function (cross-validation loss), returning the hyperparameter combination that minimizes the objective function value. Finally, the SVM model is trained using the optimal parameters found by BO, and the trained model is evaluated.
[0056] In this embodiment, hyperspectral imaging technology combined with a BOSVM model was used to accurately identify blueberry varieties. First, images and spectral data of four blueberry varieties were acquired using a hyperspectral imaging system, and a BOSVM model was constructed simultaneously. The spectral data of the blueberries are shown below. Figure 2 As shown, the trends and peak / trough positions of the spectral curves are similar, but the reflectance differs. The reflectance varies significantly within the 850-950 nm wavelength range. The troughs appearing in the 600-700 nm range are related to the anthocyanin and chlorophyll content within the blueberry, while the peaks and troughs appearing in the 800-1000 nm range may be related to the carbohydrate and water content within the blueberry. Averaging the spectral data for each blueberry variety yields average spectral data, representing the spectral characteristics of different blueberry varieties.
[0057] like Figure 3 As shown in the figure, the blueberries 'Brilliant,' 'Rika,' 'Eureka,' and 'L25' exhibit differences in reflectance within the 800-1000 nm wavelength range, demonstrating the feasibility of classifying blueberry varieties using spectral information. Texture features of blueberries were extracted from their hyperspectral images. To extract texture features more comprehensively, texture information was extracted from four directions: 0°, 45°, 90°, and 135°. Different kernel sizes (3, 5, and 7) were set to extract texture features at different scales, which helps extract detailed and large-scale texture information, enhancing feature diversity and recognition ability. Averaging the feature values from the four directions yielded 12 texture features for a single blueberry sample.
[0058] like Figure 4As shown, the figures are box plots of texture feature data (contrast, correlation, homogeneity, entropy) extracted from blueberry images after averaging different kernel sizes. The box plots display the maximum, minimum, median, mean, outliers, and dispersion of the data. Figure 4 As shown in A, Eurica has the highest mean contrast, indicating that Eurica may have a coarser texture compared to the other three varieties, and that the data dispersion between Eurica and Rica is relatively high. Figure 4 As shown in B, Eurica has the highest mean correlation, indicating that its texture may be more consistent, while L25's texture may be more irregular. Brilliant has the highest mean homogeneity, suggesting its texture may be more uniform, while Eurica has the lowest homogeneity dispersion, indicating smaller differences in homogeneity between samples. Figure 4 As shown in C, Eureka has the highest mean entropy, while the entropy values of the other three varieties are similar, indicating that Eureka may have the highest texture complexity. Figure 4 As shown in D.
[0059] Figure 4 This study fully demonstrates the differences in texture characteristics among different blueberry varieties, illustrating the feasibility of using blueberry texture features for variety identification. Color features of the RGB and HSV channels of each blueberry sample were extracted from blueberry hyperspectral images; specifically, the average value of each pixel across the six channels (R, G, B, H, S, V) was calculated, resulting in six color features for each blueberry sample. The box plots of the six color feature data are shown below. Figure 5 As shown, the box plot trends of color characteristics R, G, B, H, S, and V are basically consistent. Eurica has the lowest mean values for R, G, B, H, S, and V, indicating that Eurica's color may be duller compared to the other three blueberry varieties.
[0060] Figure 5 The H channel data in D shows little dispersion. Except for Eurica, the average H channel values of the other three blueberry varieties are relatively similar, indicating that the hues of Brilliant, Eurica, and L25 may be quite consistent. Figure 5 This study fully demonstrates the differences in color characteristics among different blueberry varieties and further illustrates the feasibility of using blueberry texture features for variety identification. Combining color and texture features allows for a more accurate description of the overall appearance of blueberry samples. Blueberry samples were randomly divided into training and test sets at a 4:1 ratio. After first-order derivative preprocessing of the hyperspectral data, as shown... Figure 6 As shown, the first-order derivative algorithm enhances the peak and trough information in the spectrum, making the spectral features more prominent and thus increasing the amount of spectral information, which helps improve model performance. Feature bands are selected using CARS, such as... Figure 7 As shown, by Figure 7As shown in Figure A, the number of characteristic wavelengths decreases continuously with the increase of sampling times, and the curve in the figure shows a trend of first dropping sharply and then gradually decreasing. This trend indicates that the process of CARS extracting characteristic wavelengths is from coarse selection to fine selection. Figure 7 From B and 7C, we know that when the number of samples is 6, the root mean square error (RMSECV) reaches its minimum value of 0.2856, at which point the number of extracted feature wavelengths is 142. The spectral data and texture color features are then fused and normalized. The fusion principle is as follows: Figure 8 As shown, the BOSVM model was trained using the training set, and then used to predict data on the test set. The BOSVM model achieved 100% classification accuracy on both the training and test sets. The confusion matrix for the test set is shown below. Figure 9 As shown in the figure, all four blueberry varieties in the test set were correctly predicted, and the kappa coefficients of both the training and test sets were 1, achieving rapid, non-destructive, and accurate identification of blueberry varieties.
[0061] In this embodiment of the application, the SOCAR-Net model is established as follows:
[0062] The SOCAR-Net model introduces a Squeeze-and-Excitation (SE) module, which utilizes a channel attention mechanism to weight and adjust the output of convolutional layers. By adaptively weighting the features of each channel, the SE module significantly improves the model's focus on important features, enhancing its feature representation capabilities. The SE module works by first inputting a feature map... Where C is the number of input channels and L is the feature length, i.e., the feature dimension of the spectrum, the Squeeze operation is then performed. The SE module first performs global average pooling on the input feature map. Since it is one-dimensional spectral data, the pooling operation is performed along the feature dimension (i.e., the wavelength dimension), rather than pooling across channels. The global average pooling formula for each channel is as follows:
[0063] (5-1)
[0064] Where xcl represents the feature Figure X The feature value of the c-th channel at position l is given by zc, where zc is the average value of that channel. Finally, an excitation operation is performed, modeling the weights of each channel using two fully connected layers, as shown in the following formula:
[0065] (5-2)
[0066] Where W1 and W2 are the weight matrices of the fully connected layer, b1 and b2 are bias terms, δ represents the ReLU activation function, σ represents the Sigmoid activation function, and yc is the weighting coefficient for each channel. Finally, the input features... Figure X By multiplying by the weighting coefficient yc of each channel for weighting adjustment, the SE module strengthens important channels and enhances the model's ability to focus on key information.
[0067] Deep convolutional networks often face gradient vanishing and information decay problems during training, making it difficult for the network to learn effective feature representations. To alleviate this problem, the SOCAR-Net model introduces residual connections. Residual connections, through skip connections, add the input features to the output of the deep network, ensuring smooth information flow between layers. The principle of residual connections is based on the assumption that the input feature is X and the output feature is F(X). The residual connection directly adds the input to the output of the convolutional layer, as shown in the following formula:
[0068] (5-3)
[0069] Where F(X) is the output calculated through several convolution operations. This is the final output. The introduction of residual connections ensures that information is not lost during deep network propagation, thereby accelerating the training process and improving the model's expressive power. Meanwhile, the SOCAR-Net model optimizes its convolutional layer design, employing three layers of convolution and pooling, with a pooling layer added after each convolutional layer to more effectively extract multi-scale features from the input data.
[0070] The SOCAR-Net model structure is as follows: Figure 10 As shown, the model contains three convolutional blocks. Each convolutional block consists of a convolutional layer, a Batch Normalization (BN) layer, a ReLU activation function, a SE module, and max pooling. After the third convolutional block, there are residual connections and fully connected layers. Regarding parameter settings, the kernel size is 3; the Adam optimizer is used; the initial learning rate is 0.001, which decays to 0.5 times the original value every 50 epochs using the StepLR scheduler; the dropout rate is 0.2; and the training epochs are 1000.
[0071] In this embodiment, hyperspectral imaging technology combined with the SOCAR-Net model was used to accurately identify the quality of blueberries. First, spectral data of the blueberries were acquired using a hyperspectral imaging system. Simultaneously, a SOCAR-Net model was constructed, which includes an SE module and residual connections. The spectral curves of the blueberry samples are shown below. Figure 11As shown, the spectral curves of different blueberry samples exhibit a consistent trend, with peaks and troughs containing quality information. Then, using the corresponding detection methods in national and industry standards, the contents of soluble solids, vitamin C, and anthocyanins in blueberries were actually measured. The spectral physicochemical co-occurrence distance algorithm (SPXY) was used to divide the sample training and test sets at a 4:1 ratio. Feature wavelengths were extracted using CARS, such as... Figure 12 As shown, by Figure 12 As shown in A, consistent with the selection of CARS characteristic wavelengths during the construction of the blueberry variety identification model, the curve in the figure also exhibits a trend of first sharp decrease and then gradual decrease. Figure 12 From B and 12C, we know that when the number of samples is 17, RMSECV reaches its minimum value of 0.4478, at which point the number of extracted feature wavelengths is 49. The SOCAR-Net model was trained using the training set, and then used to predict data on the test set. The SOCAR-Net model performed excellently in predicting soluble solids, vitamin C, and anthocyanin content. The scatter plot of the model's predicted values and the fitted values is shown below. Figure 13 , 14 As shown in Figure 15, there is a significant linear relationship between the predicted and measured values of quality, and the predicted and measured values are very close. The coefficients of determination (R2) of the training and test sets for the three indicators are all above 0.94, and the relative analysis error (RPD) of the training and test sets is above 4. This indicates that the SOCAR-Net model proposed by combining hyperspectral imaging technology can achieve non-destructive, rapid and accurate prediction of the content of soluble solids, vitamin C and anthocyanins in blueberries.
[0072] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for nondestructive detection of blueberry variety identification and quality based on hyperspectral imaging, characterized in that, The method comprises the following steps: S1, acquiring multi-dimensional feature data of the spectrum, texture and color of the blueberry sample by a hyperspectral imaging device, wherein the texture information of the blueberry is extracted from four directions of 0°, 45°, 90° and 135°, and 12 texture features of a single blueberry sample are extracted by averaging the feature values of the four directions with different kernel sizes; the average values of each pixel in R, G, B, H, S and V channels are calculated by the blueberry hyperspectral image, and each blueberry sample has six color features; the hyperspectral data is preprocessed by first derivative, and the feature wavelength is selected by CARS; S2, constructing a BOSVM model, classifying and identifying the blueberry varieties based on the multi-dimensional feature data, and outputting a variety label, wherein the variety label comprises Brilliant, Rica, Superior Rica and L25; S3, constructing a spectrally optimized SOCAR Net, based on the spectral data, simultaneously predicting the soluble solids, vitamin C, and anthocyanin contents of blueberries; S4, integrating the variety label and the multi-quality parameter to generate a detection report containing the variety identification result and the quality index; The optimization method of the BOSVM model in step S2 comprises: (1) automatically searching the kernel function parameters and the penalty factor of the support vector machine by the Bayesian optimization algorithm, and the optimization target is to maximize the classification accuracy, the best penalty parameter is 42.87, and the scale parameter of the best kernel function is 4.97; (2) the input features include the spectral band, the texture feature and the multi-dimensional data after color space conversion; The implementation of the SOCAR-Net in step S3 comprises: (1) three convolution modules, each module comprising a convolution layer, a batch normalization layer, an activation function and an SE attention module; (2) a residual connection module, which performs down-sampling on the original input through 1×1 convolution, and adds the output of the last convolution after adjusting the size by interpolation; (3) a channel attention mechanism dynamically allocates weights according to the importance of spectral features, and strengthens the key band feature extraction.
2. The nondestructive method for blueberry variety identification and quality detection based on hyperspectral imaging according to claim 1, characterized in that, The spectral coverage range of the hyperspectral imaging device in step S1 is 400-1000nm, the spectral resolution is 1.9nm, and the spectral bandwidth is 1.3nm.
3. The non-destructive method for blueberry variety identification and quality detection based on hyperspectral imaging according to claim 1, characterized in that, The classification accuracy of the training set and the test set of the variety identification model is 100%, and the kappa coefficient is 1, the determination coefficient of the soluble solids, vitamin C and anthocyanin content of the quality detection model is greater than 0.94, the root mean square error is less than 0.82, and the relative analysis error is greater than 4.
4. The blueberry variety identification and quality nondestructive detection system based on hyperspectral imaging, for realizing the blueberry variety identification and quality nondestructive detection method based on hyperspectral imaging in any one of claims 1-3, characterized in that, The hyperspectral imaging module is used for acquiring the multi-dimensional data of the spectrum, texture and color of the blueberry sample; The data preprocessing module is used for spectral data preprocessing, feature wavelength selection and feature fusion; the BOSVM processing unit realizes variety classification based on Bayesian optimization support vector machine; the SOCAR-Net processing unit predicts multi-quality parameters based on channel attention residual network; and the report generation module integrates the classification and detection results to output a comprehensive report.
5. The hyperspectral imaging based non-destructive detection system for blueberry variety identification and quality assessment according to claim 4, characterized in that, The BOSVM processing unit and the SOCAR-Net processing unit share the hyperspectral data input, and realize detection by calculation.
6. The hyperspectral imaging based non-destructive detection system for blueberry variety identification and quality assessment according to claim 4, characterized in that, The SOCAR-Net processing unit supports end-to-end training, and adopts an adaptive learning rate scheduling strategy to optimize the model parameters.
Citation Information
Patent Citations
Hyperspectral image classification method based on spectrum-space self-attention and Transform network
CN117315481A