Blueberry sugar degree hyperspectral inversion method based on improved Laplacian mapping

By combining multivariate scattering correction and preprocessing methods of fractional derivatives, improved Laplace mapping algorithm and convolutional neural network, the shortcomings of preprocessing and feature selection of blueberry hyperspectral data in the prior art are solved, and more efficient and accurate prediction of sugar levels are achieved.

CN120012044AActive Publication Date: 2025-05-16KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510500484.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing hyperspectral technology has poor generalization when predicting blueberry sugar content. Traditional preprocessing methods and feature selection algorithms are not enough to cope with blueberry hyperspectral data in complex scenarios, resulting in a degradation of model performance.

Method used

Multivariate scattering correction (MSC) and fractional derivative (FOD) are used for data preprocessing, feature bands are extracted using the improved Laplace mapping algorithm (ILE), and a lightweight sugar-predictive model is constructed through a convolutional neural network (CNN).

Benefits of technology

It improves the accuracy and speed of blueberry sugar content prediction in complex scenarios, reduces costs, and enhances the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012044A_ABST
    Figure CN120012044A_ABST
Patent Text Reader

Abstract

The invention relates to a blueberry sugar degree hyperspectral inversion method based on improved Laplacian mapping, and belongs to the technical field of blueberry hyperspectral acquisition and analysis. The method comprises the following steps: firstly, collecting blueberry samples of different varieties and different maturity degrees in a complex scene, measuring hyperspectral data of the blueberry samples and sugar content of corresponding blueberries, then dividing the blueberry sample data into a training set and a test set according to a proportion by using a statistical algorithm, and performing data preprocessing by using an MSC + FOD combination to obtain a training set and a test set; extracting characteristic wave bands from the preprocessed spectral data by using ILE provided by the invention, establishing a blueberry sugar degree prediction model by using CNN, training the model by using a training set, and finally comprehensively evaluating the model by using a test set in combination with regression indexes R2, RMSE and RPD. The method can quickly and accurately predict the sugar degrees of blueberries of different varieties and different maturity degrees in a complex scene, and provides powerful technical support for prediction of the sugar degrees of the blueberries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a blueberry sugar content hyperspectral inversion method based on improved Laplace mapping, and belongs to the technical field of blueberry hyperspectral acquisition and analysis. Background Art

[0002] As a fruit with extremely high nutritional value, blueberry is deeply loved by people. Sugar content is one of the standards for evaluating the quality of blueberries, and different sugar contents have different demands in the market. By measuring the sugar content of blueberries, agricultural producers can determine the best time to pick them, ensure the taste and quality of blueberries, and meet market demand to the greatest extent. Therefore, it is of great significance to quickly and accurately predict the sugar content in blueberries to promote the development of agricultural production. In traditional methods, determining the sugar content of blueberries involves collecting blueberries on site and a large number of chemical experimental analyses. This process is not only time-consuming and laborious, but also causes irreversible destructive damage to blueberries. In recent years, with the promotion of spectral technology in the direction of non-destructive testing, it has brought great convenience and progress to agriculture, improving the detection efficiency and accuracy without destroying the appearance and internal structure of the fruit, while reducing the various costs required for detection. At present, hyperspectral technology has been widely used in the rapid non-destructive detection research of blueberries and other fruits and vegetables, providing a powerful analytical method for quality assessment and component analysis of fruits and vegetables.

[0003] In the existing research on predicting the sugar content of blueberries using hyperspectral remote sensing technology, although there are many spectral processing methods and modeling methods, they are relatively traditional and have poor generalization. Especially when faced with blueberry hyperspectral data of different varieties and different maturity collected in complex scenes, traditional preprocessing methods such as SG smoothing, standard normal variable transformation (SNV), first-order derivatives, etc. are not enough to deal with complex spectral data, while fractional order derivatives (FOD) can not only eliminate baseline drift and suppress noise, but also amplify the subtle information hidden in the massive spectral signals, and better adapt to nonlinear spectral changes in complex scenes, thereby capturing global and local information in the spectrum at multiple scales. After preprocessing the hyperspectral data, the ability of traditional feature selection algorithms to accurately and quickly extract key information is still limited. For example, the competitive adaptive reweighting algorithm (CARS) is prone to generate redundant bands and needs to be used multiple times; the iterative information retention variable algorithm (IRIV) takes too long to calculate, affecting the modeling efficiency. Laplace mapping (LE) can analyze the local similarity between bands by constructing a Laplace matrix, adaptively assign band weights, and quickly extract characteristic wavelengths. The overall generalization ability of the algorithm is stronger. However, the blueberry spectrum after FOD processing will show nonlinear performance trends such as uneven distribution, long tail and multiple peaks due to the amplification of weak features. At this time, LE cannot better capture the complex relationship between data, and the risk of falling into local optimality and ignoring global information is greatly increased, which ultimately affects the overall performance of the model.

[0004] Therefore, the present invention proposes a blueberry sugar content hyperspectral inversion method based on improved Laplace mapping to address the above problems. Based on the inadequacy of existing research, it is proposed to use a combination of multivariate scatter correction (MSC) and fractional order derivative (FOD) to preprocess data, and to use an improved Laplace mapping algorithm (ILE) to extract the characteristic bands of blueberry hyperspectral while better adapting to FOD. Finally, a CNN network is used to construct a lightweight sugar content prediction model for blueberries of different varieties and maturity in complex scenarios and verify its applicability, providing a reference for further research on blueberry sugar content hyperspectral inversion and other related research. Summary of the invention

[0005] The purpose of the present invention is to provide a blueberry sugar content hyperspectral inversion method based on improved Laplace mapping in view of the deficiencies in the prior art. The method has higher prediction accuracy, faster prediction speed and lower cost for blueberries in complex scenarios, and has great practical application value and development potential.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0007] A blueberry sugar content hyperspectral inversion method based on improved Laplace mapping, firstly collect blueberry samples of different varieties and different maturity in complex scenes, measure the hyperspectral data of blueberry samples and the sugar content of corresponding blueberries, then select a suitable statistical algorithm to divide the blueberry sample data into a training set and a test set according to a certain ratio, use MSC+FOD combination to preprocess the data, then use the ILE proposed in the present invention to extract characteristic bands from the preprocessed spectral data, use CNN to establish a blueberry sugar content prediction model and use the training set to train the model, finally use the test set combined with the regression index R 2 , RMSE and RPD are used to comprehensively evaluate the model.

[0008] The specific steps include:

[0009] Step 1: Collect blueberry samples of different varieties and maturity in complex scenes, and measure the hyperspectral data and sugar content of the blueberry samples;

[0010] Step 2: Use statistical algorithms to divide the hyperspectral data and sugar content of blueberry samples, and divide the blueberry samples into two parts: training set and test set according to proportion;

[0011] Step 3: Preprocessing the divided blueberry sample data, wherein the preprocessing uses the MSC+FOD combination to perform spectral preprocessing on the blueberry hyperspectral data of the training set and the test set respectively to eliminate noise and highlight spectral features;

[0012] Step 4: Use the improved Laplace mapping algorithm ILE to extract characteristic bands from the blueberry sample data after MSC+FOD combination preprocessing;

[0013] Step 5: Use the convolutional neural network (CNN) in deep learning to build a blueberry sugar content prediction model in complex scenarios. Use the training set after preprocessing and feature extraction to train the model, and then use the test set to calculate the regression index R of the trained blueberry sugar content prediction model. 2 , RMSE and RPD, to validate the prediction model.

[0014] The Step 3 is specifically as follows:

[0015] Step 3.1: First, the multivariate scatter correction (MSC) method was used to preliminarily preprocess the blueberry data. The spectrum of each sample was linearly fitted with the average of all sample spectra to eliminate the scattering effect caused by external factors and enhance the correlation between the spectrum and the sugar content.

[0016] Step 3.2: Since the blueberry hyperspectral data collected in complex scenes are relatively complex, the fractional derivative FOD is introduced to process the blueberry data processed by MSC again. The measured blueberry sugar content data (m×1) and hyperspectral data (m×n) are used as input data respectively. According to the preset fractional order interval, the fractional derivative of each band of each blueberry spectral curve is calculated in 0~2 orders in sequence to obtain the processed spectral curve, where FOD is defined as the Grunwald-Letnikov function, and its formula is as follows:

[0017]

[0018] Where f(x) is the spectral signal, which represents the reflectivity at wavelength x; D represents the reflectivity f(x). Secondary derivatives; is the order; h is the derivative step size; b and a are the upper and lower limits of the derivative, respectively, and their values ​​increase dynamically with the increase of the number of bands; is the Gamma function, which is used to generalize factorials to non-integer orders. express The value of the Gamma function of order, express The Gamma function value of the order is used to scale and normalize the weights; m is a non-negative integer representing the number of steps between the current wavelength x and the historical data point x − mh; represents a series of historical data points near x, is the spectral reflectance of the function at the historical data point x − mh;

[0019] Step 3.3: By calculating the fractional derivative at point x, the local variation characteristics of the spectral signal are captured to flexibly mine the weak features in the blueberry hyperspectrum.

[0020] The Step 4 is specifically as follows:

[0021] Step 4.1: Introduce the Manhattan distance formula in the Laplace mapping algorithm to construct an undirected weighted graph, and use G(V,E) to represent the constructed graph, where V represents each vertex in the graph and E represents the edge between vertices, so as to reduce the impact of the subsequent construction of the Laplace matrix caused by the large difference in blueberry spectral distribution, long tail and multi-peak hyperspectral data, and reduce the complexity of the algorithm. The Manhattan distance The expression is as follows:

[0022]

[0023] Among them, n represents the dimension, x ik and x jk are the coordinate values ​​of two n-dimensional vectors in the kth dimension;

[0024] Step 4.2: Introduce a polynomial kernel function into the Laplace mapping algorithm to measure the similarity between each sample point, and use this to calculate the weight of each edge in the graph For each sample point, find the point that is most similar to it, then set the weight between the sample point and its most similar point to a non-zero value, and set the weights between other non-adjacent points to 0. Finally, the weight matrix W is obtained to fit the nonlinear data after FOD processing, and the intrinsic structure of the data is retained to the greatest extent. The polynomial kernel function The expression is as follows:

[0025]

[0026] Among them, x and y represent two input vectors; Indicates the dot product operation of two vectors; c is a constant term, which indicates the offset of the control function; d is the degree of the polynomial, which determines the complexity of the high-dimensional space mapped to;

[0027] The diagonal matrix D is calculated using the weight values, and the calculation formula for its diagonal elements is as follows:

[0028]

[0029] in, represents diagonal elements;

[0030] Step 4.3: Utilization Construct the Laplace matrix L and solve the generalized eigenvalue decomposition problem ,in Denote the generalized eigenvalue to obtain the eigenvector That is, the set of blueberry characteristic bands required. In order to ensure the uniqueness of the final characteristic data, combined with the constraint conditions The generalized eigenvalue decomposition problem is transformed into the −1 L performs eigenvalue decomposition, where T is the transposed matrix representation, and finally takes the eigenvectors corresponding to the d smallest eigenvalues ​​except 0 to obtain , in order to complete the extraction of blueberry characteristic spectral bands.

[0031] Specifically, Step 4.1 and Step 4.2 improve the Laplace mapping algorithm by introducing Manhattan distance and polynomial kernel function to obtain the improved Laplace mapping algorithm ILE, and Step 4.3 is performed based on ILE to extract characteristic bands in blueberry sample data.

[0032] The convolutional neural network (CNN) structure in Step 5 is composed of 1x10 and 1x3 convolution kernels and two 1x2 maximum pooling layers.

[0033] The beneficial effect of the present invention is: a blueberry sugar content prediction model MSC+FOD+ILE+CNN in complex scenarios is proposed. Compared with the traditional preprocessing, characteristic wavelength selection algorithm and machine learning model establishment, the present invention uses a combination of multivariate scattering correction (MSC) and fractional derivative (FOD) to preprocess blueberry data of different varieties and different maturity in complex scenarios, capturing complex patterns in spectral curves to better describe nonlinear changes. The Laplace mapping (ILE) is improved to extract wavelength features more quickly and accurately while adapting to FOD, and finally the advantages of the deep learning network model are used to customize the CNN network to process the complex nonlinear relationship between hyperspectral data. The ILE proposed in the present invention not only retains the original advantages of the LE algorithm, but also can better adapt to multi-peak nonlinear complex data after FOD processing, reduce the amount of calculation while improving the robustness of the model, better extract the characteristic wavelengths in the blueberry hyperspectrum, and effectively improve the existing accuracy of the sugar content prediction model based on the blueberry hyperspectrum in complex scenarios. The proposed model can quickly and accurately predict the sugar content of blueberries of different varieties and maturity levels, providing strong technical support for related research on blueberry sugar content prediction and also providing ideas for the lightweight deployment of blueberry sugar content prediction models. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Schematic diagram of the process of the method in the embodiment of the present invention;

[0035] Figure 2 A schematic diagram of the sampling environment of the research area in an embodiment of the present invention;

[0036] Figure 3 : is the high spectral reflectance curve of MSC+FOD treatment at different orders in the embodiment of the present invention;

[0037] Figure 4 It is a structural diagram of a CNN model in an embodiment of the present invention;

[0038] Figure 5 It is a scatter plot of the predicted value and actual value of the sugar content of each blueberry in the training set and the test set of the best prediction model under the three modeling methods in the embodiment of the present invention. DETAILED DESCRIPTION

[0039] The content of the present invention is further described below in conjunction with the accompanying drawings and specific implementation modes, but this does not limit the scope of the present invention.

[0040] Example 1: Figure 1 As shown, a blueberry sugar content hyperspectral inversion method based on improved Laplace mapping includes the following steps:

[0041] Step 1: Collect blueberry samples of different varieties and maturity in complex scenes, and measure the hyperspectral data and sugar content of the blueberry samples.

[0042] Specifically, in this embodiment, the blueberry samples were collected from a blueberry plantation, where three types of blueberries (L25, L11, and F6) were planted. Among them, L25 has larger fruits, high sugar content, and extremely sweet fruits; L11 has uniform fruit size, medium to high sugar content, and moderate sweetness; F6 has larger and flattened fruits, tastes sweet and sour, and has a lower sugar content than the other two. Figure 2 shown.

[0043] In this embodiment, three blueberry varieties (F6, L11, L25) were collected, and each variety was collected at three degrees of maturity (mature, semi-mature, and immature). During the collection process, the blueberry samples of the same variety were ensured to be uniform in size without any damage or defects. Among them, there were 100 L25 varieties, 36 mature samples, 32 semi-mature samples, and 32 immature samples; 88 L11 varieties, 24 mature samples, 32 semi-mature samples, and 32 immature samples; 88 F6 varieties, 24 mature samples, 32 semi-mature samples, and 32 immature samples, and finally a total of 276 blueberry samples were obtained. After that, the picked blueberry samples were boxed and stored, and the collection time, location and other information were marked, and the blueberry hyperspectral and sugar content were measured. The GaiaField-V10 model hyperspectral imager of Jiangsu Shuangli Hepu was used to collect hyperspectral data, and the band range collected by the spectrometer was 400~1000nm. The instrument used for collecting sugar content is the Atuo fruit sugar meter, model PAL-1, with a range of 0 ~ 53%, a resolution of 0.1%, and an accuracy fluctuation range of 0.2%.

[0044] Step 2: Use statistical algorithms to divide the hyperspectral data and sugar content of blueberry samples, and divide the blueberry samples into two parts: training set and test set in proportion.

[0045] Specifically, the KS algorithm was used to divide the hyperspectral data of 276 blueberry samples into two parts, a training set and a test set, in a ratio of 4:1. The training set consisted of 221 samples and the validation set consisted of 55 samples. The sugar content statistics and coefficient of variation of the blueberry samples after division are shown in Table 1.

[0046] Table 1 Statistical description of sugar content in blueberry samples

[0047] Step 3: Preprocess the divided blueberry sample data. The preprocessing uses the MSC+FOD combination to perform spectral preprocessing on the blueberry hyperspectral data of the training set and the test set respectively to eliminate noise and highlight the spectral features.

[0048] Specifically, the divided data sets are preprocessed. First, SG, MSC, SG+FOD and MSC+FOD are used to preprocess the training set and validation set data respectively. The formula of fractional derivative (FOD) is shown below. The result of blueberry hyperspectral data after MSC+FOD preprocessing is shown in the figure below. Figure 3 shown.

[0049]

[0050] Where f(x) is the spectral signal, which represents the reflectivity at wavelength x; D represents the reflectivity f(x). Secondary derivatives; is the order, here it is the fractional order in the range of 0 to 2 with an interval of 0.05; h is the derivative step size, the sampling interval of the spectrometer is 1nm, here it is set to 1; b and a are the upper and lower limits of the derivative, respectively, and their values ​​increase dynamically with the increase of the number of bands; is the Gamma function, which is used to generalize factorials to non-integer orders. express The Gamma function value of order is similar. express The Gamma function values ​​of the order are used to scale and normalize the weights; m is a non-negative integer, representing the number of steps between the current wavelength x and the historical data point x − mh; represents a series of historical data points near x, is the spectral reflectance of the function at the historical data point x − mh.

[0051] By calculating the fractional derivative at point x, the local variation characteristics of the spectral signal are captured, so as to flexibly mine the weak features in the blueberry hyperspectrum.

[0052] Under the same characteristic wavelength selection algorithm and modeling method, taking LE as the feature selection algorithm and CNN modeling method as an example, the final model indicators after the experiment are shown in Table 2.

[0053] Table 2 Comparison of model indicators of different preprocessing methods under LE + CNN mode

[0054] From the data in Table 2, it can be seen that the MSC + FOD preprocessing method has a good effect under the CNN modeling method, and it also shows that FOD has great advantages in preprocessing blueberry hyperspectral data in complex scenes. Compared with the use of SG and MSC alone, the combination of MSC + FOD method fully utilizes the advantages of the two preprocessing methods, while correcting the scattering effect and enhancing the spectral characteristics, better adapting to the complex hyperspectral data of blueberries, and providing a good foundation for the subsequent extraction of characteristic bands.

[0055] Step 4: Use the improved Laplace mapping algorithm ILE to extract characteristic bands from the blueberry sample data after MSC+FOD combination preprocessing.

[0056] Specifically, Laplace mapping (LE) and improved Laplace mapping (ILE) are used to extract feature bands from the preprocessed training set and test set. After FOD preprocessing, the blueberry hyperspectral data will have nonlinear characteristics such as uneven wavelength distribution, long tail and multi-peaks. In view of the inadaptability of LE in processing such data, the algorithm is improved by introducing the Manhattan distance formula and polynomial kernel function, while retaining the original advantages of the LE algorithm and making it more adaptable to the complex blueberry hyperspectral data after FOD preprocessing. Subsequently, a comparative experiment was carried out to verify whether the improvement of ILE is effective. Under the same modeling method, CNN modeling is still taken as an example. The preprocessing selects MSC + FOD, which performs best under CNN modeling. The final experimental results are shown in Table 3.

[0057] Table 3 Comparison of model indicators of LE and ILE under MSC + FOD + CNN mode

[0058] By comparing the data in Table 3, it can be seen that under the same combination of preprocessing and modeling methods, the performance of the final MSC+FOD+ILE+CNN model is better than that of the unimproved LE model, indicating that after the improvement, ILE successfully adapts to the data processed by FOD, and can more accurately and effectively extract the characteristic wavelengths of the blueberry hyperspectrum, thereby improving the accuracy of subsequent modeling.

[0059] Step 5: Use the convolutional neural network (CNN) in deep learning to build a blueberry sugar content prediction model in complex scenarios. Use the training set after preprocessing and feature extraction to train the model, and then use the test set to calculate the regression index R of the trained blueberry sugar content prediction model. 2 , RMSE and RPD, to validate the prediction model.

[0060] Specifically, a blueberry sugar content prediction model was established using a training set that had been preprocessed and feature bands extracted. The blueberry sugar content prediction model was constructed using the traditional machine learning algorithms of random forest (RF) and partial least squares regression (PLSR) and the convolutional neural network (CNN) algorithm in deep learning for comparative experiments. RF and PLSR were implemented using the sklearn interface in the Python 3.11 library. The CNN network consisted of two convolution kernels of 1x10 and 1x3 and two 1x2 maximum pooling layers for feature extraction. The network structure is shown in the figure below. Figure 4 As shown. The convolutional neural network model is built using the 2.16.1 version of the Tensorflow deep learning framework and Python 3.11 language in the VSCode software. For the convolutional neural network model, it is mainly composed of an input layer, a convolution layer, a pooling layer, a fully connected layer, and an output layer. The specific settings are: the output layer activation function is Linear, the activation functions of the remaining network layers are ReLu, the optimizer is Adam, the loss function of the model is RMSE, the inactivation rate of the Dropout layer is 30%, the model is uniformly trained for 350 Epochs, and each convolution layer uses a one-dimensional convolution layer. For the entire modeling process of the convolutional neural network model, taking the MSC+FOD+ILE+CNN model proposed in the present invention as an example, it is divided into the following steps:

[0061] First, the preprocessed training set data with a dimension of (221, 30) is input into the input layer, where 221 represents 221 training samples and 30 represents 30 spectral features. These features are used to establish a model to predict sugar content.

[0062] Then, the training set passes through the convolutional layer, pooling layer and other network layers of the convolutional neural network prediction model, and the blueberry sugar content prediction value is output at the output layer. During the entire training process, the model is continuously adjusted through each round of iterative training, updating the network weights to minimize the value of the loss function and find the optimal model;

[0063] Finally, the optimal model with the smallest loss function during the iterative training process under this method is saved.

[0064] Use the test set to verify and evaluate the trained prediction model and calculate the R of each prediction model on the test set. 2 , RMSE and RPD, to determine the prediction accuracy and generalization performance of each blueberry sugar content prediction model. 2 The closer the value is to 1, the better the model prediction effect is. RMSE is used to describe the difference between the predicted value and the measured value. The smaller the value, the better. RPD is used to evaluate and describe the prediction ability of the prediction model. The larger the RPD value, the better the model prediction ability. Generally speaking, when RPD < 1.0, it indicates that the model is not suitable for prediction tasks. When 2.0 < RPD < 2.5, its prediction ability is good. When RPD > 2.5, it indicates that the model has excellent prediction ability. The evaluation results of the optimal prediction models under the last three modeling methods on the training set and test set are shown in Table 4 below:

[0065] Table 4 The optimal model results under three different modeling methods

[0066] The results in Table 4 show that by comparing the three modeling methods, it can be seen that the MSC+FOD+ILE+CNN prediction model proposed in the present invention has the best effect. 2 The RPD is 2.6694, 0.8597, RMSE is 0.8552, and 0.8694, indicating that the CNN model has excellent prediction performance and has significantly improved the modeling effect compared with the other two traditional machine learning network models RF and PLSR. 2 The R of the optimal model is 0.0468, the RMSE is reduced by 0.1322, and the RPD is increased by 0.3574. 2 The results show that the blueberry sugar content of each optimal prediction model is improved by 0.012, the RMSE is reduced by 0.0362, and the RPD is improved by 0.1085. Figure 5As shown. The predicted points of the test set on the scatter plot of the CNN model are closer to the fitting line than those of RF and PLSR, and the predicted values ​​of the CNN model on the test set fit the measured values ​​best. This is mainly due to the fact that FOD + MSC combined with ILE can more effectively process and extract the characteristic wavelengths of blueberry data, and the powerful learning ability of deep learning, which together improves the accuracy and performance of modeling. Based on the above results, the effectiveness of the MSC+FOD+ILE+CNN prediction model is demonstrated, and its prediction accuracy meets the requirements of actual usage and is feasible for application.

[0067] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A blueberry sugar content hyperspectral inversion method based on improved Laplace mapping, characterized in that: The specific steps include: Step 1: Collect blueberry samples of different varieties and maturity in complex scenes, and measure the hyperspectral data and sugar content of the blueberry samples; Step 2: Use statistical algorithms to divide the hyperspectral data and sugar content of blueberry samples, and divide the blueberry samples into two parts: training set and test set according to proportion; Step 3: Preprocessing the divided blueberry sample data, wherein the preprocessing uses the MSC+FOD combination to perform spectral preprocessing on the blueberry hyperspectral data of the training set and the test set respectively to eliminate noise and highlight spectral features; Step 4: Use the improved Laplace mapping algorithm ILE to extract characteristic bands from the blueberry sample data after MSC+FOD combination preprocessing; Step 5: Use the convolutional neural network (CNN) in deep learning to build a blueberry sugar content prediction model in complex scenarios. Use the training set after preprocessing and feature extraction to train the model, and then use the test set to calculate the regression index R of the trained blueberry sugar content prediction model. 2 , RMSE and RPD, to validate the prediction model.

2. The method for inverting blueberry sugar content hyperspectral based on improved Laplace mapping according to claim 1, characterized in that: The Step 3 is specifically as follows: Step 3.1: First, the multivariate scatter correction (MSC) method was used to preliminarily preprocess the blueberry data, and a linear fit was performed between each sample spectrum and the average of all sample spectra; Step 3.2: Introduce fractional derivative FOD to process the blueberry data processed by MSC again. The measured blueberry sugar content data and hyperspectral data are used as input data respectively. According to the preset fractional order interval, the fractional derivative of each band of each blueberry spectral curve is calculated in 0~2 orders in turn to obtain the processed spectral curve, where FOD is defined as Grunwald-Letnikov function, and its formula is as follows: ; Where f(x) is the spectral signal, which represents the reflectivity at wavelength x; D represents the reflectivity f(x). Secondary derivatives; is the order; h is the derivative step size; b and a are the upper and lower limits of the derivative, respectively, and their values ​​increase dynamically with the increase of the number of bands; is the Gamma function, which is used to generalize factorials to non-integer orders. express The value of the Gamma function of order, express The Gamma function value of the order is used to scale and normalize the weights; m is a non-negative integer representing the number of steps between the current wavelength x and the historical data point x − mh; represents a series of historical data points around x, is the spectral reflectance of the function at the historical data point x − mh; Step 3.3: By calculating the fractional derivative at point x, the local variation characteristics of the spectral signal are captured.

3. The method for inverting blueberry sugar content hyperspectral based on improved Laplace mapping according to claim 1, characterized in that: The Step 4 is specifically as follows: Step 4.1: Introduce the Manhattan distance formula in the Laplace mapping algorithm to construct an undirected weighted graph. Use G(V,E) to represent the constructed graph, where V represents each vertex in the graph, E represents the edge between vertices, and Manhattan distance The expression is as follows: ; Among them, n represents the dimension, x ik and x jk are the coordinate values ​​of two n-dimensional vectors in the kth dimension; Step 4.2: Introduce a polynomial kernel function into the Laplace mapping algorithm to measure the similarity between each sample point, and use this to calculate the weight of each edge in the graph For each sample point, find the point that is most similar to it, then set the weight between the sample point and its most similar point to a non-zero value, and set the weights between other non-adjacent points to 0. Finally, the weight matrix W is obtained to fit the nonlinear data after FOD processing. The polynomial kernel function The expression is as follows: ; Among them, x and y represent two input vectors; Indicates the dot product operation of two vectors; c is a constant term, which indicates the offset of the control function; d is the degree of the polynomial, which determines the complexity of the high-dimensional space mapped to; The diagonal matrix D is calculated using the weight values, and the calculation formula for its diagonal elements is as follows: ; in, represents diagonal elements; Step 4.3: Utilization Construct the Laplace matrix L and solve the generalized eigenvalue decomposition problem ,in Represents the generalized eigenvalue to obtain the eigenvector f, which is the set of blueberry characteristic bands required, combined with the constraints The generalized eigenvalue decomposition problem is transformed into the −1 L is decomposed by eigenvalues, where T is the transposed matrix representation, and finally the eigenvectors corresponding to the d smallest eigenvalues ​​except 0 are taken to obtain f, thereby completing the extraction of the characteristic spectral band of blueberries.

4. The method for inverting blueberry sugar content hyperspectral based on improved Laplace mapping according to claim 1, characterized in that: The convolutional neural network (CNN) structure in Step 5 is composed of 1x10 and 1x3 convolution kernels and two 1x2 maximum pooling layers.

Citation Information

Patent Citations

  • Hyperspectrum-based nondestructive detection method for sugar acidity of waxberry fruits

    CN111795932A

  • Crispy pear candy precision detection method and device, cloud equipment and computer device

    CN117805024A

  • Nondestructive testing method for internal quality and maturity of strawberries based on MPCD-Mmba model

    CN119510322A

  • Method and system for hyperspectral inversion of phosphorus content of rubber tree leaves

    US20210056424A1

Cited By

  • Strawberry sugar degree efficient nondestructive testing method based on intelligent algorithm

    CN120703080A

  • Method and system for detecting maturity of red bayberry based on hyperspectrum

    CN121236611A

  • Blueberry sugar degree prediction method based on feature coding and double-branch network

    CN121834250A

  • Blueberry SSC content prediction method based on Caputo fractional derivative hyperspectral pretreatment

    CN121917497A

  • Method for predicting SSC content of blueberry based on hyperspectral pretreatment of caputo fractional derivative

    CN121917497B