A Hyperspectral Inversion Method for Blueberry Brix Based on Improved Laplacian Mapping

By combining multivariate scattering correction and preprocessing methods of fractional derivatives, the improved Laplace mapping algorithm and convolutional neural network solve the problem of poor generalization of blueberry hyperspectral data processing in the prior art, and achieve more efficient and accurate blueberry sugar level prediction.

CN120012044BActive Publication Date: 2025-06-20KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510500484.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-20
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing hyperspectral technology has poor generalization when predicting blueberry sugar content. Traditional preprocessing methods and feature selection algorithms are difficult to effectively process blueberry hyperspectral data in complex scenarios, resulting in a decline in model performance.

Method used

Multivariate scattering correction (MSC) and fractional derivative (FOD) are used for data preprocessing, feature bands are extracted using the improved Laplace mapping algorithm (ILE), and a blueberry sugar level prediction model in complex scenarios is constructed through convolutional neural network (CNN).

Benefits of technology

It improves the accuracy and speed of blueberry sugar content prediction, reduces costs, and applies to different varieties and maturity predictions in complex scenarios, improving the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012044B_ABST
    Figure CN120012044B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for hyperspectral inversion of blueberry sugar content based on improved Laplacian mapping, belonging to the technical field of blueberry hyperspectral acquisition and analysis. First, blueberry samples of different varieties and different maturities are collected in complex scenes, and the hyperspectral data of the blueberry samples and the corresponding sugar content of the blueberries are measured. Then, the blueberry sample data are divided into a training set and a test set according to a ratio using a statistical algorithm, and data preprocessing is performed using the MSC+FOD combination. Next, the ILE proposed by the present invention is used to extract characteristic bands from the preprocessed spectral data, a blueberry sugar content prediction model is established using CNN, and the model is trained using the training set. Finally, the test set is used to comprehensively evaluate the model in combination with the regression indexes R<supgt;2< / supgt>, RMSE and RPD. The present invention can quickly and accurately predict the sugar content of blueberries of different varieties and different maturities in complex scenes, providing strong technical support for the prediction of blueberry sugar content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for hyperspectral inversion of blueberry sugar content based on improved Laplacian mapping, belonging to the technical field of blueberry hyperspectral acquisition and analysis. Background Technique

[0002] Blueberries, as a fruit with extremely high nutritional value, are deeply loved by people. Sugar content, as one of the criteria for evaluating blueberry quality, has different market demands at different sugar levels. By measuring the sugar content of blueberries, agricultural producers can determine the optimal harvesting time to ensure the taste and quality of blueberries and meet market demands to the greatest extent. Therefore, quickly and accurately predicting the sugar content in blueberries is of great significance for promoting the development of agricultural production. In traditional methods, determining the sugar content of blueberries involves on-site collection of blueberries and a large number of chemical experimental analyses. This process is not only time-consuming and laborious but also causes irreparable damage to blueberries. In recent years, with the popularization of spectral technology in the field of non-destructive testing, it has brought great convenience and progress to agriculture, improving the detection efficiency and accuracy without damaging the appearance and internal structure of fruits, and at the same time reducing various costs required for detection. Currently, hyperspectral technology has been widely applied to the rapid non-destructive testing research of blueberry and other fruit and vegetable qualities, providing a powerful analysis means for the quality evaluation and component analysis of fruits and vegetables.

[0003] In the existing research on predicting blueberry sugar content using hyperspectral remote sensing technology, although there are many spectral processing methods and modeling methods, they are all relatively traditional and have poor overall generalization ability. Especially when facing the hyperspectral data of different varieties and maturities of blueberries collected in complex scenarios, traditional preprocessing methods such as SG smoothing, standard normal variate transformation (SNV), and first derivative are insufficient to handle complex spectral data. Fractional order derivative (FOD) can not only eliminate baseline drift and suppress noise but also amplify the subtle information hidden in a large amount of spectral signals, better adapt to the non-linear spectral changes in complex scenarios, and thus capture the global and local information in the spectrum at multiple scales. After preprocessing the hyperspectral data, similarly, the ability of traditional feature selection algorithms to accurately and quickly extract key information is still limited. For example, the competitive adaptive reweighted sampling algorithm (CARS) is prone to generating redundant bands and needs to be used multiple times; the iterative retained information variable algorithm (IRIV) has a long calculation time, affecting the modeling efficiency, etc. Laplacian eigenmaps (LE) can analyze the local similarity between bands by constructing a Laplacian matrix, adaptively assign band weights, and quickly extract characteristic wavelengths, and the overall generalization ability of the algorithm is stronger. However, after the blueberry spectrum is processed by FOD, due to the amplification of weak features, there will be non-linear performance trends such as uneven spectral curve distribution, long tail, and multiple peaks. At this time, LE cannot better capture the complex relationships between data, and the risk of falling into local optimum and ignoring global information is greatly increased, ultimately affecting the overall performance of the model.

[0004] Therefore, the present invention proposes a hyperspectral inversion method for blueberry sugar content based on improved Laplacian mapping in view of the above problems. On the basis of the deficiencies in existing research, it is proposed to preprocess data using a combination of multiplicative scatter correction (MSC) and fractional-order derivative (FOD), and use the improved Laplacian mapping algorithm (ILE) to extract the characteristic bands of blueberry hyperspectrum while better adapting to FOD. Finally, a lightweight sugar content prediction model for blueberries of different varieties and maturities in complex scenarios is constructed using a CNN network and its applicability is verified, providing a reference for further research on related topics such as hyperspectral inversion of blueberry sugar content. Summary of the Invention

[0005] The object of the present invention is to provide a hyperspectral inversion method for blueberry sugar content based on improved Laplacian mapping in view of the deficiencies of the prior art. This method has higher prediction accuracy, faster prediction speed and lower cost for blueberry prediction in complex scenarios, and has good practical application value and development potential.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] A hyperspectral inversion method for blueberry sugar content based on improved Laplacian mapping. First, blueberry samples of different varieties and maturities in complex scenarios are collected, and the hyperspectral data of the blueberry samples and the sugar content of the corresponding blueberries are measured. Then, a suitable statistical algorithm is selected to divide the blueberry sample data into a training set and a test set according to a certain ratio. The MSC+FOD combination is used for data preprocessing, and the ILE proposed by the present invention is used to extract the characteristic bands from the preprocessed spectral data. A blueberry sugar content prediction model is established using a CNN and the model is trained using the training set. Finally, the test set is used in combination with the regression metrics R 2 , RMSE and RPD to comprehensively evaluate the model.

[0008] Specifically, it includes the following steps:

[0009] Step1: Collect blueberry samples of different varieties and maturities in complex scenarios, and measure the hyperspectral data and sugar content of the blueberry samples;

[0010] Step2: Use a statistical algorithm to divide the hyperspectral data and sugar content of the blueberry samples, and divide the blueberry samples into a training set and a test set according to a ratio;

[0011] Step3: Preprocess the divided blueberry sample data. The preprocessing uses the MSC+FOD combination to perform spectral preprocessing on the hyperspectral data of the blueberries in the training set and the test set respectively to eliminate noise and highlight spectral features;

[0012] Step 4: Extract feature bands from the blueberry sample data preprocessed by MSC + FOD using the improved Laplace mapping algorithm ILE;

[0013] Step 5: Use the convolutional neural network CNN in deep learning to construct a blueberry sugar content prediction model in complex scenarios. Use the training set after preprocessing and feature extraction to train the model, and then use the test set to calculate the regression metrics R 2 , RMSE, and RPD for the trained blueberry sugar content prediction model to verify the prediction model.

[0014] The specific content of Step 3 is as follows:

[0015] Step 3.1: First, use the multiplicative scatter correction (MSC) method to preliminarily preprocess the blueberry data. Linearly fit each sample spectrum with the average of all sample spectra to eliminate the scattering effect caused by external factors and enhance the correlation between the spectrum and the sugar content;

[0016] Step 3.2: Since the blueberry hyperspectral data collected in complex scenarios is relatively complex, introduce the fractional order derivative (FOD) to process the blueberry data processed by MSC again. Take the measured blueberry sugar content data (m×1) and the hyperspectral data (m×n) as input data respectively. According to the preset fractional order interval, calculate the fractional order derivative of each band of each blueberry spectral curve in the range of 0 - 2 in turn to obtain the processed spectral curve, where FOD is defined as the Grunwald - Letnikov function, and its formula is as follows:

[0017]

[0018] where f(x) is the spectral signal, representing the reflectance at wavelength x; D represents taking the th derivative of the reflectance f(x); is the order; h is the derivative step size; b and a are the upper and lower limits of the derivative respectively, and their values increase dynamically with the increase of the number of bands; is the Gamma function, which is used to generalize the factorial to non - integer orders, represents the value of the Gamma function of order represents the value of the Gamma function of order, which is used for scaling and normalizing weights; m is a non - negative integer, representing the number of steps between the current wavelength x and the historical data point x - mh; represents a series of historical data points near x, is the spectral reflectance of the function at the historical data point x - mh;

[0019] Step 3.3: Calculate the fractional derivative at point x to capture the local change characteristics of the spectral signal, aiming to flexibly mine the weak features in blueberry hyperspectral data.

[0020] The specific content of Step 4 is as follows:

[0021] Step 4.1: Introduce the Manhattan distance formula into the Laplacian mapping algorithm to construct an undirected weighted graph, denoted as G(V, E), where V represents each vertex in the graph and E represents the edges between vertices, so as to reduce the impact on the subsequent construction of the Laplacian matrix due to the large difference in blueberry spectral distribution, the existence of long tails and multi-peaks in hyperspectral data, and at the same time reduce the algorithm complexity. The expression of the Manhattan distance is as follows: The expression is as follows:

[0022]

[0023] where n represents the dimension, x ik and x jk are the coordinate values of two n-dimensional vectors on the k-th dimension respectively;

[0024] Step 4.2: Introduce a polynomial kernel function into the Laplacian mapping algorithm to measure the similarity between each sample point, and calculate the weight of each edge in the graph. For each sample point, find the point most similar to it, and then set the weight between the sample point and its most similar point to a non-zero value, and set the weights between other non-neighboring points to 0. Finally, obtain the weight matrix W to fit the non-linear data after FOD processing and retain the internal structure of the data to the greatest extent. The expression of the polynomial kernel function is as follows: The expression is as follows: The expression is as follows:

[0025]

[0026] where x and y represent the two input vectors; represents the dot product operation of the two vectors; c is a constant term representing the control function offset; d is the degree of the polynomial, determining the complexity of the mapped high-dimensional space;

[0027] Calculate the diagonal matrix D using the weight values, and the formula for its diagonal elements is as follows:

[0028]

[0029] where, represents the diagonal element;

[0030] Step 4.3: Use to construct the Laplacian matrix L and solve the generalized eigenvalue decomposition problem , where represents the generalized eigenvalue to obtain the eigenvector which is the set of the desired blueberry characteristic bands. To ensure the uniqueness of the final characteristic data, combined with the constraint conditions transform the generalized eigenvalue decomposition problem into the eigenvalue decomposition of D −1 L, where T represents the transpose matrix. Finally, take the eigenvectors corresponding to the d smallest eigenvalues except 0 to obtain , thereby completing the extraction of the blueberry characteristic spectral bands.

[0031] Specifically, Step 4.1 and Step 4.2 improve the Laplacian mapping algorithm by introducing the Manhattan distance and the polynomial kernel function to obtain the improved Laplacian mapping algorithm ILE, and based on ILE, Step 4.3 is carried out to extract the characteristic bands in the blueberry sample data.

[0032] The CNN structure in the said Step5 consists of convolutional kernels of 1x10 and 1x3 and two max-pooling layers of 1x2.

[0033] The beneficial effects of the present invention are as follows: A blueberry sugar content prediction model MSC+FOD+ILE+CNN in complex scenarios is proposed. Compared with establishing a model by using traditional preprocessing, characteristic wavelength selection algorithms and machine learning, the present invention uses the combination of multiplicative scatter correction (MSC) and fractional order derivative (FOD) to preprocess blueberry data of different varieties and different maturities in complex scenarios, capturing complex patterns in the spectral curve to better describe the nonlinear changes. Improve the Laplacian mapping (ILE) to more quickly and accurately extract wavelength features while adapting to FOD. Finally, utilize the advantages of the deep learning network model and customize the CNN network to process the complex nonlinear relationship between hyperspectral data. The ILE proposed in the present invention not only retains the original advantages of the LE algorithm, but also can better adapt to the multi-peak nonlinear complex data after FOD processing, reducing the calculation amount while improving the robustness of the model, better extracting the characteristic wavelengths in blueberry hyperspectrum, and effectively improving the existing accuracy of the sugar content prediction model based on blueberry hyperspectrum in complex scenarios. The proposed model can quickly and accurately predict the sugar content of blueberries of different varieties and different maturities, providing strong technical support for the related research on blueberry sugar content prediction, and also providing ideas for the lightweight deployment of the blueberry sugar content prediction model. Description of the Drawings

[0034] Figure 1 is a schematic flow chart of the method in the embodiment of the present invention;

[0035] Figure 2 is a schematic diagram of the sampling environment of the research area in the embodiment of the present invention;

[0036] Figure 3 The hyperspectral reflectance curves of MSC+FOD processing at different orders in the embodiments of the present invention;

[0037] Figure 4 The structural diagram of the CNN model in the embodiments of the present invention;

[0038] Figure 5 The scatter plot of the predicted values and actual values of the blueberry sugar content in the training set and test set of the best prediction models under three modeling methods in the embodiments of the present invention. Specific implementation manners

[0039] The content of the present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners, but this does not limit the scope of the present invention.

[0040] Example 1: As Figure 1 shown, a hyperspectral inversion method for blueberry sugar content based on improved Laplacian mapping includes the following steps:

[0041] Step1: Collect blueberry samples of different varieties and maturities in a complex scene, and measure the hyperspectral data and sugar content of the blueberry samples.

[0042] Specifically, in this embodiment, the blueberry samples are collected from a blueberry plantation where three types of blueberries (L25, L11, F6) are planted. Among them, L25 has larger fruits, high sugar content, and extremely sweet fruits; L11 has uniform fruit sizes, medium to high sugar levels, and moderate sweetness; F6 has larger and flat fruits, a sour and sweet taste, and a relatively lower sugar content compared to the other two. The environment of the blueberry sampling area is as Figure 2 shown.

[0043] In this embodiment, a total of three blueberry varieties (F6, L11, L25) are collected, and each variety is taken at three levels of maturity (ripe, semi-ripe, unripe). During the collection process, it is ensured that the blueberry samples of the same variety have uniform sizes and no damage or defects. Among them, there are 100 samples of the L25 variety, 36 ripe samples, 32 semi-ripe samples, and 32 unripe samples; 88 samples of the L11 variety, 24 ripe samples, 32 semi-ripe samples, and 32 unripe samples; 88 samples of the F6 variety, 24 ripe samples, 32 semi-ripe samples, and 32 unripe samples, and finally a total of 276 blueberry samples are obtained. After that, the picked blueberry samples are packed and stored, marked with information such as the collection time and location, and the blueberry hyperspectrum and sugar content are measured. The hyperspectral data is collected using a GaiaField-V10 model hyperspectral imager from Jiangsu Shuangli Hupu, and the spectral range collected by the spectrometer is 400~1000nm. The instrument used for sugar content collection is an Atago fruit sugar refractometer, model PAL-1, with a measuring range of 0~53%, a resolution of 0.1%, and an accuracy fluctuation range of 0.2%.

[0044] Step 2: Use a statistical algorithm to divide the hyperspectral data and sugar content of blueberry samples, and divide the blueberry samples into two parts: a training set and a test set according to a certain proportion.

[0045] Specifically, the KS algorithm is used to divide the hyperspectral data of 276 collected blueberry samples, and they are divided into a training set and a test set according to a ratio of 4:1. Among them, the training set has 221 samples and the validation set has 55 samples. After division, the sugar content statistical results and coefficient of variation of the blueberry samples are shown in Table 1.

[0046] Table 1 Statistical description of sugar content in blueberry samples

[0047]

[0048] Step 3: Preprocess the divided blueberry sample data. The preprocessing uses the MSC+FOD combination to perform spectral preprocessing on the hyperspectral data of blueberries in the training set and the test set respectively to eliminate noise and highlight spectral features.

[0049] Specifically, preprocess the divided data set. First, use SG, MSC, SG+FOD, and MSC+FOD respectively to preprocess the data of the training set and the validation set. The formula for the fractional order derivative (FOD) is shown as follows. The result of preprocessing the blueberry hyperspectral data with MSC+FOD is as Figure 3 shown.

[0050]

[0051] Among them, f(x) is the spectral signal, representing the reflectance at wavelength x; D represents taking the th derivative of the reflectance f(x); is the order, which is in the range of 0 to 2 orders here, with each fractional order at an interval of 0.05; h is the derivative step length, and the sampling interval of the spectrometer is 1 nm, which is set to 1 here; b and a are the upper and lower limits of the derivative respectively, and their values increase dynamically with the increase in the number of bands; is the Gamma function, which is used to generalize the factorial to non-integer orders, represents the value of the Gamma function of order, and similarly represents the value of the Gamma function of order, both of which are used for scaling and normalizing weights; m is a non-negative integer, representing the number of steps between the current wavelength x and the historical data point x - mh; represents a series of historical data points near x, is the spectral reflectance of the function at the historical data point x - mh.

[0052] By calculating the fractional derivative at point x, the local change characteristics of the spectral signal are captured, aiming to flexibly mine the weak characteristics in blueberry hyperspectral data.

[0053] Under the same feature wavelength selection algorithm and modeling method, taking LE as the feature selection algorithm and CNN modeling method as an example, the final model metrics after the experiment are shown in Table 2.

[0054] Table 2 Model metrics for different preprocessing methods under the LE + CNN method

[0055]

[0056] It can be seen from the data in Table 2 that the MSC + FOD preprocessing method has good results under the CNN modeling method, and it also shows that FOD has great advantages in preprocessing blueberry hyperspectral data in complex scenarios. Compared with using SG and MSC alone, the combination of the MSC + FOD method gives full play to the advantages of the two preprocessing methods, enhances spectral features while correcting the scattering effect, better adapts to the complex blueberry hyperspectral data, and provides a good foundation for the extraction of subsequent characteristic bands.

[0057] Step4: Use the improved Laplacian mapping algorithm ILE to extract characteristic bands from the blueberry sample data preprocessed by the MSC+FOD combination.

[0058] Specifically, use Laplacian mapping (LE) and the improved Laplacian mapping (ILE) to extract characteristic bands from the preprocessed training set and test set. After FOD preprocessing, the blueberry hyperspectral data will show non-linear characteristics such as uneven wavelength distribution, long tails, and multiple peaks. In view of the inadaptability of LE in processing such data, the algorithm is improved by introducing the Manhattan distance formula and polynomial kernel function, making it more adaptable to the complex blueberry hyperspectral data preprocessed by FOD while retaining the original advantages of the LE algorithm. Subsequently, a comparative experiment is carried out to verify whether the improvement of ILE is effective. Under the same modeling method, still taking the CNN modeling method as an example, the preprocessing selects MSC + FOD, which performs the best under the CNN modeling method. The final experimental results are shown in Table 3.

[0059] Table 3 Model metrics for comparing LE and ILE under the MSC + FOD + CNN method

[0060]

[0061] By comparing the data in Table 3, it can be seen that under the same combination of preprocessing and modeling methods, the performance of the final MSC+FOD+ILE+CNN model is better than that of the unimproved LE model. This shows that after improvement, ILE successfully adapts to the data processed by FOD and can more accurately and effectively extract the characteristic wavelengths of blueberry hyperspectral, improving the subsequent modeling accuracy.

[0062] Step5: Use the convolutional neural network CNN in deep learning to construct a blueberry sugar content prediction model for complex scenarios. Use the training set after preprocessing and feature extraction to train the model, and then use the test set to calculate the regression metrics R 2 , RMSE, and RPD to verify the prediction model.

[0063] Specifically, use the training set after preprocessing and feature band extraction to establish a blueberry sugar content prediction model. Conduct comparative experiments by constructing blueberry sugar content prediction models using the random forest (RF) and partial least squares regression (PLSR) in traditional machine learning algorithms and the convolutional neural network (CNN) algorithm in deep learning respectively. RF and PLSR are implemented by calling the sklearn interface in the Python 3.11 library. The CNN network consists of two convolutional kernels of 1x10 and 1x3 and two max pooling layers of 1x2, which are used to extract features. Its network structure is as Figure 4 shown. The convolutional neural network model is built using the 2.16.1 version of the TensorFlow deep learning framework and the Python 3.11 language in the VSCode software. For the convolutional neural network model, it mainly consists of an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The specific settings are as follows: the activation function of the output layer is Linear, the activation functions of the other network layers are ReLu, the optimizer is Adam, the loss function of the model is RMSE, the inactivation rate of the Dropout layer is 30%, the model is uniformly trained for 350 Epochs, and each convolutional layer uses a one-dimensional convolutional layer. For the entire modeling process of the convolutional neural network model, taking the MSC+FOD+ILE+CNN model proposed by the present invention as an example, it is divided into the following steps:

[0064] First, input the preprocessed training set data with a dimension of (221, 30) into the input layer, where 221 represents 221 training samples and 30 represents 30 spectral features, and these features are used to establish a model to predict the sugar content;

[0065] Then, the training set passes through network layers such as the convolutional layer and pooling layer of the convolutional neural network prediction model, and the predicted value of the blueberry sugar content is output at the output layer. Throughout the training process, the model is continuously adjusted through repeated iterative training in each round, updating the network weights to minimize the value of the loss function and find the optimal model;

[0066] Finally, save the optimal model with the minimum loss function during the iterative training process under this method.

[0067] Use the test set to verify and evaluate the trained prediction model, and calculate the R 2 , RMSE, and RPD of each prediction model on the test set to determine the prediction accuracy and generalization performance of each blueberry sugar content prediction model. Among them, the closer the value of R 2 is to 1, the better the model prediction effect. RMSE is used to describe the difference between the predicted value and the measured value, and the smaller its value, the better. RPD is used to evaluate and describe the prediction ability of the prediction model. The larger the RPD value, the better the model prediction ability. Generally speaking, when RPD < 1.0, it indicates that the model is not suitable for the prediction task. When 2.0 < RPD < 2.5, its prediction ability is better. When RPD > 2.5, it indicates that the prediction ability of the model is excellent. The evaluation results of the respective optimal prediction models under the last three modeling methods on the training set and the test set are shown in Table 4 below:

[0068] Table 4 Optimal model results under three different modeling methods

[0069]

[0070] As shown in the results of Table 4, by comparing the three modeling methods, it can be seen that the MSC+FOD+ILE+CNN prediction model proposed by the present invention has the best effect. For this model, its R 2 on the validation set is 0.8597, RMSE is 0.8552, and RPD is 2.6694, indicating that the prediction performance of the CNN model is excellent, and compared with the other two traditional machine learning network models RF and PLSR, the modeling effect has a significant improvement. On the test set, the R 2 of the CNN model relative to the optimal model of PLSR has increased by 0.0468, RMSE has decreased by 0.1322, and RPD has increased by 0.3574; relative to the optimal model of RF, the R 2 has increased by 0.012, RMSE has decreased by 0.0362, and RPD has increased by 0.1085. Combining with the scatter plot of the measured values and predicted values of the blueberry sugar content of each optimal prediction model, as Figure 5As shown. Among them, the predicted points of the test set on the scatter plot of the CNN model are closer to the fitting line than those of RF and PLSR, and the predicted values and measured values of the CNN model on the test set also fit best. This is mainly due to the fact that the combination of FOD + MSC and ILE can more effectively process and extract the characteristic wavelengths of blueberry data, plus the powerful learning ability of deep learning. The three together improve the accuracy and performance of modeling. According to the above results, the effectiveness of the MSC+FOD+ILE+CNN prediction model is demonstrated, and its prediction accuracy meets the requirements of actual use, with the feasibility of application.

[0071] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. A blueberry sugar content hyperspectral inversion method based on improved Laplace mapping, characterized in that: The specific steps include: Step 1: Collect blueberry samples of different varieties and maturity in complex scenes, and measure the hyperspectral data and sugar content of the blueberry samples; Step 2: Use statistical algorithms to divide the hyperspectral data and sugar content of blueberry samples, and divide the blueberry samples into two parts: training set and test set according to proportion; Step 3: Preprocessing the divided blueberry sample data, wherein the preprocessing uses the MSC+FOD combination to perform spectral preprocessing on the blueberry hyperspectral data of the training set and the test set respectively to eliminate noise and highlight spectral features; Step 4: Use the improved Laplace mapping algorithm ILE to extract characteristic bands from the blueberry sample data after MSC+FOD combination preprocessing; Step 5: Use the convolutional neural network (CNN) in deep learning to build a blueberry sugar content prediction model in complex scenarios. Use the training set after preprocessing and feature extraction to train the model, and then use the test set to calculate the regression index R of the trained blueberry sugar content prediction model. 2 , RMSE and RPD, to validate the prediction model; The Step 4 is specifically as follows: Step 4.1: Introduce the Manhattan distance formula in the Laplace mapping algorithm to construct an undirected weighted graph. Use G(V,E) to represent the constructed graph, where V represents each vertex in the graph, E represents the edge between vertices, and Manhattan distance The expression is as follows: ; Among them, n represents the dimension, x ik and x jk are the coordinate values ​​of two n-dimensional vectors in the kth dimension; Step 4.2: Introduce a polynomial kernel function into the Laplace mapping algorithm to measure the similarity between each sample point, and use this to calculate the weight of each edge in the graph For each sample point, find the point that is most similar to it, then set the weight between the sample point and its most similar point to a non-zero value, and set the weights between other non-adjacent points to 0. Finally, the weight matrix W is obtained to fit the nonlinear data after FOD processing. The polynomial kernel function The expression is as follows: ; Among them, x and y represent two input vectors; Indicates the dot product operation of two vectors; c is a constant term, which indicates the offset of the control function; d is the degree of the polynomial, which determines the complexity of the high-dimensional space mapped to; The diagonal matrix D is calculated using the weight values, and the calculation formula for its diagonal elements is as follows: ; in, represents diagonal elements; Step 4.3: Utilization Construct the Laplace matrix L and solve the generalized eigenvalue decomposition problem ,in Represents the generalized eigenvalue to obtain the eigenvector f, which is the set of blueberry characteristic bands required, combined with the constraints The generalized eigenvalue decomposition problem is transformed into the −1 L is decomposed by eigenvalues, where T is the transposed matrix representation, and finally the eigenvectors corresponding to the d smallest eigenvalues ​​except 0 are taken to obtain f, thereby completing the extraction of the characteristic spectral band of blueberries.

2. The method for inverting blueberry sugar content hyperspectral based on improved Laplace mapping according to claim 1, characterized in that: The Step 3 is specifically as follows: Step 3.1: First, the multivariate scatter correction (MSC) method was used to preliminarily preprocess the blueberry data, and a linear fit was performed between each sample spectrum and the average of all sample spectra; Step 3.2: Introduce fractional derivative FOD to process the blueberry data processed by MSC again. The measured blueberry sugar content data and hyperspectral data are used as input data respectively. According to the preset fractional order interval, the fractional derivative of each band of each blueberry spectral curve is calculated in 0~2 orders in turn to obtain the processed spectral curve, where FOD is defined as Grunwald-Letnikov function, and its formula is as follows: ; Where f(x) is the spectral signal, which represents the reflectivity at wavelength x; D represents the reflectivity f(x). Secondary derivatives; is the order; h is the derivative step size; b and a are the upper and lower limits of the derivative, respectively, and their values ​​increase dynamically with the increase of the number of bands; is the Gamma function, which is used to generalize factorials to non-integer orders. express The value of the Gamma function of order, express The Gamma function value of the order is used to scale and normalize the weights; m is a non-negative integer representing the number of steps between the current wavelength x and the historical data point x − mh; represents a series of historical data points around x, is the spectral reflectance of the function at the historical data point x − mh; Step 3.3: By calculating the fractional derivative at point x, the local variation characteristics of the spectral signal are captured.

3. The method for inverting blueberry sugar content hyperspectral based on improved Laplace mapping according to claim 1, characterized in that: The convolutional neural network (CNN) structure in Step 5 is composed of 1x10 and 1x3 convolution kernels and two 1x2 maximum pooling layers.

Citation Information

Patent Citations

  • Crispy pear candy precision detection method and device, cloud equipment and computer device

    CN117805024A

  • Nondestructive testing method for internal quality and maturity of strawberries based on MPCD-Mmba model

    CN119510322A