A Prediction Method of Blueberry SSC Based on Fractional Derivative Coupled Optimization of Spectral Index
By combining the fractional derivative and two-dimensional optimization spectral index, the complex nonlinear relationship and band information omission in blueberry SSC prediction are solved, and fast, accurate and lossless blueberry SSC prediction is achieved, improving prediction accuracy and adaptability.
Patent Information
- Application Number
- CN202510554761.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art is difficult to quickly and accurately predict blueberry soluble solid content (SSC). Traditional methods are highly destructive, complex in operation and difficult to meet the needs of large-scale prediction. The integer derivative method is not very accurate under low signal-to-noise ratio. The traditional spectral index cannot adapt to different fruit varieties or environments, and there are band information omissions and multicollinearity problems.
The 0-2-order fractional derivative in the form of Grünwald-Letnikov was combined with the two-dimensional optimized spectral index, and the characteristic bands were screened through the Pearson correlation coefficient to construct a backpropagation neural network (BPNN) model for blueberry SSC prediction.
It realizes fast, accurate and lossless blueberry SSC prediction, improves prediction accuracy, reduces costs, adapts to different fruit varieties and environments, and solves the problems of complex nonlinear relationships and band information omissions.
Smart Images

Figure CN120067620B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the SSC of blueberries by optimizing spectral indices based on fractional-order derivative coupling, belonging to the field of visible near-infrared spectroscopy acquisition and non-destructive prediction of fruits. Background Technique
[0002] Blueberries are honored as the "king of berries" and the "queen of fruits". Their fruits contain significant medicinal values and rich nutritional and health-care effects. The prediction of the soluble solids content (SSC) of blueberries is the research focus of blueberry trade and the quality characteristics of blueberries. By predicting the SSC of blueberries, the maturity, shelf life, and damage degree can be effectively evaluated, thereby increasing the added value of blueberries. In particular, the internal sugar content is directly related to the taste and flavor of blueberries, and about 80% to 85% of the soluble solids in fruits are composed of sugars, and usually the sugar content is used to measure the SSC. Rapid and accurate estimation of the SSC of blueberries is of great significance for guiding the agricultural production of local fruit farmers and the blueberry trade activities of enterprises.
[0003] Traditionally, for the determination of the SSC of blueberries, on-site sampling is required and sent to the laboratory for chemometric analysis, and the refractometer method is commonly used. However, this method requires liquefying the sample, which is destructive, complex in operation, and time-consuming, and it is difficult to predict on a large scale and cannot meet the requirements of production management. Visible near-infrared spectral remote sensing technology, due to its high resolution and rich spectral information, brings a new approach for predicting the SSC of fruits. With the development of machine learning, many researchers have combined it with the visible near-infrared spectral technology of fruits to predict the SSC of different fruits, improving the efficiency. The backpropagation neural network (BPNN) model, as a commonly used neural network model, has obvious advantages in nonlinear processing, self-learning, and generalization ability. In the research on the prediction of the visible near-infrared SSC model of fruits, it is crucial to reasonably select the spectral feature bands closely related to the SSC prediction.
[0004] Previously, researchers have tried to use integer-order differential methods such as the first derivative (FD) and the second derivative (SD) to preprocess spectral data to screen sensitive bands and construct prediction models, which can, to a certain extent, eliminate baseline drift and background interference, thereby highlighting the characteristic peaks in the spectrum. In addition, there are also studies using other spectral preprocessing methods combined with traditional spectral indices to screen feature bands. Traditional spectral indices are usually based on the reflectance combination of specific bands (such as the normalized difference vegetation index NDVI). Research shows that spectral indices can effectively identify the bands sensitive to the inversion index and simplify the model complexity to a certain extent.
[0005] However, the existing methods still have the following limitations. Although the integer-order derivative method can screen out sensitive bands to a certain extent, its overall prediction accuracy is still not high, making it difficult to meet the requirements of high-precision fruit SSC prediction. It can only process spectral data at a fixed order (such as the first order or the second order), and it is unable to finely capture the weak changes in the spectrum. On the other hand, integer-order derivatives are prone to amplifying noise, especially when the signal-to-noise ratio of the spectrum is low, which may lead to a decrease in the accuracy of the prediction model. Traditional one-dimensional spectral indices usually rely on empirical formulas or fixed band combinations and are difficult to adapt to the spectral characteristics of different fruit varieties or growth environments, resulting in the loss of some spectral information. Secondly, traditional spectral indices ignore the possible autocorrelation characteristics between independent variables, such as the problem of multicollinearity between different bands, which may affect the stability and prediction ability of the model. When dealing with high-dimensional spectral data, traditional spectral indices are prone to introducing more noise peaks, further reducing the reliability of the model. In contrast, fractional-order derivatives can finely analyze spectral reflectance information at smaller order intervals to achieve in-depth data mining. The optimized spectral index can comprehensively capture the spectral response of target features through pairwise combination operations of all bands. The combination of fractional-order derivatives and spectral indices has achieved remarkable results in the visible near-infrared spectral modeling of soil and vegetation, such as predicting soil moisture content, salt content, and physiological and biochemical parameters of vegetation. Research shows that the fractional-order derivative model has better performance than the integer-order model, and the characteristic bands extracted by the spectral index contribute greatly to the prediction model, while simplifying the computational complexity of the model. At present, there is no report on the research of fractional-order derivative coupled with optimized spectral index for blueberry SSC, and the complex nonlinear relationship between blueberry SSC and visible near-infrared spectral data also requires a more refined processing method. Therefore, it is necessary to propose a blueberry SSC prediction method based on fractional-order derivative coupled with optimized spectral index to solve the deficiencies of the existing technology and achieve fast, accurate, and non-destructive blueberry SSC prediction. Summary of the Invention
[0006] The object of the present invention is to provide a blueberry SSC prediction method based on fractional-order derivative coupled with optimized spectral index, aiming to solve the technical problems of difficult processing of complex nonlinear relationships between visible near-infrared spectral data, redundant band variables, and omission of band information closely related to SSC. This method has high prediction accuracy, fast prediction speed, low cost, can be predicted in a large range, and has good practical application value.
[0007] To achieve the above object, the present invention is realized through the following technical solutions: A blueberry SSC prediction method based on fractional-order derivative coupled with optimized spectral index specifically includes the following steps:
[0008] Step1: Collect blueberry samples and measure the visible near-infrared spectral reflectance and SSC values of the blueberry samples indoors;
[0009] Step 2: Eliminate outliers from the measured visible near-infrared spectral reflectance data of blueberry samples and the corresponding SSC values.
[0010] Step 3: Divide the visible near-infrared spectral data and SSC data of blueberry samples after outlier elimination, and divide the blueberry samples into two parts, a training set and a validation set, according to a certain proportion.
[0011] Step 4: Preprocess the original spectral reflectance data after dividing the dataset. Use the 0-2 order fractional derivative in the Grünwald-Letnikov form for spectral transformation, with a step size of 0.2 order, and perform data normalization.
[0012] Step 5: Screen the characteristic bands of the preprocessed spectral reflectance data using the Pearson correlation coefficient between the two-dimensional optimized spectral index and the SSC value.
[0013] Step 6: Use the screened characteristic bands as model variables. Among them, the screened characteristic bands are independent variables, and the measured SSC value of blueberries in the study area is the dependent variable. Construct SSC visible near-infrared spectral prediction models of different orders for blueberry SSC prediction and conduct model tests.
[0014] Step 7: Calculate the modeling determination coefficient, prediction determination coefficient, corrected root mean square error, prediction root mean square error, and relative analysis error between the SSC predicted values and the measured values output by the prediction models of different orders to evaluate each order of prediction model, and select the optimal model for non-destructive prediction of blueberry SSC.
[0015] The specific method for outlier elimination is as follows:
[0016] Eliminate the measured abnormal data, and eliminate the data in the measured sugar content values that are higher or lower than the preset multiple of the normal range of the dataset.
[0017] The 0-2 order fractional derivative in the Grünwald-Letnikov form is specifically as follows:
[0018]
[0019] In the formula, is the independent variable band, is the reflectance of the band, is to replace the independent variable in the function f with and then the band reflectance value, is the order of the derivative. When = 0, it represents the original spectrum, When = 1, it represents the first derivative of the original spectrum. When = 2, it represents the second derivative of the original spectrum. is the function with respect to of the nth derivative, where n is the difference between the upper and lower limits of the derivative. is the gamma function and is the value of the gamma function with the independent variable
[0020] The specific data normalization process is as follows:
[0021] Perform minimum - maximum normalization on the spectral reflectance data after the original and fractional - order derivative processing. The normalization formula is:
[0022]
[0023] In the formula, the variables and represent the data values before and after normalization. and represent the maximum and minimum absorbance values of the sample at the same wavelength, or the maximum and minimum values of SSC.
[0024] Performing fractional - order derivative first can enhance the characteristic information of spectral data, highlight the change trend and details of the spectral curve, and make features such as absorption peaks and reflection peaks of the spectrum more obvious. Then performing data normalization can process these enhanced features on a unified scale, which is beneficial for the subsequent BPNN model to better learn the characteristic patterns of the data.
[0025] The specific content of Step5 is as follows:
[0026] Use the measured blueberry SSC data and spectral reflectance data of different orders as input data, and calculate 4 optimized spectral indices by combining all bands of different orders in pairs.
[0027] Use the Pearson correlation coefficient method to perform correlation analysis on the optimized spectral indices of DI, NDI, SI, and IDI of different orders and the blueberry SSC content respectively, and find the corresponding bands Band i and Band j as the optimal characteristic band combination. Obtain 4 optimal band combinations, a total of 8 bands, as the input variables of the model. The calculation formulas of the four optimized spectral indices are as follows:
[0028]
[0029]
[0030]
[0031]
[0032] In the formula, NDI is the normalized index, DI is the difference index, SI is the index, IDI is the reciprocal difference index, Band i and Band j respectively represent bands i and j, R Bandi 、R Bandj respectively represent the reflectances of two different bands.
[0033] The Pearson correlation coefficients are calculated between the four spectral index values calculated for different orders and the measured SSC values of blueberries, respectively, to obtain two-dimensional matrix data (m×m) of the correlation coefficients. The Pearson correlation coefficient The formula is:
[0034]
[0035] In the formula, and are the observed values of two variables (band reflectance and SSC value, respectively), 、 are the average values of the two variables respectively, and n is the number of samples.
[0036] Specifically, Step6 is as follows:
[0037] Using the selected spectral characteristic wavelengths as model variables, a visible near-infrared spectral prediction model for SSC value is established to predict the SSC of blueberries, and model verification is carried out. The characteristic bands selected by the characteristic band selection algorithm are used as independent variables, and the measured SSC values of blueberries in the study area are used as dependent variables. A backpropagation neural network (BPNN) model is selected to construct visible near-infrared spectral prediction models for SSC of different orders. In the BPNN model, the grid search method is used to find the optimal hyperparameter combination for each order of input variables, with the aim of finding the optimal model corresponding to each order.
[0038] Specifically, Step7 is as follows:
[0039] Calculate the modeling determination coefficient R c 2 、prediction determination coefficient R p 2, the calibration root mean square error RMSEC, the prediction root mean square error RMSEP, and the relative analysis error RPD were used to evaluate the prediction models of blueberry SSC for each order of visible near-infrared spectra, determine the prediction accuracy and generalization performance of the prediction models of blueberry SSC for each order of visible near-infrared spectra, select the optimal model, and compare the integer-order derivative model with the fractional-order derivative model. R p 2 , R c 2 , the specific calculation formulas for RMSEC, RMSEP, and RPD are as follows:
[0040] 2
[0041] 2
[0042]
[0043]
[0044]
[0045] In the formula, y mi is the measured SSC value of the i-th group of blueberries, y pi is the predicted SSC value of the i-th group of blueberries, y mean is the average SSC value in the training set or validation set, n p and n c are the numbers of blueberries in the training set and validation set, respectively.
[0046] The beneficial effects of the present invention are as follows: Compared with the method of separately using traditional integer-order derivatives for spectral data preprocessing or separately using traditional one-dimensional specific-band spectral indices for feature band extraction, the present invention innovatively combines the advantages of fractional-order derivatives and two-dimensional optimized spectral indices, effectively solving the problems of difficult handling of complex non-linear relationships between visible near-infrared spectral data, redundant feature band variables, and omission of band information closely related to SSC. And compared with the traditional refractometer measurement method, the present invention has significant advantages such as high efficiency, non-destructiveness, and low cost. At the same time, the present invention significantly improves the prediction accuracy of SSC based on blueberry visible near-infrared spectral data, can perform quantitative prediction quickly and accurately, and provides strong technical support for the non-destructive prediction research of blueberry SSC. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a schematic flow chart of the method in the embodiment of the present invention;
[0048] Figure 2 Schematic diagram of the spectral curve after fractional derivative processing in the embodiment of the present invention;
[0049] Figure 3 Thermodynamic diagram of the correlation between four optimized spectral indices of the original data (0th order) and SSC in the embodiment of the present invention;
[0050] Figure 4 Thermodynamic diagram of the correlation between four optimized spectral indices and SSC after 0.6 fractional derivative processing in the embodiment of the present invention;
[0051] Figure 5 Thermodynamic diagram of the correlation between four optimized spectral indices and SSC after 1.0 fractional derivative processing in the embodiment of the present invention;
[0052] Figure 6 Thermodynamic diagram of the correlation between four optimized spectral indices and SSC after 2.0 fractional derivative processing in the embodiment of the present invention;
[0053] Figure 7 Combined diagram of the optimal characteristic bands and the absolute value of the correlation coefficient corresponding to different orders in the embodiment of the present invention;
[0054] Figure 8 Scatter diagram of the predicted values and measured values of blueberry SSC of each order BPNN prediction model on the validation set in the embodiment of the present invention;
[0055] Figure 9 Line chart of the predicted values and measured values of blueberry SSC of each order BPNN prediction model on the validation set in the embodiment of the present invention. Detailed implementation manners
[0056] The content of the present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners, but this does not limit the scope of the present invention.
[0057] In this embodiment, blueberry samples are collected from a blueberry planting base. The base has a moderate altitude, a mild climate, and a large temperature difference between day and night, which is conducive to the accumulation of blueberry sugar; there is abundant precipitation, moderate humidity, and a moderate pH value, which is suitable for the growth of blueberries; there is sufficient light, which is beneficial to the photosynthesis of blueberries and the improvement of fruit quality.
[0058] As Figure 1 shown, a method for predicting blueberry SSC based on fractional derivative coupled with optimized spectral indices includes the following steps:
[0059] Step1: Collect blueberry samples and measure the visible near-infrared spectral reflectance and SSC value of the blueberry samples indoors.
[0060] Specifically, blueberry sample collection and processing: Manually and randomly collect blueberry samples in the research area. When sampling, classify them according to color, size, and morphology. Put 10 blueberries into a sealed sample bag, and a total of 130 groups are collected, covering three periods: mature (blue-violet, 100 groups), semi-mature (light blue-violet / pink, 15 groups), and immature (light green, 15 groups). The size, weight, and morphology of each group of blueberries are kept consistent. After the samples are transported back to the laboratory, to coordinate environmental variables, store them at 23°C (±1°C) for 12 hours first. Use the HandHeld 2 ground object spectrometer of ASD Company (band range 325 - 1075nm) to measure the visible near-infrared spectrum and image of blueberry samples indoors. In the ENVI remote sensing image processing platform, extract the visible near-infrared spectrum of blueberries. Select areas at intervals of about 90° at the equator of the fruit belly of the sample, and use the average value of 4 measurements as the single-fruit spectrum. The average spectrum of 10 samples in each group is used as the spectrum of this group of samples. Use the Atago PAL-1 fruit refractometer (sugar degree measurement range 0 - 53%) to measure the sugar degree. Measure each blueberry three times and take the average. The SSC mean value of 15 blueberries in each group is used as the SSC value of this group of samples.
[0061] Step2: Remove outliers from the measured visible near-infrared spectrum reflectance data and corresponding SSC values of blueberry samples.
[0062] Specifically, remove outliers from the measured visible near-infrared spectrum data and corresponding SSC of blueberry samples, and remove abnormal data caused by measurement errors. In this example, the removal method is to remove values that exceed three times the normal range value. After removing 0 abnormal samples, finally use the obtained 130 groups of blueberry sample data for modeling.
[0063] Step3: Divide the visible near-infrared spectrum data and SSC data of blueberry samples after outlier removal into two parts: a training set and a validation set according to a certain proportion.
[0064] Specifically, use the train_test_split algorithm to divide the visible near-infrared data of 130 groups of blueberry samples after outlier removal into a training set and a validation set according to a ratio of 8:2. Among them, the training set has 104 and the validation set has 26. The SSC statistical results after dividing the blueberry samples are shown in Table 1.
[0065] Table 1 Statistical description of SSC in blueberry samples
[0066]
[0067] Step4: Preprocess the original spectral reflectance data after dividing the dataset. Use the 0-2 order fractional derivative in the Grünwald-Letnikov form for spectral transformation, with a step size of 0.2 order, and perform data normalization.
[0068] Step4.1: Preprocess the original spectral reflectance data by fractional derivative. Use the 0-2 order fractional derivative in the Grünwald-Letnikov form for spectral transformation, with a fractional derivative interval of 0.2 order, to highlight spectral features and plot the fractional derivative curve. The formula for the Grünwald-Letnikov form is:
[0069]
[0070] In the formula, is the independent variable band, is the reflectance of the band, is the reflectance value of the band after replacing the independent variable in the function f with , is the order of the derivative. When =0, it represents the original spectrum. When =1, it represents the first derivative of the original spectrum. When =2, it represents the second derivative of the original spectrum. is the order derivative of the function with respect to . n is the difference between the upper and lower limits of the derivative. is the gamma Gamma function, is the value of the Gamma function with the independent variable . The step size of the present invention is 0.2 order. After obtaining the fractional derivative result, randomly select a group of samples to plot the fractional derivative curve in the Pycharm software and observe the change trend of spectral features.
[0071] For example, Figure 2As shown, the overall trend of the original spectral curve is as follows: The spectral curve shows a certain fluctuating trend in the wavelength range of 325 - 1075 nm, but the overall reflectance gradually increases with the increase of wavelength. The value of the original spectral reflectance is between 0 and 0.45. The peak of spectral intensity is near 720 nm, the second peak is near 550 nm, and there are two significant absorption peaks near 680 nm and 980 nm, which are related to the absorption of chlorophyll a and water respectively. In addition, the carbon - oxygen bond C - O in the water of blueberry berries can lead to different absorption peaks, thus affecting spectral analysis. In the visible light range (400 - 780 nm), pigments in blueberries (such as anthocyanins) will absorb blue light (400 - 500 nm), so the reflectance in the blue light band is relatively low, usually between 0 and 0.1. In the 500 - 600 nm (green light band), the reflectance increases, between 0.1 and 0.2, because part of the green light is reflected by the surface of blueberries; in the 600 - 700 nm (red light band), the reflectance decreases again, with the reflectance between 0.0 and 0.1, because pigments in blueberries (such as anthocyanins) will absorb red light. In the near - infrared band (780 - 1075 nm), the reflectance in the 780 - 900 nm range increases significantly, with the reflectance between 0.2 and 0.5. The epidermis and internal structure of blueberries have strong reflection on near - infrared light. The relatively high reflectance in the near - infrared band indicates that blueberries have strong reflection characteristics in this band.
[0072] Characteristics of the fractional - order derivative curve: The 0.2 - order derivative curve is similar to the original spectrum, but as the order increases, the two significant peaks near 550 nm and 720 nm gradually become prominent, the width of the reflection peak becomes narrower, and the shape becomes sharper. In addition, the reflection peak at 550 nm is still visible before the 0.8 - order, and becomes less obvious at higher orders; the reflection peak at 720 nm is still visible before the 1.6 - order and becomes less obvious at higher orders, and the derivative reflectance values at the remaining wavelengths of the curve tend to zero. There are obvious fluctuations in the original spectral curve around 325 - 400 nm, which may be due to the highlighting of noise. However, generally, the spectral curve after fractional - order derivative processing becomes smoother compared with the original spectral curve, and the reflectance value continuously decreases to tend to 0, indicating that the spectral noise is better suppressed, so that the true information of the spectrum can be displayed more clearly. By comparing the 0 - 1 - order fractional - order derivative curves, it is found that the shape of each order curve basically remains similar, retaining the shape of the two absorption peaks, which is related to the absorption of chlorophyll a and water in blueberries; by comparing the 1 - 2 - order fractional - order derivative curves, only the peak shape at 720 nm is better retained, and the shoulders and inflection points of this peak are gradually highlighted, but this is not obvious in the original spectrum. In summary, these changes in the spectral shape prove that using fractional - order derivatives can obtain fine spectral information, smooth noise to a certain extent, and retain the main characteristics of the spectrum.
[0073] Step4.2: Perform min-max normalization on the original spectral reflectance data after fractional derivative processing. The normalization formula is as follows:
[0074]
[0075] In the formula, the variables and represent the data values before and after normalization. and represent the maximum and minimum absorbance values of the samples at the same wavelength, or the maximum and minimum values of SSC.
[0076] Step5: Screen the characteristic bands of the preprocessed spectral reflectance data using the Pearson correlation coefficient between the two-dimensional optimized spectral index and the SSC value.
[0077] Step5.1: Take the measured blueberry SSC data (130×1) and the spectral data at each order (130×750) as input data, and calculate the optimized spectral index by combining every two of the 750 bands. In this example, four spectral indices, namely DI, NDI, SI, and IDI, are selected. These four indices involve ratios, differences, sums, and linear combinations. Perform the above four spectral index operations on the unprocessed original spectrum and the spectrum after fractional derivative processing respectively. The calculation formulas for the four optimized spectral indices are as follows:
[0078]
[0079]
[0080]
[0081]
[0082] In the formula, NDI is the normalized index, DI is the difference index, SI is the index, IDI is the reciprocal difference index, Band i and Band j represent bands i and j respectively, and R Bandi and R Bandj represent the reflectances of two different bands respectively. In the present invention, the case of i = j is not considered.
[0083] Step5.2: Calculate the Pearson correlation coefficients between the four spectral index values calculated at different orders and the measured blueberry SSC values respectively to obtain the two-dimensional matrix data (m×m) of the correlation coefficients. The Pearson correlation coefficient formula is as follows:
[0084]
[0085] In the formula, and are the observed values of two variables (band reflectivity and SSC value), , are the means of the two variables, and n is the number of samples.
[0086] Step 5.3: Draw the correlation matrix heat map of each order based on the normalized difference index (NDI), difference index (DI), sum index (SI), inverse difference index (IDI) and SSC, such as Figures 3 - 6 As shown in the figure, the correlation heat maps of the four indices of the original spectrum (0th order), 0.6th order, 1.0th order, and 2.0th order with SSC are plotted respectively. Due to the different calculation methods, the correlation heat maps of NDI, DI, and IDI are bounded by the diagonal line, and the relationship coefficients of the upper and lower parts are opposite, while the relationship coefficients of the upper and lower parts of the SI correlation heat map are the same, that is, they are symmetrical. Figures 3 - 6 As shown, there is a correlation coefficient (r) color bar on the right, ranging from -1 to 1, indicating that the correlation ranges from completely negative correlation to completely positive correlation. The darker the color, the greater the correlation.
[0087] from Figure 3 From the correlation heat map of the four spectral indices, the dark area is wider, indicating that the characteristic bands are redundant and there are more correlation areas. After fractional derivative processing, from 0.6 to 1 ( Figure 4 , Figure 5 ), the light-colored area gradually increases, indicating that the correlation decreases, the extracted spectral feature band information is more concentrated, and the dark-colored area decreases, which effectively solves the problem of feature band redundancy, further proving that fractional-order derivative processing can highlight spectral feature information. However, as the order increases, such as reaching the 2nd order ( Figure 6 ), important spectral feature information may be omitted, which is manifested as large-area coverage of light-colored areas and significantly reduced correlation.
[0088] From the horizontal comparison of the same spectral index, DI and SI have a good correlation with SSC, followed by NDI, and IDI is the worst. The correlation distribution patterns under the same spectral index are similar, indicating that the overall characteristics of the correlation of spectral data reflected by them are relatively consistent. Specifically, DI, IDI, NDI and SI are concentrated in the dark areas in the ranges of 500-550nm, 700-780nm, 700-850nm and 680-780nm, respectively, indicating that the spectral data in these ranges have a strong correlation with SSC, which may correspond to the specific absorption or reflection characteristics of SSC. In the 910-1000nm and other regions, the distribution of light and dark is staggered, indicating that the spectral changes in these regions are complex and may be affected by a combination of multiple factors. From the 0.6th order to the 1st order, the light-colored area gradually increases, further indicating that the correlation decreases.
[0089] The i and j wavelength positions where the maximum correlation coefficient is located are used as the optimal wavelength combination. For example, Figure 7 shows the optimal characteristic band variable combinations for each order under different exponents. As Figure 7 shown, Band i and Band j represent the optimal band combinations. The order with the highest correlation coefficient between DI and SSC is the 0.6th order, with a correlation coefficient of 0.95 and a wavelength combination of 692 nm and 703 nm; the order with the highest correlation coefficient between SI and SSC is the 0.8th order, with a correlation coefficient of 0.93 and a wavelength combination of 519 nm and 727 nm; while the highest correlation coefficients between IDI, NDI and SSC are all obtained at the 0th order, i.e., the original spectrum, with absolute values of 0.89 and 0.93 respectively, and the wavelength combinations are (484 nm, 533 nm) and (513 nm, 554 nm) respectively. The change in the correlation coefficient between the spectral index and SSC is roughly a step-like change. The correlations between SI, DI and SSC are basically close at each order, and the correlations between IDI and NDI and SSC are basically close at each order. Except for IDI, the correlation coefficients between DI, SI, NDI and SSC are stable around 0.8 from the 0.2nd to the 1.0th order, and the correlation coefficients start to gradually decrease from the 1.2th order. DI and SI are stable around 0.7 from the 1.6th to the 2.0th order.
[0090] In addition, the present invention also summarizes the characteristic bands of blueberry SSC. As Figure 7 shown, the wavelength combinations preferably selected by the DI index (0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6 orders), the IDI index (0.6, 0.8, 1.0, 1.2 orders) and NDI (0.4, 0.6, 0.8, 1.0, 1.2 orders) are all near 725 nm and 750 nm, and the wavelength combinations preferably selected by the SI index (0.2, 0.4, 0.6, 0.8, 1.0 orders) are all near 525 and 725 nm. Therefore, the range near 525 and 725 nm is considered to be the characteristic band of blueberry SSC. The high correlation of the optimal bands further confirms the effectiveness of the optimal spectral index in screening characteristic variables.
[0091] Step6: Use the selected characteristic bands as model variables. Among them, the selected characteristic bands are independent variables, and the measured values of blueberry SSC in the study area are dependent variables. Construct SSC visible near-infrared spectral prediction models of different orders to predict blueberry SSC and conduct model tests.
[0092] Specifically, the training sets of different orders after feature band extraction are input into each prediction model, and the backpropagation neural network algorithm in machine learning is used to construct a quantitative prediction model for blueberry SSC. The backpropagation neural network (BPNN) model has significant advantages in dealing with complex problems. It can accurately fit highly nonlinear relationships through multi-layer structures and non-linear activation functions, continuously adjust weights and biases with the help of the backpropagation algorithm, autonomously learn patterns and rules from a large amount of data, and can make reasonable predictions for unseen data after training. In this example, the backpropagation neural network model is implemented by calling the sklearn interface in the Pycharm software. Its network structure mainly consists of an input layer, a hidden layer, and an output layer. To compare the backpropagation neural network models of different orders in this example, especially to compare the model accuracies of integer-order derivatives, original spectra, and fractional-order derivatives, in this example, the grid search method is used to traverse all possible parameter combinations to find the optimal BPNN model hyperparameters of different orders. The entire modeling process for each order of the backpropagation neural network model is divided into the following steps:
[0093] Step6.1: Input the feature variable data (130 * 8) screened by the Pearson correlation coefficient method for each order into the input layer, as well as the dependent variable SSC value (130 * 1), where 130 represents 130 blueberry samples, and 8 represents 8 spectral feature variables (for each fractional order, four optimal feature band combinations screened out by four spectral indices). These feature variables are used to predict the SSC value of the model.
[0094] Step6.2: Subsequently, the training set passes through the network layers such as the hidden layer and output layer of the backpropagation neural network prediction model. During the entire training process, the grid search method traverses all possible parameter combinations to find the optimal hyperparameters of the BPNN model of different orders to minimize the value of the loss function. hidden_layer_sizes: represents the structure of the neural network hidden layer. Four settings are tried, such as (100) for one hidden layer with 100 neurons, (100, 200) representing two layers with 100 and 200 neurons respectively, and (100, 100) and (200, 200) representing two layers with 100 neurons and 200 neurons respectively. activation: the type of activation function for the hidden layer, and the candidate values are'relu' and 'tanh' (the output range of 'tanh' is [-1, 1]). solver: the optimization algorithm for training the neural network, and the candidate values are 'adam' (adaptive moment estimation) and'sgd' (stochastic gradient descent). max_iter: the maximum number of iterations for model training, and the candidate values are 1000 and 3000. alpha: the strength of the regularization term used to prevent overfitting, and the candidate values are 0.0001 (weak), 0.001 (medium), and 0.01 (strong). Therefore, there are a total of 4*2*2*2*3 = 96 parameter combinations. GridSearchCV will traverse these 96 parameter combinations, train the model for each combination, and evaluate the performance of the model through 5-fold cross-validation.
[0095] Step6.3: Finally, save the optimal model with the minimum loss function, that is, the minimum RMSEC, during the iterative training process, and finally make predictions on the validation set.
[0096] Step7: Calculate the modeling determination coefficient, prediction determination coefficient, corrected root mean square error, prediction root mean square error, and relative analysis error between the SSC predicted values and the measured values output by the prediction models of different orders to evaluate each order of prediction model, and select the optimal model for non-destructive prediction of blueberry SSC.
[0097] Specifically, verify and evaluate each trained prediction model. Calculate the R p 2 , RMSEP and RPD between the SSC predicted values and the measured values output by each prediction model on the validation set to evaluate the established visible near-infrared spectroscopy blueberry SSC prediction model, and determine the prediction accuracy and generalization performance of the blueberry SSC prediction model. Among them, R 2The closer the value is to 1, the better the model prediction effect. RMSE is used to describe the difference between the predicted value and the measured value, and the smaller its value, the better. RPD is used to evaluate and describe the prediction ability of the prediction model. The larger the RPD value, the better its prediction ability. When RPD < 1.0, it indicates that the model is not suitable for the prediction task. When 1.0 < RPD < 1.4, the prediction ability of the model is weak. When 1.4 < RPD < 1.8, the prediction ability of the model is average. When 1.8 < RPD < 2.0, the prediction ability of the model is good. When 2.0 < RPD < 2.5, the prediction ability of the model is very good. When RPD > 2.5, it indicates that the prediction ability of the model is excellent. Finally, the evaluation indexes of each prediction model on the training set and the test set are shown in Table 2 below:
[0098] Table 2 Modeling results of BPNN with different orders
[0099]
[0100] As shown in the results of Table 2, the fractional-order backpropagation neural network model (BPNN) has better ability in predicting the SSC of blueberries than the integer-order BPNN model and the original-order BPNN model that only uses spectral indices to screen characteristic bands without fractional-order derivative processing. It can be found that the best modeling effect is achieved by the 0.6-order BPNN model, and the modeling effects of the other 0.2-order and 0.8-order BPNN models are also excellent. The R p 2 both reach above 0.84, and the RPD both reaches above 2.5. For the 0.6-order BPNN model with the best modeling effect, its R c 2 on the training set is 0.920, RMSEC is 0.534%, RPD is 3.533. At the same time, on the validation set, R p 2 is 0.852, RMSEP is 0.482%, RPD is 2.599. According to the classification of the RPD index, it shows that the prediction performance of the 0.6-order BPNN model is excellent, and there is a significant improvement compared with the modeling effects of the integer-order derivatives (1st and 2nd orders) and the original spectrum (0th order).
[0101] Among them, on the validation set, the R p 2 of the 0.6-order BPNN model relative to the optimal BPNN model of the 0th order (original spectrum) is increased by 0.051, RMSEP is decreased by 0.077%, and RPD is increased by 0.356. Secondly, the R p 2They were increased by 0.082 and 0.22 respectively, the RMSEP was decreased by 0.044% and 0.278% respectively, and the RPD was increased by 0.216 and 0.95 respectively. Combining with the scatter plots of the measured and predicted values of blueberry SSC on the validation set for each prediction model, as Figure 8 shown. Among them, the scatter plots of the 0.6-order backpropagation neural network model (BPNN) for the training set and the validation set are closer to the 1:1 line than those of the 0-order, 1-order, and 2-order BPNN models, and the predicted values and the measured values on the validation set fit best. This is mainly due to the powerful information capture of the fractional derivative coupled optimized spectral index.
[0102] Step8: Combining with the line graph of the validation set of the backpropagation neural network model (BPNN), as Figure 9 shown, it can be concluded that the optimal 0.6-order BPNN model reduces the gap between the predicted values and the measured values for some validation set data.
[0103] Based on the above results, it shows the effectiveness of establishing a prediction model using the fractional derivative coupled optimized spectral index, and its prediction accuracy is relatively high, meeting the requirements of actual use, and having the feasibility of application.
[0104] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A blueberry SSC prediction method based on fractional derivative coupled optimization of spectral indices, characterized in that, It includes the following specific steps: Step1: Collect blueberry samples and measure the visible near-infrared spectral reflectance and SSC value of the blueberry samples indoors; Step2: Eliminate outliers from the measured visible near-infrared spectral reflectance data and the corresponding SSC values of the blueberry samples; Step3: Divide the visible near-infrared spectral data and SSC data of the blueberry samples after outlier elimination, and divide the blueberry samples into two parts, a training set and a validation set, according to a ratio; Step4: Preprocess the original spectral reflectance data after dividing the dataset. Use the 0-2 order fractional derivative in the Grünwald-Letnikov form for spectral transformation, with a step size of 0.2 order, and perform data normalization; Step5: Use the Pearson correlation coefficient between the two-dimensional optimized spectral index and the SSC value to screen the characteristic bands from the preprocessed spectral reflectance data; Step6: Use the selected characteristic bands as model variables. Among them, the selected characteristic bands are independent variables, and the measured SSC value of blueberries in the study area is the dependent variable. Construct SSC visible near-infrared spectral prediction models of different orders to predict the SSC of blueberries, and conduct model tests; Step7: Calculate the modeling determination coefficient, prediction determination coefficient, corrected root mean square error, prediction root mean square error, and relative analysis error between the SSC predicted values and the measured values output by the prediction models of different orders to evaluate each order of prediction model, and select the optimal model for non-destructive prediction of blueberry SSC.
2. A blueberry SSC prediction method based on fractional derivative coupling optimization of spectral indices according to claim 1, characterized in that The specific method for eliminating outliers is as follows: Eliminate the measured abnormal data, and eliminate the data in the measured sugar content values that are higher or lower than a preset multiple of the normal range of the dataset.
3. A blueberry SSC prediction method based on fractional derivative coupled optimization of spectral indices according to claim 1, characterized in that The 0-2 order fractional derivative in the Grünwald-Letnikov form is specifically as follows: ; In the formula, is the independent variable band, is the reflectance of the band, is the value of the band reflectance after replacing the independent variable in the function f with , is the order of the derivative. When = 0, it represents the original spectrum, = 1 represents the first derivative of the original spectrum, = 2 represents the second derivative of the original spectrum, is the function with respect to of the order derivative, where n is the difference between the upper and lower limits of the derivative, is the gamma function, is the value of the gamma function with the independent variable .
4. A blueberry SSC prediction method based on fractional derivative coupled optimization of spectral indices according to claim 1, characterized in that, The specific content of Step5 is as follows: Use the measured blueberry SSC data and spectral reflectance data of different orders as input data, and calculate 4 optimized spectral indices by combining all bands of different orders in pairs. The Pearson correlation coefficient method was used to analyze the correlation between the optimized spectral indices of DI, NDI, SI, and IDI at different orders and the blueberry SSC content, respectively, to find the band i and Band j , which were used as the optimal characteristic band combinations. Four optimal band combinations were obtained, totaling 8 bands, which were used as the input variables of the model. The calculation formulas of the four optimized spectral indices are shown as follows: ; ; ; ; In the formula, NDI is the Normalized Difference Index, DI is the Difference Index, SI is the Index, IDI is the Inverse Difference Index, Band i and Band j represent bands i and j respectively, and R Bandi and R Bandj represent the reflectances of two different bands respectively.
Citation Information
Patent Citations
Fig maturity evaluation method based on spectral data and mathematical model
CN119044077A
Fruit maturity comprehensive evaluation and shelf life prediction method based on near infrared spectrum technology
CN119413757A