Wheat grain moisture content lossless prediction method based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement
By combining hyperspectral imaging technology with the Wasserstein generative adversarial network, the destructive and time-consuming problems of wheat moisture detection were solved, and rapid, accurate and non-destructive detection of wheat grain moisture content was achieved, which improved the level of agricultural management and promoted the intelligent and green development of agriculture.
Patent Information
- Application Number
- CN202510716941.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional wheat moisture detection methods are destructive, time-consuming, and difficult to meet the high-efficiency needs of modern agriculture. Hyperspectral imaging technology faces problems of noise interference and poor model robustness in wheat moisture detection, and lacks an accurate visualization method for the internal moisture distribution of wheat grains.
Hyperspectral imaging technology combined with Wasserstein generative adversarial network data enhancement is used. Data is acquired through visible light-near infrared and shortwave infrared imaging systems. The t-distributed random neighbor embedding dimensionality reduction algorithm is used to verify the data quality. A regression model based on extreme learning machine, back propagation neural network and convolutional neural network is constructed. The internal moisture distribution of wheat grains is visualized through model inversion technology.
It realizes rapid, accurate and non-destructive detection of wheat grain moisture content, improves prediction accuracy, supports scientific decision-making in agricultural management, optimizes processing technology and storage management, reduces resource consumption, and promotes intelligent and green development of agriculture.
Smart Images

Figure CN120629028A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of agricultural information technology, and specifically provides a method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement. Background Art
[0002] Wheat moisture content is a key indicator of its storage safety and processing quality. Changes in this indicator are directly related to wheat stability during storage and its performance in subsequent processing steps. Wheat with excessively high moisture content is susceptible to mold infestation during storage, and may even develop problems such as overheating and clumping. In severe cases, this can lead to a significant decline in grain quality or even complete loss. Conversely, excessively low moisture content can affect wheat processing performance, such as reducing flour yield or affecting the taste and texture of the final product. Therefore, how to quickly and accurately measure wheat moisture content has become a key research topic in modern agricultural production and grain storage.
[0003] Traditional detection methods, such as the oven drying method, are currently the most widely used method for moisture determination. This method calculates the moisture content by heating the sample and weighing it, and has high accuracy. However, the oven drying method also has obvious limitations. Its detection process requires destructive treatment of the sample, and the entire process is time-consuming, usually taking several hours or even longer, which makes it difficult to meet the needs of modern agriculture for high efficiency and rapid response. In addition, this method also requires specialized experimental equipment and personnel to operate, which lacks convenience in practical application. Therefore, finding an efficient, non-destructive and accurate detection method has become the key to solving the current practical problems in agricultural production.
[0004] In this context, hyperspectral imaging has emerged as an emerging nondestructive testing technology. Compared with traditional methods, hyperspectral imaging can obtain rich spectral and spatial distribution information from samples without damaging them. This technology allows analysis of the sample's spectral characteristics at different wavelengths, thereby inferring its internal composition and properties. This method is not only nondestructive and rapid, but also provides more comprehensive information about the sample, thus showing significant potential for application in agricultural product quality testing. However, the current application of hyperspectral imaging technology in wheat moisture detection still faces several challenges that need to be addressed. Hyperspectral data typically contains information from a large number of spectral bands, resulting in a massive amount of data, which places high demands on data processing capabilities. Furthermore, hyperspectral data inevitably contains noise interference, which can arise from environmental fluctuations, instrument errors, or sample complexity, all of which can affect the accuracy and robustness of the model. Furthermore, existing research on wheat moisture detection has mostly focused on overall moisture content measurement, lacking precise visualization methods for the internal moisture distribution of wheat kernels. The moisture distribution inside wheat grains is of great significance for understanding the laws of moisture migration, optimizing processing technology and improving storage management.
[0005] In summary, while hyperspectral imaging technology offers a promising solution for wheat moisture detection, numerous challenges remain in its practical application. Future research should focus on effectively removing noise during data processing and improving the generalization of models. Furthermore, developing precise visualization methods for the internal moisture distribution of wheat kernels will be a key area for advancing the application of this technology. These improvements will help further advance hyperspectral imaging technology in wheat moisture detection, thereby better serving the practical needs of modern agricultural production and grain storage management. Summary of the Invention
[0006] The purpose of the present invention is to provide a non-destructive prediction method for wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement, which can effectively combine spectral and moisture content data, improve the prediction accuracy of wheat moisture content, and provide a scientific basis for agricultural management; at the same time, the present invention aims to solve the problems existing in traditional prediction methods when processing high-dimensional spectral data, such as insufficient feature extraction and poor model stability, thereby realizing rapid, accurate and non-destructive detection of wheat grain moisture content.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement, which includes the following steps:
[0008] Step 1: Collect wheat grain samples;
[0009] Step 2: Using a visible-near infrared and short-wave infrared hyperspectral imaging system to obtain hyperspectral image data of wheat grains and measure the moisture content of the wheat grains;
[0010] Step 3: Enhance the spectral and moisture content data based on the Wasserstein generative adversarial network, generate data, and verify the data quality through the t-distributed random neighbor embedding dimensionality reduction algorithm;
[0011] Step 4: Use spectral derivative preprocessing to remove noise and feature selection algorithm to select characteristic wavelengths, and build a regression model based on extreme learning machine, back propagation neural network and convolutional neural network to predict wheat grain moisture content;
[0012] Step five: Use model inversion technology to visualize the spatial distribution of water inside wheat grains.
[0013] Furthermore, in step 1, different varieties of wheat are selected for sampling to ensure the diversity and representativeness of the data; at the same time, the samples are collected under the same environmental conditions to reduce the impact of external factors on the moisture content and ensure the accuracy of subsequent analysis.
[0014] Furthermore, the wheat grain samples cover at least 5 provinces across the country and include 10 varieties, namely "Hengmai 29", "Malan No. 1", "Zhengmai 379", "Xinmai 26", "Jimai 22", "Hemai 29", "Huamai 21", "Lianmai 186", "Wankenmai 22" and "Gushenmai 19", with a total sample size of 700 grains and a moisture content range of 7%-16%.
[0015] Furthermore, in step 2, the parameters of the visible light-near infrared hyperspectral imaging system are: spectral range 382.67-1010.64nm, resolution 2.8nm, CCD camera pixel 804×440, and moving platform speed 7mm / s; the parameters of the shortwave infrared hyperspectral imaging system are: spectral range 982.38-2562.36nm, resolution 6.5nm, charge coupled device camera pixel 320×256, and moving platform speed 17mm / s.
[0016] Furthermore, both the visible-near-infrared hyperspectral imaging system and the short-wave infrared hyperspectral imaging system were enclosed in a dark box. Before collecting data, the hyperspectral imaging system was preheated, and then the wheat grains were arranged in a 5×7 array on a mobile platform for imaging; the original hyperspectral image was corrected using white and black reference objects to reduce the impact of uneven light distribution and eliminate redundant information.
[0017] Furthermore, a white reference is obtained by using a white board with a reflectivity of 99%, and a black reference is obtained by covering the camera lens with a lid; the correction is done using the following formula:
[0018] R Cal =(R Raw -R Dark ) / (R white -R Dark )
[0019] Among them, R Cal is the calibrated hyperspectral image of wheat, R Raw is the original hyperspectral image of wheat, R Dark is the reference image with 0% reflectivity, R White is a reference image with a reflectivity of 99%.
[0020] Furthermore, in step 2, after the hyperspectral image is acquired, the moisture content of the wheat grains is determined by a direct drying method.
[0021] Furthermore, in step three, the Wasserstein generative adversarial network adopts a differentiated network structure for visible light-near infrared and short-wave infrared data: the visible light-near infrared band generator contains 7 layers of deconvolution, and the discriminator contains 8 layers of convolution; the short-wave infrared band generator contains 6 layers of deconvolution, and the discriminator contains 8 layers of convolution; the distribution consistency of the generated data and the real data is verified by the t-distributed random neighbor embedding dimensionality reduction algorithm, and 8000 iterations of synthetic data are selected for model training.
[0022] Furthermore, in the step four, the regression model constructed in the step four is: in the visible light-near infrared band, the first-order derivative-continuous projection algorithm-convolutional neural network modeling is used. Specifically, the original spectrum is preprocessed with the first-order derivative to enhance the feature resolution, and then the characteristic bands with strong correlation are screened out by the continuous projection algorithm. Finally, the screened data is input into the convolutional neural network for training and prediction. The determination coefficient of the model prediction set is 0.9371, the root mean square error is 0.5717, and the residual prediction deviation is 3.8955; in the short-wave near-infrared band, the first-order derivative-ReliefF algorithm-convolutional neural network modeling is used. Specifically, the original spectrum is first preprocessed with the first-order derivative, and then the feature selection is performed by the ReliefF algorithm. Finally, the convolutional neural network is input for training. The determination coefficient of the model prediction set is 0.8095, the root mean square error is 1.0009, and the residual prediction deviation is 2.3201.
[0023] Furthermore, in step five, after the model training is completed, the newly acquired hyperspectral image is processed using model inversion technology to predict the moisture content of wheat at each pixel point; specifically, the single-pixel spectrum is input into the optimal model through model inversion to generate a moisture distribution heat map, in which different color gradients are used to represent high and low moisture content (such as blue for low and red for high).
[0024] Furthermore, the method is applied in wheat processing quality monitoring, storage safety assessment and rapid grain quality detection scenarios.
[0025] The present invention proposes a nondestructive wheat grain moisture content prediction method based on hyperspectral imaging and Wasserstein generative adversarial network data augmentation. This method innovatively constructs a differentiated Wasserstein generative adversarial network data augmentation network to generate high-quality spectral and moisture data tailored to the characteristics of the visible-near-infrared and shortwave infrared bands. The generated data is then verified for distributional consistency with real samples using a t-distributed random neighbor embedding dimensionality reduction algorithm, effectively expanding the training set size. Furthermore, spectral preprocessing methods are used to eliminate noise interference. Feature bands are selected using the SPA and ReliefF algorithms. A convolutional neural network regression model is constructed, and the nonlinear correlation between spectral and moisture content is deeply extracted through convolutional layers. Experimental results show that this method exhibits excellent prediction performance in both the visible-near-infrared and shortwave infrared bands, significantly outperforming traditional extreme learning machine and backpropagation neural network methods. This method has broad applications in agricultural monitoring, food safety, and precision agriculture. It can monitor moisture content in real time, optimize irrigation and harvesting, and improve crop yield and quality while reducing resource consumption. Furthermore, this technology can be extended to other crops, promoting agricultural scientific and technological progress. By optimizing water resource utilization, it supports the sustainable development of agriculture, reduces environmental impact, and promotes intelligent and green agricultural development. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 The present invention is a flowchart of the method for non-destructive prediction of moisture content in wheat grains.
[0027] Figure 2 The structure of the Wasserstein generative adversarial network in the visible-near-infrared band.
[0028] Figure 3 The structure of the Wasserstein generative adversarial network in the shortwave near-infrared band.
[0029] Figure 4Spectral data generated for different training time periods in the (a) visible-near-infrared band and (b) shortwave infrared band; moisture data generated for different training time periods in the (c) visible-near-infrared band and (d) shortwave infrared band; visualization of real data and generated data using the t-distributed random neighbor embedding dimensionality reduction algorithm in the (e) visible-near-infrared band and (f) shortwave near-infrared band.
[0030] Figure 5 Visualization of the moisture content distribution in a single wheat kernel in (a) the visible-NIR band and (b) the shortwave-NIR band. DETAILED DESCRIPTION
[0031] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] like Figure 1 As shown, a non-destructive prediction method for wheat grain moisture content based on hyperspectral imaging and data enhancement includes the following steps:
[0033] Step 1: Collect wheat grain samples. Select different varieties of wheat for sampling to ensure data diversity and representativeness. Samples are collected under the same environmental conditions to reduce the impact of external factors on moisture content and ensure the accuracy of subsequent analysis.
[0034] Step 2: Use a visible-near-infrared and shortwave infrared hyperspectral imaging system to obtain hyperspectral image data of wheat grains. In this process, the collected wheat samples are scanned using a hyperspectral imaging system to obtain hyperspectral images from the visible light to the near infrared and shortwave infrared bands. Then, the moisture content of the wheat grains is measured;
[0035] Step 3: Enhance the spectral and moisture content data using a Wasserstein generative adversarial network. Data is generated and verified for quality using a t-distributed random neighbor embedding dimensionality reduction algorithm. In this step, the Wasserstein generative adversarial network is trained to generate synthetic spectral data and corresponding moisture content data, expanding the dataset and improving the robustness of the model. The t-distributed random neighbor embedding dimensionality reduction algorithm is applied to visualize and verify the quality of the synthetic data, ensuring consistency with the real data in a high-dimensional space, thereby improving the reliability of subsequent analysis.
[0036] Step 4: Spectral preprocessing methods are used to remove noise and feature selection algorithms are used to select characteristic wavelengths. Spectral preprocessing methods can effectively remove noise from spectral data and enhance the performance of signal features. Feature selection algorithms are used to select characteristic wavelengths that are highly correlated with wheat moisture content. These selected characteristic wavelengths will be used to construct regression models, and predictions will be made using various models such as extreme learning machines, backpropagation neural networks, and convolutional neural networks to improve the model's predictive power and accuracy.
[0037] Step five uses model inversion technology to visualize the spatial distribution of moisture within wheat kernels. After model training is complete, the newly acquired hyperspectral imagery is processed using model inversion technology to predict the moisture content of wheat at each pixel. This step generates a spatial distribution map of moisture within the wheat kernel, visualizing the distribution of moisture within the wheat and providing a valuable basis for subsequent wheat storage and processing decisions.
[0038] Step 1 of this embodiment specifically includes:
[0039] China's main wheat-producing regions are located in Hebei, Henan, Shandong, Jiangsu, and Anhui. To ensure comprehensive representation, this study selected ten wheat varieties from these five provinces. Consequently, 700 wheat kernels from ten different varieties were used: Hengmai 29, Malan 1, Zhengmai 379, Xinmai 26, Jimai 22, Hemai 29, Huamai 21, Lianmai 186, Wankenmai 22, and Gushenmai 19. The wheat kernels were stored at room temperature to ensure consistent quality.
[0040] Step 2 of this embodiment specifically includes:
[0041] Hyperspectral images were acquired using a visible-near-infrared (VIS) system and a shortwave-near-infrared (SWNIR) system. The VIS-NIR system consisted of an ICLB1620CCD camera with an 804×440 pixel resolution and an ImSpectorV10E imaging spectrometer. The system had a spectral resolution of 2.8 nm and a wavelength range of 382.67–1010.64 nm. The setup also included a halogen light source, a mobile platform, and a Dell computer. The mobile platform moved at a speed of 7 ms / s, covering a range of 80–280 mm, with an exposure time of 3 ms. The SWNIR system was equipped with an EM285CL camera with a 320×256 pixel resolution and an ImSpectorN25E imaging spectrometer. The system had a wavelength range of 982.38–2562.36 nm and a spectral resolution of 6.5 nm. Similar to the VIS-NIR band, it also includes a halogen light source, a mobile platform that moves at a speed of 17 mm / s within a distance of 80 to 300 mm, and a Dell computer. The exposure time was set to 1.5 ms and the intensity was adjusted to 250. Both systems were enclosed in a dark box to eliminate light interference. Before collecting data, the hyperspectral imaging system was preheated for 30 minutes to stabilize the light source and reduce interference. Wheat grains were arranged in a 5×7 array on the mobile platform for imaging. The raw hyperspectral image was corrected using white and black reference objects to reduce the impact of uneven light distribution and eliminate redundant information. The white reference was obtained using a white board with a reflectivity of 99%, while the black reference was obtained by covering the camera lens with a lid. The correction uses the following formula (1):
[0042] R Cal =(R Raw -R Dark ) / (R white -R Dark (1)
[0043] Among them, RCal is the calibrated hyperspectral image of wheat, RRaw is the original hyperspectral image of wheat, RDark is the reference image with 0% reflectance, and RWhite is the reference image with 99% reflectance.
[0044] After hyperspectral image acquisition, the moisture content of wheat grains was determined using the direct drying method, in accordance with the standards of the American Society of Agricultural Engineers. The specific procedure was as follows: the sample to be tested was placed in a covered aluminum box. Each box contained wheat grains, and after sealing, the initial weight was recorded (accurate to 0.001 g). The boxes were then placed in a constant-temperature forced-air drying oven and dried continuously at 130 ± 2°C for 19 hours. After drying, the boxes were immediately cooled to room temperature and weighed again to a constant weight. The moisture content of the wheat grains was then calculated using Equation (2).
[0045] Moisture content (%) = (m1-m2) / m1×100% (2)
[0046] Where m1 is the weight of the wheat grains and m2 is the weight of the wheat grains after drying.
[0047] Step 3 of this embodiment specifically includes:
[0048] The present invention uses the Wasserstein generative adversarial network to improve data quality and optimize the accuracy of model prediction. Figure 2 and Figure 3 The Wasserstein generative adversarial network architecture for the visible-near-infrared (VIS) and shortwave-near-infrared (SWNIR) bands is shown. The inherent data size differences between the VIS and SWNIR bands necessitate the development of a network architecture that can effectively handle these differences. The VIS spectra typically have higher spectral resolution and larger data volumes than the SWNIR. In the VIS-NIR band, the generator consists of seven deconvolutional layers, while the discriminator consists of eight convolutional layers. This architecture is designed to effectively process the high-resolution data in the VIS-NIR band, resulting in accurate spectral and moisture content data generation. In contrast, the SWNIR band requires a different architecture, with the generator consisting of six deconvolutional layers and the discriminator consisting of eight convolutional layers. This modified architecture is optimized for the lower spectral resolution and smaller data size of the SWIR, enabling the network to effectively process and generate spectral and moisture content data in this band. Furthermore, the implementation of additional layers in the discriminator facilitates the learning and representation of more complex functions, thereby enhancing the discriminator's ability to distinguish between generated and real data. Furthermore, a more robust discriminator is less prone to mode collapse, where the same output produced by the generator has limited variations, thus encouraging the generator to generate more diverse and richer data. Therefore, the Wasserstein generative adversarial network is able to effectively handle different data sizes by adopting a specialized architecture for each band, achieving higher performance.
[0049] Figure 4 (a) shows spectra generated in the visible-near-infrared spectrum at 500, 2000, 4000, and 8000 iterations, respectively. As time increases, the generated spectra become smoother and more consistent with the true spectra. After approximately 500 iterations, the generated spectra exhibit similar profiles. However, they still have considerable noise. At 8000 iterations, the generated spectra show a high degree of similarity to the true spectra. Figure 4 (b) shows the shortwave infrared spectra generated at 500, 2000, 4000 and 8000 iterations. Figure 4Similarly, the generated spectral curve (a) exhibits its initial shape at 500 iterations and reaches equilibrium at 8000 iterations. However, compared to the generated visible-NIR spectral curve, residual noise still exists. Overall, the spectral data generated by the Wasserstein GAN reaches a stable state at 8000 iterations. Figure 4 (c)-(d) Block diagrams depict the real and generated moisture content data for the VIS-NIR and SWIR spectra. At 500 iterations, the generated moisture data differed significantly from the real moisture data. After 2000 iterations, the generated moisture data were nearly identical to the real moisture data, demonstrating that the moisture content data generated by the Wasserstein GAN are of high quality and suitable for subsequent regression model development. Therefore, the spectral and moisture data generated after 8000 iterations were selected for subsequent analysis, as they showed the highest similarity to the real data and demonstrated strong convergence, providing a reliable foundation for further research.
[0050] To further investigate the credibility of the spectral and moisture data generated by the Wasserstein GAN, we employed the t-distributed random neighbor embedding (T-DNE) dimensionality reduction algorithm. This is a nonlinear dimensionality reduction and visualization technique that preserves the local and global structure of the original dataset by initially calculating pairwise similarities between any two data points in a high-dimensional space. Figure 4 (e)-(f) show visualizations of real and generated spectral and moisture content data in the visible-near-infrared and shortwave infrared bands using the t-distributed random neighbor embedding dimensionality reduction algorithm. The real and generated spectral and moisture data exhibit similar clustering structures in both three-dimensional space and two-dimensional mapping (including the XY, YZ, and XZ planes), indicating significant similarity between the two datasets and demonstrating the similarity between the generated spectral and moisture content data and the real ones. Therefore, the generated spectral and moisture data are of extremely high quality, demonstrating strong potential and making a significant contribution to the subsequent construction of robust and accurate regression models. The high-quality data generated can provide a reliable foundation for model development, enabling the created regression model to effectively capture the complex relationships between spectral and moisture data. This, in turn, improves model performance and interpretability, ultimately enhancing the overall accuracy and reliability of the regression model.
[0051] Step 4 of this embodiment specifically includes:
[0052] During the acquisition of sample spectral data, noise and significant interference are often encountered. These interferences can be attributed to various sources, including dark current, light scattering, and potential human error. Effectively removing noise from spectral data is crucial to prevent data distortion and ensure the accuracy of subsequent analysis. Therefore, the present invention adopts spectral preprocessing methods to eliminate the effects of noise. Spectral preprocessing methods can be roughly divided into four categories: baseline correction, scatter correction, smoothing, and scaling. To investigate the effects of four different preprocessing methods on the prediction performance of the regression model, first-order derivatives, standard normalized variables, SG smoothing, and automatic scaling were used. When evaluating spectral preprocessing methods, it is crucial to adopt a multifaceted approach that considers both the noise reduction effect and the overall performance of the model, as this is the only way to determine the optimal preprocessing method. By adopting this comprehensive perspective, the selected method is able to effectively minimize noise while providing accurate and reliable results, thereby ensuring the validity and robustness of the prediction model.
[0053] Given the inherent redundancy of hyperspectral data, filtering for characteristic bands is crucial to remove redundant information. Generally speaking, characteristic band screening methods can be categorized into three main categories: filtering, wrapping, and embedding. Filtering methods, when assessing variable importance, ignore potential interdependencies or synergies between variables and are therefore unsuitable for the purposes of this study. In contrast, wrapping and embedding methods can account for interactions between variables and are more suitable for selecting characteristic bands for hyperspectral data. The Successive Projections Algorithm (SPA) is a classic wrapping method that minimizes collinearity among variables through direct operations in vector space. The SPA algorithm uses an iterative process to progressively select spectral vectors that are as orthogonal as possible. These vectors serve as the final elements for decomposing and interpreting mixed spectral data. The ReliefF algorithm, on the other hand, is a classic embedding method that relies on a user parameter, k, called the "number of neighbors." This parameter uses the k most recent hits and misses in the score update for each target instance, thereby improving the reliability of the weight estimate. Therefore, by selecting an appropriate parameter, k, the ReliefF algorithm can effectively handle noise and interference in hyperspectral data.
[0054] The extreme learning machine (ELM), a single-hidden-layer feedforward neural network, has the core feature that the weights of the hidden layer nodes are randomly or manually assigned and do not require updating. Based on this feature, the ELM significantly improves computational efficiency while maintaining model accuracy through a single-step learning process that only calculates the output weights. The backpropagation neural network (BPNN), a multi-layer feedforward network, uses the backpropagation algorithm for network training and weight optimization. During the forward propagation phase, information is passed layer by layer to the output layer; during the backpropagation phase, the weights are iteratively updated by backpropagating the error gradient, thereby continuously narrowing the deviation between the predicted and actual values. The one-dimensional convolutional neural network (1D-CNN), a variant of the convolutional network specifically designed to process one-dimensional time series data, typically consists of alternating convolutional and pooling layers, with a fully connected layer at the end to map the feature space to the output target.
[0055] Although neural network models demonstrate efficient feature learning capabilities in specific scenarios, insufficient data often hinders their full performance in real-world applications. To address this, generative adversarial networks (GANs) are introduced for data augmentation. By expanding limited samples and optimizing data distribution, they provide a more robust input foundation for subsequent modeling. It has been shown that tripling the dataset generated using GANs can improve modeling results by expanding the limited dataset and better representing the underlying patterns. Therefore, to augment the dataset, the Wasserstein GAN was used to augment the data in the visible-near-infrared band and the shortwave-near-infrared band, tripling the sample size of both datasets. For regression modeling, the dataset was partitioned into a calibration set and a prediction set in a 7:3 ratio to ensure robustness in model training and evaluation. This partitioning strategy enables reliable estimation of model parameters and comprehensive evaluation of predictive performance, more accurately reflecting the relationship between predictors and the response variable. In order to comprehensively evaluate the predictive ability of the model, the determination coefficient of the calibration set (R2 C) and root mean square error of the calibration set (RMSEC), as well as the determination coefficient of the prediction set (R2 P) and root mean square error of the prediction set (RMSEP), performance to deviation ratio (RPD) and range error ratio (RER) of the prediction set were used, and the evaluation was performed using formulas (3)-(6).
[0056]
[0057] RPD=SD / RMSEP (5)
[0058] RER=Range / RMSEP (6)
[0059] When evaluating spectral preprocessing methods, a multifaceted approach is crucial, considering both noise reduction effectiveness and overall model performance, as the synergy of these two factors allows for a more comprehensive determination of the optimal preprocessing method. By employing this multidimensional evaluation framework, the selected method not only effectively reduces noise interference but also significantly improves the model's prediction accuracy and robustness. Table 1 shows the performance of a regression model for moisture content prediction using the full visible-near-infrared spectrum. A higher R² indicates a better model fit, while a lower RMSE indicates a smaller discrepancy between the predicted and actual results. First-order derivative preprocessing combined with an extreme learning machine performed well in moisture prediction, achieving an R² C of 0.7398, a RMSEC of 1.1456, an R² P of 0.6851, and an RMSEP of 1.2518. Autoscaling preprocessing combined with a back-propagation neural network demonstrated even better performance, with R² C and R² P reaching 0.8071 and 0.7791, respectively, and RMSEC and RMSEP decreasing to 0.9683 and 1.0919. The first-order derivative preprocessing combined with the convolutional neural network showed the best performance, with its R2 C and R2 P increased to 0.9388 and 0.9215, and RMSEC and RMSEP significantly decreased to 0.5531 and 0.6308.
[0060] Table 1: Performance of the regression model for moisture content prediction based on four VIS-NIR pretreatment methods
[0061]
[0062]
[0063] Table 2: Performance of the regression model for moisture content prediction based on four shortwave near-infrared pretreatment methods
[0064]
[0065] Table 2 illustrates the performance of the regression models for predicting SWIR moisture content using different preprocessing methods. Following the same principles used in VIS-NIR analysis, the R 2 The best model in short-wave infrared was selected by comparing the results of RMSE. The best performing models in moisture content prediction include the 1st-derivative-ELM model, the 1st-derivative-BPNN model, and the 1st-derivative-CNN model. Specifically, the 1st-derivative-ELM model has the highest R in the calibration set and the prediction set. 2 C 、R 2 PThe 1st-derivative-BPNN model showed better prediction performance. and The 1st-derivative-CNN model is improved to 0.8449 and 0.7611 respectively, while RMSEC and RMSEP are reduced to 0.9407 and 1.1530. The RMSEP value of the proposed method is slightly lower than that of the former (0.8327), but its RMSEP value (1.1056) shows better prediction stability.
[0066] Taken together, these results indicate that first-order derivative and autoscaling preprocessing, through a synergistic dual optimization mechanism of noise suppression and feature enhancement, effectively improves the generalization performance of multi-component prediction models in both the visible-near-infrared and shortwave near-infrared bands. This finding is highly consistent with previous research conclusions and further validates the universal advantages of first-order derivative and autoscaling preprocessing for spectral data feature extraction. Not only does this validate the technical advantages of first-order derivative preprocessing in eliminating baseline drift, but it also confirms the universal value of autoscaling preprocessing for optimizing the distribution of spectral response values. The adaptability of different spectral preprocessing methods to machine learning and deep learning regression models varies significantly. First-order derivative preprocessing is more suitable for CNN models due to their autonomous feature extraction capabilities. Autoscaling preprocessing, on the other hand, exhibits stronger synergy with models such as ELM and BPNN, as these algorithms are more sensitive to scale changes in the input data.
[0067] Eliminating redundant variables is a widely adopted strategy for simplifying models and enhancing versatility. Its predictive performance is comparable to or even exceeds that of full-wavelength models. Based on the selection of characteristic bands in the visible-near-infrared region, the SPA and ReliefF algorithms exhibit differentiated feature extraction properties for different preprocessed data.
[0068] The SPA algorithm uses a root mean square error (RMSE) optimization strategy to select characteristic bands for both auto-scaled and first-order derivative preprocessed spectral data. For the water content prediction task, seven characteristic bands (404.00, 424.20, 425.55, 661.92, 879.6, 978.79, and 994.74 nm) were selected for the auto-scaled preprocessed spectral data, while eight characteristic bands (399.99, 401.32, 405.34, 766.79, 970.08, 971.54, 983.15, and 988.95 nm) were required for the first-order derivative preprocessed spectral data to maintain model accuracy. This further reveals the distribution of feature importance for the ReliefF algorithm. The algorithm realizes band screening by setting the importance threshold. For water content prediction, 21 characteristic bands (399.99, 401.32, 402.66, 404.00, 406.69, 408.03, 409.37, 412.06, 413.41, 414.75, 418.8, 420.15, 422.85, 425.55, 428.26, 430.97, 9 84.6, 991.84, 993.29, 997.63, 999.08 nm), while the first-order derivative preprocessed spectral data can still select 11 characteristic bands (406.69, 410.72, 412.06, 413.41, 416.10, 425.55, 437.77, 441.85, 990.39, 991.84, 993.29 nm) under the condition of a higher threshold of 0.06, indicating that the first-order derivative preprocessing effectively improves the feature discrimination ability.
[0069] The results of the optimization of short-wave near-infrared characteristic bands showed significant differences. Analysis of the SPA algorithm based on the root mean square error criterion showed that 16 characteristic bands (1001.04, 1007.32, 1019.95, 1178.85, 1333.71, 1348.10, 1362.53, 1391.51, 1449.83, 1648.37, 1670.40, 1844.61, 1851.75, 1858.88, 1915.43, and 1977.8 nm) were selected for moisture content prediction using the first-order derivative preprocessing of spectral data. The ReliefF algorithm uses wavelength importance assessment and a threshold of 0.02 to screen out 17 characteristic bands (1051.98, 1064.96, 1421.91, 1457.15, 1486.48, 1493.82, 1552.57, 1699.72, 1873.10, 1880.19, 1922.44, 1936.40, 1950.30, 1964.13, 1977.88, 1984.74, and 1991.57 nm) from the first-order derivative preprocessed spectral data for moisture prediction.
[0070] Further analysis revealed that spectral preprocessing profoundly impacted feature selection results. Spectral data preprocessed with first-order derivatives generally produced more characteristic bands than those preprocessed with autoscaling, suggesting that first-order derivative preprocessing enhanced the availability of detailed information by eliminating baseline drift.
[0071] The selection of feature bands is crucial for building accurate regression models. Screening feature bands can reduce data dimensionality, minimize the impact of noise, and enhance model interpretability, thereby helping to discover underlying relationships and patterns, thereby improving prediction accuracy and reducing the risk of overfitting.
[0072] Table 3 shows the regression performance of the prediction models for various physical and chemical indicators in the visible-near infrared band. Through comparative analysis, it was found that the optimal models for different indicators showed significant differences. In terms of moisture content prediction, the extreme learning machine based on first-order derivative preprocessing combined with the SPA feature band selection algorithm showed the best performance, and its correction set The prediction set The Autoscale-SPA-BPNN model performs better in terms of prediction stability. The 1st-derivative-SPA-CNN model showed the strongest comprehensive performance. The results reached 0.9403 and 0.9371 respectively, RMSEC and RMSEP were as low as 0.5438 and 0.5717, and RPD and RER were improved to 3.8955 and 15.2713. Table 4 further shows the modeling results of the short-wave near-infrared band. In moisture prediction, the 1st-derivative-SPA-ELM model maintains its basic performance advantage. is 0.8004, RMSEC is 1.0723, and the 1st-derivative-ReliefF-CNN model is optimized by feature selection algorithm. It increased to 0.8095 and RMSEP decreased to 1.0009.
[0073] The results of the study confirmed that the feature band selection based on the SPA algorithm can effectively improve the model performance. This is due to the SPA algorithm's ability to accurately extract potential spectral features, thereby optimizing data representation and model fitting. This conclusion is highly consistent with the research findings that feature selection improves the generalization ability of the model, verifying the scientific nature of the method system of the present invention. From the perspective of model architecture, the CNN model performs best in predicting moisture content. Its advantage stems from its ability to automatically extract spatial hierarchical spectral features, especially when processing high-dimensional spectral data. It is significantly better than the ELM and BPNN models. Although the BPNN model is close to the performance of the CNN model in some scenarios, its RPD value is generally lower than 3, and its robustness is insufficient. Due to the limitations of its shallow structure, the ELM model only performs well in specific preprocessing combinations.
[0074] Table 3: Performance of the regression model for predicting moisture content based on the visible-near infrared characteristic band
[0075]
[0076] Table 4: Performance of the regression model for predicting moisture content based on the shortwave near-infrared characteristic band
[0077]
[0078]
[0079] Step 5 of this embodiment specifically includes:
[0080] Hyperspectral imaging technology offers significant advantages over traditional spectral analysis techniques. Its unique spatial resolution, combined with continuous spectral information, enables spatial visualization of physical and chemical parameters within a sample. To systematically assess the three-dimensional distribution of physical and chemical parameters within wheat grains, a model inversion approach was employed. The optimal prediction model was applied pixel by pixel to the original region-of-interest image data, resulting in a chemical composition distribution map that intuitively reflects the gradients of physical and chemical parameters.
[0081] Hyperspectral imaging offers distinct advantages by integrating spectral and spatial information, enabling visualization of moisture content distribution within individual wheat kernels. Figure 5 The results of moisture content inversion models constructed based on visible-near-infrared and short-wave near-infrared bands for single wheat kernel detection are presented. The moisture distribution characteristics are characterized by a color gradient, where cold tones (blue) correspond to the lowest values and warm tones (red) represent the highest values. 35 independent samples were used to generalize and verify the optimal models of the two types of bands. The results showed that although the inversion image of the visible-near-infrared band presented an overall blue background with a high pixel density, the main area of the sample was still mainly distributed in red, while the inversion image of the short-wave near-infrared band more significantly showed the clustering characteristics of high-value red areas. The detection results of the two bands were consistent with the distribution law of moisture gradient inside the wheat kernel, that is, the high-moisture area was concentrated in the core of the endosperm while the surface was relatively dry.
[0082] In summary, in order to improve the accuracy and robustness of prediction, the present invention introduces the Wasserstein generative adversarial network data enhancement technology, designs a differentiated network structure to generate high-quality synthetic spectral data and corresponding wheat moisture content data, expands the training data set, and improves the robustness of the model; utilizes model inversion technology to analyze the spectral information corresponding to each pixel point in the hyperspectral image, and combines the trained prediction model to realize the visualization of the spatial distribution of moisture inside the wheat grain, providing guidance for wheat storage and processing.
[0083] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the scope of protection of the present invention in any form. All technical solutions obtained by equivalent substitution, etc., fall within the scope of protection of the present invention. Parts not covered by the present invention are the same as the existing technology or can be implemented using existing technology.
Claims
1. A non-destructive prediction method for wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement, characterized in that: The following steps are involved: Step 1: Collect wheat grain samples; Step 2: Using a visible-near infrared and short-wave infrared hyperspectral imaging system to obtain hyperspectral image data of wheat grains and measure the moisture content of the wheat grains; Step 3: Enhance the spectral and moisture content data based on the Wasserstein generative adversarial network, generate data, and verify the data quality through the t-distributed random neighbor embedding dimensionality reduction algorithm; Step 4: Use spectral derivative preprocessing to remove noise and feature selection algorithm to select characteristic wavelengths, and build a regression model based on extreme learning machine, back propagation neural network and convolutional neural network to predict wheat grain moisture content; Step five: Use model inversion technology to visualize the spatial distribution of water inside wheat grains.
2. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: In the step 1, different varieties of wheat are selected for sampling to ensure the diversity and representativeness of the data; at the same time, the samples are collected under the same environmental conditions to reduce the impact of external factors on the moisture content and ensure the accuracy of subsequent analysis; specifically, the wheat grain samples cover at least 5 provinces in the country and include 10 varieties, namely "Hengmai 29", "Malan No. 1", "Zhengmai 379", "Xinmai 26", "Jimai 22", "Hemai 29", "Huamai 21", "Lianmai 186", "Wankenmai 22" and "Gushenmai 19", with a total sample size of 700 grains and a moisture content range of 7%-16%.
3. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: In step 2, the parameters of the visible light-near infrared hyperspectral imaging system are: spectral range 382.67-1010.64nm, resolution 2.8nm, CCD camera pixel 804×440, and moving platform speed 7mm / s; the parameters of the shortwave infrared hyperspectral imaging system are: spectral range 982.38-2562.36nm, resolution 6.5nm, charge coupled device camera pixel 320×256, and moving platform speed 17mm / s.
4. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1 or 3, characterized in that: Both the visible-near-infrared hyperspectral imaging system and the shortwave infrared hyperspectral imaging system were enclosed in a dark box. Before collecting data, the hyperspectral imaging system was preheated, and then wheat grains were arranged in a 5×7 array on a mobile platform for imaging. The original hyperspectral image was corrected using white and black reference objects to reduce the impact of uneven light distribution and eliminate redundant information.
5. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 4, characterized in that: The white reference is obtained by using a white board with a reflectivity of 99%, and the black reference is obtained by covering the camera lens with a lid; the correction is made using the following formula: R Cal =(R Raw -R Dark ) / (R white -R Dark ) Among them, R Cal is the calibrated hyperspectral image of wheat, R Raw is the original hyperspectral image of wheat, R Dark is the reference image with 0% reflectivity, R White is a reference image with a reflectivity of 99%.
6. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: In the step 2, after the hyperspectral image is collected, the moisture content of the wheat grains is determined by a direct drying method.
7. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: In step three, the Wasserstein generative adversarial network adopts a differentiated network structure for visible light-near infrared and short-wave infrared data: the visible light-near infrared band generator contains 7 layers of deconvolution, and the discriminator contains 8 layers of convolution; the short-wave infrared band generator contains 6 layers of deconvolution, and the discriminator contains 8 layers of convolution; the distribution consistency of the generated data and the real data is verified by the t-distributed random neighbor embedding dimensionality reduction algorithm, and 8000 iterations of synthetic data are selected for model training.
8. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: The regression model constructed in step 4 is: In the visible-near-infrared band, a first-order derivative-continuous projection algorithm-convolutional neural network model was used. Specifically, the original spectrum was preprocessed with the first-order derivative to enhance feature resolution. The continuous projection algorithm was then used to filter out characteristic bands with strong correlation. Finally, the filtered data was input into the convolutional neural network for training and prediction. The model's coefficient of determination on the prediction set was 0.9371, the root mean square error was 0.5717, and the residual prediction bias was 3.8955. In the shortwave near-infrared band, the first-order derivative-ReliefF algorithm-convolutional neural network modeling is adopted. Specifically, the original spectrum is first preprocessed with the first-order derivative, then feature selection is performed using the ReliefF algorithm, and finally the convolutional neural network is input for training. The determination coefficient of the model prediction set is 0.8095, the root mean square error is 1.0009, and the residual prediction deviation is 2.3201.
9. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: In the step five, the single-pixel spectrum is input into the optimal model through model inversion to generate a visualized thermal map of the internal moisture distribution of the wheat grains, in which different color gradients are used to represent the moisture content.
10. The method for non-destructive prediction of wheat grain moisture content based on hyperspectral imaging and Wasserstein generative adversarial network data enhancement according to claim 1, characterized in that: The method is applied in wheat processing quality monitoring, storage safety assessment and rapid grain quality detection scenarios.
Citation Information
Cited By
An online nondestructive detection method for water content of agricultural products based on multispectral imaging
CN122448770A