A peanut biomass inversion method based on feature extraction and screening

Through the drone, feature screening and model construction are carried out through feature acquisition of high-spectral image data, combined with wavelet transformation and vegetation index, the problem of low inversion accuracy of peanut biomass is solved and higher inversion accuracy and accuracy are achieved.

CN119295972BActive Publication Date: 2025-05-09HENAN UNIV OF ECONOMICS & LAW +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411310618.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-05-09
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify peanut biomass characteristics, resulting in low inversion accuracy.

Method used

The drone is equipped with a hyperspectral camera to collect peanut hyperspectral image data, and multi-scale features are extracted through wavelet transformation and mathematical transformation. The peanut vegetation index and variable projection importance methods are combined for feature screening. The random forest model, backpropagation neural network and support vector machine model after particle swarm optimization are constructed for peanut biomass inversion.

Benefits of technology

It improves the accuracy and accuracy of peanut biomass inversion, and reduces the data redundancy and dimensional disaster problems caused by the improvement of hyperspectral resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119295972B_ABST
    Figure CN119295972B_ABST
Patent Text Reader

Abstract

The present invention discloses a peanut biomass inversion method based on feature extraction and screening, comprising the following steps: S1: using an unmanned aerial vehicle equipped with a hyperspectral camera to collect hyperspectral image data of sample peanuts; S2: obtaining the measured biomass of the sample peanuts; S3: preprocessing the collected hyperspectral image data of the sample peanuts to obtain initial spectral reflectance data; S4: extracting and screening the features of the initial spectral reflectance data; S5: using the screened spectral features, constructing a peanut biomass inversion model through a random forest model after particle swarm optimization, a back propagation neural network and a support vector machine, and using the measured biomass of the sample peanuts as a data set to train and test the peanut biomass inversion model. The present invention realizes the screening and feature combination of effective features of peanut hyperspectral images, thereby improving the accuracy and precision of the peanut biomass inversion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of crop growth monitoring, and in particular relates to a peanut biomass inversion method based on feature extraction and screening. Background Art

[0002] As an important agricultural product, peanut is one of the most widely grown oil crops in the world. my country is the world's largest peanut producer, consumer and importer. With the improvement of people's living standards, the demand for oil has increased significantly. In 2023, my country's peanut production reached 16.4 million tons, accounting for 37% of the global total. However, many grain and oil exporting countries have strengthened export control, which poses a serious challenge to my country's oil supply security. It can be seen that peanut production is of great significance to my country's economy and agriculture. Rapid, accurate and non-destructive monitoring of peanut growth and estimation of its biomass are crucial to ensuring the security of national oil supply and improving agricultural production efficiency.

[0003] At present, there are relatively few studies on remote sensing estimation of peanut biomass at home and abroad, especially the research on the reflectance of the canopy of peanut plants and the influence of feature screening and feature combination on the accuracy of the inversion model is still insufficient. The Chinese invention patent with announcement number CN105115910B discloses a method for detecting the distribution of protein content in peanuts based on hyperspectral imaging technology, including: collecting spectral images of peanut samples at characteristic wavelengths, inputting the spectral reflectance values ​​at characteristic wavelengths after preprocessing into the quantitative model of peanut protein content distribution, and obtaining the distribution of protein content in peanut samples. The method for establishing a quantitative model of protein content in peanuts includes collecting hyperspectral images of peanuts and determining their protein content using conventional methods; the hyperspectral images are subjected to image correction and background deletion to extract the average spectrum; the average spectrum after preprocessing is used as the independent variable and the protein content is used as the dependent variable to establish a mathematical model of full-band protein content, and on this basis, the regression coefficient is used to determine the characteristic wavelength, and the quantitative model is established and tested. Although the above scheme achieves the inversion estimation of peanut biomass, the quantitative model only uses the average spectrum as the independent variable, which easily ignores important spectral information, cannot guarantee the accurate identification of target features, and is difficult to guarantee the accuracy of the model. Therefore, in order to improve the inversion estimation accuracy of peanut biomass, effective screening of the spectral characteristics of peanut samples is the main research direction of peanut biomass monitoring. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art, solve the problem that peanut biomass characteristics cannot be accurately identified and the inversion accuracy is low, and provide a peanut biomass inversion method based on feature extraction and screening.

[0005] In order to achieve the above object, the present invention adopts the following technical solution:

[0006] A peanut biomass inversion method based on feature extraction and screening comprises the following steps:

[0007] S1: Use a drone equipped with a hyperspectral camera to collect hyperspectral image data of sample peanuts;

[0008] S2: Obtain the measured biomass of sample peanut;

[0009] S3: preprocessing the collected hyperspectral image data of the sample peanuts to obtain initial spectral reflectance data;

[0010] S4: feature extraction and screening of initial spectral reflectance data;

[0011] S4.1: Perform wavelet transform on the initial spectral reflectance data to extract multi-scale features in the initial spectral reflectance data to obtain multi-scale spectral reflectance data;

[0012] S4.2: Perform mathematical transformations on multiscale spectral reflectance data;

[0013] S4.3: Based on historical data, select the peanut vegetation index and calculate the peanut vegetation index data based on the multi-scale spectral reflectance data;

[0014] S4.4: Combined with the measured biomass of the sample peanut, the initial spectral reflectance data, the multi-scale spectral reflectance data after mathematical transformation, and the peanut vegetation index data are subjected to feature screening using variable projection importance to obtain the screened spectral features;

[0015] S5: Using the screened spectral features, a peanut biomass inversion model was constructed through a random forest model optimized by particle swarm optimization, a back propagation neural network, and a support vector machine. The measured biomass of sample peanuts was used as a data set to train and test the peanut biomass inversion model.

[0016] The present invention extracts the spectral reflectance data of ground sampling points through unmanned aerial vehicle hyperspectral images, and constructs the spectrum and vegetation index after mathematical transformation as feature input. Subsequently, the variable projection importance (VIP) method is used to screen the input features. The peanut biomass inversion model is established by using the machine learning methods of support vector machine (SVM), back propagation neural network (BPNN) and random forest (RF), and the model is optimized by particle swarm optimization algorithm to realize the inversion model for fast and accurate prediction of peanut biomass.

[0017] Preferably, the step of preprocessing the hyperspectral image data of the collected peanut samples in S3 includes: S3.1: performing correction preprocessing on the hyperspectral image data of the peanut samples, including radiation correction, atmospheric correction and geometric correction; S3.2: using a Savitzky-Golay filter to perform smoothing preprocessing on the hyperspectral impact data after the correction preprocessing to obtain initial spectral reflectance data. The Savitzky-Golay filter is used for smoothing to improve the signal-to-noise ratio of the spectrum and reduce the influence of external interference, so as to eliminate the unevenness or burrs on the spectral curve caused by the hyperspectral sensor being affected by external environmental factors or the sensor itself during the collection of spectral data.

[0018] Preferably, the wavelet transform in S4.1 adopts a continuous wavelet transform whose wavelet mother function is a Gaussian-4 function, and performs a convolution operation on the initial spectral reflectance data based on the translation and scaling of the wavelet mother function, and its expression is:

[0019]

[0020] Where: W f (a, b) are wavelet coefficients, f(λ) is the initial spectral reflectance data, λ is the spectral band, ψ a,b The Gaussian-4 function is used, where a is the scale factor and b is the translation factor.

[0021] The initial spectral reflectance data is decomposed into multiple scales by continuous wavelet transform to extract the implicit effective information in the initial spectral reflectance.

[0022] Preferably, the mathematical transformation in S4.2 comprises a first order differential.

[0023] Preferably, the peanut vegetation index includes a modified ground chlorophyll index and a bimodal canopy nitrogen index, and the calculation formula is:

[0024]

[0025] Where MMTCI stands for Modified Ground Chlorophyll Index, DCNI stands for Bimodal Canopy Nitrogen Index, and R 750 , R 710 , R 700 , R 680 , R 670 Represents the multi-scale spectral reflectance data at wavelengths of 750nm, 710nm, 700nm, 680nm, and 670nm respectively.

[0026] Preferably, the step of using variable projection importance to perform feature screening on the initial spectral reflectance data, the multi-scale spectral reflectance data after first-order differentiation, and the peanut vegetation index data in S4.4 includes:

[0027] S4.4.1: Use the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index as independent variables, and the measured biomass of the sample peanut as the dependent variable, and perform partial least squares regression to extract the principal components.

[0028] S4.4.2: Calculate the correlation coefficient between each principal component and the dependent variable.

[0029] S4.4.3: Obtain the importance of the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index for the peanut biomass inversion model fitting through the importance calculation in the variable projection importance. The importance VIP j The calculation formula is:

[0030]

[0031] In the formula: k is the number of independent variables; j is the index of the independent variable, y is the dependent variable, m is the number of principal components extracted, c is the h is the hth extracted principal component; r(y,c h ) is the correlation coefficient between the dependent variable and the hth principal component, w hj is the weight of the independent variable on the principal component.

[0032] The present invention uses the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differentiation, and the reflectance data of the peanut vegetation index as independent variables, so as to more comprehensively and meticulously depict the key information of the peanut biomass reflectance spectrum and improve the accuracy of model establishment. The independent variables (initial spectral reflectance data, multi-scale spectral reflectance data after the first-order differentiation, and reflectance data of the peanut vegetation index) are weighed and calculated through the variable projection importance to explain the dependent variable (the measured biomass of the sample peanut), and the variables are screened.

[0033] By screening features and making full use of hyperspectral data, the dimensions that affect the data are reduced, and the efficiency and accuracy of data processing are improved, thereby achieving accurate identification and quantitative inversion of target features and reducing the problems of data redundancy and dimensionality disaster caused by the improvement of hyperspectral resolution.

[0034] S4.4.4: Use importance values ​​to perform feature screening on the initial spectral reflectance data, the multi-scale spectral reflectance data after first-order differentiation, and the reflectance data of the peanut vegetation index to obtain the screened spectral features.

[0035] Preferably, the particle swarm optimization algorithm in S5 sets the learning factors to 1.5 and 1.7, the inertia weight to 0.7, and uses the root mean square error function as the fitness function.

[0036] Preferably, the step of using the measured biomass of the sample peanut as a data set to train and test the peanut biomass inversion model in S5 includes: randomly selecting n sample points from the measured biomass of the sample peanut as a test set and the remaining sample points as a training set, using the data in the test set to test the trained model, and the evaluation results are output as the determination coefficient and the root mean square error, and the calculation formula is as follows:

[0037]

[0038] Where: R 2 is the determination coefficient, RMSE is the root mean square error, n is the number of sample points, i is the index of the sample point, is the predicted value of peanut biomass, is the average value of peanut biomass.

[0039] Preferably, the step of obtaining the measured biomass of the sample peanuts in S2 is: according to the position of the sample peanuts in the peanut field, the peanut plants within a single sampling area are harvested, the dry weight of the peanut plants is obtained after pre-processing the peanut plants, and then the dry weight of the peanut plants per unit area is obtained to obtain the measured biomass of the sample peanuts.

[0040] The present invention extracts spectral reflectance data from hyperspectral images, constructs the first-order differential of reflectance and vegetation index as feature input, and uses the variable projection importance method to screen feature bands. A random forest model, a back propagation neural network and a support vector machine model are constructed based on the combined features of initial spectral reflectance data, multi-scale spectral reflectance data after first-order differential and reflectance data of peanut vegetation index, so as to achieve the screening and feature combination of effective features of hyperspectral images, thereby improving the accuracy and precision of peanut biomass inversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention is further described in detail below with reference to the accompanying drawings.

[0042] Figure 1 is a method block diagram of the present invention;

[0043] Figure 2 is a flow chart of the present invention;

[0044] Figure 3 is the initial spectral reflectance spectrum of the present invention;

[0045] Figure 4 It is a multi-scale spectral reflectance spectrum after first-order differentiation of the present invention;

[0046] Figure 5 It is the importance value ranking result of the present invention. DETAILED DESCRIPTION

[0047] like Figure 1 , Figure 2 As shown, the present invention provides a peanut biomass inversion method based on feature extraction and screening, comprising the following steps: S1: using an unmanned aerial vehicle equipped with a hyperspectral camera to collect hyperspectral image data of sample peanuts; S2: obtaining the measured biomass of the sample peanuts; S3: preprocessing the collected hyperspectral image data of the sample peanuts to obtain initial spectral reflectance data; S4: extracting and screening the features of the initial spectral reflectance data; S5: using the screened spectral features, a random forest model after particle swarm optimization, a back propagation neural network and a support vector machine to construct a peanut biomass inversion model, and using the measured biomass of the sample peanuts as a data set to train and test the peanut biomass inversion model.

[0048] In this embodiment, a UAV equipped with a hyperspectral imager is used for data collection, and the lateral and directional overlap rates of the images are set to 80% to ensure full coverage and high-quality stitching of the data. The image band range of the hyperspectral imager collected in real time is 400-1000nm, including 176 spectral channels, a spectral resolution of 3.5nm, and a spatial resolution of about 1.8cm / pixel.

[0049] Before the UAV takes off, a hyperspectral imager is used to capture standard black-and-white images to eliminate radiation errors and lens geometric distortion. After takeoff, gray cloth areas with reflectivity of 30% and 50% are captured, and all images are calibrated for radiation.

[0050] In this embodiment, the step of obtaining the measured biomass of the sample peanut in S2 is: according to the position of the sample peanut in the peanut field, the peanut plants in a single sampling area are harvested, and after the peanut plants are pre-treated, the dry weight of the peanut plants is obtained, and then the dry weight of the peanut plants per unit area is obtained to obtain the measured biomass of the sample peanut. Specifically, N ground sampling points are selected in different varieties of test fields, and the spatial position coordinates of each point are located and recorded using a GNSS receiver. At the sampling point position, all peanut plants in a single sampling area are harvested. In this embodiment, the single sampling area is set to 0.5*0.5m. The harvested peanut plants are pre-treated, and the pre-treatment includes removing the soil, sealing the peanut plants in bags and taking them back to the laboratory for cleaning. The cleaned sample peanut plants are naturally air-dried and oven-dried. After the weight does not change significantly, they are weighed and converted to a unit area as the measured biomass of the sample peanuts at each sampling point.

[0051] In this embodiment, the step of preprocessing the collected hyperspectral image data of the sample peanuts in S3 is: S3.1: Correction preprocessing is performed on the hyperspectral image data of the sample peanuts, including radiation correction, atmospheric correction and geometric correction. Specifically, after the hyperspectral image acquisition is completed, the data preprocessing software in the imager is used to preprocess all image data, and the spatial resolution of the preprocessed image is about 5cm. Subsequently, the canopy spectral reflectance of the image is extracted using remote sensing image processing software, such as ENVI software. The region of interest (ROI for short) is divided according to the position coordinates of each ground sampling point, the spectral reflectance of all pixels in the region of interest is extracted and the average value is calculated, thereby obtaining the canopy spectral reflectance of each sampling position.

[0052] S3.2: Use Savitzky-Golay filter to smooth the hyperspectral impact data after correction preprocessing to obtain initial spectral reflectance data. Specifically, the SG filter parameters with a window size of 13 and an order of 5 are used. Convolution smoothing is performed using the Savitzky-Golay filter (SG filter for short) to improve the signal-to-noise ratio of the spectrum and reduce the impact of external interference. The reflectance images before and after smoothing preprocessing are shown in Figure 2. Figure 3 and Figure 4 shown.

[0053] In this embodiment, the steps of feature extraction and screening of the initial spectral reflectance data in S4 are as follows: S4.1: Perform wavelet transform on the initial spectral reflectance data, extract multi-scale features in the initial spectral reflectance data, and obtain multi-scale spectral reflectance data. Specifically, the wavelet transform uses a continuous wavelet transform whose wavelet mother function is a Gaussian-4 function, performs a convolution operation on the initial spectral reflectance data based on the translation and scaling of the wavelet mother function, and decomposes the hyperspectral information into wavelet energy coefficients at different scales and different band lengths without changing the original band range and position. The expression is:

[0054]

[0055] Where: W f (a, b) are wavelet coefficients, f(λ) is the initial spectral reflectance data, λ is the spectral band, ψ a,b The Gaussian-4 function is used, where a is the scale factor and b is the translation factor.

[0056] In this embodiment, at a scale of 1 to 10 (2 1 , 2 2 , 2 3 , 2 4 , 2 5 , 2 6 , 27 , 2 8 , 2 9 , 2 10 ) is decomposed to extract the deep implicit information in the initial spectral reflectance data.

[0057] S4.2: Perform mathematical transformation on the multi-scale spectral reflectance data; the mathematical transformation includes first-order differential. Specifically, the calculation formula of the first-order differential is as follows:

[0058]

[0059] Among them, BF i refers to the reflectivity value of the i-th band, B i refers to the wavelength value of the i-th band, DF i It refers to the first-order differential eigenvalue of the i-th band.

[0060] S4.3: Based on historical data, the peanut vegetation index is selected and the peanut vegetation index data is calculated based on the multi-scale spectral reflectance data; the peanut vegetation index includes the modified ground chlorophyll index and the bimodal canopy nitrogen index, and its calculation formula is:

[0061]

[0062] Where MMTCI stands for Modified Ground Chlorophyll Index, DCNI stands for Bimodal Canopy Nitrogen Index, and R 750 , R 710 , R 700 , R 680 , R 670 Represents the multi-scale spectral reflectance data at wavelengths of 750nm, 710nm, 700nm, 680nm, and 670nm respectively.

[0063] In this embodiment, by constructing the vegetation index, not only the interference information is reduced, but also the key information of the vegetation is highlighted, which can reflect the deep-level characteristics of the vegetation.

[0064] S4.4: Combined with the measured biomass of the peanut sample, the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the peanut vegetation index data are subjected to feature screening using variable projection importance to obtain the screened spectral features.

[0065] The specific steps include: S4.4.1: taking the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differentiation, and the reflectance data of the peanut vegetation index as independent variables, and taking the measured biomass of the sample peanut as the dependent variable, performing partial least squares regression to extract the principal components. In this embodiment, each principal component includes each independent variable and the weight of each independent variable in the principal component.

[0066] S4.4.2: Calculate the correlation coefficient between each principal component and the dependent variable.

[0067] S4.4.3: Obtain the importance value (referred to as VIP value) of the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index for the peanut biomass inversion model fitting through the importance calculation in the variable projection importance. j The calculation formula is:

[0068]

[0069] In formula (6), k is the number of independent variables; j is the index of the sampling point, y is the dependent variable, m is the number of principal components extracted, and c is the h is the hth extracted principal component; r(y,c h ) is the correlation coefficient between the dependent variable and the hth principal component, w hj is the weight of the independent variable on the principal component.

[0070] The present invention uses the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index as independent variables, so as to more comprehensively and meticulously characterize the key information of the peanut biomass reflectance spectrum and improve the accuracy of model establishment.

[0071] The explanatory power of independent variables (initial spectral reflectance data, multi-scale spectral reflectance data after first-order differential, and reflectance data of peanut vegetation index) on the dependent variable (measured biomass of sample peanuts) was measured and calculated through variable projection importance, and the variables were screened to construct a high-precision peanut biomass inversion model.

[0072] In this embodiment, when calculating the VIP value of the initial spectral reflectance data, the number of independent variables k is 176, and the maximum value of j is the total number of sampling points N. When calculating the VIP value of the multi-scale spectral reflectance data after the first-order differential, the number of independent variables k is 175, and the maximum value of j is the total number of sampling points N. When calculating any vegetation index in the peanut vegetation index, the number of independent variables k is 175, and the maximum value of j is the total number of sampling points N.

[0073] S4.4.4: Use importance values ​​to perform feature screening on the initial spectral reflectance data, the multi-scale spectral reflectance data after first-order differentiation, and the reflectance data of the peanut vegetation index, respectively, to obtain the screened spectral features. Specifically, select features with higher importance values ​​or remove features with lower importance values. By screening the features, the hyperspectral data is fully utilized, the dimensions that affect the data are reduced, and the efficiency and accuracy of data processing are improved, thereby achieving accurate identification and quantitative inversion of target features, and reducing the problems of data redundancy and dimensionality disaster caused by the improvement of hyperspectral resolution.

[0074] In this embodiment, in order to screen out key characteristic variables, the variable projection importance method is used to evaluate the correlation and sort the VIP values ​​of the three data features. The peanut biomass inversion model is constructed using three machine learning algorithms: support vector machine (SVM), back propagation neural network (BPNN) and random forest (RF), and the hyperparameters of these three models are optimized by particle swarm optimization (PSO) algorithm. The inversion estimation of peanut biomass in the test area is realized by using the most relevant feature combination and model.

[0075] In this embodiment, the particle swarm optimization algorithm in S5 sets the learning factors to 1.5 and 1.7, the inertia weight to 0.7, and uses the root mean square error function as the fitness function. By using the particle swarm optimization algorithm, the optimal parameter combination of the model is found to improve the performance and prosperity of the model.

[0076] The present invention extracts spectral reflectance data from hyperspectral images, constructs the first-order differential of reflectance and vegetation index as feature input, and uses the variable projection importance method to screen feature bands. A random forest model, a back propagation neural network and a support vector machine model are constructed based on the combined features of initial spectral reflectance data, multi-scale spectral reflectance data after first-order differential and reflectance data of peanut vegetation index, so as to achieve the screening and feature combination of effective features of hyperspectral images, thereby improving the accuracy and precision of peanut biomass inversion.

[0077] In this embodiment, the step of using the measured biomass of the sample peanut as a data set to train and test the peanut biomass inversion model in S5 includes: randomly selecting n sample points from the measured biomass of the sample peanut as a test set and the remaining sample points as a training set, using the data in the test set to test the trained model, and the evaluation results are output as the determination coefficient and the root mean square error. The calculation formula is as follows:

[0078]

[0079] Where: R 2 is the determination coefficient, RMSE is the root mean square error, n is the number of sample points, i is the index of the sample point, is the predicted value of peanut biomass, is the average value of peanut biomass. 2 The closer it is to 1 and the smaller the RMSE is, the higher the model accuracy is.

[0080] In this embodiment, when testing the peanut biomass inversion model, the VIP values ​​of the 176 original reflectance bands in the hyperspectral image, the 175 characteristic bands after the first-order differential processing, and the vegetation index were respectively calculated with the measured peanut biomass to screen out effective features with higher VIP values. The VIP value ranking results are as follows: Figure 5 The calculated features with greater correlation and VIP values ​​are shown in Table 1.

[0081] Table 1 Features with high correlation and VIP values

[0082]

[0083] In the initial spectral reflectance data, the VIP values ​​of the two bands with wavelengths of 748.8nm and 745.3nm are the highest, reaching 1.42 and 1.40 respectively. In the multi-scale spectral reflectance data after the first-order differential, the VIP values ​​of DF99 and DF100 are the highest, reaching 2.57 and 2.56 respectively; wherein DF99 and DF100 are calculated according to formula (3) from the adjacent bands 734.8nm and 731.3nm and 731.3nm and 727.8nm in the initial spectral reflectance data. In the peanut vegetation index data, the VIP values ​​of the bimodal canopy nitrogen index (DCNII) and the modified ground chlorophyll index (MMTCI) are the highest, reaching 1.59 and 1.52 respectively. Through correlation analysis, the bands and spectral reflectance ranges of the effective characteristics of peanut biomass are determined, and the bands in this embodiment are concentrated at 700-750nm.

[0084] The initial spectral reflectance data after screening, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index were used to construct the peanut biomass inversion model based on three machine learning methods of particle swarm optimization. Specifically, according to the VIP calculation results, the top 1% of features were selected as input factors, or 2 vegetation indices were selected, and the peanut biomass inversion models based on three machine learning methods of particle swarm optimization were constructed respectively. The accuracy results of training and testing are shown in Table 2. Among them, the particle swarm optimized random forest model PSO-RF constructed by the multi-scale spectral reflectance data after the first-order differential and the reflectance data of the peanut vegetation index has the highest accuracy.

[0085] Table 2. Comparison of the accuracy of peanut biomass inversion models constructed by three machine learning methods based on particle swarm optimization

[0086]

[0087] The initial spectral reflectance data after screening, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index were combined, and the peanut biomass inversion model was constructed using the machine learning method after particle swarm optimization. The two data with the highest VIP value in each type of feature were combined as model input. The accuracy results of training and testing are shown in Table 3. Among them, the random forest model RF constructed after the feature combination of the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index has the highest accuracy, and the determination coefficient R on the training set and the test set is 1. 2 The RSE and RMSE were 0.846 and 0.800, respectively, and the RMSE were 0.073 and 0.076, respectively. Therefore, compared with the biomass inversion model based on a single type of feature, the performance of the three machine learning models constructed using combined features was significantly improved.

[0088] By using the particle swarm optimization algorithm (PSO) to optimize the hyperparameters of the three machine learning models, these models can search for the optimal solution globally, thereby improving the accuracy of the model and alleviating the negative impact of feature redundancy to a certain extent.

[0089] Table 3. Comparison of the accuracy of peanut biomass inversion models based on three machine learning methods based on particle swarm optimization

[0090]

[0091] The above are only preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope defined by the claims.

Claims

1. A peanut biomass inversion method based on feature extraction and screening, characterized in that: The following steps are involved: S1: Use a drone equipped with a hyperspectral camera to collect hyperspectral image data of sample peanuts; S2: Obtain the measured biomass of sample peanut; S3: preprocessing the collected hyperspectral image data of the sample peanuts to obtain initial spectral reflectance data; S4: feature extraction and screening of initial spectral reflectance data; S4.1: Perform wavelet transform on the initial spectral reflectance data to extract multi-scale features in the initial spectral reflectance data to obtain multi-scale spectral reflectance data; S4.2: Perform mathematical transformations on multiscale spectral reflectance data; S4.3: Based on historical data, select the peanut vegetation index and calculate the peanut vegetation index data based on the multi-scale spectral reflectance data; S4.4: Combined with the measured biomass of the sample peanut, the initial spectral reflectance data, the multi-scale spectral reflectance data after mathematical transformation, and the peanut vegetation index data are subjected to feature screening using variable projection importance to obtain the screened spectral features; S5: Using the screened spectral features, a peanut biomass inversion model was constructed through a random forest model, back propagation neural network and support vector machine after particle swarm optimization. The measured biomass of sample peanuts was used as a data set to train and test the peanut biomass inversion model.

2. The peanut biomass inversion method based on feature extraction and screening according to claim 1 is characterized in that: The step of preprocessing the collected hyperspectral image data of the sample peanuts in S3 includes: S3.1: Correction preprocessing is performed on the hyperspectral image data of the sample peanuts, including radiation correction, atmospheric correction and geometric correction; S3.2: The Savitzky-Golay filter is used to smooth the hyperspectral impact data after correction preprocessing to obtain the initial spectral reflectance data.

3. The peanut biomass inversion method based on feature extraction and screening according to claim 1, characterized in that: The wavelet transform described in S4.1 uses a continuous wavelet transform whose wavelet mother function is a Gaussian-4 function, and performs a convolution operation on the initial spectral reflectance data based on the translation and scaling of the wavelet mother function. The expression is: Where: W f (a, b) are wavelet coefficients, f(λ) is the initial spectral reflectance data, λ is the spectral band, ψ a,b The Gaussian-4 function is used, where a is the scale factor and b is the translation factor.

4. The peanut biomass inversion method based on feature extraction and screening according to claim 1, characterized in that: The mathematical transformation described in S4.2 includes first-order differentiation.

5. The peanut biomass inversion method based on feature extraction and screening according to claim 1, characterized in that: The peanut vegetation index includes the modified ground chlorophyll index and the bimodal canopy nitrogen index, and its calculation formula is: Where MMTCI stands for Modified Ground Chlorophyll Index, DCNI stands for Bimodal Canopy Nitrogen Index, and R 750 , R 710 , R 700 , R 680 , R 670 Represents the multi-scale spectral reflectance data at wavelengths of 750nm, 710nm, 700nm, 680nm, and 670nm respectively.

6. The peanut biomass inversion method based on feature extraction and screening according to claim 4, characterized in that: The steps of using variable projection importance to perform feature screening on the initial spectral reflectance data, the multi-scale spectral reflectance data after mathematical transformation, and the peanut vegetation index data in S4.4 include: S4.4.1: The initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index are used as independent variables, and the measured biomass of the sample peanut is used as the dependent variable to perform partial least squares regression to extract the principal components; S4.4.2: Calculate the correlation coefficient between each principal component and the dependent variable; S4.4.3: Obtain the importance of the initial spectral reflectance data, the multi-scale spectral reflectance data after the first-order differential, and the reflectance data of the peanut vegetation index for the peanut biomass inversion model fitting through the importance calculation in the variable projection importance. VIP j The calculation formula is: In the formula: k is the number of independent variables; j is the index of the independent variable, y is the dependent variable, m is the number of principal components extracted, c is the h is the hth extracted principal component; r(y,c h ) is the correlation coefficient between the dependent variable and the hth principal component, w hj is the weight of the independent variable on the principal component; S4.4.4: Use importance values ​​to perform feature screening on the initial spectral reflectance data, the multi-scale spectral reflectance data after first-order differentiation, and the reflectance data of the peanut vegetation index to obtain the screened spectral features.

7. The peanut biomass inversion method based on feature extraction and screening according to claim 1, characterized in that: The particle swarm optimization algorithm described in S5 sets the learning factors to 1.5 and 1.7, the inertia weight to 0.7, and uses the root mean square error function as the fitness function.

8. The peanut biomass inversion method based on feature extraction and screening according to claim 1, characterized in that: The steps of using the measured biomass of sample peanuts as a data set to train and test the peanut biomass inversion model in S5 include: Randomly select n sample points from the measured biomass of sample peanuts as the test set and the remaining sample points as the training set. Use the data in the test set to test the trained model. The evaluation results are output as the determination coefficient and root mean square error. The calculation formula is as follows: Where: R 2 is the determination coefficient, RMSE is the root mean square error, n is the number of sample points, i is the index of the sample point, is the predicted value of peanut biomass, is the average value of peanut biomass.

9. The peanut biomass inversion method based on feature extraction and screening according to claim 1, characterized in that: The steps for obtaining the measured biomass of the sample peanuts in S2 are: according to the location of the sample peanuts in the peanut field, the peanut plants within a single sampling area are harvested, the dry weight of the peanut plants is obtained after pre-processing, and then the dry weight of the peanut plants per unit area is obtained to obtain the measured biomass of the sample peanuts.

Citation Information

Patent Citations

  • Method for detection of protein content distribution in peanut based on hyperspectral imaging technology

    CN105115910B