Canopy chlorophyll content inversion method based on feature extraction and machine learning model
Through the combination of drone and ground data, a chlorophyll content inversion model based on machine learning was constructed, which solved the problem of insufficient research on chlorophyll content in different growth periods, and achieved efficient and accurate chlorophyll monitoring, which was suitable for the growth trend assessment of potatoes and other crops.
Patent Information
- Application Number
- CN202510092587.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing technology lacks research on the inversion of chlorophyll content for different growth periods during potato growth, resulting in insufficient applicability and accuracy of the model. The traditional measurement methods are costly and greatly affected by environmental factors, making it difficult to apply on a large scale.
The UAV equipped with a multi-spectral camera combined with ground measured data, and features are selected through Pearson's correlation analysis and competitive adaptive reweighted sampling algorithm, a chlorophyll content inversion model based on machine learning is constructed, and the model parameters are optimized using gray wolf optimization and sparrow optimization algorithm, and the inversion is combined with random forest, support vector regression and extreme gradient enhancement algorithm.
It significantly improves the accuracy of inversion of chlorophyll content in different growth periods, enhances the applicability and accuracy of the model, reduces costs, and is suitable for large-scale crop monitoring.
Smart Images

Figure CN120014295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chlorophyll content inversion, and in particular to a canopy chlorophyll content inversion method based on feature extraction and machine learning models. Background Art
[0002] As the world's fourth-largest food crop, potatoes possess significant economic and social value. my country, with a potato planting area of 4.218 million hectares, holds a significant position in global potato production. Chlorophyll plays a crucial role in potato growth, as it is a key component of photosynthesis. Chlorophyll is fundamental to the growth and development of crops like potatoes, making it a key indicator for assessing potato growth.
[0003] Traditionally, chlorophyll content is typically measured using chemical instruments. While this method is flexible, it's costly and affected by factors like light and temperature, making it unsuitable for large-scale crop monitoring. Chlorophyll value meters are widely used in outdoor field environments to measure the relative chlorophyll content of leaf canopies. In recent years, they have garnered increasing attention for crop growth monitoring. Unmanned aerial vehicles (UAVs) equipped with multispectral cameras have red-edge and near-infrared detection capabilities. Compared to conventional cameras, they offer efficient and rapid data processing and are relatively inexpensive, making them widely used in crop yield prediction.
[0004] With the development of artificial intelligence, the application of machine learning (ML) algorithms has become more and more extensive. Machine learning models can provide support for the study of canopy chlorophyll content and improve the inversion accuracy. Although previous studies have shown that the combination of drone multispectral images and machine learning models significantly improves the accuracy of chlorophyll content inversion and has been widely used in the study of various crops, most studies have only focused on a single growth period and lack comparative studies for different growth periods. Different potato varieties have significant differences in their impact on agricultural research models. Therefore, the present invention collected drone and ground data of sixteen varieties of potatoes and conducted comparative studies on the single growth period and the full growth period of potatoes. Summary of the Invention
[0005] The purpose of the present invention is to provide a canopy chlorophyll content inversion method based on feature extraction and machine learning models, construct an efficient chlorophyll content inversion model, combine the characteristics of each developmental stage with measured data, develop a model combination of different feature selection methods and parameter optimization, and significantly improve the accuracy of model inversion; conduct a comparative study on the single growth period and the entire growth period of potatoes, determine the optimal inversion model suitable for the single growth period and the entire growth period, and ensure the accuracy and applicability of the research.
[0006] To achieve the above objectives, the present invention provides a canopy chlorophyll content inversion method based on feature extraction and machine learning model, comprising the following steps:
[0007] Step S1: using a combination of drone-mounted multispectral cameras and ground measurements to collect and process potato canopy chlorophyll content data in the study area, and construct a vegetation index;
[0008] Step S2, using Pearson correlation analysis and competitive adaptive reweighted sampling CARS algorithm to perform feature selection on vegetation index;
[0009] Step S3: constructing a chlorophyll content inversion model based on machine learning;
[0010] Step S4: Optimizing the chlorophyll content inversion model constructed in step S3 based on the gray wolf optimization algorithm and the sparrow optimization algorithm;
[0011] Step S5: perform accuracy evaluation on the chlorophyll content inversion model optimized in step S4.
[0012] Preferably, in step S1, a combination of a multispectral camera mounted on an unmanned aerial vehicle and ground measurement is used to collect and process the chlorophyll content data of potato canopies in the study area, and a vegetation index is constructed. The specific process is as follows:
[0013] Step S11: The UAV carries a multispectral camera to acquire and process image data;
[0014] First, an unmanned aerial vehicle (UAV) was used to collect image data of the study area. The UAV was equipped with four multispectral sensor probes, corresponding to the spectral bands of green light (GREEN), red light (RED), red edge light (REG), and near-infrared light (NIR).
[0015] Then, the collected images were input into Pix4Dmapper software for multispectral image synthesis and calibration, thereby obtaining orthorectified multispectral reflectance images;
[0016] Step S12: collecting and measuring ground data;
[0017] First, the study area was divided into 16 plots, each corresponding to one potato variety. Ground data for a total of 16 potato varieties were collected. Each plot was divided into 20m × 100m, and data from 10 sample points were collected evenly in each plot. Leaf canopy chlorophyll content was collected at the potato planting bases in the study area at both the seedling and mature stages, with 160 ground data points collected for each growth period.
[0018] Then, the 16 sample plots divided above were used as objects. Ten well-growing potato canopies were evenly selected as sample points in each sample plot. When measuring each leaf canopy, the veins were avoided. Ten sample points were evenly collected at the mesophyll using a handheld chlorophyll meter. The average value was taken as the true value of one sample point. The readings obtained from the potato plants selected at the sample points of each sample plot represented the relative chlorophyll content (SPAD) of the sample plot.
[0019] Step S13: constructing a vegetation index based on the acquired and processed data;
[0020] The geometrically corrected drone orthophoto images were divided into 16 corresponding areas according to potato varieties. Ten sample points were taken in each area. Combined with the latitude and longitude coordinates of the ground measurement points, the reflectance of the four bands of GREEN, RED, REG and NIR corresponding to each sample point was extracted on ArcGIS. Vegetation indices were constructed based on these reflectances to enhance the physiological characteristics of vegetation.
[0021] Preferably, in step S11, 20 reference points of a uniform white plate are selected to perform geometric correction on the image, and the reflectance of the reference points is converted into the reflectance of the multispectral image using a pseudo-standard geometric correction method, so as to facilitate the extraction of accurate reflectance of the multispectral band. The correction formula is as follows:
[0022]
[0023] Among them, DN i is the reflectance value of the uniform white plate area; is the original mean before correction; is the average value of a uniform whiteboard area; DN i,corrected is the reflectance value after correction.
[0024] Preferably, in step S2 vegetation index feature selection, the Pearson feature selection method is used to measure the linear relationship between vegetation index and potato canopy SPAD, and the Pearson correlation coefficient r X,Y The calculation formula is as follows:
[0025]
[0026] Where n is the number of samples; x i and y i is the i-th data value of variables X and Y; and are the means of variables X and Y respectively.
[0027] Preferably, in step S2 vegetation index feature selection, a competitive adaptive reweighted sampling algorithm is used to evaluate the relevance of each vegetation index to the potato canopy SPAD prediction model through multiple sampling and analysis. The specific process is as follows:
[0028] Step S221, Monte Carlo sampling: Assume that the original vegetation index data set has n samples, the number of samples for each sampling is m (m≤n), and the number of sampling times is k; for the i-th sampling (i=1,2,…,k), the obtained sampling subset S i Contains m samples randomly drawn from the original data set;
[0029] Step S222: Establish a partial least squares regression PLS model: for the subset S obtained by the i-th sampling i , establish the PLS model; let X i is the vegetation index matrix m×p in the subset, where p is the number of vegetation indices; Y i is the corresponding target variable vector m×1; the PLS model is implemented by i and Y i Decompose and regress to obtain the regression coefficient vector β i , as shown below:
[0030] X=TP T +E (3);
[0031] Y=UQ T +F (4);
[0032] Where T and U are the score matrices of X and Y respectively; P and Q are the loading matrices; E and F are the residual matrices;
[0033] Regression coefficient β i , as shown below:
[0034] β i =W(P T W) -1 (5);
[0035] Where W is the weight matrix, which is determined during the iteration process of the PLS algorithm;
[0036] Step S223: Calculate the vegetation index importance index;
[0037] Assume ω i,j is the weight of the jth vegetation index in the PLS model of the i-th sampling, and the importance index I of the jth vegetation index is calculated. j , as shown below:
[0038]
[0039] Where n is the number of sampling times; a is the empirical index, where a = 1, and the vegetation index is selected according to the importance index.
[0040] Preferably, in step S3, a chlorophyll content inversion model is constructed based on machine learning, and the specific process is as follows:
[0041] Step S31, random forest (RF): An RF model is established to conduct an inversion experiment on the chlorophyll content of potato canopies. A ten-fold cross-validation method is used to continuously change the selection of the test set so that each subset has the opportunity to become a test set, and the performance of the model on different data subsets is evaluated;
[0042] The parameters of the RF model were optimized using different optimization algorithms. In each round of cross-validation, the model was trained using different parameter combinations, and the model performance was evaluated on the test set to obtain an ideal chlorophyll content inversion model.
[0043] Step S32, support vector regression SVR: establish an SVR model to conduct an inversion study on the chlorophyll content of potato canopy, and use the RBF kernel function to fit the data. Its core expression is as follows:
[0044] K(x i ,x j )=exp(-γ||x i -x j || 2 ) (7);
[0045] Among them, x i and x j is the input sample; γ is the parameter of the kernel function, which determines the distribution range of the sample in the feature space;
[0046] During the training process, a parameter optimization algorithm is used to continuously optimize parameters based on training data, and different hyperparameter combinations are evaluated to achieve chlorophyll content prediction and obtain a chlorophyll content inversion model;
[0047] Step S33, extreme gradient boosting XGBoost: The potato chlorophyll dataset is divided into a training set and a test set in a ratio of 8:2. During the training phase, the XGBoost model uses the training set as the basis, the input features, and the target variable, and uses the gradient boosting strategy to iteratively train the decision tree.
[0048] In each iteration, the XGBoost model first calculates the gradient of the loss function with respect to the predicted value, and then fits a new decision tree based on the obtained gradient information. The newly generated decision tree learns and corrects the errors generated by the previous decision tree in the prediction process, gradually improving the prediction performance and accuracy of the entire model, and obtaining the optimal chlorophyll content inversion model.
[0049] Preferably, in step S4, the parameters of the RF, SVR and XGBoost chlorophyll content inversion models are optimized based on the Gray Wolf Optimization Algorithm to invert the chlorophyll content of the potato canopy. The specific steps are as follows:
[0050] Step S411: inputting drone image data and measured potato canopy chlorophyll content;
[0051] Step S412: using the Pearson correlation analysis method and the CARS algorithm to perform feature selection on the constructed vegetation index;
[0052] Step S413: Establish RF, SVR and XGBoost chlorophyll inversion models;
[0053] Step S414: Initialize the gray wolf population and parameters, and define the fitness function of the gray wolf optimization algorithm to evaluate the quality of each parameter combination;
[0054] Step S415: Divide the social hierarchy of the gray wolves and calculate the distances between them to determine the update direction and position, and update the optimal position and fitness of the gray wolves;
[0055] Step S416: Determine whether the stop condition is met, if so, execute step S417, otherwise execute step S415;
[0056] Step S417: output the optimal parameters of RF, SVR and XGBoost chlorophyll inversion models;
[0057] Step S418: train the chlorophyll inversion model to obtain the results of the chlorophyll content inversion model, and evaluate the model using evaluation indicators.
[0058] Preferably, in step S4, the parameters of the RF, SVR and XGBoost chlorophyll inversion models are optimized based on the sparrow optimization algorithm, and the chlorophyll content is inverted. The specific steps are as follows:
[0059] Step S421: inputting drone image data and measured chlorophyll content data;
[0060] Step S422: select and construct vegetation indices using the Pearson and CARS algorithms;
[0061] Step S423: Establish RF, SVR and XGBoost chlorophyll content inversion models;
[0062] Step S424: Initialize the position of the sparrow population and define a fitness function;
[0063] Step S425: Calculate the fitness of each sparrow according to the fitness function. The sparrow with the best fitness is considered the discoverer and leads the search direction of the population. The remaining sparrows are considered followers and update their positions based on the positions of the discoverer and themselves.
[0064] Step S426: Randomly select some sparrows as scouts to explore the new search area and update their own positions, thereby updating the optimal position and fitness of the global sparrows;
[0065] Step S427: Determine whether the termination condition is met. If so, execute step S428; otherwise, execute step S425.
[0066] Step S428: output the optimal model parameters optimized by the sparrow algorithm;
[0067] Step S429: training the potato chlorophyll content inversion model, outputting the results of the chlorophyll content inversion model, and evaluating the accuracy of the model.
[0068] Preferably, in step S5, the chlorophyll content inversion model optimized in step S4 is evaluated for accuracy, and the specific process is as follows:
[0069] The data set was randomly divided into training set and test set with a ratio of 8:2. The ten-fold cross validation method was used to evaluate the generalization ability of the model during training. The coefficient of determination R was used to 2 , root mean square error RMSE and mean absolute error MAE parameters to measure the accuracy of the model, and the calculation formula is as follows:
[0070]
[0071]
[0072]
[0073] Where n is the number of samples; x i and y i is the measured value; is the predicted value; and is the average of the measured values.
[0074] Therefore, the present invention adopts the above-mentioned canopy chlorophyll content inversion method based on feature extraction and machine learning model, and uses data from different growth periods, multiple feature extraction methods and different optimization algorithms to optimize and improve the canopy SPAD inversion model. It is found that different feature extraction methods have significant differences in inversion accuracy in different growth periods. The CARS algorithm is more effective in extracting vegetation index in two single growth periods, and the Pearson correlation coefficient method in the entire growth period is more accurate. The accuracy of the optimization models of different optimization algorithms is significantly improved, and the SSA algorithm is better in all growth periods. These application methods significantly improve the applicability and accuracy of canopy SPAD inversion research in different growth periods of potatoes.
[0075] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a flow chart of the canopy chlorophyll content inversion method based on feature extraction and machine learning model of the present invention;
[0077] Figure 2 This is the UAV image preprocessing process of the present invention;
[0078] Figure 3 It is the flow chart of the gray wolf optimization algorithm of the present invention;
[0079] Figure 4 It is the flow chart of the sparrow optimization algorithm of the present invention;
[0080] Figure 5 is a descriptive statistical graph of SPAD of potato canopies at different growth stages in an embodiment of the present invention;
[0081] Figure 6 is the CARS selection result of the RF model of the original image in the mature stage in the embodiment of the present invention;
[0082] Figure 7 : These are scatter plots of the SPAD actually measured during the seedling stage and the SPAD predicted by the random forest (RF) model in an embodiment of the present invention; wherein, (a)-(f) are scatter plots corresponding to the random forest (RF) model; (g)-(l) are scatter plots corresponding to the support vector regression (SVR) model; (m)-(r) are scatter plots corresponding to the extreme gradient boosting (XGBoost) model;
[0083] Figure 8 : These are scatter plots of the SPAD actually measured in the mature stage and the SPAD predicted by each model in the embodiment of the present invention; (a)-(f) are scatter plots corresponding to the random forest (RF) model; (g)-(l) are scatter plots corresponding to the support vector regression (SVR) model; (m)-(r) are scatter plots corresponding to the extreme gradient boosting (XGBoost) model;
[0084] Figure 9 These are scatter plots of the SPAD actually measured during the entire growth period and the SPAD predicted by each model in the embodiment of the present invention; (a)-(f) are scatter plots corresponding to the random forest (RF) model; (g)-(l) are scatter plots corresponding to the support vector regression (SVR) model; and (m)-(r) are scatter plots corresponding to the extreme gradient boosting (XGBoost) model. DETAILED DESCRIPTION
[0085] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0086] like Figure 1 As shown in Figure 2, the canopy chlorophyll content inversion method based on feature extraction and machine learning model includes the following steps:
[0087] Step S1: Using an unmanned aerial vehicle (UAV) equipped with a multispectral camera and ground measurement, the potato chlorophyll content data in the study area are collected and processed, and a vegetation index is constructed.
[0088] Step S11: The unmanned aerial vehicle (UAV) carries a multispectral camera to acquire and process image data.
[0089] First, we used an unmanned aerial vehicle (UAV) to collect image data from the study area. The UAV has a takeoff weight of 1.356 kg, a maximum horizontal flight speed of 20 m / s, and a maximum takeoff altitude of 6,000 m. The UAV is equipped with four multispectral sensors, corresponding to the green (GREEN), red (RED), red-edge (REG), and near-infrared (NIR) spectral bands. The image sensor has 12 million effective pixels.
[0090] Data collection was conducted in clear, windless weather to prevent image distortion caused by strong light and temperature. Each flight took off from the same pre-set location. A calibration standard was placed before each takeoff to ensure accurate image data collection. The altitude was set to 30m, the speed to 5m / s, the lateral angle to 70°, and the heading to 80°.
[0091] Then, the collected images are input into Pix4Dmapper software for multispectral image synthesis and calibration, so as to obtain orthorectified multispectral reflectance images. The UAV image preprocessing process is as follows: Figure 2 shown.
[0092] Twenty reference points of a uniform white plate are selected to perform geometric correction on the image. The pseudo-standard geometric correction method is used to convert the reflectance of the reference points into the reflectance of the multispectral image, which facilitates the extraction of accurate reflectance of the multispectral band. The correction formula is as follows:
[0093]
[0094] Among them, DN i is the reflectance value of the uniform white plate area; is the original mean before correction; is the average value of a uniform whiteboard area; DN i,corrected is the reflectance value after correction.
[0095] Step S12: collecting and measuring ground data.
[0096] First, the study area was divided into 16 plots, each plot corresponding to one potato variety. A total of 16 ground data of potato varieties were collected. The area of each plot was divided into 20m×100m, and data from 10 sample points were evenly collected in each plot. The leaf canopy chlorophyll content was collected from the potato planting base in the study area at the seedling stage and mature stage respectively. A total of 160 ground data were collected in each growth period.
[0097] Then, the correlation between the image data obtained by the UAV and the ground measured data was studied. After the UAV image collection, the 16 sample plots divided above were used as objects, and 10 well-growing potato canopies were evenly selected as sample points in each sample plot. When measuring each leaf canopy, the veins were avoided. A handheld chlorophyll meter was used to evenly collect 10 sample points in the mesophyll, and the average value was taken as the true value of a sample point. The readings obtained from the potato plants selected at the sample point of each sample plot represented the relative chlorophyll content of the sample plot (Soil-Plant Analysis Development, SPAD).
[0098] Step S13: constructing a vegetation index based on the acquired and processed data.
[0099] Vegetation indices can significantly reflect the changes in the true physiological characteristics of vegetation. Generally, potato canopy image data collected by drones are affected by factors such as atmospheric interference, light intensity, and internal physiological characteristics. Vegetation indices can significantly reduce these effects. Therefore, 18 vegetation indices were initially selected, as shown in Table 1.
[0100] Table 1 18 vegetation indices used for SPAD estimation in this invention
[0101] Vegetation Index name formula GRVI Green-Red Vegetation Index GRVI=(GR) / (G+R) MCARI Modified Chlorophyll Absorption Ratio Index MCARI=(REG-R)-(0.2*(REG-G))*(REG / R) DVI Difference Vegetation Index DVI=NIR-R MTVI Modified Triangular Vegetation Index 1.5*(1.2*(REG-G)-2.1*(RG)) WDRVI Wide Dynamic Range Vegetation Index WDRVI=(0.12*NIR-R) / (0.12*NIR+R) EVI2 Two-band enhanced vegetation index EVI2=2.5*(NIR-R) / (NIR+2.4*R+1) RECI Red-edge chlorophyll index RECI=(NIR / REG)-1 GCI Green chlorophyll index GCI=(NIR / G)-1 NDVI Normalized Difference Vegetation Index NDVI = (NIR-R) / (NIR+R) GNDVI Green Normalized Difference Vegetation Index GNDVI=(NIR-G) / (NIR+G) RVI Ratio Vegetation Index RVI=NIR / R NDGI Normalized Difference Greenness Index NDGI=(REG-G) / (REG+G) MSRI Modified Simple Ratio Index MSR=(NIR / R-1) / (NIR / R+1) OSAVI Optimizing Soil Adjustment Vegetation Index OSAVI = (NIR-R) / (NIR+R+0.16) SRI Simple ratio index SR=NIR / REG NDRE Normalized difference red edge index NDRE=(NIR-REG) / (NIR+REG) NLI Nonlinear vegetation index NLI = (NIR*NIR-R) / (NIR*NIR+R) TVI Triangular Vegetation Index TVI=0.5*(120*(NIR-REG)-200*(R-REG))
[0102] Among them, G represents the green band value, R represents the red band value, REG represents the red edge band value, and NIR represents the near infrared band value.
[0103] The geometrically corrected drone orthophoto images were divided into 16 corresponding areas according to potato varieties. Ten sample points were taken in each area. Combined with the latitude and longitude coordinates of the ground measurement points, the reflectance of the four bands of GREEN, RED, REG and NIR corresponding to each sample point was extracted on ArcGIS. Vegetation indices were constructed based on these reflectances to enhance the physiological characteristics of vegetation.
[0104] Step S2: Use Pearson correlation analysis and competitive adaptive reweighted sampling (CARS) algorithm to perform feature selection on the index.
[0105] Since characteristic variables such as vegetation indices usually show complexity and diversity, which leads to information redundancy and increases research errors, it is necessary to select the initially selected vegetation characteristics to determine the optimal variable characteristics for inverting potato chlorophyll content.
[0106] Step S21: Pearson feature selection method.
[0107] The Pearson feature selection method is a statistical indicator method based on the Pearson correlation coefficient, which measures the degree of linear correlation between two variables. In vegetation index selection, it can be used to measure the linear relationship between the vegetation index and the SPAD of potato canopy. If the Pearson correlation coefficient between the vegetation index and the SPAD measured in the potato canopy is close to +1 or -1, it means that there is a strong linear relationship between them. This vegetation index has important value in estimating SPAD. The Pearson correlation coefficient r X,Y The calculation formula is as follows:
[0108]
[0109] Where n is the number of samples; x i and y i is the i-th data value of variables X and Y; and are the means of variables X and Y respectively.
[0110] The vegetation index is used as the free variable X and the potato canopy SPAD is used as the target variable Y to calculate the vegetation index. X,Y The vegetation index is selected by the value of X,Y When the absolute value of is greater than a certain threshold, the variable is considered to be significantly correlated and the corresponding vegetation index is selected.
[0111] In the present invention, the threshold value selected in the seedling stage and the mature stage is 0.3, and the threshold value selected in the whole growth period is 0.7, and the selected vegetation index is used for research.
[0112] Step S22: Competitive Adaptive Reweighted Sampling (CARS) algorithm.
[0113] CARS is a feature selection method that combines Monte Carlo sampling and partial least squares regression (PLS). When selecting vegetation indices, its core concept is to evaluate the relevance of each vegetation index to the potato canopy SPAD prediction model through multiple sampling and analysis. It simulates a "competition" process, with each vegetation index vying for inclusion in the model. Vegetation indices are selected based on correlation metrics, eliminating feature variables with weak correlations to select vegetation indices with strong correlations. The specific process is as follows:
[0114] Step S221: Monte Carlo sampling.
[0115] Assume that the original vegetation index dataset has n samples, the number of samples for each sampling is m (m≤n), and the number of sampling times is k. For the i-th sampling (i=1,2,…,k), the obtained sampling subset S i It contains m samples randomly selected from the original data set. This process is mainly achieved through random sampling algorithm.
[0116] Step S222: Establish a partial least squares regression (PLS) model.
[0117] For the subset S obtained by the i-th sampling i , establish a PLS model. Let X i is the vegetation index matrix in the subset (m×p, p is the number of vegetation indices), Y i is the corresponding target variable vector (m×1). The PLS model is constructed by i and Y i Decompose and regress to obtain the regression coefficient vector β i This process involves iterative decomposition using a complex PLS algorithm, as shown below:
[0118] X=TP T +E (3);
[0119] Y=UQ T +F(4);
[0120] Where T and U are the score matrices of X and Y respectively; P and Q are the loading matrices; E and F are the residual matrices.
[0121] Regression coefficient β i , as shown below:
[0122] β i =W(P T W) -1 (5);
[0123] Among them, W is the weight matrix, which is determined during the iteration process of the PLS algorithm.
[0124] Step S223: Calculate the vegetation index importance index.
[0125] Assume ω i,j is the weight (or regression coefficient or other related indicators) of the jth vegetation index in the PLS model of the i-th sampling, and the importance index I of the jth vegetation index is calculated. j , as shown below:
[0126]
[0127] Where n is the number of sampling times; a is the empirical index, where a = 1, and the vegetation index is selected according to the importance index.
[0128] Step S3: construct a chlorophyll content inversion model based on machine learning.
[0129] Step S31: Random Forest (RF).
[0130] Random Forest (RF) is an ensemble learning algorithm that constructs multiple independent decision trees. When all decision trees are constructed, a new sample is input into each decision tree for prediction, and then the average of these predicted values is calculated to obtain the final chlorophyll content prediction result.
[0131] In the present invention, an inversion experiment of potato canopy chlorophyll content was carried out by establishing an RF model. A ten-fold cross-validation method was adopted to continuously change the selection of the test set so that each subset had the opportunity to become a test set, thereby comprehensively evaluating the performance of the model on different data subsets.
[0132] The RF model parameters were optimized using different optimization algorithms. In each round of cross-validation, the model was trained using different parameter combinations, and model performance was evaluated on the test set to obtain an ideal chlorophyll content inversion model. The tuning parameters of the RF model used in this invention are shown in Table 2.
[0133] Step S32: Support Vector Regression (SVR).
[0134] Support Vector Regression (SVR) is a regression algorithm based on the principle of Support Vector Machine (SVM). The key to this model lies in the selection of kernel function. Commonly used kernel functions include linear kernel function, polynomial kernel function, radial basis function (RBF), etc.
[0135] In this paper, an SVR model is established to invert the chlorophyll content of potato canopy, and the RBF kernel function is used to fit the data. Its core expression is as follows:
[0136] K(x i ,x j )=exp(-γ||x i -x j || 2 ) (7);
[0137] Among them, x i and x j is the input sample; γ is the parameter of the kernel function, which determines the distribution range of the sample in the feature space.
[0138] During the training process, a parameter optimization algorithm was used to continuously optimize parameters based on the training data, and different hyperparameter combinations were evaluated to enable the model to more accurately predict chlorophyll content. The parameters tuned for the SVR model in this invention are shown in Table 2.
[0139] Step S33: Extreme Gradient Boosting (XGBoost).
[0140] Extreme Gradient Boosting (XGBoost) is a machine learning algorithm based on Gradient-Boosted Decision Trees (GBDT). The XGBoost model has significantly improved learning efficiency over the former. Its core idea is to gradually add new learners (usually decision trees) to improve the predictive ability of the overall model by minimizing the loss function on the basis of the previous model.
[0141] In this paper, a potato chlorophyll dataset was divided into a training set and a test set in an 8:2 ratio. During the training phase, the XGBoost model used the training set as the basis, the input features, and the target variable, and iteratively trained the decision tree using a gradient boosting strategy. In each iteration, the XGBoost model first calculated the gradient of the loss function with respect to the predicted value, and then fitted a new decision tree based on the obtained gradient information. The purpose was to effectively reduce the loss value. The newly generated decision tree can learn and correct the errors generated by the previous decision tree during the prediction process, thereby gradually improving the prediction performance and accuracy of the entire model.
[0142] Since the XGBoost model has many parameters, the present invention uses two optimization algorithms to optimize its parameters to obtain the optimal chlorophyll content inversion model. The parameters of the XGBoost model improved in the present invention are shown in Table 2.
[0143] Table 2 Parameters and significance of RF, SVR and XGBoost models
[0144]
[0145]
[0146] Step S4: Optimize the chlorophyll content inversion machine learning model constructed in step S3 based on the gray wolf optimization algorithm and the sparrow optimization algorithm.
[0147] Step S41: Parameter optimization is performed using the Gray Wolf Optimization Algorithm.
[0148] The GreyWolfOptimizer (GWO) is a heuristic optimization algorithm based on the predation behavior of grey wolf packs. Grey wolf packs have a strict hierarchy, and the algorithm uses this hierarchy to guide the pack's search direction.
[0149] In this invention, the positions of the gray wolf population are randomly initialized. The position of each gray wolf represents a set of parameters of the machine learning model. Based on the selected evaluation indicators and machine learning model, a fitness function is defined, and the wolf pack is divided into levels and the positions are continuously updated. After the positions are updated, the fitness and relative optimal solution of each gray wolf are repeatedly calculated. Each solution can be regarded as the position of a wolf. The wolf pack continuously updates its position to approach the optimal parameters. The flowchart of the gray wolf optimization algorithm to improve the chlorophyll content inversion model is as follows: Figure 3 As shown in the figure, by improving the model, the accuracy of RF, SVR and XGBoost models is improved, and their complex parameter problems are also significantly improved.
[0150] The Gray Wolf Optimization Algorithm optimizes the parameters of the RF, SVR, and XGBoost models. The specific steps for SPAD inversion of potato canopy are as follows:
[0151] Step S411: inputting drone image data and measured potato canopy chlorophyll content.
[0152] Step S412: Use the Pearson correlation analysis method and the CARS algorithm to perform feature selection on the constructed vegetation index.
[0153] Step S413: Establish RF, SVR and XGBoost chlorophyll inversion models.
[0154] Step S414: Initialize the gray wolf population and parameters, and define the fitness function of the gray wolf optimization algorithm to evaluate the quality of each parameter combination.
[0155] Step S415: Divide the social levels of the gray wolves and calculate the distances between them to determine the update direction and position, and update the optimal position and fitness of the gray wolves.
[0156] Step S416: Determine whether the stop condition is met. If so, execute step S417; otherwise, execute step S415.
[0157] Step S417: Output the optimal parameters of the RF, SVR, and XGBoost models.
[0158] Step S418: train the inversion model to obtain the results of the chlorophyll content inversion model, and evaluate the model using evaluation indicators.
[0159] Step S42: Parameter optimization is performed using the sparrow optimization algorithm.
[0160] Sparrow Search Algorithm (SSA) is an optimization algorithm proposed to simulate the foraging and anti-predation behavior of sparrow groups.
[0161] In the chlorophyll inversion model of the present invention, a group of "sparrows" are randomly initialized, that is, corresponding to a set of model parameter combinations, and sparrows with different roles such as discoverers, followers and scouts are set. The discoverer is responsible for exploring the search space, and the follower conducts a local search around the discoverer. At the same time, a warning mechanism is introduced. When encountering danger, some sparrows will randomly move to other locations to avoid falling into the local optimum. The scout is responsible for exploring the information of the new area and providing updated information on the population position to improve the search efficiency of the population, thereby optimizing and improving the parameter accuracy of the entire chlorophyll content inversion model. The flow chart of the SSA algorithm to improve the chlorophyll content inversion model is as follows: Figure 4 shown.
[0162] The specific steps of using the Sparrow Optimization Algorithm to optimize the parameters of the RF, SVR, and XGBoost models and invert chlorophyll content are as follows:
[0163] Step S421: input drone image data and measured chlorophyll content data.
[0164] Step S422: Use the Pearson and CARS algorithms to select and construct vegetation indices.
[0165] Step S423: Establish RF, SVR and XGBoost chlorophyll content inversion models.
[0166] Step S424: Initialize the position of the sparrow population and define a fitness function.
[0167] Step S425: Calculate the fitness of each sparrow according to the fitness function. The sparrow with the best fitness is regarded as the discoverer, which leads the search direction of the population. The remaining sparrows are regarded as followers, which update their positions according to the positions of the discoverer and themselves.
[0168] Step S426: Randomly select some sparrows as scouts to explore the new search area and update their own positions, thereby updating the best position and fitness of the global sparrows.
[0169] Step S427: Determine whether the termination condition is met. If so, execute step S428; otherwise, execute step S425.
[0170] Step S428: Output the optimal model parameters optimized by the sparrow algorithm.
[0171] Step S429: training the potato chlorophyll content inversion model, outputting the results of the chlorophyll content inversion model, and evaluating the accuracy of the model.
[0172] Step S5: perform accuracy evaluation on the chlorophyll content inversion model optimized in step S4.
[0173] The dataset was randomly partitioned into a training set and a test set ratio of 8:2. To enhance model performance and prevent overfitting, a ten-fold cross-validation method was used during training to evaluate the model's generalization ability. This method allows for more comprehensive data analysis to assess model performance.
[0174] In each test, the accuracy and other evaluation indicators are calculated, and the average value of the evaluation indicators is calculated after multiple tests to evaluate the generalization ability of the model. The determination coefficient R is used 2 , root mean square error RMSE and mean absolute error MAE parameters to measure the accuracy of the model. 2 The closer it is to 1, the better the stability and accuracy of the model, and the higher the degree of fit; the smaller the RMSE value, the higher the prediction accuracy of the model; MAE is used to test the prediction ability of the model, and a lower MAE value indicates that the model has better prediction accuracy; its calculation formula is as follows:
[0175]
[0176]
[0177]
[0178] Where n is the number of samples; x i and y i is the measured value; is the predicted value; and is the average of the measured values.
[0179] Example
[0180] The potato planting base research area of this embodiment is located at 41°9'36"N, 111°36'10"E. Figure 1 As shown, the region has a temperate continental monsoon climate, with a diverse topography of plains, hills, mountains, and hills and plains, providing a diverse soil type and climatic environment. Influenced by the southeast monsoon in summer and the Siberian High in winter, the region boasts an average annual temperature of around 6°C, over 3,000 hours of sunshine per year, and up to 360 mm of annual precipitation. These specific geographical and climatic characteristics provide excellent light conditions for potato growth in the region.
[0181] The study area was divided into 16 plots, each corresponding to one potato variety. Ground data were collected for a total of 16 potato varieties. Each plot measured 20 m × 100 m, and data was collected from 10 sample points evenly distributed across each plot. A total of 160 ground data points were collected during each growing season. The drone collected data between 8:30 AM and 11:30 AM on July 10 and between 9:00 AM and 11:00 AM on August 12, under clear, windless weather conditions to protect the images from strong light and temperature. Ground measurements were conducted at potato planting sites in the study area during the seedling stage (July 10, 2024) and the mature stage (August 12, 2024).
[0182] Research image data was acquired using a multispectral camera mounted on an unmanned aerial vehicle (UAV). Pix4Dmapper software was used to stitch the multispectral images together. The resulting orthophotos were then calibrated and imported into ENVI 5.3. The collected data points were annotated and used as regions of interest (RIOs), representing the sampling points closest to the measured chlorophyll content of the potato canopy. The longitude and latitude of these RIOs were then extracted. The longitude and latitude were then entered into Arcgis software, and the reflectance of the four bands (GREEN, RED, REG, and NIR) was extracted. Based on this, a vegetation characteristic index was constructed to reflect structural changes and serve as a characteristic variable for model input.
[0183] 1. SPAD descriptive statistics of potato canopy.
[0184] The values of potato canopy SPAD at different growth stages varied slightly, as shown in Table 3. The range of canopy SPAD at the seedling stage was 40.30 to 56.00, with an average of 48.47, a standard deviation of 3.70, and a coefficient of variation of 7.63%. The range of canopy SPAD at the mature stage was 28.20 to 54.20, with an average of 43.20, a standard deviation of 5.40, and a coefficient of variation of 12.50%. The range of canopy SPAD throughout the entire growth period was 28.20 to 56.00, with an average of 45.84, a standard deviation of 5.33, and a coefficient of variation of 11.63%.
[0185] Table 3 Descriptive statistics of SPAD of potato canopy at different growth stages
[0186] Growth period Number of samples Minimum Maximum average value Standard deviation Coefficient of variation Seedling stage 162 40.30 56.00 48.47 3.70 7.63% Maturity 162 28.20 54.20 43.20 5.40 12.50% Full reproductive period 324 28.20 56.00 45.84 5.33 11.63%
[0187] The difference in coefficient of variation reflects the changes in SPAD of potato canopy at different developmental stages. It is precisely because of these differences that the feasibility and necessity of studying the inversion of SPAD of potato canopy at different developmental stages are established. The descriptive statistics of SPAD of potato canopy at different growth stages are as follows: Figure 5 shown.
[0188] 2. Selection of vegetation index.
[0189] Orthophoto fusion and correction were performed on UAV images of the seedling stage, maturity stage and full growth period to remove noise and redundant information to obtain the orthophoto images for research. The characteristic vegetation indices constructed by corresponding to the measured ground SPAD in the three growth periods and each band were selected using the Pearson correlation analysis method and the CARS algorithm respectively.
[0190] (1) Pearson correlation analysis method.
[0191] For the Pearson correlation analysis, different developmental stages have different correlation characteristics. From the seedling stage to the mature stage, the correlation gradually decreases. This is related to the characteristics of the SPAD and vegetation index. The correlation coefficients between the various vegetation indices and SPAD at each developmental stage are shown in Table 4.
[0192] Table 4 Pearson correlation coefficient between chlorophyll content and vegetation index
[0193]
[0194] During the seedling stage, the vegetation indices with the strongest correlation with SPAD were EVI2, NDVI, and OSAVI, with correlation coefficients of 0.40. There were 10 indices with correlations greater than the threshold of 0.3. The vegetation features selected in the dataset were: WDRVI, EVI2, RECI, NDVI, GNDVI, MSRI, OSAVI, SRI, NDRE, and NLI.
[0195] In the mature stage, the vegetation index with the strongest correlation with SPAD is TVI, with a correlation coefficient of 0.35. There are 7 vegetation indices with correlations greater than the threshold of 0.3. The vegetation features selected in the dataset are: DVI, EVI2, NDVI, GNDVI, MSRI, OSAVI and TVI.
[0196] During the entire growth period, the vegetation index with the strongest correlation with SPAD is WDRVI, with a correlation coefficient of 0.80. There are 10 vegetation indices with correlations greater than the threshold of 0.7. The vegetation features selected in the dataset are: GRVI, DVI, MTVI, WDRVI, EVI2, NDVI, NDGI, MSRI, OSAVI and TVI.
[0197] The above coefficients are all correlation coefficients between vegetation indices and SPAD at the 0.01 significance level, indicating significant reliability. In summary, the correlation for the entire growth period is significantly improved compared to that for a single growth period. This is because the characteristics of the entire growth period can capture the seasonal patterns and cyclical changes in time series data, improving the correlation between the two.
[0198] (2)CARS algorithm.
[0199] For the CARS algorithm, an exponential decay function is used to select variables, and the performance of the variable subset obtained by each sampling is evaluated through cross-validation. It can quickly select variables that have important contributions to the model from a large number of feature variables to improve the predictive ability and efficiency of the model.
[0200] In this embodiment, the number of Monte Carlo (MC) samples is set to 50, and the number of extracted feature variables is determined based on the minimum value of the root mean square error (RMSE). Taking the CARS-RF model of the original image of the mature stage as an example, Figure 6 As shown in Figure 2, the number of selected feature variables gradually decreases, and the change rate is faster in the early stage and slower in the later stage. There are two stages: "primary selection" and "fine selection". This is mainly due to the characteristics of the algorithm. Figure 6It can be seen that as the number of Monte Carlo samples gradually increases, the value of the root mean square error also changes accordingly. When the number of samples is 35, the RMSE reaches the minimum value of 3.12. Four characteristic variables are selected, corresponding to the vegetation indices MCARI, DVI, GCI and NLI. The vegetation indices of the corresponding models for each growth period obtained after selection by the CARS algorithm are shown in Tables 5, 6 and 7.
[0201] Table 5 CARS algorithm selection results at seedling stage
[0202]
[0203] Table 6 CARS algorithm selection results in the mature stage
[0204]
[0205]
[0206] Table 7 CARS algorithm selection results for the entire growth period
[0207]
[0208]
[0209] In summary, the optimal vegetation indices for model research were selected using two feature selection methods: Pearson correlation analysis and the CARS algorithm. Table 5 shows that the majority of vegetation indices selected by the different optimization models using the CARS feature selection method at each growth stage contain near-infrared (NIR) wavelengths. This is because the reflection, absorption, and scattering properties of objects in this wavelength range produce specific spectral characteristics, which are beneficial for SPAD inversion studies in potato canopies.
[0210] For the two single growth periods, seedling stage and maturity stage, the vegetation features selected by the CARS algorithm are more effective than those selected by the Pearson correlation analysis method. This is because the CARS algorithm can better explore the relationship between variables in a single growth period and has a higher sensitivity to features that are important indicators of vegetation growth status. The Pearson algorithm usually performs linear correlation analysis on the overall data. Because the amount of data in a single growth period is small, the Pearson method is not as sensitive to such local, smaller changes as the CARS algorithm.
[0211] In the long time series data of the entire growth period, the results of the Pearson algorithm have high stability and interpretability. The linear relationship between the obtained vegetation index and canopy SPAD can intuitively reflect the growth trend of vegetation throughout the growth period. The CARS algorithm will be affected by factors such as overfitting during the entire growth period. Because it pays too much attention to the variable selection of a certain stage, it will ignore the important variable information of other growth stages, resulting in its vegetation index extraction effect during the entire growth period being inferior to the Pearson algorithm.
[0212] 3. Select the best optimization model for inverting potato canopy SPAD.
[0213] (1) The best model for inverting canopy SPAD during the seedling stage.
[0214] This example uses multispectral image bands of potato seedlings, maturity, and the entire growth period (seedling and maturity) as the basis to construct a vegetation index to perform an inversion study on its canopy SPAD. Pearson correlation analysis and the CARS algorithm are used to select the vegetation index to remove redundant and interfering feature information. The gray wolf algorithm and the sparrow algorithm are used to optimize and improve the three machine learning models, respectively. The two feature selection methods are combined with the model improved by the optimization algorithm, and the accuracy is compared with that of the original model. The results are shown in Table 8.
[0215] Table 8 SPAD inversion results of potato canopy during seedling stage
[0216]
[0217]
[0218] For the Pearson correlation coefficient method, the GWO-RF model has the highest inversion accuracy, and its test set determination coefficient R 2 The value of root mean square error (RMSE) was 2.75, and the value of mean absolute error (MAE) was 2.05. For the CARS algorithm, the SSA-RF model has the highest inversion accuracy, and its test set determination coefficient R 2 The SPAD values measured during the seedling stage and the SPAD values predicted by each model are shown in the scatter plots. Figure 7 shown.
[0219] Therefore, in the seedling stage, the effect of the CARS feature extraction method is slightly better than that of the Pearson correlation analysis method, and the model with the highest inversion accuracy is the CARS-SSA-RF model.
[0220] (2) The best model for inverting canopy SPAD at maturity.
[0221] During the potato maturity period, the previously selected vegetation indices were input into different machine learning models as feature variables, and the models were optimized and improved using the gray wolf optimization algorithm and the sparrow optimization algorithm. The results are shown in Table 9.
[0222] Table 9 SPAD inversion results of potato canopy at maturity
[0223]
[0224]
[0225] For the Pearson correlation coefficient method, the SSA-RF model has the highest inversion accuracy, and its test set determination coefficient R 2 The value of root mean square error (RMSE) reached 0.57, the value of mean absolute error (MAE) was 3.59, and the value of inversion accuracy was 2.97. For the CARS algorithm, the SSA-XGBoost model has the highest inversion accuracy, and its test set determination coefficient R 2 The SPAD value is 0.66, the root mean square error RMSE is 3.23, and the mean absolute error MAE is 2.75. The scatter plot of the actual SPAD measured in the mature stage and the SPAD predicted by each model is as follows: Figure 8 shown.
[0226] Therefore, in the mature stage, the effect of the CARS feature extraction method is slightly better than that of the Pearson correlation analysis method, and the model with the highest inversion accuracy is the CARS-SSA-XGBoost model.
[0227] (3) The best model for inverting canopy SPAD during the entire growth period.
[0228] In order to study the growth process of potatoes throughout their entire growth period and thus evaluate the situation of potatoes throughout their entire growth cycle, it is necessary to fuse the data from the two growth periods and combine the characteristics of the vegetation index with the optimization algorithm to invert the potato SPAD. The results are shown in Table 10.
[0229] Table 10 SPAD inversion results of potato canopy during the whole growth period
[0230]
[0231]
[0232] For the Pearson correlation coefficient method, the SSA-XGBoost model has the highest inversion accuracy, and its coefficient of determination of the test set is R 2The value of root mean square error (RMSE) reached 0.87, the value of mean absolute error (MAE) was 2.39, and the value of inversion accuracy was 1.99. For the CARS algorithm, the SSA-XGBoost model has the highest inversion accuracy, and its coefficient of determination R 2 The SPAD values for the whole growth period were measured and the SPAD values for each model were predicted. Figure 9 shown.
[0233] Therefore, during the entire reproductive period, the features selected by the Pearson correlation analysis method were better than those selected by the CARS algorithm, and the model with the highest inversion accuracy was the Pearson-SSA-XGBoost model.
[0234] In summary, in this invention, two optimization algorithms are combined with two feature selection methods to study three machine learning models, and multiple groups of comparative studies are set up, which is more reliable and convincing. The results show that in the seedling stage, the best results are achieved by using the sparrow optimization algorithm (SSA). The model with the highest research accuracy is the CARS-SSA-RF model, with a determination coefficient R 2 The results reached 0.60, the root mean square error RMSE was 2.63, and the mean absolute error MAE was 2.00. In the mature stage, the best results were achieved by using the sparrow optimization algorithm. The model with the highest accuracy was the CARS-SSA-XGBoost model, with a determination coefficient R 2 The results reached 0.66, the root mean square error RMSE was 3.23, and the mean absolute error MAE was 2.75. In the whole reproductive period, the best effect was achieved by using the sparrow optimization algorithm. The model with the highest accuracy was the Pearson-SSA-XGBoost model, with a determination coefficient R 2 It reached 0.87, the root mean square error RMSE was 2.39, and the mean absolute error MAE was 1.99.
[0235] In each growth period, the sparrow optimization algorithm (SSA) has a more flexible search strategy. The behaviors of the discoverers and followers and the vigilance mechanism of the scouts make the information exchange between individuals more frequent and efficient, and can converge to the optimal solution more quickly. The gray wolf optimization algorithm is prone to falling into local optimality during the search process, especially when dealing with complex functions, and may not be able to find the global optimal solution. According to the inversion results of the potato canopy SPAD mentioned above, if Figure 7 、 Figure 8 and Figure 9As shown in the figure, the accuracy of the model improved by the optimization algorithm has been significantly improved compared with the unimproved model. Different growth periods and research models must be combined with specific feature extraction methods and optimization algorithms to obtain the best prediction effect. This is of great significance for the inversion of potato canopy SPAD by UAV multispectral technology.
[0236] Therefore, the present invention adopts the above-mentioned canopy chlorophyll content inversion method based on feature extraction and machine learning model, and uses data from different growth periods, multiple feature extraction methods and different optimization algorithms to optimize and improve the canopy SPAD inversion model. It is found that different feature extraction methods have significant differences in inversion accuracy in different growth periods. The CARS algorithm is more effective in extracting vegetation index in two single growth periods, and the Pearson correlation coefficient method in the entire growth period is more accurate. The accuracy of the optimization models of different optimization algorithms is significantly improved, and the SSA algorithm is better in all growth periods. These application methods significantly improve the applicability and accuracy of canopy SPAD inversion research in different growth periods of potatoes.
[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A canopy chlorophyll content inversion method based on feature extraction and machine learning model, characterized in that: The following steps are involved: Step S1: using a combination of drone-mounted multispectral cameras and ground measurements to collect and process potato canopy chlorophyll content data in the study area, and construct a vegetation index; Step S2, using Pearson correlation analysis and competitive adaptive reweighted sampling CARS algorithm to perform feature selection on vegetation index; Step S3: Based on machine learning, a chlorophyll content inversion model is constructed. The specific process is as follows: Step S31, random forest RF: by establishing an RF model, an inversion experiment is conducted on the chlorophyll content of potato canopy. A ten-fold cross-validation method is used to continuously change the selection of the test set so that each subset has the opportunity to become a test set, and the performance of the model on different data subsets is evaluated; The parameters of the RF model were optimized using different optimization algorithms. In each round of cross-validation, the model was trained using different parameter combinations, and the model performance was evaluated on the test set to obtain an ideal chlorophyll content inversion model. Step S32, support vector regression SVR: establish an SVR model to conduct an inversion study on the chlorophyll content of potato canopy, and use the RBF kernel function to fit the data. Its core expression is as follows: (7); in, and is the input sample; is the parameter of the kernel function, which determines the distribution range of the sample in the feature space; During the training process, a parameter optimization algorithm is used to continuously optimize parameters based on the training data, and different hyperparameter combinations are evaluated to achieve chlorophyll content prediction and obtain a chlorophyll content inversion model; Step S33, extreme gradient boosting XGBoost: The potato chlorophyll dataset is divided into a training set and a test set in a ratio of 8:
2. During the training phase, the XGBoost model uses the training set as the basis, the input features, and the target variable, and uses the gradient boosting strategy to iteratively train the decision tree. In each iteration, the XGBoost model first calculates the gradient of the loss function with respect to the predicted value, and then fits a new decision tree based on the obtained gradient information. The newly generated decision tree learns and corrects the errors generated by the previous decision tree during the prediction process, gradually improving the prediction performance and accuracy of the entire model, and obtaining the optimal chlorophyll content inversion model; Step S4: Optimize the chlorophyll content inversion model constructed in step S3 based on the gray wolf optimization algorithm and the sparrow optimization algorithm. The specific process is as follows: Step S41: Optimize the parameters of the RF, SVR, and XGBoost chlorophyll content inversion models based on the Gray Wolf Optimization Algorithm to invert the chlorophyll content of the potato canopy. The specific steps are as follows: Step S411: inputting drone image data and measured potato canopy chlorophyll content; Step S412: using the Pearson correlation analysis method and the CARS algorithm to perform feature selection on the constructed vegetation index; Step S413: Establish RF, SVR and XGBoost chlorophyll inversion models; Step S414: Initialize the gray wolf population and parameters, and define the fitness function of the gray wolf optimization algorithm to evaluate the quality of each parameter combination; Step S415: Divide the social hierarchy of the gray wolves and calculate the distances between them to determine the update direction and position, and update the optimal position and fitness of the gray wolves; Step S416: Determine whether the stop condition is met, if so, execute step S417, otherwise execute step S415; Step S417: output the optimal parameters of RF, SVR and XGBoost chlorophyll inversion models; Step S418: training the chlorophyll inversion model to obtain the results of the chlorophyll content inversion model, and evaluating the model using evaluation indicators; Step S42: Optimize the parameters of the RF, SVR, and XGBoost chlorophyll inversion models based on the Sparrow Optimization Algorithm, and perform chlorophyll content inversion. The specific steps are as follows: Step S421: inputting drone image data and measured chlorophyll content data; Step S422: select and construct vegetation indices using the Pearson and CARS algorithms; Step S423: Establish RF, SVR and XGBoost chlorophyll content inversion models; Step S424: Initialize the position of the sparrow population and define a fitness function; Step S425: Calculate the fitness of each sparrow according to the fitness function. The sparrow with the best fitness is considered the discoverer and leads the search direction of the population. The remaining sparrows are considered followers and update their positions based on the positions of the discoverer and themselves. Step S426: Randomly select some sparrows as scouts to explore the new search area and update their own positions, thereby updating the optimal position and fitness of the global sparrows; Step S427: Determine whether the termination condition is met. If so, execute step S428; otherwise, execute step S425. Step S428: output the optimal model parameters optimized by the sparrow algorithm; Step S429: training a potato chlorophyll content inversion model, outputting the results of the chlorophyll content inversion model, and evaluating the accuracy of the model; Step S5: perform accuracy evaluation on the chlorophyll content inversion model optimized in step S4.
2. The canopy chlorophyll content inversion method based on feature extraction and machine learning model according to claim 1 is characterized in that: In step S1, a combination of drone-mounted multispectral cameras and ground measurements was used to collect and process potato canopy chlorophyll content data in the study area, and a vegetation index was constructed. The specific process is as follows: Step S11: The UAV carries a multispectral camera to acquire and process image data; First, an unmanned aerial vehicle (UAV) was used to collect image data of the study area. The UAV was equipped with four multispectral sensor probes, corresponding to the spectral bands of green light (GREEN), red light (RED), red edge light (REG), and near-infrared light (NIR). Then, the collected images were input into Pix4Dmapper software for multispectral image synthesis and calibration, thereby obtaining orthorectified multispectral reflectance images; Step S12: collecting and measuring ground data; First, the study area was divided into 16 plots, each corresponding to one potato variety. Ground data for a total of 16 potato varieties were collected. Each plot was divided into 20m × 100m, and data from 10 sample points were collected evenly in each plot. Leaf canopy chlorophyll content was collected at the potato planting bases in the study area at both the seedling and mature stages, with 160 ground data points collected for each growth period. Then, the 16 sample plots divided above were used as objects. Ten well-growing potato canopies were evenly selected as sample points in each sample plot. When measuring each leaf canopy, the veins were avoided. Ten sample points were evenly collected at the mesophyll using a handheld chlorophyll meter. The average value was taken as the true value of one sample point. The readings obtained from the potato plants selected at the sample points of each sample plot represented the relative chlorophyll content (SPAD) of the sample plot. Step S13: constructing a vegetation index based on the acquired and processed data; The geometrically corrected drone orthophoto images were divided into 16 corresponding areas according to potato varieties. Ten sample points were taken in each area. Combined with the latitude and longitude coordinates of the ground measurement points, the reflectance of the four bands of GREEN, RED, REG and NIR corresponding to each sample point was extracted on ArcGIS. Vegetation indices were constructed based on these reflectances to enhance the physiological characteristics of vegetation.
3. The canopy chlorophyll content inversion method based on feature extraction and machine learning model according to claim 2 is characterized in that: In step S11, 20 reference points of a uniform white plate are selected to perform geometric correction on the image. The pseudo-standard geometric correction method is used to convert the reflectance of the reference points into the reflectance of the multispectral image, which facilitates the extraction of accurate reflectance of the multispectral band. The correction formula is as follows: (1); in, is the reflectance value of the uniform white plate area; is the original mean before correction; is the average value of a uniform whiteboard area; is the corrected reflectance value.
4. The canopy chlorophyll content inversion method based on feature extraction and machine learning model according to claim 1 is characterized in that: In step S2 vegetation index feature selection, the Pearson feature selection method was used to measure the linear relationship between vegetation index and potato canopy SPAD, and the Pearson correlation coefficient The calculation formula is as follows: (2); in, is the sample size; and is a variable and No. data values; and The variables are and The mean of the vegetation index as a free variable , taking potato canopy SPAD as the target variable to calculate, and then according to Select the vegetation index by using the value of 5. The canopy chlorophyll content inversion method based on feature extraction and machine learning model according to claim 1 is characterized in that: In step S2 vegetation index feature selection, a competitive adaptive reweighted sampling algorithm was used to evaluate the relevance of each vegetation index to the potato canopy SPAD prediction model through multiple sampling and analysis. The specific process is as follows: Step S221, Monte Carlo sampling: Assume that the original vegetation index data set has n samples, and the number of samples for each sampling is , , the number of sampling is k; for the i-th sampling, i=1,2,…,k, the obtained sampling subset Si contains m samples randomly selected from the original data set; Step S222: Establish a partial least squares regression PLS model: for the subset S obtained by the i-th sampling i , establish the PLS model; let X i is the vegetation index matrix m×p in the subset, where p is the number of vegetation indices; Y i is the corresponding target variable vector m×1; the PLS model is implemented by i and Y i Perform decomposition and regression to obtain the regression coefficient vector , as shown below: (3); (4); in, and They are and The scoring matrix of and is the loading matrix; and is the residual matrix; Regression coefficient , as shown below: (5); in, is the weight matrix, which is determined during the iteration of the PLS algorithm; Step S223: Calculate the vegetation index importance index; set up It is The vegetation index is The weights in the PLS model of the sampling times are calculated. Importance index of vegetation index , as shown below: (6); in, is the number of sampling times; is the experience index, here we set =1, vegetation index is selected according to importance index.
6. The canopy chlorophyll content inversion method based on feature extraction and machine learning model according to claim 1, characterized in that: In step S5, the accuracy of the chlorophyll content inversion model optimized in step S4 is evaluated. The specific process is as follows: The data set is randomly divided into training set and test set with a ratio of 8:
2. The ten-fold cross validation method is used to evaluate the generalization ability of the model during training. The coefficient of determination is used , root mean square error and mean absolute error The parameter is used to measure the accuracy of the model, and its calculation formula is as follows: (8); (9); (10); Where, is the number of samples; and is the measured value; is the predicted value; and is the average of the measured values.