A method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index
By calculating the cotton boll index using UAV remote sensing imagery, and combining threshold segmentation and correlation analysis, the problems of low resolution in satellite remote sensing and UAV remote sensing focusing on plants rather than bolls were solved, thus achieving high-precision cotton yield estimation.
Patent Information
- Application Number
- CN202211105043.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-09-09
AI Technical Summary
In current cotton yield estimation technologies, satellite remote sensing has low spatial resolution, making it impossible to accurately estimate at the field scale. Furthermore, UAV remote sensing methods primarily focus on cotton plants rather than bolls, resulting in low estimation accuracy.
By using UAV remote sensing imagery to highlight the spectral characteristics of cotton bolls, and employing the cotton boll index for modeling, combined with threshold segmentation and correlation analysis, a cotton yield estimation model is established.
It improved the accuracy of cotton yield estimation, with a correlation coefficient of 0.84, achieving unbiased estimation at the plot scale, with an average R² of 0.77 and an average rRMSE of 7.5%.
Smart Images

Figure CN116630827B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of telemetry, in particular to a method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index. BACKGROUND
[0002] Traditional cotton yield estimation methods require manual sampling, which is labor-intensive and time-consuming. This method is increasingly unsuitable for practical applications. Because remote sensing technology can quickly obtain extensive information from ground objects, it is widely used for cotton yield estimation. Common remote sensing platforms include satellites and unmanned aerial vehicles.
[0003] Due to the rich data sources and complete data processing procedures, satellite remote sensing is first applied to cotton yield estimation. Under the prior art, some people use time-series enhanced vegetation index obtained from HJ-1A / B satellites to rank the cotton yield of a county. Some people use time-series images obtained from SPOT and Landsat series satellites to calculate vegetation index, and combine measured digital photos and meteorological data to estimate cotton yield according to an empirical formula. Some people use the measured yield data of cotton in the sampling grid and the vegetation index calculated by satellites to build a regression model, and the correlation coefficient reaches 0.87 in the middle and late growth stages of crops. Based on the normalized vegetation index, a quadratic regression model is also applied to remote sensing data and estimated cotton yield. Some people use 16-day synthetic NDVI data of MODIS satellites to extract seasonal parameters, and use a multiple regression model constructed by the seasonal parameters to estimate cotton yield. Although satellite data have a wide range of applications, their spatial resolution is low. This makes satellite remote sensing unsuitable for accurate yield estimation at the field scale. In addition, satellite remote sensing is limited by the revisit period and may not be able to obtain data during the critical growth period of crops.
[0004] Unmanned aerial vehicle remote sensing has the advantages of strong flexibility, high spatial resolution, and customizable data acquisition time. These features overcome some limitations of satellite remote sensing, leading to a rapid increase in the use of unmanned aerial vehicle images in cotton yield estimation.
[0005] Many indicators related to the physiological and biochemical state of crops can be extracted from unmanned aerial vehicle images. These mainly include vegetation index, canopy coverage, and plant height information. These indicators are used by scholars to build models to estimate cotton yield. Some people use the ratio vegetation index and cotton yield to build a model, and the R 2The value reached 0.47. Someone used the vegetation index obtained from time series images in a cotton crop growth model to estimate yield. Someone used two-stage drone images to construct 8 indicators, then used a multivariate regression model based on these indicators to estimate cotton yield, and also evaluated the performance of a drone-based remote sensing system with a low-cost RGB camera to estimate cotton yield based on plant height. Someone used cotton plant height and canopy cover information to calculate cotton yield according to an empirical formula retrieved from point cloud-based digital surface models and orthomosaic images. Someone used vegetation indices and texture features extracted from ultra-high resolution RGB images obtained from drones to estimate cotton yield.
[0006] The above methods mainly focus on the state of cotton plants, but cotton yield is mainly produced by bolls as cotton organs, not the whole plant. Yield estimation methods based on boll extraction are more direct than focusing on the whole cotton plant. As cotton production continues to improve in mechanization, leaf fall ripening has become a necessary condition for cotton production, so bolls are obvious enough in drone images to be extracted. Based on this, many studies have focused on the extraction of bolls. Laplace image transformation was used to establish the boll coverage of a plot to estimate yield, and R 2 reached 0.83 after removing outliers. Threshold segmentation and morphological filtering were also used to extract boll regions in drone images, and someone used this method to establish a linear regression model of boll area and yield for different irrigated plots, R 2 between 0.63 and 0.65. Someone used RGB color thresholds to separate each RGB component of the image. The white part of the image can be masked as bolls. Someone used threshold segmentation in the HSV color space to obtain bolls. Classification algorithms were also used to identify bolls. Someone used a robotic platform equipped with a digital camera to detect and count the number of bolls, which was based on the operation of dividing a single boll into two or more disjoint regions. Someone collected near-infrared drone images of cotton fields. An object-oriented method was used to segment bolls and background. Someone established a cotton yield estimation model based on time series drone remote sensing indices combined with the percentage of boll opening pixels extracted by U-Net. The time series data of various vegetation indices and the proportion of open bolls were input into the backpropagation neural network as training data, R 2 reached 0.853.
[0007] Previous studies mainly used spectral reflectance and commonly used indices sensitive to vegetation, rather than highlighting bolls, and rarely considered the significant features of bolls in yield estimation. When the spectral characteristics of bolls are not as strong as those of leaves, this leads to a lot of uncertainty. At the same time, too many spectral band indices are involved, resulting in the need for complex calculations. SUMMARY
[0008] To solve the problems of the prior art, the present application provides a method for obtaining a cotton boll index to estimate cotton yield by using unmanned aerial vehicle remote sensing images, which aims to develop a new spectral index that emphasizes the spectral characteristics of cotton bolls to improve the accuracy of yield modeling, and the index takes into account the spectral differences between cotton bolls and other field targets, which can support the use of naive models or as input to contribute to deep learning models and cotton yield estimation using original wavebands as input.
[0009] The technical scheme adopted by the present application is as follows:
[0010] A method for obtaining a cotton boll index to estimate cotton yield by using unmanned aerial vehicle remote sensing images, comprising the following steps:
[0011] A. Randomly collecting orthographic images of ground cotton plants by unmanned aerial vehicles to obtain multispectral unmanned aerial vehicle orthographic mosaic images;
[0012] B. Calculating a cotton boll index that highlights the spectral characteristics of cotton bolls according to the reflectance wavebands from the orthographic mosaic images;
[0013] C. Extracting bolls by threshold segmentation, comparing the performance of different cotton boll indices using correlation analysis, and modeling the cotton boll index with the best performance and field survey yield;
[0014] D. Modeling and mass production mapping.
[0015] The technical scheme provided by the present application has the following beneficial effects:
[0016] The present application calculates remote sensing indices from multispectral unmanned aerial vehicle images to highlight the spectral characteristics of cotton bolls, extracts bolls by threshold segmentation, compares the performance of different indices using correlation analysis, and models the index with the best performance and field survey yield.
[0017] The present application has the following conclusions through experiments:
[0018] 1. Bgr&nir_n is an index that normalizes the NIR waveband by the sum of the blue, green, and red wavebands, and it shows the best performance in highlighting boll information. The correlation coefficient between the DCP extracted from this index and the measured yield is 0.84.
[0019] 2. The DCP extracted from bgr&nir_n combined with the random forest method realizes unbiased estimation of cotton yield at the sample plot scale; the average R 2 based on 5-fold cross-validation is 0.77, and the average rRMSE is 7.5%.
[0020] The above results show that the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index has higher estimation accuracy compared with the prior method of using only cotton boll extraction to estimate yield. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0022] Figure 1 A method flow chart of the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0023] Figure 2 A test area schematic diagram of the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0024] Figure 3 A data display diagram of unmanned aerial vehicle collection of the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0025] Figure 4 A quantitative analysis result box plot of the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0026] Figure 5 A statistical process schematic diagram of DCP of the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0027] Figure 6 A 34 boll mask schematic diagram of the center area of the test field of the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0028] Figure 7 A performance schematic diagram of evaluation index in highlighting cotton bolls in the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0029] Figure 8 A correlation schematic diagram of DCP and yield based on ground investigation in the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0030] Figure 9 A five-fold cross-validation result schematic diagram of the model in the method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index;
[0031] Figure 10 For the method of the present application for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index, a yield map is generated by using a random forest method;
[0032] Figure 11 For the method of the present application for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index, a schematic diagram of the correlation between DCP and DUC;
[0033] Figure 12 For the method of the present application for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index, a schematic diagram of the correlation between DUC and yield;
[0034] Figure 13 For the method of the present application for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index, a comparison diagram of the relationship between DCP and yield and the relationship between DTC and yield. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the drawings.
[0036] The present embodiment provides a method for estimating cotton yield by using unmanned aerial vehicle remote sensing image to obtain cotton boll index, which comprises the following technical scheme.
[0037] I. Selection of research area
[0038] The research area of the present embodiment is located in XX City, Hebei Province. XX City is one of the main cotton production areas in China. A total of 125 test plots are set up in the test field, including three test areas with different planting areas, with theoretical planting areas of 68.00 square meters (test area A), 54.72 square meters (test area B), and 18.24 square meters (test area C). The three test areas with different areas are shown in Figure 2 As shown in the figure, the representative cotton variety Guoxin-26 is planted in the research area, which is provided by Hebei Cotton Seed Engineering Technology Research Center.
[0039] The climate conditions in XX City are suitable for cotton planting. XX City has distinct seasons and belongs to a continental semi-arid monsoon climate, with an average annual temperature of 12.2℃, a minimum temperature of -25℃, and a maximum temperature of 43℃. The average annual rainfall is 549.5 mm. The annual rainfall distribution is uneven, with 80% of the rainfall in summer.
[0040] II. Data acquisition
[0041] 2.1 Unmanned aerial vehicle data acquisition
[0042] The DJI M300 aircraft produced by DJI was equipped with ALTUM sensors produced by Micasense, which were used to collect UAV data. The DJI M300 aircraft integrated an RTK module, had strong anti-magnetic interference ability and precise positioning ability. The camera collected 1280x960 pixels in each image band. The spectral bands were approximately 475nm for the blue band, 560nm for the green band, 668nm for the red band, 717nm for the red edge band (RE), 840nm for the near-infrared red band, and 8-14 microns for the thermal infrared band.
[0043] The data collection time was from 11:00 to 13:00 on October 22, 2021, local time. The selected flight height was 50 meters, and the ground sampling distance was 2.16 cm. The heading overlap rate was 80%, and the side overlap rate was 75%. Radiometric correction and image stitching of the data were performed in the three-dimensional reconstruction software Agisoft Metashape. A sun sensor and a reflector plate were used for radiometric correction. Blue / green / red / RE / NIR reflectance images and thermal infrared emissivity images of the test area were obtained. The sensor parameters are shown in Table 1:
[0044] Table 1. Sensor parameters
[0045]
[0046] The data collected by the UAV was displayed in true color (red, green, and blue synthesis) in Figure 3 .
[0047] 2.2 Ground data acquisition
[0048] The total number of bolls (NTC) of 10 plants (18.24 square meters, 54.72 square meters) or 15 plants (63.00 square meters) of cotton were obtained by random sampling and manual counting. Because the total number of cotton plants in the test plot was known, the NTC of the entire plot was calculated.
[0049] All cotton in the plots was hand-picked, and the corresponding yield of each test plot was obtained by actual measurement. The field survey data is shown in the appendix, which is divided into three tables corresponding to three different sizes of plots.
[0050] All ground observation data and UAV data were obtained on October 22, 2021.
[0051] III. Specific method flow
[0052] As shown in the appendix Figure 1 , the method for estimating cotton yield by using UAV remote sensing image to obtain boll index in this embodiment includes three parts:
[0053] (1) Cotton boll index calculation. Reflectance bands (blue, green, red, red edge, and near infrared) from the orthomosaic image were used to calculate the cotton boll index that highlights the spectral features of cotton bolls.
[0054] (2) Cotton boll extraction and density index construction. Adaptive threshold segmentation was applied to extract cotton boll pixels.
[0055] (3) Modeling and performance evaluation. Correlation analysis was performed on the extracted results and field survey data, and a model was established based on the correlation analysis results.
[0056] 3.1 Cotton boll index calculation
[0057] Spectral analysis was first performed using drone data from the pre-harvest period. The spectral differences between cotton bolls and other ground objects at harvest time were analyzed. Then, a cotton boll index was designed based on the results of the spectral analysis.
[0058] 3.1.1 Spectral analysis
[0059] The visual features of cotton bolls were first analyzed. Cotton bolls are visually white, indicating that they have high reflectance in the visible light bands (blue, green, and red). If an index is designed to highlight the high reflectance feature in the visible light bands, it will have better performance in extracting cotton bolls. As shown in FIG. 1, the reflectance distribution of 65,603 cotton boll pixels and 118,996 non-cotton boll pixels extracted from the multispectral image through visual interpretation, the quantitative analysis results of the box plot confirm this. Box plots are widely used to visualize the analysis of continuous unimodal data. Figure 4
[0060] Cotton bolls have different reflectance in the five bands than non-cotton boll objects. From short to long waves, cotton boll pixels show a consistent downward trend, with slightly higher reflectance in the near-infrared band. However, non-cotton boll objects show an upward trend from short to long wave radiation. Different reflectance characteristics can be used to design a cotton boll index to distinguish bolls from the background. Cotton bolls have high reflectance in the visible light band and low reflectance in the infrared band. Conversely, non-cotton boll objects have high reflectance in the infrared band and low reflectance in the visible light band. The reflectance data obtained from the box plot analysis is shown in Table 2:
[0061] Table 2. Reflectance data from box plot analysis, C represents cotton and O represents non-cotton
[0062]
[0063] 3.1.2. Index design
[0064] Remote sensing indices are mathematical combinations of several spectral bands designed to highlight the spectral features of the target object. The most widely used type of indices is the mathematical combination of reflectance bands. Many scientists have proposed different forms of indices, including normalization, ratio, difference, and summation. These methods are also used in this implementation to construct remote sensing indices.
[0065] From spectral analysis, it can be seen that there is a significant difference in reflectance between bolls and other targets. Bolls have high reflectance in the visible light band (red, green, and blue) and low reflectance in the infrared (red edge and near infrared). On the contrary, other targets have high reflectance in the infrared band and low reflectance in the visible light band. In order to highlight this feature, three methods of operation in the visible and infrared bands are difference (d), ratio (r), and normalization (n). Through index calculation, the pixel value of bolls is higher, and the pixel value of other objects is lower.
[0066] To explore whether the red edge band has an impact on the index performance, the red edge band is not involved in the construction of all indices
[0067] In order to explore whether it is possible to highlight bolls only in the visible light band, some indices are constructed using the red band instead of infrared information. This is because the reflectance performance of the red band in spectral analysis is closer to the infrared information than the blue / green band.
[0068] In the visible band, the difference between bolls and other targets in the blue and green bands is greater than that in the red band. Some indices aim to explore whether it is possible to achieve the target without using the red band.
[0069] This implementation designs 34 indices. Considering the band acquisition capabilities of different sensors, it is divided into three combinations. Four methods of establishing cross-band combination tables are established. The three combinations are: only visible light band (C1), combination of visible light band and near infrared band (C2), and combination including visible light band, near infrared band, and red edge band (C3). The four methods of combination band are difference (d), ratio (r), normalization (n), and summation (sum)
[0070] Table 3 lists the calculation formulas and abbreviations of all indices.
[0071] Table 3: All indices, including three combinations (C1, C1, and C3). The four methods of combination band are difference (d), ratio (r), normalization (n), and summation (sum). Blue, Green, Red, RE, NIR represent the reflectance of blue, green, red, red edge, and near infrared bands, respectively.
[0072]
[0073]
[0074] 3.2 Cotton boll extraction and density index construction
[0075] Band math operations were applied to the UAV ortho image to obtain the cotton boll index introduced in section 3.1 and highlight the spectral reflectance characteristics of open cotton bolls on the index image. A Gaussian filter was used on the cotton boll index to remove speckle noise. Gaussian filters are linear smoothing filters commonly used to reduce noise in image processing. This process uses a pixel-level weighted average method. The value of each pixel is calculated as a weighted average of that pixel and its neighboring pixels, which effectively suppresses normal noise. A Gaussian filter was applied to the index image to produce a smoothed image to combat the distributed noise present in the cotton boll index grayscale image. The equation of a two-dimensional Gaussian filter is:
[0076]
[0077] where x and y represent the distance between the center pixel and its neighboring pixels, and σ represents the standard deviation. In this embodiment, a (3, 3) kernel with a standard deviation of 0.8 was chosen.
[0078] Otsu's method was then applied for global thresholding. Otsu's method was proposed by Japanese scholar Otsu in 1979 and is known as Otsu's method and the maximum between-class variance method. The basic principle is to divide the image into foreground and background parts according to the gray level characteristics of the image. For the best threshold, the difference between the two parts should be the largest. In this embodiment, the maximum between-class variance was used as the criterion in Otsu's algorithm to evaluate the difference. The result of the segmentation was binarized to obtain a boll mask.
[0079] Correlation analysis was used to quantitatively evaluate the boll mask generated from the 34 cotton boll indices. Correlation is commonly used to determine whether there is an association between two observed variables and to estimate the strength of this relationship. The Pearson correlation coefficient is a measure of the linear association between two normally distributed random variables, commonly denoted as r. It is defined as follows:
[0080]
[0081] where X and Y are two random variables, and are their mean values.
[0082] Density measures were constructed prior to performing correlation analysis. Since the test plot areas were different, the number of open boll pixels (NCP) for each plot was converted into an index representing NCP per square meter, which was intended to evaluate the average boll density of different plots. The DCP can be obtained by dividing the NCP of different plots by the area of the test plot.
[0083]
[0084] DCP is involved in correlation analysis with field survey data, DTC, and yield. The better the correlation of DCP with field survey data, the more prominent the indicators for producing boll face maps highlight the characteristics of bolls. The statistical process of DCP is shown in FIG. 8, where the yellow box is the boundary of the test area, the red pixels are generated by the boll shedding index, and the green pixels are the pixels that DCP needs to be counted. The boundary of the test plot is applied to the area statistics of the boll face map to obtain the DCP of each plot. Figure 5
[0085] 3.3. Modeling and yield mapping
[0086] Since the objective of the present application is to estimate cotton yield using only the boll index, the yield data from field surveys is used as the output data of the model. DCP represents the average shedding of the test area, and the extracted DCP indicates how many boll pixels are in the map, and the larger the value, the better the harvesting effect of the plot. DCP, as input data for the model, is an indicator generated by the index.
[0087] Cross-validation (CV) is used to evaluate the model and hyperparameter selection. It can generate a reliable evaluation of the performance of the model and reduce the chance of chance. The idea of k-fold cross-validation is to omit a certain number of samples at a time, perform validation, and repeat the process k times. Our data is manually sampled, so the amount of data is small (125 groups). The value of k is set to 5 so that 25 groups of data are reserved for training at a time.
[0088] This embodiment implements five algorithms commonly used in yield estimation, as shown in Table 4, linear regression (LR), support vector regression (SVR), classification and regression tree (CART), random forest (RF), and K-nearest neighbor (KNN). All model hyperparameters are determined by grid search cross-validation.
[0089] Table 4. Best model parameters after grid search
[0090]
[0091] 3.3.1. Linear regression algorithm
[0092] Many scholars use the linear regression algorithm to establish a yield estimation model, which has good estimation performance, and the simple linear regression equation is:
[0093] Y = a0+ a t x + ε (4)
[0094] where "Y" is the dependent variable, "X" is the independent variable, a0and a1are the regression coefficients or regression parameters, and ε is the difference between the predicted data and the observed data that explains equation (4).
[0095] For a dataset (x i , y i ), i = 1, 2,..., n, the optimal parameters are usually estimated by least squares, which aims to minimize the sum of squares of ε:
[0096]
[0097] 3.3.2. Support Vector Regression (SVR)
[0098] SVR is an extension of Support Vector Classification (SVC). SVR is an algorithm that estimates the relationship between system inputs and outputs based on available samples or training data. Input data and output data are jointly represented by an n-dimensional space. The goal of the support vector machine algorithm is to identify a hyperplane in the space that clearly represents the relationship between input data and output data.
[0099] A crucial step in SVR is the selection of a kernel function, which is usually used to transform a dataset from a low-dimensional space to a high-dimensional space to better express the relationship between input data and output data. Various kernel functions have been proposed and used in various applications, such as linear functions, polynomial functions, radial basis functions, and sigmoid functions.
[0100] In this embodiment, a linear kernel is used to construct a linear regression model. SVR and simple linear regression described in Section 3.3.1 represent the prediction performance of linear models.
[0101] 3.3.3. Classification and Regression Tree (CART)
[0102] CART is a machine learning method for constructing an estimation model from data. CART uses a non-parametric regression technique to develop a decision tree based on a binary algorithm that divides the current sample data into a left subtree and a right subtree until each leaf node is homogeneous or a function that measures one quality is minimized.
[0103] The CART method is widely used in the field of parametric regression. Therefore, it has been integrated into many machine learning libraries.
[0104] 3.3.4. Random Forest (RF)
[0105] RF is another popular ensemble learning method for regression and classification problems. Ensemble learning is a method of machine learning. Multiple models are trained for the same problem, and the average of multiple models is used to improve prediction accuracy and control overfitting.
[0106] RF is a bagging ensemble learning method. It uses bootstrap technique to randomly sample from the original training sample data to generate new training sample combinations, and combines the decision trees generated from each sample set into a decision forest. The decision result is determined by the number of votes from all decision trees. The key parameter of this model is n, the number of trees in the forest, which is set to 60 to distinguish from CART.
[0107] 3.3.5. The nearest neighbor algorithm (KNN)
[0108] KNN predicts the value of the target variable based on the similarity between the target value and its spatial neighbors. It is a non-parametric model, and the most critical hyperparameters are the number of neighbors (n) and the number of adjacent data points.
[0109] The value of n is set to 14, and the Euclidean distance is used as the distance judgment standard. The average value of the nearest 14 sample points is assigned to the regressor as the predicted value.
[0110] 3.3.6. Precision evaluation
[0111] In this embodiment, two different indicators are used to evaluate the performance of the model: the relative root mean square error (rRMSE) and the determination coefficient (R2)
[0112]
[0113]
[0114] where is the predicted value, y i is the actual value, is the average value of the actual value.
[0115] A smaller rRMSE value indicates that the prediction model has better performance. R 2 The value of R 2 ranges between 0 and 1. The closer the value of R
[0116] 4. Results and analysis
[0117] The cotton boll extraction results of all the indexes are shown in Section 4.1, and visual analysis is performed. The extraction effect of the indexes is quantitatively evaluated in Section 4.2. The correlation analysis of the extracted DCP and DTC and yield is performed, and the results show that our indexes have stronger correlation with yield, with a Pearson correlation coefficient of 0.84. The modeling effect of the best index extraction result and the actual yield is shown in Section 4.3.
[0118] 4.1. Cotton boll extraction
[0119] The boll index combined with a Gaussian filter and the Otsu method can be used to extract bolls from the cotton field background. The 34 boll masks of the boll index in the center area of the test field are shown in Figure 6 Fig. 3, which shows a raw image of a small area in the test field, and 34 boll mask images generated from the 34 indices are superimposed on the raw image. C1, C2, and C3 are included in the figure to represent three band combinations.
[0120] Through visual analysis, the index in C1 has the worst performance in extraction. There are many false negatives in b&r_d, b&r_r, and bg&r_r. The bolls in the raw image are not covered. There are many false positives in g&r_r. Parts of the bolls that are not opened in the raw image are also covered. The performance of bgr_sum is better than the others. The white bolls in the raw image are covered by the red mask.
[0121] The extraction based on the indices in C2 has significantly improved. The extraction of the indices of all C2 combinations is relatively stable. The white boll pixels in the raw image are covered by the red mask. The extraction performance improves as the number of bands involved in the construction of the index increases.
[0122] The extraction performance of the indices in C3 is similar to that of the indices in C2. However, the performance of the difference method is poor, and false positives are obvious. The performance of the other combinations is consistent with that of the extraction indices in C2.
[0123] The extraction performance of the normalized method is better than that of the difference and ratio methods. For example, in the combination of bgr&renir, the difference method has many false negatives, the ratio method is slightly better, and the normalized method almost completely covers the boll pixels.
[0124] The performance of the indices with multiple bands (e.g., bgr&nir_n, bgr&nir_r, and bgr&renir_n) is better than that of the indices with fewer bands (e.g., bg&nir_n, bg&nir_r, and bg&renir_n)
[0125] 4.2. Correlation analysis
[0126] The correlation between the results of different index extraction and the field survey data was analyzed at the sample plot scale. 4.2.1 and 4.2.2 show the correlation results of different indices with DTC and yield, respectively.
[0127] 4.2.1. Correlation analysis of DCP and DTC
[0128] The correlation analysis introduced in Example 3.2 was used to evaluate the performance of the indices in highlighting bolls. The results are shown in Figure 7 Fig. 6: the correlation between DTC and DCP was extracted from each index, and the figure includes three combinations and four calculation methods.
[0129] b&r is not recommended; it performs the worst among all three combinations: difference, ratio, and normalization. The results show that this band combination cannot be used for highlighting bolls. However, several indices such as g&nir_n, bg&nir_n, bgr&nir_n, bgr&renir_n, and so on perform well with R2 greater than 0.7.
[0130] All indices have some significant effects on boll characteristics. The C1 indices perform the worst with R2 less than 0.4. The indices in C2 perform the best with R2 of 0.75. The indices in C3 perform similarly to those in C2 after introducing the red edge band.
[0131] The four calculation methods present a regularity in indices constructed with the infrared band: the normalization calculation method has the highest correlation coefficient, the ratio method is second, and the difference method is the lowest.
[0132] bgr&nir_n performs the best with R2 of 0.75. It normalizes the near-infrared band by the sum of blue, green, and red bands, thereby highlighting the high reflectance of bolls in the visible band. bgr&renir_n also performs well with R2 of 0.71. Vegetation shows great changes in the red edge band, so including the red edge band introduces uncertainty, which can explain why bgr&renir_n is not as good as bgr&nir_n.
[0133] 4.2.2. Correlation between DCP and yield
[0134] In this section, the correlation between DCP extracted from each index and yield based on ground survey is analyzed, as shown in Fig. 4.2.2.1: the correlation between yield and DCP from each index, which contains three combinations and four calculation methods. Figure 9
[0135] All indices have a higher correlation with yield than DTC. Among them, the C1 indices perform the worst with R2 less than 0.53. However, once the infrared information is added, the correlation coefficient increases significantly. The correlation coefficients of C2 and C3 indices are as high as 0.84. The results of indices in C2 and C3 are similar.
[0136] Among the three combinations, the indices that apply red, green, and blue bands simultaneously perform the best. In the C1 combination, bgr_sum performs the best with R2 of 0.53. In the C2 combination, bgr&nir_n has the highest R2 of 0.84. In the C3 combination, bgr&renir_n performs the best with R2 of 0.84.
[0137] In this analysis, similar patterns were observed as in the DTC correlation analysis: the normalized calculation method had the highest correlation coefficient, the ratio form was second, and the difference method had the lowest value.
[0138] 4.3. Results of model establishment and accuracy evaluation
[0139] Bgr&nir_n contributed to the construction and selection of yield estimation models. According to the results in Section 4.2, the correlation coefficient of bgr&nir_n and DTC was 0.75, which was the highest among all indicators, and the correlation coefficient of bgr&nir_n and yield reached 0.84, which was also the highest. At the same time, considering the waveband acquisition ability of the sensor, the indicators using near-infrared waveband and visible light waveband may have more extensive application. Therefore, bgr&nir_n is recommended.
[0140] Five-fold cross-validation was used to determine the R2 and rRMSE of the training set and the test set, as shown in Figure 5 for the five-fold cross-validation results of the five models. Black dots represent the training set results, and green dots represent the test set results. Figure 9
[0141] The method of this embodiment provides an unbiased estimate of cotton yield. The predicted values and actual values of the five models are distributed on both sides of the 1:1 line, and the distribution is relatively tight.
[0142] The results of model evaluation are summarized in Table 5.
[0143] Table 5: Evaluate model fitting performance based on R2 and rRMSE. For each column of data, use the red-yellow-green scale to represent the evaluation index. Green is good, red is poor, and yellow is in between.
[0144]
[0145]
[0146] The best yield estimation was achieved by the random forest model. In the 5-fold cross-validation, the average R2 of the training set reached 0.767, and the average R2 of the test set reached 0.704. The maximum R2 in the training set was 0.781, and the maximum R2 in the test set was 0.755. The average rRMSE of the training set and the test set was 7.478% and 8.367%, respectively.
[0147] The linear regression model and the support vector regression model achieved similar performance. The average rRMSE of the training set reached 8.3%, and the average rRMSE of the test set reached 8.4%. The average R2 in the training set was about 0.71, and the average R2 in the test set was over 0.70.
[0148] The CART model has a clear overfitting. While the average R2 reaches 0.784 in the training set, it only reaches 0.678 in the test set. The average rRMSE reaches 7.213% in the training set, but only 8.719% in the test set, which can prove the high risk of overfitting for a single decision tree
[44] . The random forest method solves the overfitting problem by increasing the number of trees.
[0149] The KNN model shows the worst performance. In 5-fold cross-validation, the average R2 of the training set only reaches 0.706, and the average R2 of the test set reaches 0.664. The average rRMSE of the training set and test set is 8.411% and 8.898%, respectively.
[0150] For linear models, the performance difference between the training set and the test set is not obvious. The average R2 of the linear regression and SVR models is about 0.7 in the training and test sets. But the average R2 of the nonlinear models in the training set and test set has a certain difference, which is less than 0.11.
[0151] The yield map is generated by applying the random forest method ( Figure 10 ). The high yield area is concentrated in the northern part of the cotton field, and the yield is low in the southern part. One possible reason is that better farmland management measures are implemented in the northern part of the cotton field.
[0152] In addition to the technical solutions of the above embodiments, the question of whether the remote sensing of the cotton field by the unmanned aerial vehicle will be blocked by the canopy is raised. A conjecture is proposed that when the cotton is close to harvesting, the cotton plant naturally ages, the canopy shielding effect is weakened, and the unmanned aerial vehicle can monitor the whole plant, especially in the cotton field after leaf shedding and ripening treatment.
[0153] This embodiment proves this conjecture. In the later stage of boll shedding, the method of this embodiment can extract the bolls of the whole plant, which helps the estimation of yield.
[0154] Furthermore, this embodiment aims to prove that the unmanned aerial vehicle can detect the bolls of the whole cotton plant, not just the canopy. First, the upper bolls of shedding are defined as the bolls harvested from the upper six branches of the cotton plant. The number of upper bolls (NUC) of all sample plots comes from the ground survey. NUC is divided by the area of the test area to obtain the density of upper bolls (DUC). DUC represents the canopy information of the cotton plant. DTC represents the information of the whole cotton plant. We also analyze the correlation between DCP and DUC ( Figure 11 ). All the DCP extracted by the index is irrelevant to DUC, and the correlation coefficient is less than 0.3. Compared with the correlation analysis result of DCP and DTC in section 4.2.1 (the correlation coefficient reaches 0.74), the correlation of DCP and DTC is higher; in other words, the unmanned aerial vehicle remote sensing is more relevant to the information about the whole cotton plant, not just the information about the canopy.
[0155] There is also evidence (e.g. Figure 12 ) that UAV detection has some penetration ability before harvest. Figure 12 In (a) the relationship between DUC and actual yield, (b) the relationship between DTC and actual yield, (c) the relationship between DCP and actual yield. R2 is achieved by RF.
[0156] The relationship between DUC and yield is weak. This is because yield is harvested from all bolls, not just the upper bolls. The relationship between DTC and yield is slightly stronger. The method of the present embodiment is closer to fitting boll weight using DTC than DUC. This also shows that UAV remote sensing can obtain information about the whole plant.
[0157] Because UAV remote sensing has good detection ability for bolls of defoliated and ripened cotton fields, there are many applications of UAV in cotton production guidance in the future, such as boll opening rate inversion, defoliation effect evaluation, yield prediction and estimation.
[0158] In the correlation analysis, we found that the method of the present embodiment is more suitable for yield than DTC. We believe that boll size is a possible reason for this phenomenon.
[0159] In Figure 13 , four pixels (black squares) are occupied by bolls of different sizes (red circles). All four pixels are identified as open boll pixels, resulting in the same DCP corresponding to different DTC. However, the yield of the red circles should be similar as shown in Figure 13 (a). Figure 13 (b).
[0160] Figure 13 In (a), the four pixels are occupied by one large boll. Figure 13 In (b), the four pixels are occupied by four small bolls. Their yields are similar, but the number of bolls is very different.
[0161] This can also explain the phenomenon seen in Figure 13 ; the relationship between DCP and yield is better than the relationship between DTC and yield. DCP represents the number of boll pixels in the cotton field, and DTC represents the number of bolls. DCP reduces the impact of boll size on yield estimation to some extent. This shows that the method of the present embodiment is more accurate than the traditional method of using boll number.
[0162] The method of the embodiment is more accurate than the existing method of extracting cotton bolls for yield estimation using only threshold segmentation. OTSU and morphological filtering can extract the cotton boll area of each cell, which is used to estimate the cotton yield, with an R2 of up to 0.65. The Laplacian threshold is used to calculate the cotton unit coverage (CUC), which is used for yield modeling. Without removing outliers, the R2 reaches 0.34. The two methods described above apply threshold segmentation to the original reflectance. The present embodiment calculates the DCP before threshold segmentation. The cotton boll index does not directly use the original band, but highlights the characteristics of the cotton boll. The spectral information is fully utilized, so the average R2 of five-fold cross-validation obtained by the method of the present embodiment is 0.77.
[0163] In summary, previous studies on cotton boll extraction directly used the reflectance of the original waveband. A more direct method is to design an index based on spectral features to extract cotton bolls. The present embodiment calculates remote sensing indexes from multispectral unmanned aerial images to highlight the spectral features of cotton bolls and extracts the lint cotton bolls through threshold segmentation. Correlation analysis is used to compare the performance of different indexes. The index with the best performance is modeled with the yield of the field survey.
[0164] The main conclusions are as follows:
[0165] 1. Bgr&nir_n is an index that normalizes the NIR band by the sum of the blue, green, and red bands, and it shows the best performance in highlighting the information of the lint cotton bolls. The correlation coefficient between the DCP extracted from this index and the measured yield is 0.84.
[0166] 2. The DCP extracted from bgr&nir_n combined with the random forest method achieves unbiased estimation of cotton yield at the plot scale; the average R2 based on 5-fold cross-validation is 0.77, and the average rRMSE is 7.5%.
[0167] The new method of the present embodiment for extracting cotton bolls by calculating remote sensing indexes has higher estimation accuracy compared to the existing method of using only cotton boll extraction for yield estimation.
[0168] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for estimating cotton yield using unmanned aerial vehicle remote sensing imagery, comprising the following steps: A. Randomly collecting the orthographic image of the ground cotton plant by unmanned aerial vehicle, obtaining the multispectral unmanned aerial vehicle orthographic mosaic image; B. According to the reflectivity band from the orthographic mosaic image, calculating the cotton boll index highlighting the spectral characteristics of the cotton boll; The specific method of step B includes: B1. Using the unmanned aerial vehicle data from the pre-harvest period for spectral analysis, i.e. calculating the cotton boll index highlighting the spectral characteristics of the cotton boll according to the different reflectivity of the five reflectivity bands: blue, green, red, infrared and near-infrared from the orthographic mosaic image; B2. Designing an index index: the specific method is to divide into three combinations, establishing four cross-band combination table methods, three combinations are: visible light band C1, combination of visible light band and near-infrared band C2, and combination including visible light band, near-infrared band and red edge band C3, and four methods of combination band are difference d, ratio r, normalization n and summation sum; The index is bgr&nir_n, i.e. ; C. Extracting the boll by threshold segmentation, comparing the performance of different cotton boll indexes using correlation analysis, and modeling the cotton boll index with the best performance and field survey yield; D. Modeling and mass mapping.
2. The method for estimating cotton yield by using UAV remote sensing image to obtain cotton boll index according to claim 1, characterized in that, The specific method of step C is: Using a Gaussian filter on the cotton boll index to remove speckle noise, this process uses a pixel-level weighted average method, the value of each pixel is calculated as the weighted average value of the pixel and its adjacent pixels, and the Gaussian filter is applied to the index image to produce a smooth image to resist the distribution noise existing in the cotton boll index grayscale image, the equation of the two-dimensional Gaussian filter is: (1) Where x and y represent the distance between the center pixel and its adjacent pixels, and σ represents the standard deviation; Then apply Otsu method to global threshold segmentation, according to the gray characteristics of the image, divide the image into foreground and background two parts; The result of segmentation is binarized to obtain a boll mask; Using correlation analysis to quantitatively evaluate the boll mask generated by 34 cotton boll indexes, obtaining the Pearson correlation coefficient r: (2) wherein and are two random variables, and are their means; The calculation method of density measure is: converting the number of bolls NCP of each plot into an index representing NCP per square meter, aiming to evaluate the average boll density of different plots, dividing the NCP of different plots by the area of the test plot to obtain DCP: (3) DCP represents the average boll of the test area, applying the boundary of the test plot to the area statistics of the boll mask to obtain the DCP of each plot.
3. The method for estimating cotton yield by using UAV remote sensing image to obtain boll index according to claim 1, characterized in that, The specific method of step D is: The modeling method is to use the field survey yield data as the output data of the model, use DCP as the input data of the model, and use cross-validation for model evaluation and hyperparameter selection.
4. The method for estimating cotton yield by using UAV remote sensing image to obtain boll index according to claim 3, characterized in that, The evaluation model algorithm includes linear regression, support vector regression, classification and regression tree, random forest and nearest neighbor algorithm.
5. The method for estimating cotton yield by using UAV remote sensing image to obtain boll index according to claim 4, characterized in that, The evaluation model algorithm is random forest algorithm.
Citation Information
Patent Citations
SENP cotton yield estimation method based on unmanned aerial vehicle image and estimation model construction method
CN111275567A
Cotton vegetation index remote sensing detection method based on Sentinel-2 satellite
CN111521562A