A method, system, electronic device and medium for automatically identifying crop lodging
Through the automated crop lodging recognition method, multi-source remote sensing data and machine learning algorithms are used to achieve high-precision lodging recognition without manual participation, solving the problems of high cost and low precision in existing methods.
Patent Information
- Application Number
- CN202310137820.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-02-20
AI Technical Summary
Existing crop lodging recognition methods require manual participation in obtaining samples, resulting in high time and economic costs. The optical data of a single phase is easily disturbed by crop growth factors, reducing identification accuracy.
An automated crop lodging recognition method is adopted to obtain multi-source remote sensing data before and after the disaster, calculate the band difference, exponential difference and texture features, and use the recursive feature elimination method to screen features. Combining the isolated forest algorithm and the random forest supervision classifier, lodging and non-lodging samples are automatically extracted to achieve lodging recognition without manual intervention.
It improves the accuracy and automation of crop lodging recognition, reduces the identification cost, and can conduct low-cost and high-precision lodging recognition of large-scale crops without manual participation.
Smart Images

Figure CN116310864B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of crop lodging identification, and in particular to a crop lodging automatic identification method, system, electronic equipment and medium. Background Art
[0002] Crop lodging refers to the phenomenon that upright-growing crops become crooked in patches or even fall to the ground. Lodging is a common problem in crop production and has become one of the important limiting factors for high and stable yields. Unreasonable cultivation measures such as fertilization, irrigation, and planting density may cause lodging. If management measures such as field drainage and crop propping up are taken in time after lodging occurs, the yield loss will be significantly reduced. Therefore, rapid identification of the scope of crop lodging is of great significance for optimizing disaster prevention and mitigation measures, formulating reasonable agricultural management policies, and ensuring stable crop yields.
[0003] Multispectral remote sensing data can well reflect the growth status and canopy structure of vegetation, and is widely used in lodging identification. However, under severe weather conditions, the acquisition of high-quality visible light multispectral remote sensing data is limited, while radar remote sensing images have the advantages of all-day, all-weather, and not affected by weather conditions, making them particularly suitable for lodging monitoring. To this end, a large number of studies have used both optical remote sensing images and radar remote sensing images for lodging identification, thereby better integrating the advantages of the data and improving identification accuracy. In addition, several researchers have pointed out that the use of single-phase optical data is difficult to avoid interference from factors such as crop growth during lodging identification, and time series images or multi-phase images are often required for monitoring.
[0004] At present, the methods used for lodging identification are basically based on relatively mature machine learning and deep learning algorithms, such as random forest algorithm, support vector machine SVM algorithm, maximum likelihood classification (MLC) algorithm, XGBoost algorithm, transfer learning, Unet, SegNet, etc. The above algorithms can all achieve lodging identification well, but they all require manual selection and labeling of samples. In the process of monitoring crop growth abnormalities such as crop lodging based on remote sensing images, manual participation in obtaining samples will greatly increase the time cost and economic cost of lodging identification. Therefore, exploring a lodging identification method that does not require manual participation, has strong generalization ability, and has a high degree of automation is the direction of researchers' efforts.
[0005] As an abnormal crop event caused by extreme weather events such as typhoons and rainstorms, lodging events can be identified by anomaly detection technology. The isolation forest algorithm is an unsupervised anomaly detection algorithm that defines anomalies as sample points that are easier to separate from all sample points. It does not need to specify the range and pattern of normal sample points, nor does it need to calculate indicators related to distance and density. Anomaly detection can be achieved by calculating the anomaly score of sample points. It is widely used in attack detection in network security, financial transaction fraud detection, noise data filtering and other fields. At present, the application of the isolation forest algorithm in the field of remote sensing is still relatively small, and no research has applied the isolation forest algorithm to lodging identification. In addition, the existing research on the discrimination of binarization (abnormal and normal) based on the isolation forest algorithm has introduced the threshold segmentation algorithm. In actual application, the sample points on both sides of the threshold often have similar characteristics, and it is difficult to reliably distinguish the data into two categories with completely opposite attributes only by the threshold, which will increase the uncertainty of the model and reduce the accuracy of crop lodging identification. Summary of the invention
[0006] The purpose of the present invention is to provide a crop lodging automatic identification method, system, electronic equipment and medium to improve the accuracy of crop lodging identification.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A crop lodging automatic identification method, comprising:
[0009] Acquire pre-disaster Sentinel-1 data, pre-disaster Sentinel-2 data, post-disaster Sentinel-1 data, and post-disaster Sentinel-2 data of the area to be identified; the area to be identified is a crop growing area where a disaster occurs;
[0010] Calculate a feature set based on the pre-disaster Sentinel 1 data, the pre-disaster Sentinel 2 data, the post-disaster Sentinel 1 data, and the post-disaster Sentinel 2 data; the feature set includes a band difference feature set, an index difference feature set, and a texture feature set;
[0011] Using a recursive feature elimination method to perform feature screening on the feature set to obtain a screened feature set;
[0012] According to the screened feature set, the lodging samples and non-lodging samples in the crop-covered area that is not flooded by water bodies after the disaster are determined using the isolation forest algorithm; the crop-covered area that is not flooded by water bodies after the disaster is determined in the area to be identified based on the post-disaster Sentinel-2 data and the Dynamic World dataset;
[0013] According to the lodging samples and the non-lodging samples, a random forest supervised classifier is used to extract the lodging range in the crop-covered area that is not flooded by water after the disaster.
[0014] Optionally, the process of determining the crop-covered area that is not flooded by water after the disaster specifically includes:
[0015] Acquire the DynamicWorld dataset and the feature dataset of the post-disaster Sentinel-2 data; the feature dataset includes the second band data and second index data of the post-disaster Sentinel-2 data; the second index data includes the Normalized Difference Vegetation Index, the Enhanced Vegetation Index, the Red Edge Position Index, the Modified Red Edge Normalized Difference Vegetation Index and the Surface Water Index; the second band data is multispectral remote sensing data;
[0016] Determine initial cultivated land coverage data according to the DynamicWorld dataset; the initial cultivated land coverage data is data with a label of 4 in the DynamicWorld dataset;
[0017] According to the initial cultivated land coverage data, an initial cultivated land coverage range is obtained;
[0018] Using the ISODATA algorithm to perform unsupervised classification on the feature data set to obtain a classification result;
[0019] In the classification results, the category with the smallest average value of the normalized vegetation index and the largest average value of the bipolar water index is selected as the water body type;
[0020] The water body type is removed from the initial cultivated land coverage to obtain the crop coverage area that is not submerged by the water body after the disaster.
[0021] Optionally, calculating a feature set according to the pre-disaster Sentinel No. 1 data, the pre-disaster Sentinel No. 2 data, the post-disaster Sentinel No. 1 data, and the post-disaster Sentinel No. 2 data specifically includes:
[0022] Calculating the first index data of the pre-disaster Sentinel-1 data and the first index data of the post-disaster Sentinel-1 data; the first index data includes a radar vegetation index and a dual-polarization water index;
[0023] Calculating the second index data of the pre-disaster Sentinel-2 data and the second index data of the post-disaster Sentinel-2 data;
[0024] Calculate a first difference between the first index data of the post-disaster Sentinel No. 1 data and the first index data of the pre-disaster Sentinel No. 1 data, and a second difference between the second index data of the post-disaster Sentinel No. 2 data and the second index data of the pre-disaster Sentinel No. 2 data, and form an index difference feature set;
[0025] Calculate the third difference between the first band data of the post-disaster Sentinel 1 data and the first band data of the pre-disaster Sentinel 1 data, and the fourth difference between the second band data of the post-disaster Sentinel 2 data and the second band data of the pre-disaster Sentinel 2 data, and form a band difference feature set; the first band data is radar remote sensing data; the second band data is multispectral remote sensing data;
[0026] The texture features of the index difference feature set and the band difference feature set are calculated to obtain a texture feature set; the texture features include mean, variance, homogeneity, contrast, dissimilarity, entropy, angular second moment and correlation.
[0027] Optionally, the band difference feature set, the exponent difference feature set and the texture feature set are subjected to feature screening using a recursive feature elimination method to obtain a screened feature set, specifically comprising:
[0028] Using the recursive feature elimination method to screen the texture feature set to obtain an optimal texture feature set;
[0029] Recombining the optimal texture feature set with the band difference feature set and the exponential difference feature set to obtain a recombined feature set;
[0030] The recursive feature elimination method is used to screen the recombined feature set to obtain the screened feature set.
[0031] Optionally, according to the filtered feature set, an isolation forest algorithm is used to determine the lodging samples and non-lodging samples in the crop-covered area that is not flooded by water after the disaster, specifically including:
[0032] Inputting the values corresponding to the pixels of the crop-covered area that is not flooded by water after the disaster in the filtered feature set into the isolation forest algorithm to obtain a normalized anomaly score for each pixel;
[0033] Determining a normalized anomaly score histogram according to the normalized anomaly score of each pixel;
[0034] According to the normalized anomaly score, the pixels of the crop-covered area that is not flooded by water after the disaster are sorted in ascending order and divided into n groups of pixel sets;
[0035] Calculating the coefficient of variation of each group of pixel sets, and fitting the n coefficients of variation into a coefficient of variation curve;
[0036] Calculate the coefficient of variation contribution rate of the pixel according to the coefficient of variation curve;
[0037] Remove the pixel sets whose cumulative coefficient of variation contribution rate is less than 0.99 from the n pixel sets to obtain the purified crop pixel set;
[0038] Determine the lodging samples and non-lodging samples in the purified crop pixel group according to the contribution rate of the coefficient of variation; the lodging samples are the samples whose histogram percentiles are in [x p=0.99 +5%,x p=0.99 +10%] interval, and the non-lodging samples are the pixels whose histogram percentiles are in the [95%, 100%] interval; p=0.99 is the percentile of the normalized anomaly score histogram where the coefficient of variation contribution is 0.99.
[0039] A crop lodging automatic identification system, comprising:
[0040] A data acquisition module, used to acquire pre-disaster Sentinel 1 data, pre-disaster Sentinel 2 data, post-disaster Sentinel 1 data and post-disaster Sentinel 2 data of an area to be identified; the area to be identified is a crop growing area where a disaster occurs;
[0041] A feature calculation module, used to calculate a feature set based on the pre-disaster Sentinel 1 data, the pre-disaster Sentinel 2 data, the post-disaster Sentinel 1 data and the post-disaster Sentinel 2 data; the feature set includes a band difference feature set, an index difference feature set and a texture feature set;
[0042] A data screening module, used for performing feature screening on the band difference feature set, the exponential difference feature set and the texture feature set by using a recursive feature elimination method to obtain a screened feature set;
[0043] A lodging identification module is used to determine the lodging samples and non-lodging samples in the crop-covered area that is not flooded by water bodies after the disaster using the isolation forest algorithm according to the filtered feature set; the crop-covered area that is not flooded by water bodies after the disaster is determined in the area to be identified based on the post-disaster Sentinel-2 data and the DynamicWorld dataset;
[0044] The lodging range determination module is used to extract the lodging range in the crop-covered area that is not flooded by water after the disaster by using a random forest supervised classifier based on the lodging samples and the non-lodging samples.
[0045] An electronic device comprises: a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the above-mentioned crop lodging automatic identification method.
[0046] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned crop lodging automatic identification method is implemented.
[0047] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0048] The automatic crop lodging identification method of the present invention obtains Sentinel 1 data and Sentinel 2 data before and after the disaster in the area to be identified, processes the obtained data to form a feature set, uses a recursive feature elimination method to perform feature screening on the feature set, and uses an isolation forest algorithm to determine the lodging samples and non-lodging samples in the crop coverage area that is not flooded by water after the disaster based on the screened feature set; and uses a supervised classifier to extract the lodging range in the crop coverage area that is not flooded by water after the disaster based on the lodging samples and the non-lodging samples. The present invention proposes an algorithm that first automatically extracts lodging samples and non-lodging samples based on the isolation forest algorithm, and then determines the lodging range based on the random forest supervised classifier, so as to realize low-cost, high-precision, and highly automated lodging identification of a large range of crops based on remote sensing images without human intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0050] Figure 1 A flow chart of the crop lodging automatic identification method provided by the present invention;
[0051] Figure 2 It is a flow chart of the crop lodging automatic identification method of the present invention in practical application;
[0052] Figure 3 This is the first characteristic screening curve diagram in this embodiment;
[0053] Figure 4 This is the second characteristic screening curve diagram in this embodiment;
[0054] Figure 5 Schematic diagram of normalized anomaly score in this embodiment;
[0055] Figure 6 Schematic diagram of the pixel extraction process based on the coefficient of variation in this embodiment; wherein, Figure 6 (a) is the coefficient of variation curve; Figure 6 (b) is a schematic diagram of the contribution rate of the coefficient of variation; Figure 6 (c) is a schematic diagram of the position of the finally extracted pixels in the normalized anomaly score histogram;
[0056] Figure 7Schematic diagram of lodging and non-lodging samples in this embodiment;
[0057] Figure 8 Schematic diagram of the lodging degree in this embodiment. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] The purpose of the present invention is to provide a crop lodging automatic identification method, system, electronic equipment and medium, which improve the accuracy of crop lodging identification.
[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] Embodiment 1
[0062] The method of the present invention is described by taking Bohetai Township, Zhaoyuan County, Daqing City, Heilongjiang Province, which was hit by Typhoon Mesak on September 3, 2020 and suffered a lodging disaster, as a specific embodiment.
[0063] Super Typhoon Mesak No. 9 in 2020 arrived at Bohetai Township at 23:00 on September 3. At this time, the maximum wind speed in the center of the typhoon was 8 (18 meters per second) and continued to move northwest. Heilongjiang Province initiated a flood control level IV emergency response at 14:00 on September 2, and a meteorological disaster (strong winds, heavy rains) level III emergency response at 14:30 on September 2. The typhoon caused large-scale power outages and water shortages in urban areas of Heilongjiang, Jilin and other places, damaged some houses in townships, and large-scale lodging of crops such as corn.
[0064] like Figure 1 and Figure 2 As shown, the crop lodging automatic identification method provided by the present invention comprises:
[0065] Step 101: Acquire pre-disaster Sentinel 1 data, pre-disaster Sentinel 2 data, post-disaster Sentinel 1 data, and post-disaster Sentinel 2 data of an area to be identified; the area to be identified is a crop growing area where a disaster occurs.
[0066] The data used in the present invention are mainly divided into aerospace image data (Sentinel data), aerial image data (UAV data), and surface type distribution products.
[0067] The aerospace image data includes Sentinel 1 data and Sentinel 2 data provided by the public data archive of Google Earth Engine, wherein Sentinel 1 data can be obtained on the GEE platform by linking ee.ImageCollection("COPERNICUS / S1_GRD"), and the present invention selects the dual-frequency bands VV and VH in the IW (interference wideband) mode of Sentinel 1 data; Sentinel 2 data is multispectral instrument surface reflectance data, which can be obtained on the GEE platform by linking ee.ImageCollection("COPERNICUS / S2_SR_HARMONIZED"), and the present invention selects the 2-8A band and the 11th band of Sentinel 2 data. The specific band information and time phase information of the above data are shown in Tables 1 and 2.
[0068] Table 1 Sentinel 1 data table
[0069]
[0070] Table 2 Sentinel 2 data table
[0071]
[0072] The drone data used in this paper was obtained by aerial photography of Bohetai Township, Zhaoyuan County, Daqing City, Heilongjiang Province on September 28, 2020, covering part of the area within the range of 124.0°E-125.5°E, 45.48°N-45.54°N in Bohetai Township, with a coverage area of about 5km 2 The drone used for shooting is FK-001 vertical take-off, equipped with SONYa7r camera, ExmorCMOS sensor model, sensor size of 35.9mm×24mm, effective pixel of 36.4 million, image size of 7360×4912 pixels, pixel size of 4.9μm×4.9μm, aerial photography altitude of 410m, lateral overlap rate of 70%, heading overlap rate of 70%, and the ground resolution of the obtained image is 0.06m. The processing of drone images includes data quality inspection, image feature point extraction, image matching, aerial triangulation and regional network adjustment, generation of digital elevation model, orthorectification to generate digital orthophotos, etc. The drone data has two functions: one is to obtain the sample points required for feature screening and accuracy verification through visual interpretation; the other is to obtain the distribution data of different types of crop plots in the study area through digital means.
[0073] The surface type data used in this paper comes from the DynamicWorld database released by Google, which can be obtained through "GOOGLE / DYNAMICWORLD / V1" on the GEE platform. Dynamic World is a near real-time 10m resolution global land use and land cover dataset, which is produced using deep learning based on 10m resolution Sentinel-2 images (Sentinel 2 data), including probability information and label information of 9 surface use types. This paper extracts the cultivated land pixels in the study area based on the DynamicWorld dataset on August 21, 2020.
[0074] Step 102: Calculate a feature set based on the pre-disaster Sentinel 1 data, the pre-disaster Sentinel 2 data, the post-disaster Sentinel 1 data, and the post-disaster Sentinel 2 data; the feature set includes a band difference feature set, an exponential difference feature set, and a texture feature set.
[0075] Furthermore, the step 102 specifically includes:
[0076] Calculate the first index data of the pre-disaster Sentinel 1 data and the first index data of the post-disaster Sentinel 1 data. The first index data includes a radar vegetation index and a dual-polarization water index.
[0077] Calculate the second index data of the pre-disaster Sentinel-2 data and the second index data of the post-disaster Sentinel-2 data; the second index data include the normalized vegetation index, the enhanced vegetation index, the red edge position index, the modified red edge normalized vegetation index and the surface water index.
[0078] Calculate a first difference between the first index data of the post-disaster Sentinel No. 1 data and the first index data of the pre-disaster Sentinel No. 1 data, and a second difference between the second index data of the post-disaster Sentinel No. 2 data and the second index data of the pre-disaster Sentinel No. 2 data, and form an index difference feature set;
[0079] Calculate the third difference between the first band data of the post-disaster Sentinel 1 data and the first band data of the pre-disaster Sentinel 1 data, and the fourth difference between the second band data of the post-disaster Sentinel 2 data and the second band data of the pre-disaster Sentinel 2 data, and form a band difference feature set.
[0080] The texture features of the index difference feature set and the band difference feature set are calculated to obtain a texture feature set; the texture features include mean, variance, homogeneity, contrast, dissimilarity, entropy, angular second moment and correlation.
[0081] In practical applications, the present invention preprocesses Sentinel-1 data and Sentinel-2 data based on the GEE platform. For Sentinel-1 data, first refer to the framework for preprocessing Sentinel-1SAR backscatter data in Google Earth Engine given by Mullissa et al., and perform boundary noise correction, speckle filtering and radiation terrain normalization on each Sentinel-1 data. After that, the median of each period of images before and after the disaster is calculated respectively, and then the images with the smallest average difference with the median data are further selected for mosaicking to obtain pre-disaster synthetic radar images and post-disaster synthetic radar images. For Sentinel-2 data, firstly, each Sentinel-2 data is de-clouded. The de-clouding process is performed using the quality assessment band QA60 in the Sentinel-2 data, and the pixels with pixel values of bits 10 and 11 of the band are masked as 0, which represent cloud and cirrus pixels respectively. After that, all bands are resampled to a resolution of 10 meters. Finally, the median of each period of images before and after the disaster is calculated respectively, and then the images with the smallest average difference with the median data are further selected for mosaicking to obtain the synthetic multispectral images before and after the disaster.
[0082] The indexes selected in lodging identification mainly include vegetation indexes used to reflect vegetation growth conditions and water body indexes used to reflect surface moisture content. By investigating the indices commonly used in previous studies and comprehensively considering the band characteristics of the data source of the present invention, the present invention calculates 7 water body and vegetation indices for lodging identification, specifically including radar vegetation index (RVI) and Sentinel-1 dual-polarization water index (SDWI) calculated based on Sentinel-1 data, normalized vegetation index (NDVI), enhanced vegetation index (EVI), red edge position index (REP), modified red edge normalized vegetation index (MNDVI705) and surface water body index (LSWI) calculated based on Sentinel-2 data, see Table 3. Among them, RVI is more sensitive to vegetation moisture content and plant structure, SDWI has better performance in extracting large-scale water body information in complex surface environments, REP and MNDVI705 are more sensitive to biochemical parameters of green plant growth conditions, and LSWI is more sensitive to soil and vegetation liquid water content.
[0083] Table 3 Index calculation data and formula statistics
[0084]
[0085] Among them, VH is the cross-polarization band, VV is the vertical polarization band, NIR is the near-infrared band, RED is the red band, BLUE is the blue band, RE1, RE2, RE3 are the three red edge bands of Sentinel-2, and SWIR is the shortwave infrared band of Sentinel-2.
[0086] Using multi-period data differences for lodging identification can better avoid misjudgments caused by factors such as crop variety differences, mixed pixels, and crop sowing time differences, and improve the accuracy of lodging identification. The present invention calculates the difference between the radar remote sensing data (first band data) of Sentinel 1 data after a disaster and before the disaster and the difference between the first index data, the difference between the multispectral remote sensing data (second band data) of Sentinel 2 data after a disaster and before the disaster and the difference between the second index data, and obtains a difference feature set containing 18 difference features.
[0087] Texture is a visual feature that reflects the homogeneity phenomenon in an image. It reflects the surface structure organization and arrangement properties of an object with slow or periodic changes. Many studies have pointed out that texture features are very important in lodging identification, and the texture features of lodging and non-lodging crops show significant numerical differences. Therefore, the present invention further calculates the texture of the above 18 difference features based on the gray level co-occurrence matrix (GLCM). The window size for calculating the texture is 5×5, the step size is 1, and the direction is 45°. The types of texture include 8 categories: mean, variance, homogeneity, contrast, dissimilarity, entropy, angular second moment, and correlation. Finally, a total of 144 texture feature data are obtained to constitute a texture feature set.
[0088] Step 103: Using a recursive feature elimination method, the band difference feature set, the exponent difference feature set and the texture feature set are subjected to feature screening to obtain a screened feature set.
[0089] Furthermore, the step 103 specifically includes:
[0090] The texture feature set is screened using the recursive feature elimination method to obtain an optimal texture feature set.
[0091] The RFE algorithm based on random forest classifier screened 144 texture features, and the results are as follows Figure 3 As shown in the figure. With the increase of the number of features, the classification accuracy is significantly improved. When the number of features is 7, the classification accuracy reaches the highest, which is 0.8755. After that, with the increase of the number of features, the classification accuracy of the model fluctuates in a small range and continues to decline. When the number of features is 7, the optimal texture feature set selected includes the mean texture features corresponding to RedEdge1, RedEdge2, RedEdge4, SWIR band, NDVI, mNDVI705, and LSWI index.
[0092] The optimal texture feature set is recombined with the band difference feature set and the exponential difference feature set to obtain a recombined feature set.
[0093] The recursive feature elimination method is used to screen the recombined feature set to obtain the screened feature set.
[0094] The recombined feature set is screened based on the RFE algorithm. The results are as follows: Figure 4 shown. Figure 4 The middle X-axis represents the number of features input to the RFE algorithm; the left Y-axis represents the features, and the features selected by the RFE algorithm are lit in blue, and the depth of blue indicates the importance of the feature in the RFE algorithm; the right Y-axis represents the accuracy of the RFE algorithm. For example, when the feature of the RFE algorithm is 1, only SWIR is selected, and the accuracy of the RFE algorithm is 0.7845. As the number of features in the RFE algorithm increases, the accuracy of the RFE algorithm improves significantly; when the number of features is 5, the accuracy of the RFE algorithm begins to fluctuate in a small range and slowly improves until the number of features reaches 18, when the RFE algorithm has the highest accuracy of 0.8932. The features selected at this time are the most suitable features for lodging identification.
[0095] In practical applications, the present invention uses a recursive feature elimination (RFE) algorithm based on a random forest classifier to select the optimal feature combination suitable for lodging identification and evaluate the feature contribution. The RFE algorithm is a greedy optimization algorithm that trains feature subsets multiple times to find the feature subset with the best performance. The specific steps of the RFE algorithm based on the random forest classifier are as follows: 1) k features are input into the random forest classifier as the initial feature data set, the importance of each feature is calculated, and the accuracy of the initial feature subset is calculated by cross-validation; 2) a feature with the lowest feature importance is removed from the current feature subset to obtain a new feature subset, which is input into the random forest classifier again, the importance of each feature in the new feature subset is calculated, and the accuracy of the new feature subset is calculated by cross-validation; 3) step 2 is repeated iteratively until all features are exhausted, and finally k feature subsets with different numbers of features are obtained, and the feature subset with the highest accuracy is selected as the optimal feature combination.
[0096] The present invention uses post-disaster drone data and selects a total of 579 samples as sample data for the random forest classifier in the RFE algorithm through visual interpretation. In order to comprehensively use various features, avoid feature collinearity problems, and improve the accuracy of lodging identification, the feature screening of the present invention is divided into two steps. First, 144 texture feature sets are screened to obtain the optimal texture feature set for lodging identification; then, the optimal texture feature set is recombined with the band difference feature set and the exponential difference feature set, and they are screened again to obtain the final feature set (screened feature set) for lodging identification.
[0097] By analyzing the selected feature types, it can be found that the remote sensing band difference feature reflects the highest importance, followed by the texture feature, while the importance of the remote sensing index difference feature is significantly lower than the first two. Among the remote sensing band difference features, the most important one is the SWIR band, which has stronger atmospheric penetration than other bands and has more obvious absorption characteristics for vegetation moisture content; secondly, the Red and Green bands also reflect high importance in the classification; all RedEdge bands are selected, and the red edge band can accurately reflect the growth status of crops, which is very important for identifying lodged crops. The most important of the remote sensing index difference features is the mNDVI705 index calculated based on the RedEdge band, followed by the LSWI index calculated based on the SWIR band, and finally the NDVI index. The importance of the index difference feature basically inherits the importance of the band used to calculate it, and the index difference feature does not reflect the information advantage over the single-band difference feature. Among the texture features obtained after the first step of feature screening, the mean texture features except the Green band were all selected, among which the most important ones were the mean texture features of mNDVI705 and RedEdge1.
[0098] Table 4 Schematic diagram of texture feature set after screening
[0099]
[0100] Step 104: Based on the filtered texture feature set, the isolation forest algorithm is used to determine the fallen samples and non-fallen samples in the crop-covered area that is not flooded by water bodies after the disaster; the crop-covered area that is not flooded by water bodies after the disaster is determined within the area to be identified based on the post-disaster Sentinel-2 data and the DynamicWorld dataset.
[0101] Furthermore, the process of determining the crop-covered area that is not flooded by water after the disaster specifically includes:
[0102] The DynamicWorld data set and the characteristic data set of the post-disaster Sentinel-2 data are obtained. The characteristic data set includes the second band data and the second index data of the post-disaster Sentinel-2 data.
[0103] Initial farmland coverage data is determined according to the DynamicWorld dataset; the initial farmland coverage data is data with a label of 4 in the DynamicWorld dataset.
[0104] The initial cultivated land coverage range is obtained according to the initial cultivated land coverage data.
[0105] The ISODATA algorithm is used to perform unsupervised classification on the feature data set to obtain a classification result.
[0106] In the classification results, the category with the smallest average value of the normalized vegetation index and the largest average value of the bipolar water index is selected as the water body type.
[0107] The water body type is removed from the initial cultivated land coverage to obtain the crop coverage area that is not submerged by the water body after the disaster.
[0108] In practical applications, since lodging disasters are often accompanied by heavy rainfall, crops suffer from the dual damage of lodging and flooding, resulting in a part of the cultivated land pixels before the disaster turning into water pixels after the disaster. If the cultivated land coverage before the disaster is still used as the research area, the flooded farmland area will often be classified as lodging farmland in the subsequent lodging identification, thus affecting the identification accuracy. Therefore, lodging identification should be carried out within the cultivated land coverage area that has not been flooded after the disaster.
[0109] The present invention extracts the scope of cultivated land coverage after the disaster (the scope of cultivated land coverage that is not flooded after the disaster) based on the DynamicWorld data set on August 21 and the post-disaster Sentinel-2 data. First, extract the data with label information 4 (cultivated land) in the DynamicWorld data set as the initial cultivated land coverage data; then use the ISODATA algorithm to perform unsupervised classification on the data set composed of the Sentinel-2 band data (second band data) and the second index data after the disaster, and define the surface cover types contained in the data as 4 categories, namely the most common 4 categories of water bodies, bare land, vegetation, and buildings in the image of the area to be identified in the present invention. In the classification results, the category with the lowest NDVI average value and the highest SDWI average value in each category is selected as the water body type, and this part of the water body type is removed from the initial cultivated land coverage to obtain the cultivated land coverage area that is not flooded after the disaster.
[0110] Furthermore, the step 104 specifically includes:
[0111] Step 1041: Input the values corresponding to the pixels of the crop-covered area that is not flooded by water after the disaster in the filtered feature set into the isolation forest algorithm to obtain a normalized anomaly score for each pixel.
[0112] In practical applications, the present invention uses the isolation forest algorithm in the Scikit-learn toolkit, takes each pixel in the cultivated land coverage area that is not flooded after the disaster as a sample, inputs the numerical value corresponding to the sample into the isolation forest algorithm after the filtered texture feature set, and calculates the normalized anomaly score of each pixel. The normalized anomaly score interval is [-1,1], the closer to -1, the higher the degree of abnormality of the pixel, and the closer to 1, the lower the degree of abnormality of the pixel, that is, the smaller the normalized anomaly score, the more likely it is to be a lodging pixel.
[0113] The present invention further extracts lodging and non-lodging samples in the crop coverage area that is not flooded by water after the disaster. Limited by the limited spatial resolution of remote sensing data and the influence of the classification error of the DynamicWorld dataset itself, the pixels in the extracted cultivated land coverage area may contain some mixed pixels (one pixel contains multiple types of land objects), misclassified non-cultivated land pixels, etc. Secondly, the extraction of lodging and non-lodging samples should be carried out in crop pixels in principle, and the pixels in the cultivated land coverage area extracted in the previous step may include some non-crop pixels, such as bare land pixels, mixed shrub pixels, etc. Both of the above situations will interfere with the extraction of lodging and non-lodging samples. In the currently extracted cultivated land coverage area, the number of pixels of other land object types is a minority, and compared with the difference between lodging crop pixels and non-lodging crop pixels in cultivated land pixels, the difference between other land object type pixels and crop pixels is greater, that is, the abnormal score values calculated by other land object type pixels are closer to -1. At the same time, the characteristic data of different types of samples must have the characteristics of small intra-class differences and large inter-class differences. If the coefficient of variation is used to characterize the degree of data dispersion, the coefficient of variation of the characteristic data of the same category of pixels should be smaller, while the coefficient of variation of the characteristic data of different categories of pixels should be larger. Therefore, the pixels in the cultivated land coverage area are sorted and grouped according to the anomaly score. The group with a large coefficient of variation and a high degree of anomaly (that is, a low anomaly score value, close to -1) is more likely to be a mixture of non-crop pixels or a mixture of non-crop pixels (or mixed pixels) and crop pixels. On the contrary, the group with a small coefficient of variation and a low degree of anomaly (that is, a high anomaly score value, close to 1) is more likely to be a pure crop pixel.
[0114] Based on the above analysis, the coefficient of variation of the normalized anomaly score of pixels in the cultivated land coverage area is calculated to extract the fallen and non-fallen pixels. Specifically, it includes four steps: histogram statistics and pixel grouping, coefficient of variation calculation and curve fitting, coefficient of variation contribution rate calculation and pixel extraction.
[0115] Step 1042: Determine a normalized anomaly score histogram according to the normalized anomaly score of each pixel.
[0116] The normalized anomaly score of cultivated land calculated based on isolation forest is as follows: Figure 5 As shown in the figure, the depth of the color band indicates the degree of pixel abnormality. The darker the color (the closer the value is to -1), the greater the degree of abnormality.
[0117] Step 1043: sorting the pixels of the crop-covered area that is not flooded by water after the disaster in ascending order according to the normalized anomaly score and dividing them into n groups of pixel sets.
[0118] In practical applications, the histogram of the normalized anomaly score is calculated, and the pixels are divided into 20 groups of pixel sets in the order of the normalized anomaly score from small to large, that is, each group of pixel sets contains 5% of the total number of pixels.
[0119] Step 1044: Calculate the coefficient of variation of each group of pixel sets, and fit the n coefficients of variation into a coefficient of variation curve.
[0120] In practical applications, the coefficient of variation of each set of pixels is calculated to obtain 20 sample points of the coefficient of variation. The coefficient of variation is a normalized measure of the degree of discreteness of the probability distribution. Its value is the ratio of the standard deviation (σ) to the mean (μ), which can be calculated by formula (1). In formula (1), x refers to the pixel at the percentile of the histogram x%. The natural exponential function (formula (2)) is used to fit the coefficient of variation of the discrete 20 groups into a continuous coefficient of variation function. In formula (2), x is the pixel at the percentile of the histogram x%, y is cv is the fitted coefficient of variation corresponding to the pixel with the histogram percentile of x%, and a and b are the function coefficients.
[0121]
[0122] y cv =a×e -bx (2)
[0123] Step 1045: Calculate the coefficient of variation contribution rate of the pixel according to the coefficient of variation function.
[0124] In practical applications, based on the coefficient of variation function y cv , use formula (3) to calculate the coefficient of variation contribution rate p. The coefficient of variation contribution rate p represents the cumulative coefficient of variation value of the pixels in the top x% of the histogram percentile. When p>0.99, it means that the top x% of the histogram percentile p=0.99 % of the pixels contribute 99% of the coefficient of variation, and the histogram percentile x p=0.99 % to 100% of the pixels have almost no contribution to the coefficient of variation. The smaller the coefficient of variation of a group of pixels, the more homogeneous the group of pixels is, that is, the easier it is to eliminate the influence of misclassified pixels. Therefore, the present invention determines the histogram percentile x p=0.99 % to 100% pixels are the cultivated land pixel groups that can be used to extract classification samples.
[0125]
[0126] Step 1046: Remove the pixel sets with cumulative coefficient of variation contribution rate less than 0.99 from the n pixel sets to obtain the purified crop pixel set. In practical applications, further remove 5% of the pixels on the lower side of the normalized anomaly score in the cultivated land pixel set to better eliminate the influence of mixed pixels (such as a mixture of crops and bare land).
[0127] Step 1047: Determine the lodging samples and non-lodging samples in the purified crop pixel group according to the contribution rate of the coefficient of variation. In practical applications, in the crop pixel group, the pixels whose normalized anomaly score is closer to -1 (smaller) are more likely to be lodging pixels, and the pixels whose anomaly score is closer to 1 (smaller) are more likely to be non-lodging pixels. Therefore, the histogram percentile is selected in [x p=0.99 +5%,x p=0.99 The pixels in the interval [95%, 100%] were selected as lodging samples, and the pixels in the interval [95%, 100%] of the histogram percentile were selected as non-lodging samples.
[0128] The coefficient of variation curve obtained by fitting is as follows Figure 6 As shown in (a), the calculated contribution rate of the coefficient of variation is Figure 6 As shown in (b), the position of the finally extracted pixel in the normalized anomaly score histogram is as follows: Figure 6 (c) When the coefficient of variation contribution rate is equal to 0.99, the histogram percentile is at 18%, that is, the pixels in the histogram percentile interval [18%, 100%] are crop pixels that can be used for classification. At this time, remove 5% of the pixels on the side (left) with lower abnormal scores in the crop pixel histogram, take the pixels in the interval [23%, 28%] as lodging pixels, and take the samples in the interval [95%, 100%] as non-lodging pixels, including 104549 pixels evenly distributed in the study area, as shown in Figure 7 shown.
[0129] Step 105: Based on the lodging samples and the non-lodging samples, a random forest supervised classifier is used to extract the lodging range in the crop-covered area that is not flooded by water after the disaster.
[0130] Furthermore, step 105 specifically includes:
[0131] The lodging samples and the non-lodging samples are input into the random forest supervised classifier, and the obtained lodging range is used as the lodging range in the crop-covered area that is not flooded by water after the disaster.
[0132] Based on the automatically extracted training samples, the present invention uses the random forest classifier provided in the ArcGIS 10.6 image classification toolkit to classify the lodging pixels and non-lodging pixels, extracts the normalized anomaly score data within the lodging range obtained by the random forest classification, and uses the geometric interval classification method to classify the lodging pixels into 4 levels in the order of normalized anomaly scores from small to large, representing different degrees of lodging severity, namely, extremely severe lodging, severe lodging, moderate lodging, and mild lodging. Figure 8 shown.
[0133] The characteristics of the present invention are explained below by comparing the method of the present invention with the method in the prior art.
[0134] 1. Superiority of Random Forest Classifier
[0135] In order to verify the superiority of the random forest classifier, the present invention uses the random forest classifier, SVM classifier, and maximum likelihood classifier provided in the ArcGIS 10.6 image classification toolkit to classify the fallen pixels and non-fallen pixels based on the automatically extracted training samples, and extract the fallen range in the crop-covered area that was not flooded by water after the disaster. Afterwards, 200 fallen points and 200 non-fallen points were selected as points for accuracy verification based on drone images, and the overall accuracy and Kappa coefficient of the fallen range obtained by the above three classifiers were calculated respectively, and the extraction accuracy of random forest and the other two methods was compared.
[0136] The accuracy of the extraction of the range of the lodging pixel of the present invention is shown in Table 5. The random forest classifier has the highest accuracy, with an overall accuracy of 0.7775 and a kappa accuracy of 0.5550, followed by the SVM classifier, with an overall accuracy of 0.7443 and a kappa accuracy of 0.4890, and finally the MLC classifier, with an overall accuracy of 0.6741 and a kappa accuracy of 0.3457.
[0137] Table 5 Accuracy statistics of lodging recognition
[0138] Classifier Kappa Accuracy Overall accuracy RF 0.5550 0.7775 SVM 0.4890 0.7443 MLC 0.3457 0.6741
[0139] 2. Importance of feature screening
[0140] In order to analyze whether feature screening can help improve the accuracy of lodging extraction, the accuracy of lodging extraction using all features and using screened features was compared, as shown in Table 6. By comparison, it can be seen that the overall accuracy of lodging extraction with feature screening is significantly higher than that without feature screening. Among them, the SVM algorithm has the most obvious improvement, with an overall accuracy improvement of 20.46%, followed by the MLC algorithm, with an overall accuracy improvement of 20.38%. In addition, the supervised classification operation time is significantly reduced after feature screening, especially the random forest algorithm, with the operation time reduced by nearly four-fifths. Therefore, in the process of lodging extraction, feature screening first helps to improve the accuracy of the algorithm, greatly reduce the amount of data calculation, and save calculation time.
[0141] Table 6 Comparison results of lodging recognition accuracy
[0142]
[0143] 3. Feasibility of automatic sample selection
[0144] In order to analyze the advantages and disadvantages of the lodging extraction algorithm based on automatic sample selection, the lodging samples and the selected features based on the visual interpretation of UAV data were used to extract the lodging range using random forest classifier, SVM classifier, and maximum likelihood classifier, and the accuracy of the lodging range identified by the automatically extracted samples was compared with that of the algorithm based on the automatically extracted samples. The results are shown in Table 7. The accuracy of the algorithm based on the automatic sample selection is lower than that of the algorithm based on the manually selected samples. The SVM algorithm has the smallest accuracy difference, with an overall accuracy difference of only 0.0962, and the MLC algorithm has the largest difference, with an overall accuracy difference of 0.1531. However, the acquisition of high-quality samples often requires additional data (such as the UAV data used in this study) or field surveys, which increases the workload and cost. Although the automated sample extraction method sacrifices some accuracy, it has low manual participation, high degree of freedom, strong business degree, and fast lodging identification speed, which greatly improves efficiency and reduces costs, and has good business application potential. In addition, considering that the automated lodging extraction method combining isolation forest and random forest has the highest accuracy and is least affected by feature selection, it is recommended to use the method combining isolation forest and random forest first in the actual application of lodging extraction.
[0145] Table 7 Comparison results of lodging recognition accuracy
[0146]
[0147] 4. Applicability of the algorithm in identifying lodging of different crop types
[0148] In order to explore the applicability of this algorithm for lodging identification of different types of crops, the present invention extracts corn and rice plots according to the distribution data of each type of crop plots, and performs lodging identification of each crop based on the combination method of isolated forest and random forest, and compares the recognition accuracy of the lodging identification results in the corn and rice plots of the entire study area. As shown in Table 8, when the corn lodging range is extracted in the corn covered pixel, the overall accuracy is 0.8025, which is 1.29% higher than the overall accuracy when the lodging range is extracted in the cultivated land pixels of the entire study area; when the lodging range is extracted in the rice covered pixel, the overall accuracy is 0.7375, which is 0.83% higher than the overall accuracy when the lodging range is extracted in the cultivated land pixels of the entire study area. By comparison, it can be seen that the lodging extraction accuracy for a single crop is similar to that for the entire study area, which is related to the choice of using the difference data before and after the disaster for lodging identification in the present invention, and the difference calculation can reduce the influence of crop type differences to a certain extent. In summary, considering that there is little difference in the accuracy of lodging identification between distinguishing and not distinguishing crop types, and obtaining reliable distribution data of various types of crops requires additional workload and data support, lodging extraction of the entire cultivated land pixels is a more economical and efficient method.
[0149] Table 8 Statistics of lodging recognition accuracy
[0150]
[0151] The present invention uses Sentinel 1 and Sentinel 2 data to explore the practicality of radar remote sensing data and optical remote sensing data in lodging detection, and comprehensively analyzes the characteristic bands suitable for lodging detection through the feature screening process, providing reference for related research on lodging detection or other crop disaster detection using Sentinel data. The present invention proposes a set of automated sample extraction methods based on isolation forests, which improves the current sample extraction methods that require a large amount of ground data and manual participation to support. The method greatly improves the efficiency of sample extraction and has universal application value. It is also applicable to the detection of other disasters such as floods, straw burning, and hail encountered by crops. The present invention is the first study to apply the isolation forest algorithm to the automatic extraction of remote sensing image classification samples, and provides a reference for applying machine learning algorithms in the field of anomaly recognition to the remote sensing field. The present invention proposes a sequential integration algorithm that first automatically extracts lodging or non-lodging samples based on the isolation forest algorithm, and then performs lodging recognition based on a supervised classifier, and recommends the use of a random forest classifier in the supervised classification process. The lodging distribution map produced based on the present invention can serve the rapid identification of the lodging range, and provide first-hand information for field damage assessment, insurance claim preparation, disaster rescue, etc. At the same time, the method proposed by the present invention provides a reference for applying the machine learning algorithm in the field of anomaly recognition to the field of remote sensing.
[0152] Based on the isolation forest algorithm, this paper proposes a method for automatic and rapid identification of crop lodging in a large area using remote sensing images without human intervention (crop lodging automatic identification method), and draws the following conclusions: 1) The RFE algorithm can well screen out features suitable for lodging identification, greatly improving the accuracy and efficiency of the algorithm. The feature screening results show that the remote sensing band difference features and texture difference features have a high contribution rate in lodging identification, among which the SWIR band and the red edge band are the most important; 2) The normalized anomaly score of cultivated land pixels calculated based on the isolation forest algorithm reflects the The abnormal degree of each cultivated land pixel affected by the disaster after the disaster was obtained. Among the lodging disasters that occurred at the same time, the lodging degree of soybean was the most serious, followed by corn; 3) The coefficient of variation curve was fitted using the intra-group coefficient of variation of the normalized anomaly score histogram, and the contribution rate of the coefficient of variation was calculated, which better realized the extraction of lodging and non-lodging sample pixels and avoided the influence of cultivated land extraction errors and mixed pixels; 4) The accuracy of lodging identification by random forest classifier was higher than that of SVM classifier and MLC classifier, and the overall accuracy of lodging identification could reach 78%, and the Kappa accuracy could reach 77%. In summary, the algorithm proposed in this study has a high degree of automation, low data cost, and strong applicability in actual business. It has high application value in agricultural disaster emergency management, agricultural insurance underwriting and loss assessment, and can also provide a reference for the identification of other sudden disaster events of crops.
[0153] Embodiment 2
[0154] In order to execute the method corresponding to the above embodiment 1 to achieve corresponding functions and technical effects, a crop lodging automatic identification system is provided below, including:
[0155] The data acquisition module is used to acquire pre-disaster Sentinel 1 data, pre-disaster Sentinel 2 data, post-disaster Sentinel 1 data and post-disaster Sentinel 2 data of the area to be identified; the area to be identified is the crop growing area where the disaster occurs.
[0156] The feature calculation module is used to calculate a feature set based on the pre-disaster Sentinel No. 1 data, the pre-disaster Sentinel No. 2 data, the post-disaster Sentinel No. 1 data and the post-disaster Sentinel No. 2 data.
[0157] The data screening module is used to perform feature screening on the band difference feature set, the exponent difference feature set and the texture feature set by using a recursive feature elimination method to obtain a screened feature set.
[0158] A lodging identification module is used to determine the lodging samples and non-lodging samples in the crop-covered area that is not flooded by water bodies after the disaster based on the filtered feature set; the crop-covered area that is not flooded by water bodies after the disaster is determined in the area to be identified based on the post-disaster Sentinel-2 data and the Dynamic World dataset.
[0159] The lodging range determination module is used to extract the lodging range in the crop-covered area that is not flooded by water after the disaster by using a random forest supervised classifier based on the lodging samples and the non-lodging samples.
[0160] Embodiment 3
[0161] This embodiment provides an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the crop lodging automatic identification method of the first embodiment.
[0162] Embodiment 4
[0163] This embodiment provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the crop lodging automatic identification method of the first embodiment is implemented.
[0164] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0165] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for automatically identifying crop lodging, It is characterized in that include: Acquire pre-disaster Sentinel-1 data, pre-disaster Sentinel-2 data, post-disaster Sentinel-1 data, and post-disaster Sentinel-2 data of the area to be identified; the area to be identified is a crop growing area where a disaster occurs; Calculate a feature set based on the pre-disaster Sentinel 1 data, the pre-disaster Sentinel 2 data, the post-disaster Sentinel 1 data, and the post-disaster Sentinel 2 data; the feature set includes a band difference feature set, an index difference feature set, and a texture feature set; Using a recursive feature elimination method to perform feature screening on the band difference feature set, the index difference feature set and the texture feature set to obtain a screened feature set; Based on the screened feature set, the isolated forest algorithm is used to determine the lodging samples and non-lodging samples in the crop-covered area that is not flooded by water after the disaster, including: Inputting the values corresponding to the pixels of the crop-covered area that is not flooded by water after the disaster in the filtered feature set into the isolation forest algorithm to obtain a normalized anomaly score for each pixel; Determining a normalized anomaly score histogram according to the normalized anomaly score of each pixel; According to the normalized anomaly score, the pixels of the crop-covered area that is not flooded by water after the disaster are sorted in ascending order and divided into n groups of pixel sets; Calculating the coefficient of variation of each group of pixel sets, and fitting the n coefficients of variation into a coefficient of variation curve; Calculate the coefficient of variation contribution rate of the pixel according to the coefficient of variation curve; Remove the pixel sets whose cumulative coefficient of variation contribution rate is less than 0.99 from the n pixel sets to obtain the purified crop pixel set; Determine the lodging samples and non-lodging samples in the purified crop pixel group according to the contribution rate of the coefficient of variation; the lodging samples are the ones whose histogram percentiles are in [ %, %] interval, and the non-lodging samples are the pixels whose histogram percentile is in [ , %] interval pixels; is the percentile of the normalized anomaly score histogram at the coefficient of variation contribution rate of 0.99; The process of determining the crop-covered area that is not flooded by water after the disaster specifically includes: Acquire the Dynamic World data set and the characteristic data set of the post-disaster Sentinel-2 data; the characteristic data set includes the second band data and second index data of the post-disaster Sentinel-2 data; the second index data includes the Normalized Difference Vegetation Index, the Enhanced Vegetation Index, the Red Edge Position Index, the Modified Red Edge Normalized Difference Vegetation Index and the Surface Water Index; the second band data is multispectral remote sensing data; Determining initial cultivated land coverage data according to the Dynamic World dataset; According to the initial cultivated land coverage data, an initial cultivated land coverage range is obtained; Using the ISODATA algorithm to perform unsupervised classification on the feature data set to obtain a classification result; In the classification results, the category with the smallest average value of the normalized vegetation index and the largest average value of the bipolar water index is selected as the water body type; Removing the water body type from the initial cultivated land coverage to obtain a crop coverage area that is not submerged by water bodies after the disaster; According to the lodging samples and the non-lodging samples, a random forest supervised classifier is used to extract the lodging range in the crop-covered area that is not flooded by water after the disaster.
2. The method for automatically identifying crop lodging according to claim 1, It is characterized in that According to the pre-disaster Sentinel No. 1 data, the pre-disaster Sentinel No. 2 data, the post-disaster Sentinel No. 1 data and the post-disaster Sentinel No. 2 data, a feature set is calculated, specifically including: Calculating the first index data of the pre-disaster Sentinel-1 data and the first index data of the post-disaster Sentinel-1 data; the first index data includes a radar vegetation index and a dual-polarization water index; Calculating the second index data of the pre-disaster Sentinel-2 data and the second index data of the post-disaster Sentinel-2 data; Calculate a first difference between the first index data of the post-disaster Sentinel No. 1 data and the first index data of the pre-disaster Sentinel No. 1 data, and a second difference between the second index data of the post-disaster Sentinel No. 2 data and the second index data of the pre-disaster Sentinel No. 2 data, and form an index difference feature set; Calculate the third difference between the first band data of the post-disaster Sentinel 1 data and the first band data of the pre-disaster Sentinel 1 data, and the fourth difference between the second band data of the post-disaster Sentinel 2 data and the second band data of the pre-disaster Sentinel 2 data, and form a band difference feature set; the first band data is radar remote sensing data; the second band data is multispectral remote sensing data; The texture features of the index difference feature set and the band difference feature set are calculated to obtain a texture feature set; the texture features include mean, variance, homogeneity, contrast, dissimilarity, entropy, angular second moment and correlation.
3. The crop lodging automatic identification method according to claim 1, It is characterized in that The band difference feature set, the exponent difference feature set and the texture feature set are subjected to feature screening by using a recursive feature elimination method to obtain a screened feature set, which specifically includes: Using the recursive feature elimination method to screen the texture feature set to obtain an optimal texture feature set; Recombining the optimal texture feature set with the band difference feature set and the exponential difference feature set to obtain a recombined feature set; The recursive feature elimination method is used to screen the recombined feature set to obtain the screened feature set.
4. A crop lodging automatic identification system, It is characterized in that include: A data acquisition module, used to acquire pre-disaster Sentinel 1 data, pre-disaster Sentinel 2 data, post-disaster Sentinel 1 data and post-disaster Sentinel 2 data of an area to be identified; the area to be identified is a crop growing area where a disaster occurs; A feature calculation module, used to calculate a feature set based on the pre-disaster Sentinel 1 data, the pre-disaster Sentinel 2 data, the post-disaster Sentinel 1 data and the post-disaster Sentinel 2 data; the feature set includes a band difference feature set, an index difference feature set and a texture feature set; A data screening module, used for performing feature screening on the band difference feature set, the exponential difference feature set and the texture feature set by using a recursive feature elimination method to obtain a screened feature set; The lodging identification module is used to determine the lodging samples and non-lodging samples in the crop-covered area that is not flooded by water after the disaster by using the isolation forest algorithm according to the filtered feature set, and specifically includes: Inputting the values corresponding to the pixels of the crop-covered area that is not flooded by water after the disaster in the filtered feature set into the isolation forest algorithm to obtain a normalized anomaly score for each pixel; Determining a normalized anomaly score histogram according to the normalized anomaly score of each pixel; According to the normalized anomaly score, the pixels of the crop-covered area that is not flooded by water after the disaster are sorted in ascending order and divided into n groups of pixel sets; Calculating the coefficient of variation of each group of pixel sets, and fitting the n coefficients of variation into a coefficient of variation curve; Calculate the coefficient of variation contribution rate of the pixel according to the coefficient of variation curve; Remove the pixel sets whose cumulative coefficient of variation contribution rate is less than 0.99 from the n pixel sets to obtain the purified crop pixel set; Determine the lodging samples and non-lodging samples in the purified crop pixel group according to the contribution rate of the coefficient of variation; the lodging samples are the ones whose histogram percentiles are in [ %, %] interval, and the non-lodging samples are the pixels whose histogram percentile is in [ , %] interval pixels; is the percentile of the normalized anomaly score histogram at the coefficient of variation contribution rate of 0.99; The process of determining the crop-covered area that is not flooded by water after the disaster specifically includes: Acquire the Dynamic World data set and the characteristic data set of the post-disaster Sentinel-2 data; the characteristic data set includes the second band data and second index data of the post-disaster Sentinel-2 data; the second index data includes the Normalized Difference Vegetation Index, the Enhanced Vegetation Index, the Red Edge Position Index, the Modified Red Edge Normalized Difference Vegetation Index and the Surface Water Index; the second band data is multispectral remote sensing data; Determining initial cultivated land coverage data according to the Dynamic World dataset; According to the initial cultivated land coverage data, an initial cultivated land coverage range is obtained; Using the ISODATA algorithm to perform unsupervised classification on the feature data set to obtain a classification result; In the classification results, the category with the smallest average value of the normalized vegetation index and the largest average value of the bipolar water index is selected as the water body type; Removing the water body type from the initial cultivated land coverage to obtain a crop coverage area that is not submerged by water bodies after the disaster; The lodging range determination module is used to extract the lodging range in the crop-covered area that is not flooded by water after the disaster by using a random forest supervised classifier based on the lodging samples and the non-lodging samples.
5. An electronic device, It is characterized in that include: A memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the crop lodging automatic identification method according to any one of claims 1 to 3.
6. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for automatically identifying crop lodging according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Corn lodging area extraction system and method based on maximum likelihood method
CN111091052A
Remote sensing extraction method and device for lodging corn
CN112766036A