Vegetation index and cascaded K-means clustering fused soybean remote sensing extraction method

By integrating vegetation index and cascaded K-means clustering into a soybean remote sensing extraction method, and utilizing Sentinel-2 and Sentinel-1 image data and elevation models, combined with GSCI and cascaded K-means clustering algorithms, high-precision and automated extraction of soybean planting areas was achieved. This solved the problem of balancing cost and accuracy in existing technologies and improved the automation and stability of remote sensing extraction.

CN120913080AActive Publication Date: 2025-11-07JILIN UNIVERSITY

Patent Information

Application Number
CN202511421283.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-07
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

Existing soybean remote sensing extraction methods struggle to balance cost, accuracy, and automation. They rely on ground samples and manually set thresholds, making it difficult to achieve high-precision, automated extraction over large soybean growing areas.

Method used

A remote sensing extraction method for soybeans that integrates vegetation index and cascaded K-means clustering is proposed. A feature set is constructed using Sentinel-2 multispectral imagery, Sentinel-1 radar imagery, and elevation model data. The cascaded K-means clustering algorithm and the Greenness and Structure Index (GSCI) are combined to achieve unsupervised classification and automated extraction on the Google Earth Engine platform.

Benefits of technology

It achieves high-precision and automated extraction in large-scale soybean planting areas, eliminates human intervention, improves cross-regional application capabilities and classification stability, and reduces the cumbersomeness of sample collection and threshold parameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913080A_ABST
    Figure CN120913080A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing extraction, and particularly relates to a soybean remote sensing extraction method fusing vegetation indexes and cascaded K-means clustering. Comprising the following steps: S1, constructing a soybean extraction original feature set; s2, extracting a farmland area based on the soybean extraction original feature set, and generating a dry farmland mask in the farmland area; s3, performing unsupervised classification on the soybean extraction optimization feature set by adopting a cascade K-means clustering algorithm to obtain an initial clustering category in a dry farmland mask limited region; s4, constructing a green degree and a structure comprehensive index, and combining the green degree and the structure comprehensive index with a natural discontinuity method to obtain a soybean reference region; and S5, constructing a soybean consistency scoring formula, and obtaining a soybean remote sensing extraction result based on the soybean consistency scoring formula. The method does not depend on ground samples or empirical thresholds, and can realize automatic extraction of high-precision and extensible large-range soybean planting distribution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of remote sensing extraction, and particularly relates to a soybean remote sensing extraction method fusing a vegetation index and cascade K-means clustering. BACKGROUND

[0002] Soybean is an important oil and protein crop in the world, and plays an irreplaceable role in ensuring food security, providing livestock feed and producing vegetable oil. With the growth of population and the increasing demand for sustainable agriculture, accurately grasping the distribution and yield information of soybean has become a key link for agricultural management, policy making, international trade and scientific research.

[0003] Traditional acquisition of soybean planting information relies on artificial census and statistics, which has certain accuracy, but has problems of low efficiency, high cost, long cycle, etc., especially cannot adapt to crop identification and area assessment in large-scale regions, which restricts the informatization and fine management of modern agriculture. The rapid development of remote sensing technology provides a new path for soybean planting information extraction. By using the spatial, spectral and temporal characteristics of remote sensing images, crop monitoring and identification can be realized in a large range, quickly and automatically. This non-contact and efficient means has become an important direction of agricultural remote sensing research and application.

[0004] However, the current mainstream soybean remote sensing extraction method still faces many challenges. Supervised classification-based algorithms, such as random forest and support vector machine, usually rely on a large number of ground samples for training, which not only has high sample collection cost and long time consumption, but also has poor model generalization ability due to regional differences in samples, and needs to be frequently updated to adapt to different regions or years of extraction tasks. Unsupervised classification methods such as K-means and ISODATA are not supported by samples, but are highly sensitive to initial parameters and noise, and lack semantic labels, which need to be determined by artificial judgment, and are prone to classification errors and unstable clustering problems.

[0005] In addition, the method of using vegetation index (such as NDVI) for soybean identification usually relies on empirical or statistical threshold, and extracts soybean area by binaryzation segmentation, but this method is sensitive to regional differences and seasonal changes, and lacks adaptability and universality. Even on the Google Earth Engine (GEE) platform, by calling time-series remote sensing data and classification algorithms, the processing efficiency can be improved, but it still relies on ground samples or manual threshold setting, and still needs human intervention in key steps, which makes it difficult to realize truly end-to-end automatic extraction, and restricts its stability and consistency in large-scale, multi-region and long-term continuous mapping.

[0006] In summary, the existing soybean remote sensing extraction method is difficult to achieve effective balance among cost, precision and automation degree, and a new method of soybean remote sensing extraction is needed, which can realize high precision, automation and generalization of large-scale soybean planting area extraction on the GEE platform without manual sample selection and threshold setting, to break through the current technical bottleneck and promote the development of agricultural remote sensing intelligence. SUMMARY

[0007] Therefore, the present application aims to provide a soybean remote sensing extraction method combining vegetation index and cascade K-means clustering to solve the problem that the existing technology is difficult to achieve balance among cost, precision and automation degree. The present application does not rely on ground samples or empirical thresholds and can realize high-precision, scalable and automated extraction of large-scale soybean planting distribution.

[0008] To achieve the above-mentioned purpose, the technical scheme of the present application is as follows: A soybean remote sensing extraction method combining vegetation index and cascade K-means clustering, specifically comprising the following steps: S1: constructing a soybean extraction original feature set based on Sentinel-2 multispectral images, Sentinel-1 radar images and elevation model data; S2: extracting farmland areas based on the soybean extraction original feature set and generating a dry farmland mask in the farmland areas; S3: cropping the soybean extraction original feature set based on the dry farmland mask to obtain a soybean extraction optimized feature set, and using a cascade K-means clustering algorithm to perform unsupervised classification on the soybean extraction optimized feature set to obtain initial clustering categories in the dry farmland mask limited area; S4: constructing a greenness and structure comprehensive index and performing index calculation in the dry farmland mask limited area using the greenness and structure comprehensive index, and using a natural discontinuity method to perform hierarchical processing on the index calculation results to obtain a soybean reference area; S5: constructing a soybean consistency scoring formula and inputting the initial clustering categories in the dry farmland mask limited area and the soybean reference area into the soybean consistency scoring formula for discrimination to obtain a soybean remote sensing extraction result.

[0009] Further, in step S1, the soybean remote sensing extraction features included in the soybean extraction original feature set include chlorophyll vegetation index, land surface moisture index, enhanced vegetation index, vertical emission-vertical reception, vertical emission-horizontal reception, near-infrared mean value and slope direction.

[0010] Further, the chlorophyll vegetation index, the land surface water index, and the enhanced vegetation index are preprocessed based on the Sentinel-2 multispectral image; the vertical emission-vertical reception and the vertical emission-horizontal reception are preprocessed based on the Sentinel-1 radar image; the near-infrared mean value is preprocessed based on the Sentinel-2 multispectral image; and the slope direction is preprocessed based on the elevation model data. The preprocessed results of the chlorophyll vegetation index, the land surface water index, the enhanced vegetation index, the vertical emission-vertical reception, the vertical emission-horizontal reception, the near-infrared mean value, and the slope direction are synthesized by using the cat function to obtain a soybean extraction original feature set.

[0011] Further, in step S2, the soybean extraction original feature set is masked by using the farmland grid of the WorldCover v200 dataset to remove the non-farmland area, and the farmland area extraction is completed. The threshold value is set, and the B11 band mean value of different areas in the Sentinel-2 multispectral image is calculated in the farmland area. The area with a B11 band mean value greater than the threshold value is regarded as a dry field area, and a dry field mask is generated.

[0012] Further, the pixel value of the farmland grid is 40, and the threshold value is 0.15.

[0013] Further, the specific method of unsupervised classification of the soybean extraction optimized feature set by using the cascade K-means clustering algorithm includes: In the ee.Clusterer.wekaCascadeKMeans function built in the GEE platform, the scale parameter is set to 20, the numPixels parameter is set to 3000, the seed parameter is set to 42, and the geometries parameter is set to true. The cascade K-means clustering classifier is realized by calling the ee.Clusterer.wekaCascadeKMeans function built in the GEE platform. The input of the cascade K-means clustering classifier is the soybean extraction optimized feature set, and the output of the cascade K-means clustering classifier is the initial clustering class in the dry field mask limited area.

[0014] Further, the calculation formula of the greenness and structure integrated index GSCI is: ; wherein, represents the reflectivity value of the near-infrared band, represents the reflectivity value of the red edge second band.

[0015] Further, the highest level region obtained by using the natural break method for hierarchical processing is regarded as the soybean reference area.

[0016] Further, the soybean consistency score formula is: ; Wherein, is the consistency score of the i-th category of soybean, is the number of overlapping pixels of the i-th category and the soybean reference region, is the total number of pixels of the soybean reference region.

[0017] Compared with the prior art, the present application can achieve the following beneficial effects: (1) The soybean remote sensing extraction method fusing vegetation index and cascade K-means clustering of the present application realizes an end-to-end automated process from feature extraction to classification determination by integrating unsupervised cascade K-means clustering, greenness and structure comprehensive index (GSCI) and consistency score, which eliminates human intervention, saves time-consuming sample collection and tedious threshold parameter adjustment, significantly improves the automation level, and greatly enhances the objectivity of the method.

[0018] (2) The soybean remote sensing extraction method fusing vegetation index and cascade K-means clustering of the present application is completely executed in the GEE cloud, realizes automatic collection of source region samples through GSCI grading, has strong cross-regional application ability, and eliminates the model migration problem caused by data deviation in different regions through the automatic sampling mechanism based on image features, so that efficient extraction of large-scale soybean planting areas can be completed in a short time.

[0019] (3) Compared with the commonly used unsupervised classification algorithm (such as K-means clustering and furthest priority clustering) in GEE, the cascade K-means clustering can effectively alleviate the problems of initial value sensitivity and noise interference through multi-level progressive clustering and dynamic center updating mechanism, and significantly improve the stability and classification accuracy of clustering. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings, which form a part of the present application, are used to provide further understanding of the present application, and the schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute improper limitations on the present application. In the drawings: Figure 1 is a flowchart of the soybean remote sensing extraction method fusing vegetation index and cascade K-means clustering according to the embodiment of the present application; Figure 2 is a structural block diagram of the soybean remote sensing extraction method fusing vegetation index and cascade K-means clustering according to the embodiment of the present application; Figure 3 is a time series curve diagram of the greenness and structure comprehensive index according to the embodiment of the present application; Figure 4 The soybean remote sensing extraction area and the statistical area fitting schematic diagram are shown in the embodiments of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not constitute a limitation on the present application.

[0022] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0023] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and therefore cannot be understood as indicating or implying that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" and the like are only used for description purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features limited by "first", "second" and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0024] In the description of the present application, it should be noted that unless otherwise specified and limited, the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication between two elements inside. For those skilled in the art, the specific meaning of the above terms in the present application can be understood through specific circumstances.

[0025] Google Earth Engine (GEE) is a cloud computing-based remote sensing and geospatial analysis platform developed by Google, mainly used for large-scale environmental monitoring, geospatial modeling and spatio-temporal change analysis. The platform integrates dozens of mainstream remote sensing data such as Landsat, MODIS and Sentinel, which can be called at any time; GEE uses Google cloud server to process data, supporting large-scale spatio-temporal analysis tasks without local image download. The related operations of the invention are realized based on the GEE platform.

[0026] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0027] As Figure 1 shown, the present invention proposes a soybean remote sensing extraction method combining vegetation index and cascade K-means clustering, specifically including the following steps: S1: constructing a soybean extraction original feature set based on Sentinel-2 multispectral image, Sentinel-1 radar image and elevation model data; S2: extracting farmland area based on the soybean extraction original feature set, and generating a dry farmland mask in the farmland area; S3: cropping the soybean extraction original feature set based on the dry farmland mask, eliminating non-target area feature information, obtaining a soybean extraction optimized feature set, and using cascade K-means clustering algorithm to perform unsupervised classification on the soybean extraction optimized feature set, obtaining the initial clustering class in the dry farmland mask limited area; S4: constructing a greenness and structure comprehensive index, and using the greenness and structure comprehensive index to perform index calculation in the dry farmland mask limited area, and using the natural discontinuity method to perform hierarchical processing on the index calculation result, obtaining a soybean reference area; S5: constructing a soybean consistency scoring formula, and inputting the initial clustering class in the dry farmland mask limited area and the soybean reference area into the soybean consistency scoring formula for discrimination, obtaining a soybean remote sensing extraction result.

[0028] As Figure 2As shown, the present application firstly acquires Sentinel-2 multispectral images, Sentinel-1 radar images and elevation data, and after preprocessing, constructs a soybean extraction original feature set containing vegetation index, polarization, texture and terrain, and based on auxiliary data and threshold method, extracts farmland area and generates farmland mask in turn; then two parallel processes are executed respectively, one is to perform farmland mask cutting on the soybean extraction original feature set based on the farmland mask, to eliminate non-target area feature information, to construct a soybean extraction optimized feature set, and to use a cascade K-means clustering algorithm to perform unsupervised clustering on the multi-source feature set to obtain an initial cluster class of ground objects, and the other is to construct a greenness and structure comprehensive index (GSCI) based on near-infrared and red edge bands, and to divide it into five levels using the natural discontinuity method, and to select the highest value level as the soybean reference area; by consistency scoring on the initial cluster class and the soybean reference area, the cluster class with the highest score is identified as soybean, and the soybean distribution map is extracted, and statistical analysis and classification accuracy evaluation are performed, and finally high-precision and low-sample-dependent automatic remote sensing extraction of soybean planting area is realized.

[0029] Further, the natural discontinuity method is a grading method for dividing continuous data into several discrete levels, which is widely used in remote sensing image grading and mapping. The core idea is to minimize intra-class difference and maximize inter-class difference, so as to obtain the most reasonable data grouping.

[0030] In some embodiments, in step S1, the soybean remote sensing extraction features contained in the soybean extraction original feature set include chlorophyll vegetation index, land surface moisture index, enhanced vegetation index, vertical emission-vertical reception, vertical emission-horizontal reception, near-infrared mean value, and slope direction.

[0031] It should be noted that the chlorophyll vegetation index, the land surface moisture index, the enhanced vegetation index, the vertical emission-vertical reception, the vertical emission-horizontal reception, the near-infrared mean value, and the slope direction are originally numerical indicators representing the physical or vegetation state of the target area under a specific remote sensing band or polarization mode, but in remote sensing data processing, they are usually expressed and applied in the form of raster images with geographical spatial resolution. The present application uniformly inputs the above parameters in the form of images to facilitate subsequent image processing and feature extraction operations.

[0032] In some embodiments, the chlorophyll vegetation index, the land surface moisture index, and the enhanced vegetation index are preprocessed based on Sentinel-2 multispectral images; the vertical emission-vertical reception and the vertical emission-horizontal reception are preprocessed based on Sentinel-1 radar images; the near-infrared mean value is preprocessed based on Sentinel-2 multispectral images; and the slope direction is preprocessed based on elevation model data. The preprocessed results of chlorophyll vegetation index, land surface water index, enhanced vegetation index, vertical transmission-vertical reception, vertical transmission-horizontal reception, near-infrared average and slope direction are synthesized by using the cat function to obtain the original feature set of soybean extraction.

[0033] The soybean remote sensing extraction feature refers to: in the classification of remote sensing images, the multi-dimensional input information of the distribution of soybean planting area is extracted, and the band information in the multi-source remote sensing data and the derived index layer are spliced to form a multi-dimensional classification feature layer, which is convenient for subsequent clustering processing.

[0034] The vegetation index (including chlorophyll vegetation index (GCVI), land surface water index (LSWI) and enhanced vegetation index (EVI)) is calculated using the Sentinel-2 image data set of the GEE platform. The calculation formulas of the chlorophyll vegetation index (GCVI), land surface water index (LSWI) and enhanced vegetation index (EVI) are as follows: ; ; ; Among them, 、 、 、 、 respectively represent the reflectivity values of near-infrared (usually corresponding to B8 of Sentinel-2), red (usually corresponding to B4 of Sentinel-2), green (usually corresponding to B3 of Sentinel-2), blue (usually corresponding to B2 of Sentinel-2) and short wave 1 band (usually corresponding to B11 of Sentinel-2) in the remote sensing image.

[0035] The specific process of preprocessing the vegetation index is: based on the growth period of soybean, the time window of Sentinel-2 image synthesis is limited to July 1 to September 30, and 10% is used as the cloud threshold for preliminary filtering, and then the maskS2clouds function is used to further remove the cloud interference pixels, so as to ensure the stability of the features; finally, the median function is used to synthesize single-scene data according to the time sequence, which effectively enhances the representativeness of the data.

[0036] Polarization features: polarization features include vertical transmission-vertical reception (VV) and vertical transmission-horizontal reception (VH). VV means that the radar transmits and receives the echo signal in the vertical polarization mode. VH means that the radar transmits vertically and receives horizontally.

[0037] The specific process of preprocessing the polarization features is: the two polarization features W and VH are calculated using the Sentinel-1 image dataset of the GEE platform. Based on the soybean growth period, the time window of July 1 to September 30 is set as the Sentinel-1 image synthesis time window. First, the threshold mask is used to remove the edge or low signal-to-noise ratio area (less than -30 dB) to improve the stability of the radar features. Then, the max function is used to calculate the maximum value of the time series dataset of the Sentinel-1 image (the Sentinel-1 image dataset arranged in time sequence) to synthesize single-scene polarization feature data and enhance the structural expression capability.

[0038] Texture features: A 3x3 moving window is used in GEE to select the near-infrared band to calculate the gray matrix, and the near-infrared mean (B8_savg) feature is selected as the texture feature for soybean classification.

[0039] The specific process of preprocessing the texture features is: the dataset of the Sentinel-2 image from July 1 to September 30 is extracted, and the near-infrared band B8 of the Sentinel-2 image dataset is multiplied by 1000 to convert it to an integer type for texture calculation. The glcmTexture function (gray level co-occurrence matrix texture analysis) is used to extract the savg texture feature of the near-infrared band B8. In order to better reflect the complexity of the spatial structure of vegetation, the median function is used to calculate the B8_savg dataset between July 1 and September 30, and then synthesize single-scene data.

[0040] Terrain features: The slope direction (elevation) information calculated from the elevation (DEM) data of the study area is used as the terrain feature.

[0041] The global SRTM digital elevation model data (USGS / SRTMGL1_003) is used to calculate the slope factor, and the single-scene elevation data is obtained as the terrain feature to reflect the influence of terrain changes on the planting area.

[0042] The above steps represent the integration of full-range remote sensing features from spectrum, structure, radar reflectivity, spatial texture, and terrain factors. The cat function is used to synthesize a multi-dimensional single-scene classification image, and a multi-source remote sensing feature set is constructed, i.e. the soybean extraction original feature set, which provides sufficient feature input for the subsequent cascade K-means clustering.

[0043] In some embodiments, in step S2, the soybean extraction original feature set is masked using the farmland raster of the WorldCover v200 dataset to remove non-farmland areas and complete farmland area extraction. A threshold is set, the B11 band mean value of different regions in the Sentinel-2 multispectral image in the farmland area is calculated, and the region with a B11 band mean value greater than the threshold is regarded as a dry farmland region to generate a dry farmland mask.

[0044] It should be noted that the pixel value of the grid data of 40 in the 2020 global land cover data set (WorldCover v200) published by the European Space Agency (ESA) is used to mask the multi-source remote sensing feature set of the research area, to remove non-farmland areas such as forests, water areas, built-up areas, etc. and complete farmland coverage extraction; then, using the threshold method, the region less than 0.15 is regarded as a rice planting area based on the mean value synthesis data of short-wave infrared 1 (B11 of Sentinel-2) from April to June, and the region with a B11 band mean value greater than 0.15 is regarded as a dry farmland crop planting area containing soybeans within the farmland coverage range, to complete the dry farmland mask.

[0045] In some embodiments, the updateMask function is used to crop the soybean extraction original feature set to remove non-target area feature information, and obtain a soybean extraction optimized feature set. The specific method of unsupervised classification of the soybean extraction optimized feature set by using the cascade K-means clustering algorithm includes: in the ee.Clusterer.wekaCascadeKMeans function built in the GEE platform, the scale parameter is set to 20, the numPixels parameter is set to 3000, the seed parameter is set to 42, and the geometries parameter is set to true. The cascade K-means clustering classifier is realized by calling the ee.Clusterer.wekaCascadeKMeans function built in the GEE platform. The input of the cascade K-means clustering classifier is the soybean extraction optimized feature set, and the output of the cascade K-means clustering classifier is the initial clustering class in the limited area of the dry farmland mask.

[0046] It should be noted that the scale parameter is set to 20, that is, the pixel is collected at a resolution of 20 meters; the numPixels parameter is set to 3000, which specifies that the maximum number of sampled pixels is 3000; the seed parameter is set to 42, which specifies the random number seed 42 to ensure that the sampling is reproducible; the geometries parameter is set to true to retain the geometric spatial information of each sampled pixel. The parameters are input into the ee.Clusterer.wekaCascadeKMeans function and are waiting for calling.

[0047] The ee.Clusterer.wekaCascadeKMeans function is used to call the cascade K-means clustering classifier of the GEE platform, and then the optimal clustering structure is automatically explored in the range of 4 to 8 clusters, that is, the minimum cluster number (minClusters) is set to 4, and the maximum cluster number (maxClusters) is set to 8.

[0048] Further, the cascade K-means clustering is a multi-level unsupervised clustering algorithm, which is improved on the basis of the traditional K-means algorithm. Through the cascade iteration of the clustering result layer by layer, the stability of the classification and the preservation ability of the local structure are enhanced. The ee.Clusterer.wekaCascadeKMeans in Google Earth Engine is a specific implementation function of the algorithm, which has the characteristics of automatic optimization of initial center and enhancement of clustering diversity, and is suitable for scenes with high heterogeneity within the class but similar spectrum between pixels in remote sensing data. The basic principle is: initially, the K-means clustering algorithm is used for data clustering, and then the original clustering result is re-clustered (i.e. the clustering center or the data within the cluster is re-clustered) in each layer, forming a "hierarchical subdivision", and finally generating a clustering result with more hierarchical structure and separation ability.

[0049] In some embodiments, the formula for calculating the greenness and structure composite index GSCI is: ; Among them, represents the reflectance value of the near-infrared band, represents the reflectance value of the red edge second band.

[0050] It should be noted that the greenness and structure composite index (Greenness and Structure Composite Index, GSCI) proposed in the present application is a new type of remote sensing vegetation index specially constructed for extracting soybean planting areas. The original intention of the index is to optimize the remote sensing extraction task of soybean planting areas. Its construction logic is derived from two key features of soybeans in remote sensing spectrum: 1. Significant greenness performance, which is characterized by high reflectance in the near-infrared band; 2. Relatively complex vertical canopy structure, which has a clear response in the red edge band. By combining the reflectance of NIR and VRE2 bands in a simple product form, GSCI gives soybean crops unique spectral response characteristics. This composite index not only improves the spectral separability of soybeans from other vegetation or background objects, but also enhances its temporal recognition stability at different time nodes. Therefore, GSCI can reflect the leaf green level and canopy structure complexity of crops, and has good adaptability and potential for promotion in soybean remote sensing classification, growth period monitoring, and growth analysis tasks.

[0051] As Figure 3 shown, the time series curves of greenness and structure composite index of soybean, corn, rice, water body, grassland and built-up area are shown, which shows that the GSCI index has good soybean separation in the main growth period of soybean.

[0052] Next, in order to facilitate subsequent automatic extraction of typical soybean planting area samples from the source area and build an automatic classification process, in order to realize automatic classification processing based on remote sensing index image, the application takes GSCI (greenness and structure composite index) image as input data source, and programs natural breaks method on Google Earth Engine (GEE) platform, and according to the natural distribution characteristics of the pixel value in the image, it is unsupervised optimal segmentation, automatically determine the segmentation threshold, and divide the continuous numerical value into five non-overlapping numerical intervals. Specifically, the following steps are included: 1. Image array processing First, extract all valid pixel values in the GSCI index image, eliminate invalid values (such as cloud mask, boundary null value), and convert them into a one-dimensional numerical array, denoted as: ; Wherein, represents the index value of the th valid pixel, is the number of valid pixels.

[0053] 2. Natural breaks method modeling calculation Take the data set D as input, and use the natural breaks method to perform optimal classification on the data set D. The basic principle of this classification method is: Under the premise of given classification number k, find a set of segmentation points, so that the difference between each class is the smallest (that is, the within-group variance is the smallest), and the difference between classes is the largest (that is, the inter-class difference is the largest). The optimization goal is to minimize the within-group variance (Within-Class Variance, WCV): ; Wherein, WCV (Within-Class Variance) is the total within-group variance (that is, the total sum of squared errors); is the j th classification, which contains a plurality of data points ; is a data value of the j th class; is the average value of all data in the th class; j is the class index, and the value range is 1, 2, 3, 4, 5.

[0054] Meanwhile, the total sum of squares (TSS) of the whole set is defined as follows: ; ; Under the optimal grouping condition, the Goodness of Variance Fit (GVF) index can be further calculated: ; Wherein, the closer the GVF is to 1, the better the grouping effect is.

[0055] 3. Threshold extraction and hierarchical mapping Through the dynamic programming algorithm, all possible combinations of segmentation positions are calculated to obtain a set of optimal hierarchical threshold values : Wherein, , , wherein is the optimal breakpoint obtained by calculation, and each class interval is defined as , a total of 5 intervals.

[0056] 4. Automatic hierarchical classification of images Finally, according to the above threshold, the original index image is classified pixel by pixel: ; Wherein, f(x) is the class number of each pixel, and the output is a hierarchical index image, and the pixel value represents the classification to which it belongs.

[0057] Subsequently, according to the order of GSCI value from high to low, the above five hierarchical intervals are labeled as "very high, high, medium, low, and very low" five probability levels, to represent the possibility of the region as a soybean target object. The interval corresponding to "very high" level is selected as the high confidence soybean reference sample area.

[0058] This hierarchical method does not need to rely on artificial experience to set the threshold, and can automatically divide the optimal classification interval according to the numerical distribution of the image itself, which not only retains the physical meaning of the GSCI index, but also improves the accuracy and stability of the class division in subsequent unsupervised classification. It is suitable for remote sensing automatic recognition and classification modeling tasks of regional scale crop planting areas.

[0059] In some embodiments, the soybean consistency score formula is: ; Wherein, is the soybean consistency score of the i-th class, is the number of overlapping pixels of the i-th class and the soybean reference area, Total number of pixels of the soybean reference region.

[0060] The soybean consistency score formula is used to determine the pixel consistency between the supervised clustering categories and the reference map, and the soybean coincidence degree of each category and the reference region is calculated respectively, and the category with the highest soybean coincidence degree is obtained by comparison.

[0061] Taking the main soybean planting areas in Heilongjiang Province in 2021 as an example, the soybean remote sensing extraction of 9 cities in Heilongjiang Province is carried out respectively, the soybean planting area of each city is calculated according to the remote sensing extraction result, and the determination coefficient (R 2 ) of the soybean planting area obtained by the remote sensing extraction result and the statistical data is calculated, as shown in Figure 4 In the fitting results of the two kinds of data, the determination coefficient (R²) reaches 0.86, which shows that the present application has good fitting ability and classification accuracy in the soybean remote sensing identification task, and verifies the effectiveness and practical value of the present application in the large-scale crop identification.

[0062] It should be understood that the various forms of the flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which is not limited herein.

[0063] The above specific embodiments do not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A remote sensing extraction method for soybeans that integrates vegetation indices and cascaded K-means clustering, characterized in that: Specifically comprising the following steps: S1: constructing a soybean extraction original feature set based on Sentinel-2 multispectral images, Sentinel-1 radar images and elevation model data; S2: extracting farmland areas based on the soybean extraction original feature set, and generating a dry farmland mask in the farmland areas; S3: cropping the soybean extraction original feature set based on the dry farmland mask to obtain a soybean extraction optimized feature set, and using a cascade K-means clustering algorithm to perform unsupervised classification on the soybean extraction optimized feature set to obtain initial clustering categories in the dry farmland mask limited area; S4: constructing a greenness and structure comprehensive index, and performing index calculation in the dry farmland mask limited area using the greenness and structure comprehensive index, and using the natural discontinuity method to perform hierarchical processing on the index calculation results to obtain a soybean reference area; S5: constructing a soybean consistency scoring formula, and inputting the initial clustering categories in the dry farmland mask limited area and the soybean reference area into the soybean consistency scoring formula for discrimination to obtain a soybean remote sensing extraction result.

2. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 1, characterized in that: In step S1, the soybean remote sensing extraction features contained in the soybean extraction original feature set include chlorophyll vegetation index, land surface moisture index, enhanced vegetation index, vertical emission-vertical reception, vertical emission-horizontal reception, near-infrared mean, and slope direction.

3. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 2, characterized in that: The chlorophyll vegetation index, land surface moisture index and enhanced vegetation index are preprocessed based on Sentinel-2 multispectral images; the vertical emission-vertical reception and vertical emission-horizontal reception are preprocessed based on Sentinel-1 radar images; the near-infrared mean is preprocessed based on Sentinel-2 multispectral images; and the slope direction is preprocessed based on elevation model data; The preprocessing results of the chlorophyll vegetation index, land surface moisture index, enhanced vegetation index, vertical emission-vertical reception, vertical emission-horizontal reception, near-infrared mean and slope direction are synthesized using the cat function to obtain the soybean extraction original feature set.

4. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 1, characterized in that: In step S2, the farmland area extraction is completed by using the farmland raster of the WorldCover v200 data set to mask the soybean extraction original feature set and remove the non-farmland area; A threshold is set, and in the farmland area, the B11 band mean of different areas in the Sentinel-2 multispectral image is calculated, and the area with a B11 band mean greater than the threshold is taken as a dry farmland area to generate a dry farmland mask.

5. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 4, characterized in that: The pixel value of the farmland raster is 40; and the threshold is 0.

15.

6. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 1, characterized in that: The specific method of using a cascade K-means clustering algorithm to perform unsupervised classification on the soybean extraction optimized feature set includes: In the ee.Clusterer.wekaCascadeKMeans function built in the GEE platform, the scale parameter is set to 20, the numPixels parameter is set to 3000, the seed parameter is set to 42, and the geometries parameter is set to true; The cascade K-means clustering classifier is implemented by calling the ee.Clusterer.wekaCascadeKMeans function built in the GEE platform, the input of the cascade K-means clustering classifier is the soybean extraction optimized feature set, and the output of the cascade K-means clustering classifier is the initial clustering category in the limited area of the dry field mask.

7. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 1, characterized in that: The calculation formula of the greenness and structure comprehensive index GSCI is: ; wherein, represents a reflectance value in the near-infrared band, represents a reflectance value in the red edge second band.

8. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 1, characterized in that: The highest level region obtained by using the natural discontinuity method for hierarchical processing is taken as the soybean reference area.

9. The soybean remote sensing extraction method of fusion vegetation index and cascade K-means clustering according to claim 1, characterized in that: The soybean consistency score formula is: ; wherein, is the soybean consistency score of the i-th category, is the number of pixels of the i-th category that overlap with the soybean reference region, is the total number of pixels of the soybean reference region.

Citation Information

Patent Citations

  • Remote sensing monitoring method and device for water quality parameters of shallow aquatic plant lake

    CN108020511A

  • Soybean remote sensing mapping method based on improved GWCCI index

    CN117523412A

  • Image processing based advisory system and a method thereof

    US20210365683A1

  • System and Method for Mapping Land Cover Types with Landsat, Sentinel-1, and Sentinel-2 Images

    US20220392215A1

Cited By

  • Image classification method, system and device based on small-batch K-means, and medium

    CN121600405A