A scene self-adaption based ground sample placement method for crop remote sensing classification

By constructing a stratified sampling base map using a data-driven unsupervised clustering algorithm, the problems of sample class balance and insufficient accuracy in remote sensing crop classification are solved, achieving efficient and low-cost sample acquisition and high-precision classification.

CN116824355BActive Publication Date: 2025-12-12INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310080765.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-19
Publication Date
2025-12-12
Estimated Expiration
2043-01-19

AI Technical Summary

Technical Problem

Existing remote sensing crop classification methods struggle to ensure sample class balance and accuracy under complex planting patterns, resulting in insufficient classification accuracy and high sample acquisition costs.

Method used

A data-driven approach is adopted, using an unsupervised clustering algorithm to perform spectral clustering on the remote sensing data to be classified, constructing a stratified sampling base map, determining the optimal sample allocation method, ensuring balance between sample classes and diversity within class classes, and reducing sample redundancy.

Benefits of technology

Achieving high-precision classification with a smaller number of samples saves sample preparation time and costs, improves remote sensing classification efficiency, and avoids class imbalance problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824355B_ABST
    Figure CN116824355B_ABST
Patent Text Reader

Abstract

A kind of scene self-adapting-based crop remote sensing classification ground sample layout method, comprising: S1, the image data set of the region to be classified research is acquired, and total sample size and expected user accuracy are set;S2, the characteristics of image data are analyzed, and the classification target area is extracted using unsupervised classification method;S3, based on the classification target area extracted in the foregoing step, clustering analysis is carried out on the target area, the optimal clustering class number is determined, each clustering class is taken as a layer, and the number of samples in each layer is allocated according to sample allocation principle, to finally determine the number of samples and the position of samples;S4, output the sample layout result to be collected.The method of the present application can significantly reduce sample redundancy, avoid class imbalance and other problems compared with other sample point selection methods, improve the representativeness of sample points, and further improve the classification accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sample layout, more particularly, to a crop remote sensing classification ground sample layout method based on scene self-adaptation. BACKGROUND

[0002] Generally, samples for remote sensing data classification are obtained by field investigation sampling method or visual interpretation based on existing data, which usually requires sample labeling personnel to have relevant professional knowledge and collect enough samples to ensure classification accuracy. In recent years, the combination of spatial sampling theory based on geostatistics and remote sensing technology has improved the accuracy and efficiency of remote sensing classification through "prior optimization" method, and improved sample class balance and alleviated local sample redundancy.

[0003] Stratified sampling method can improve sampling efficiency to some extent, which divides the population into different layers to increase the representativeness of a certain class. Some studies try to use existing or historical distribution data of remote sensing classification objects to construct sampling base map, so as to improve the pertinence and spatial sampling efficiency of sample acquisition.

[0004] In the absence of available survey object distribution data, some scholars collect a large number of samples to reduce the standard error of the results, and use easy-to-implement methods such as systematic sampling and simple random sampling to quickly layout samples. Some other scholars use various auxiliary information data to indirectly construct stratified base map. For example, when mapping farmland and crop types, stratified base map is constructed according to soil type distribution information, administrative division information, or crop planting frequency extracted based on historical information.

[0005] Remote sensing classification research based on stratified sampling often uses existing data sets (such as land use products) as stratification basis, but in remote sensing crop classification application, especially in China, complex planting patterns such as rotation, multiple cropping, fallow, and frequent farmer field activities will cause the size of the field to change and the type of crops to change, at which time the effect of historical crop distribution data will be very limited.

[0006] Simple random sampling and systematic sampling method often do not consider the spatial correlation between survey objects, although they can ensure the balance of sample spatial distribution to some extent, but they cannot guarantee the balance of sample classes, which may lead to insufficient rare class samples, so that the accuracy cannot meet the research needs.

[0007] Soil type and administrative division data cannot reflect the characteristic difference between crops, and when historical information is incomplete or new crops and planting patterns appear, the sampling base map extracted by crop planting frequency cannot well reflect the distribution of crops in the season due to its lag.

[0008] Although the idea of clustering hierarchy has been widely applied in other disciplines, such as processing non-equilibrium data, text processing, land supervision, geological exploration and path planning, cloth pattern classification, etc., in the layout of remote sensing classification sample points, due to the different feature gaps between land types, such as the feature difference between cultivated land and residential land is much greater than the feature difference between crop types in farmland, it is necessary to preprocess in combination with the local land type distribution before clustering analysis of image data to exclude obvious interference categories. At the same time, according to the classification demand and actual cost, a suitable sample layout scheme is selected. SUMMARY

[0009] In view of the problems in the background art, in order to overcome the above-mentioned difficulties, the present application adopts a data-driven method, and proposes a method for constructing a hierarchical sampling base map using spectral, vegetation index and other features of the image to be classified and judging the hierarchical sampling sample allocation method. The core idea is to estimate the number of categories in advance by using unsupervised clustering algorithm for spectral clustering of remote sensing data to be classified when collecting samples, and to use the optimal sample allocation method of each layer according to the given total sample number, so as to ensure the balance between sample categories and the diversity within the category, reduce sample redundancy, and achieve high-precision classification effect with fewer samples, greatly saving the time cost of sample making and improving the efficiency of remote sensing classification.

[0010] Therefore, the present application proposes a scene-adaptive crop remote sensing classification ground sample layout method, comprising: S1, acquiring image data set of the region to be classified, and setting total sample quantity and expected user accuracy; S2, analyzing the characteristics of the image data, and extracting the classification target region by using an unsupervised classification method; S3, based on the classification target region extracted in the foregoing step, performing clustering analysis on the target region to determine the optimal number of clustering categories, taking each clustering category as a layer, and allocating the number of samples in each layer according to the sample allocation principle, and finally determining the number and position of the samples; and S4, outputting the sample layout result to be collected.

[0011] The present application has the following beneficial effects: in the crop classification based on remote sensing data, in the training sample making link, aiming at the problems of "where to collect" and "how many to collect", based on the characteristics of the image data to be classified, a data-driven scene-adaptive sample point selection method is proposed, which can calculate the optimal sample point spatial distribution and sample point quantity allocation in advance, compared with other sample point selection methods, can significantly reduce sample redundancy, avoid category imbalance and other problems, improve the representativeness of the sample points, and further improve the classification accuracy of the model.

[0012] 1) The simple and applicable multi-clustering and hierarchical method is designed in the application, and the hierarchical base map can be obtained according to the time sequence spectral characteristics of the image without the aid of other auxiliary data. Since the classification object in the experiment is crops, the farmland area can be automatically extracted through the decision tree, thereby reducing the interference of other land types in the classification process.

[0013] 2) After the sampling work guided by the method, the accuracy of classification and extraction can be ensured under the condition of fewer samples, the cost is saved, and the situation that the overall accuracy is high but the classification accuracy of the rare class is low due to the lack of enough samples of the rare class caused by the uneven distribution of samples between classes is avoided.

[0014] 3) The best clustering number of kmeans obtained by combining the elbow method can be inconsistent with the actual crop class number, because the phenomenon of'same object and different spectrum' on the remote sensing image is reflected in the hierarchical sampling base map as one crop class corresponding to multiple layers. By analyzing the characteristics of the remote sensing image data, the sampling base map generated by clustering divides the overall into different categories and ensures that each category can allocate sample points, ensuring the intra-class diversity of the sample set to increase the representativeness of the sample set.

[0015] 4) The total sample number set in advance is compared with the theoretical total sample number calculated, and the optimal sample allocation mode of each layer is selected. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to make the application easier to understand, the application will be described in more detail by referring to the specific embodiments shown in the accompanying drawings. These drawings only depict typical embodiments of the application and should not be considered as limiting the scope of protection of the application.

[0017] Figure 1 Flowchart for one embodiment of the method of the application.

[0018] Figure 2 Flowchart for one embodiment of step S2 of the method of the application.

[0019] Figure 3 Flowchart for another embodiment of step S2 of the method of the application.

[0020] Figure 4 Flowchart for one embodiment of step S3 of the method of the application.

[0021] Figure 5 Flowchart for another embodiment of the method of the application.

[0022] Figure 6 Map of the study area location and the processed April image.

[0023] Figure 7Fig. for the relationship between the cluster number k of land use and SSE and the second derivative of SSE.

[0024] Figure 8 Fig. for the clustering result of surface type when k = 3.

[0025] Figure 9 Fig. for the NDVI time series curve of the clustering category.

[0026] Figure 10 Fig. for the relationship between the cluster number k of farmland and SSE and the second derivative of SSE.

[0027] Figure 11 Fig. for the clustering result of farmland area and k = 4.

[0028] Figure 12 Fig. for the classification result of equal stratified sampling. DETAILED DESCRIPTION

[0029] Embodiments of the present application will be described below with reference to the accompanying drawings so as to be better understood by those skilled in the art and to be implemented, but the listed embodiments are not intended to limit the present application, and the embodiments described below and the technical features in the embodiments can be combined with each other without conflict, wherein the same components are denoted by the same reference numerals.

[0030] The technical solution of the present application will be described in detail below taking the application scenario of crop remote sensing classification as an example. The application scenarios of the technical solution of the present application include but are not limited to the application scenario of crop remote sensing classification.

[0031] The technical idea of the present application is that, first, the image data set of the region to be classified is acquired, the features of the image data are extracted and analyzed, the target region of classification is extracted by using an unsupervised classification method, then secondary clustering is performed on the target region to determine the optimal number of clustering categories, and then each clustering category is taken as a layer, the number of samples in each layer is allocated according to the sample allocation principle, and finally the number and position of the samples are determined.

[0032] In one embodiment, as shown in FIG. 1, the method of the present application comprises S1-S4. Figure 1

[0033] S1: Acquire the image data set of the region to be classified, and set the total sample size N and the expected user accuracy. Step S1 comprises S11-S13.

[0034] S11: Preprocess the acquired remote sensing data set to be classified.

[0035] ​The remote sensing dataset is image data obtained by satellites, unmanned aerial vehicles, aerial photography and other means. It includes but is not limited to data of optical satellites and radar satellites, such as landsat, sentinel 1, sentinel 2, gaofen 1, gaofen 2, and other single-period, multi-period, single-band and multi-band remote sensing satellite data. The time phase range should consider the growth and phenology time of different crop types in the study area, and also need to consider the influence of cloud cover on data quality. For example, when selecting remote sensing data sets, data with cloud cover less than 10% or cloud cover less than 5% in the study area from April to September can be selected as input.

[0036] Then, batch data preprocessing is performed on the input remote sensing dataset, and the preprocessing steps include orthorectification, radiation correction, geometric registration, image cropping and splicing, etc.

[0037] S12: Set the total number of samples N.

[0038] According to experience and investigation needs, the total number of samples N to be collected is determined, such as N = 300.

[0039] S13: Set the expected classification accuracy.

[0040] In theory, the higher the classification accuracy of remote sensing image, the higher the requirement for samples and images to be classified, so the classification product meets the expected classification accuracy requirement, for example, setting the expected overall accuracy to 85%, or setting the user accuracy to 85%.

[0041] S2: Target area extraction. Separate the non-target ground type area such as buildings and water bodies from the farmland (or other vegetation cover) land class, extract the crop classification object area, and reduce the interference of non-target ground objects on sample layout. As shown in Figure 2 S2 includes S21-S26.

[0042] S21: Build classification features and extract features.

[0043] Extract the spectral features and vegetation indices RVI, GCVI, NDVI, NDBI and MNDWI of the remote sensing dataset processed in S1 as classification features. The calculation method is:

[0044] RVI = NIR / RED

[0045] GCVI = NIR / GREEN-1

[0046] NDVI = (NIR-RED) / (NIR+RED)

[0047] NDBI = (SWIR-NIR) / (SWIR+NIR)

[0048] MNDWI=(GREEN-MIR) / (GREEN+MIR)

[0049] Among them, RED is the red band, GREEN is the green band, NIR is the near-infrared band, MIR is the mid-infrared band, and SWIR is the short-wave infrared band.

[0050] S22: Unsupervised clustering. Based on the multidimensional feature data extracted in step S21, a preliminary classification of the remote sensing dataset for the study area is performed.

[0051] Preferably, an unsupervised classification method is used to perform preliminary classification processing on the remote sensing dataset of the study area. Unsupervised classification methods include, but are not limited to, k-means clustering and ISODATA. Unsupervised classification methods typically require specifying the number of clusters M as input beforehand. In this invention, within a certain range of cluster numbers [L1, L2], unsupervised clustering is performed sequentially on the values ​​of the cluster number L. For example, the land surface types existing in a study area are typically water bodies, cultivated land, forest land, and buildings; the range of L can be [4, 8], with L1 = 4 and L2 = 8.

[0052] S23: Determine the optimal number of clusters K1.

[0053] Preferably, a common method for determining the K1 value is the elbow method. The specific steps are as follows:

[0054] (1) Calculate the sum of squared errors (SSE) of the unsupervised clustering results corresponding to each cluster number L obtained in step S22, and plot the k-SSE curve in a two-dimensional Cartesian coordinate system.

[0055] The formula for calculating SSE is as follows:

[0056]

[0057] In the formula, C i The number of clusters L corresponds to the i-th cluster in the unsupervised clustering results, m i It is cluster C i The mean (centroid) of all sample features in the cluster, p is the cluster C. i Features of the sample points, count(C) i ) is a cluster C i The number of samples in the sample.

[0058]

[0059] In the formula, SSE M It is the sum of squared errors of the number of clusters L, |pm i | 2 It is cluster Ci The distance square of the sample point p and the centroid m. i

[0060] (2) Find the point with the maximum slope change rate (d2_max) in the k-SSE curve and take it as the inflection point. The number of clusters corresponding to the inflection point is the optimal number of clusters K1.

[0061] The calculation formula of the slope change rate is as follows:

[0062]

[0063]

[0064] In the formula, is the slope when the number of clusters is L, is the slope change rate when the number of clusters is L, and d2_max is the maximum value in the range of the number of clusters [L1, L2].

[0065] S24: Select the unsupervised clustering result.

[0066] According to the optimal number of clusters K1 obtained in S23, the result corresponding to the number K1 is selected from the unsupervised clustering result obtained in S22.

[0067] S25: Construct the threshold function rule of the decision tree.

[0068] Before constructing the decision tree, divide the current feature space of each node by the threshold function rule and determine the contained categories.

[0069] (1) For each image of step S1, calculate the NDVI values of all pixels in each cluster category in step S24 and take the average.

[0070] (2) Find the maximum value NDVI_Max and the minimum value NDVI_Min in the 5-10 month (set time range) from the NDVI values of the cluster category images calculated in step (1), and calculate the difference NDVI_Diff between the maximum and minimum values of NDVI in the 5-10 month. Compare NDVI_Max and the threshold value to distinguish the vegetation coverage area and the non-vegetation coverage area; compare NDVI_Diff and the threshold value to distinguish the farmland area and other vegetation coverage areas. The threshold function rule selects the vegetation index and its default value as shown in Table 1.

[0071] Table 1 Threshold function rule selects the vegetation index and its default value

[0072]

[0073] ​S26: According to the selected decision number of the classified features and the clustered object features of the image, a decision tree is constructed, an object-oriented method is used to identify different land surface types in the study area, and a target land area is extracted step by step. The farmland (or other vegetation cover) area extracted is the target area to be extracted in the study.

[0074] As shown in Figure 3 , the constructed decision tree is as follows:

[0075] (1): Each class of the unsupervised clustering result is regarded as an object. According to the time series image used, the NDVI maximum value NDVI_max in May-October is calculated. If NDVI_max is greater than the first threshold value (for example, the default value is 0.5), the class is a vegetation cover area. If NDVI_max is less than the first threshold value (for example, the default value is 0.4), the class is a non-vegetation cover area. More preferably, if NDVI_max is less than the second threshold value (for example, the default value is 0.4), the class is a non-vegetation cover area.

[0076] (2): For the vegetation cover area, the NDVI maximum minimum difference NDVI_Diff in May-October is calculated. If NDVI_Diff is greater than the set threshold value (for example, the default value is 0.4), the land class is a farmland area. If NDVI_Diff is less than the set threshold value (for example, the default value is 0.4), the class is other vegetation cover area.

[0077] S3: Based on the classified target area extracted in the preceding step, a clustering analysis is performed on the target area to determine the optimal clustering class number. Each clustering class is taken as a layer, and the number of samples in each layer is allocated according to the sample allocation principle to finally determine the number of samples and the location of the samples. As shown in Figure 1 , step S3 includes S31-S33.

[0078] S31: A hierarchical sampling base map is generated for the target area extracted in S2.

[0079] S31-1: The target area extracted in step S2 is used to mask the image, and classification features are constructed and extracted in the target area. This step is the same as step S21.

[0080] S31-2: Unsupervised clustering processing is performed. Within a certain clustering number value range [M1, M2], the clustering number M value is input in sequence to perform unsupervised clustering. This step is the same as step S22.

[0081] S31-3: The elbow method is used to determine the optimal clustering number K2. For example, in the northern region where the planting system is relatively simple, the value range of M can be [2, 8], M1=2, M2=8, and the best clustering number obtained according to the elbow method is K2=6. This step is the same as step S23.

[0082] S31-4: Obtain stratified sampling base map. Take the optimal cluster number K2 of S31-3 and the classification features of S31 as the input of the unsupervised clustering again, and the output of the unsupervised clustering is the stratified sampling base map. This step is the same as step S24.

[0083] S32: Assign the number of samples in each layer. As shown in Figure 4

[0084] S32-1: Calculate the theoretical total sample number N'.

[0085] After the stratified base map of step S3 and the expected classification accuracy requirement set in step S13, the reference theoretical total sample number N' for different classification accuracy requirements is calculated according to the formula proposed by Cochran, as follows:

[0086] 1) Calculate the theoretical total sample number N' according to the expected overall accuracy of simple random sampling method:

[0087] N' = z 2 O(1-O) / d 2 (5)

[0088] In the formula, N' is the theoretical total sample number; O is the expected overall accuracy set in step S13; d is the expected half-width of the confidence interval (usually 0.05); z is the percentile of the standard normal distribution (usually 1.96).

[0089] 2) Calculate the theoretical total sample number N' according to the expected user accuracy of each category of stratified sampling:

[0090]

[0091] Where S i i is the standard deviation of the i-th category, and U i is the expected user accuracy of the i-th category set in step S13.

[0092] When the image elements in the study area are very large, we have:

[0093]

[0094] In the formula, N' is the theoretical total sample number, is the standard error when estimating the overall accuracy (usually 0.05); W i is the area proportion of the i-th category.

[0095] S32-2: Set the evaluation rule: N < N'

[0096] S32-3: Select the appropriate sample allocation method in each layer: equal allocation method or area ratio allocation method. ​

[0097] The total sample number N set in the step S12 is compared with the size of the theoretical total sample number N' calculated in the step S32-1, and a suitable intra-layer sample distribution mode is selected through the judgment rule in the step S32-2.

[0098] (1) When N < N', the area ratio distribution mode is selected, that is, the total sample number is distributed to each layer according to the proportion of the area of each layer, and the sample number of each layer is obtained. The calculation formula is:

[0099]

[0100] In the formula, n h is the sample point number of the hth layer; S h is the pixel number of the hth layer; and N is the total sample point number.

[0101] (2) When N > N', the equal distribution mode is selected, that is, the total sample number is evenly distributed to each layer, and the sample number of each layer is the same. The calculation formula is:

[0102]

[0103] In the formula, n h is the sample point number of the hth layer; h is the layer number; and N is the total sample point number.

[0104] S33: stratified sampling.

[0105] On the basis of the sample number n h of each layer obtained in the step S32 and the sample base map obtained in the step S31, the spatial positions of the intra-layer samples are randomly distributed in the intra-layer.

[0106] S4: sample layout result. The sample layout map to be collected in the crop remote sensing classification task, the spatial positions and the number of the samples come from the step S3.

[0107] The present application has been experimentally verified. It is proved that the method of the present application is feasible, and can be used for crop sample layout and remote sensing image classification in a typical agricultural research area in Liaoyuan City, Liaoning Province.

[0108] The experimental steps are as follows: (1) After preprocessing the Sentinel-2 remote sensing image covering the study area, crop classification features are extracted from multiple period image data; (2) The elbow method is used to determine the optimal number of kmeans clustering, and the farmland area and the best sampling base map are obtained through secondary clustering; (3) According to the comparison of the calculated total sample quantity and the actual sample quantity, the allocation method of the in-layer sample quantity is determined, and the sample space distribution map is obtained; (4) The support vector machine classifier is used to classify the same image respectively, and finally a set of real ground samples is used to compare the classification accuracy of different sampling strategies.

[0109] The classification accuracy obtained by the experiment is 86.25%, 88.75%, 88.75%, 90.00% and 91.25% respectively, which is close to the classification accuracy (92.5%) of 388 samples obtained by traditional visual interpretation. It is proved that less sampling can also obtain the same and higher classification accuracy.

[0110] The specific implementation steps are as follows: Figure 5 ):

[0111] First, download the Sentinel-2 image product. Based on the Google Earth Engine platform, 6 Sentinel-2 images with less cloud coverage covering the study area are selected, and then the image is cropped and spliced according to the boundary of the study area. After processing, a set of multi-temporal images is obtained. Figure 6

[0112] Second, screen the band, calculate the vegetation index and extract the feature. According to the difference of the spectral characteristics of crops, the index is calculated: NDVI (Normalized Difference Vegetation Index, Normalized Difference Vegetation Index), NDWI (Normalized Difference Water Index, Normalized Vegetation Water Index), NDBI (Normalized Difference Built-up Index, Normalized Building Index), The specific calculation formula is as follows:

[0113] NDVI = (NIR-RED) / (NIR+RED) (10)

[0114] NDWI = (GREEN-NIR) / (GREEN+NIR) (11)

[0115] NDBI = (SWIR-NIR) / (SWIR+NIR) (12)

[0116] ​In the formula, GREEN, RED, SWIR and NIR respectively correspond to the image brightness values of the 3rd, 4th, 6th and 8th bands of the Sentinel-2 image band. Ten bands (B2-B8, B8A, B11, B12) and three vegetation indexes (NDVI, NDWI, NDBI) are extracted for each period of sentinel-2 satellite images, and a total of 78 features are extracted for six periods of sentinel-2 images.

[0117] Third step, kmeans unsupervised classification and elbow method to determine the number of surface types. First, use the kmeans clustering algorithm to perform rough clustering on the time series remote sensing image, and each cluster category is a surface type. The true surface type of each cluster category is temporarily unknown. The core index of the elbow method is SSE (sum of the squared errors), which is the clustering error of all samples, representing the good or bad of the clustering effect. Let k start from 1 until a suitable upper limit is taken, cluster for each k value and record the corresponding SSE, then draw the relationship diagram of k and SSE, and finally select the elbow corresponding k as the best clustering number. When the elbow is difficult to distinguish, the second derivative is usually used to get the optimal value. The calculation formula of SSE is as formula (1).

[0118] The k value (k=3) corresponding to the first elbow is selected as the optimal clustering number ( Figure 7 ), and the corresponding clustering result diagram ( Figure 8 ) is obtained. Since rough clustering is performed on the entire study area, the cluster category may not correspond to the true surface type, but it has little effect on the final classification accuracy.

[0119] Fourth step, select threshold function rule and object-oriented decision tree classification. In this experiment, the default threshold of the decision tree is not changed, each cluster category is regarded as an object, and the NDVI time series curve of the three cluster categories is calculated ( Figure 9 ), and the classification index value and decision tree classification result are obtained (Table 2).

[0120] Fifth step, extract farmland area. According to the classification result of the decision tree, the farmland area is extracted ( Figure 11 ).

[0121] Sixth step, kmeans unsupervised classification and elbow method to determine the best layer number. Repeat the operation of the third step on the farmland area, and each cluster category corresponds to a layer, and the true surface type of each cluster category is temporarily unknown. Select the k value (k=4) corresponding to the first elbow as the best layer number ( Figure 10 ), and obtain the corresponding clustering result ( Figure 11). Due to the existence of "same object different spectrum" and "different spectrum of the same object" phenomenon in crop classification process, the clustering category and the real category may exist one-to-one, one-to-many, many-to-one and many-to-many.

[0122] Step 7, calculation of the theoretical total sample size. According to formula (2) and formula (3), when the confidence level is 95%, the expected half-width of the confidence interval is 0.05, the expected overall accuracy of simple random sampling is 85%, the theoretical total sample size is calculated to be 195; when the confidence level is 95%, the standard error is 0.01, and the user accuracy in the sample classification result of traditional artificial visual interpretation is used (Table 3), the theoretical total sample size of stratified random sampling is calculated to be 508.

[0123] Step 8, stratified sampling sample allocation method within the layer. On the basis of the total sample size of 25, 49, 100, 169 and 225, the sample size within each layer is allocated equally (Table 4). The sample space position of each layer is obtained by simple random sampling method.

[0124] Step 9, support vector machine (Support Vector Machine, SVM) supervised classification and classification accuracy evaluation. The true value of the sample obtained in step 6 is given by field sampling or professional visual interpretation, and the obtained sample is input into the SVM classifier. Using a certain number of validation sample data, the overall accuracy (overall accuracy, OA) and Kappa coefficient are calculated as evaluation indexes by using confusion matrix. The sample and crop distribution result graph are output Figure 12 ), Figure 12 , in which (a) is the classification result graph of 388 samples obtained by traditional visual interpretation; (b)-(f) are the classification results and sample distribution graphs of different total sample sizes of equal stratified sampling.

[0125] Table 2 Classification index values and decision tree classification results

[0126]

[0127] Table 3 Classification results of traditional artificial visual interpretation

[0128]

[0129] Table 4 Equal stratified sampling method within the layer sample size and its classification accuracy

[0130]

[0131] The above-described embodiments are merely preferred specific embodiments of the present application. The phrase "in an embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments" used in the specification can refer to one or more of the same or different embodiments according to the present disclosure. Common variations and replacements made by those skilled in the art within the technical scope of the present application should be included in the protection scope of the present application.

Claims

1. A method for deploying ground samples for remote sensing classification of crops based on scene adaptation, characterized in that, include: S1, Obtain the image dataset of the study area to be classified, and set the total sample size N and the expected classification accuracy; S2, extract and analyze the features of the image data, and use an unsupervised classification method to extract the classification target region, including: S21, constructing classification features and extracting features; S22, performing unsupervised clustering processing on the multidimensional feature data extracted in step S21; S23, determining the optimal number of clusters K1; S24, selecting the corresponding unsupervised clustering results based on the optimal number of clusters K1; S25, constructing a decision tree threshold function rule; S26, based on the aforementioned decision tree, using an object-oriented method to identify different land surface types within the study area to be classified, and extracting the classification target region; S3, Based on the classification target region extracted in the previous steps, perform cluster analysis on the target region to determine the optimal number of clusters K2. Using each cluster as a layer, allocate the number of samples within each layer according to the sample allocation principle, ultimately determining the number and location of samples. This includes: S31: Generating a stratified sampling base map: Using the classification target region extracted in step S2, mask the image of the study area to be classified, construct and extract classification features within the classification target region; perform unsupervised clustering processing, sequentially performing unsupervised clustering on the number of clusters M within a certain range of cluster size values; use the elbow method to determine the optimal number of clusters K2; using the optimal number of clusters K2 as input, the output unsupervised clustering result is the stratified sampling base map; S32: Allocating the number of samples within a layer: Calculate the theoretical total number of samples N'; set the evaluation rule: N < N'; select the sample allocation method within the layer: when N < N', select the allocation method based on area ratio; when N > N', select the equal allocation method; S33: Stratified sampling.

2. The deployment method according to claim 1, characterized in that, In step S21, the spectral features and vegetation indices of the remote sensing dataset processed in S1 are extracted, including: red band, green band, blue band, near-infrared band and red edge band; RVI, GCVI, NDVI, NDBI and MNDWI.

3. The deployment method according to claim 2, characterized in that, Step S23 includes: Calculate the sum of squared errors SSE corresponding to each cluster number L obtained in step S22 and plot the k-SSE curve in a two-dimensional Cartesian coordinate system. Find the point in the k-SSE curve with the largest rate of change of slope and take it as the inflection point. The number of clusters corresponding to the inflection point is the optimal number of clusters K1.

4. The deployment method according to claim 3, characterized in that, Step S25 includes: For each image in step S1, calculate the NDVI value of all pixels in one cluster category in step S24 and take the average. Find the maximum value NDVI_Max and minimum value NDVI_Min within a set time range from the calculated NDVI values ​​of each image in the cluster category, and calculate the difference between the maximum and minimum NDVI values ​​within the set time range, NDVI_Diff. By comparing NDVI_Max with a threshold, vegetation-covered areas and non-vegetation-covered areas are distinguished; by comparing NDVI_Diff with a threshold, farmland areas and other vegetation-covered areas are distinguished.

5. The deployment method according to claim 4, characterized in that, Step S26 includes: Each class of the unsupervised clustering results is treated as an object. Based on the time series image used, the maximum NDVI value NDVI_max within a set time range is calculated. If NDVI_max is greater than the first threshold, the class is a vegetation-covered area; if NDVI_max is less than the first threshold, the class is a non-vegetation-covered area. For vegetated areas, the maximum and minimum difference between NDVI and NDVI, NDVI_Diff, is calculated within a set time range. If NDVI_Diff is greater than a set threshold, the land type is farmland; if NDVI_Diff is less than a set threshold, the type is other vegetated areas.

Citation Information

Patent Citations

  • County-level crop planting area remote sensing statistic sampling survey scheme design method

    CN107527014A

  • Polarized radar remote sensing image-based flood inundation range automatic extraction method

    CN110570462A