Field-level crop extraction method based on SAR and optical images

By fusing SAR and optical remote sensing image data, using GAN and semantic segmentation networks, the problem of difficult to obtain high-precision field-scale crop spatial distribution information in the prior art is solved, and efficient and high-precision field-level crop extraction is achieved.

CN118397412BActive Publication Date: 2025-05-06INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI

Patent Information

Application Number
CN202410505675.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-05-06
Estimated Expiration
2044-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-precision acquisition of field-scale crop spatial distribution information, especially when optical images cannot achieve high-precision mapping due to interference from clouds and fog.

Method used

By fusing SAR remote sensing image data with optical remote sensing image data, using a generative adversarial network (GAN) and semantic segmentation network, image fusion and field extraction are performed to achieve efficient and high-precision field-level crop extraction.

Benefits of technology

It realizes efficient and high-precision field-level crop extraction, which can effectively supplement the shortcomings of optical images and provide more accurate field information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397412B_ABST
    Figure CN118397412B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of agricultural remote sensing, and relates to a field-level crop extraction method integrating SAR and optical images, including: SAR remote sensing image data, optical remote sensing image data, ground crop sample data and field data; filtering processing; cloud removal processing; cropping and clipping processing after superposition and registration, and establishing a sample data set of fused images; obtaining a fusion learning generation network of fused images based on SAR remote sensing image data and optical remote sensing data; generating a fusion data set; constructing a field extraction semantic segmentation network model, and obtaining a field identification semantic segmentation network; extracting field information using the field identification semantic segmentation network, and obtaining a field mask area; and performing supervised learning on crops in the field mask area to obtain crop classification results. The present invention realizes efficient and high-precision mapping in the field crop extraction process by integrating SAR remote sensing image data and optical image data and utilizing dense time series remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of agricultural remote sensing, and in particular relates to a field-level crop extraction method integrating SAR and optical images. Background Art

[0002] Accurate information on farmland use is of great significance for rationally planting crop types, improving farmland utilization, and ensuring food security. Traditional agricultural monitoring methods such as manual surveys and aerial photography have problems such as low efficiency, high cost, and long cycles. With the development of high-temporal and spatial resolution remote sensing data and information technology, agricultural informatization urgently needs to solve the problem of how to obtain high-precision field-scale crop spatial distribution information in a timely manner.

[0003] Seasonality is one of the most prominent characteristics of crops. The phenological evolution of each type of crop produces a unique temporal distribution of spectral reflectance. With the development of remote sensing technology, many image data with different spatial resolutions have emerged, which carry rich spectral information and can quickly obtain surface information. Therefore, multi-temporal remote sensing data has become an effective data source for monitoring crop growth dynamics and classifying them.

[0004] At present, crop identification methods based on machine learning mainly rely on processes such as optical image feature extraction and data fusion. However, invalid pixels generated by interference such as clouds and fog in optical images make it impossible to achieve high-precision field-scale mapping. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a field-level crop extraction method integrating SAR and optical images, comprising:

[0006] Collect SAR remote sensing image data, optical remote sensing image data, ground crop sample data and field data;

[0007] The amplitude pixel information of SAR remote sensing image data is filtered and resampled after coordinate conversion;

[0008] Calculate remote sensing index data based on optical remote sensing image data, use remote sensing index data to enhance optical remote sensing image features, perform cloud removal on the optical remote sensing image after enhanced image features, and resample after coordinate conversion;

[0009] The resampled SAR remote sensing image data, optical remote sensing image data and field data are superimposed and registered, and then the clippings are processed to establish the sample data set required for image fusion network training;

[0010] Based on the generative adversarial network, the data in the sample data set is used to train the generative adversarial network, and a fusion learning generative network based on the fusion image of SAR remote sensing image data and optical data is obtained;

[0011] A fusion learning generative network is used to generate a fusion dataset of SAR remote sensing image data and optical remote sensing image data based on a sample dataset;

[0012] Based on the fusion data set of SAR remote sensing image data and optical remote sensing image data and field data, a field extraction semantic segmentation network model is constructed. Field feature information is learned through the field extraction semantic segmentation network model to obtain a field extraction semantic segmentation network.

[0013] The field information is extracted using the field extraction semantic segmentation network to obtain the field mask area;

[0014] Using the data in the sample dataset, supervised learning is performed on the crops in the masked area of ​​the field to obtain the crop classification results.

[0015] Based on the above technical solution, the present invention can also be improved as follows.

[0016] Furthermore, the SAR remote sensing image data are Sentinel-1GRD amplitude image data during the crop growth period, the scanning mode is IW, and the polarization mode is VV polarization; the optical remote sensing image data are several spectral band data of Sentinel-2MS I L2A optical image data; the field data are labeled raster data of field distribution; and the ground crop sample data are ground sample points containing latitude and longitude information and crop types.

[0017] Furthermore, the remote sensing index data include the Normalized Difference Vegetation Index NDVI, the Chlorophyll Vegetation Index GCVI, the Land Surface Water Index LSWI, the Normalized Difference Building Index NDBI, the Tassel Cap Brightness Coefficient TCB and the Modified Normalized Difference Water Index MNDWI;

[0018] The optical remote sensing image data is declouded according to the remote sensing index data, including: using a declouding algorithm to perform cloud projection through dark pixels to produce a cloud mask for declouding, setting a fixed time domain and time step, performing median synthesis within the time domain and time step, and reducing the hole values ​​generated by the declouding algorithm.

[0019] Furthermore, the resampled SAR remote sensing image data, optical remote sensing image data and field data are superimposed and registered, and then clipped to establish a sample data set, including:

[0020] Set the slice image size, retain the set ratio of overlapping pixels for rectangular sliding window slicing, and set N W is the number of image slice rows, N H is the number of image slice columns, Width is the image length, Height is the image width, Clip x is the pixel length of the image slice, Clipy is the pixel width of the image slice, Step x is the pixel size of the overlap between adjacent slices, Step y is the pixel size of adjacent slice overlap, the image slice formula is:

[0021]

[0022]

[0023] Let the number of final slice rows be N Wend , the final number of slice columns is N Hend , by adding the first slice of the reverse sliding window to preserve data integrity, the extended slice calculation formula is:

[0024] N Wend =math.floor(N W )+1;

[0025] N Hend =math.floor(N H )+1.

[0026] Furthermore, the fusion learning generation network is an FGAN network; the FGAN network includes two branch architectures, each of which is based on the Unet model and modifies the feature extraction network. After extracting low-dimensional semantic information with two layers of two-dimensional convolution with a kernel size of 5×5, three multi-layer convolutions with a kernel size of 3×3 are superimposed to obtain feature maps of different scales; the ASPP spatial pyramid pooling structure is introduced to obtain feature maps of several scales, and then four CatConv2d convolutions with a kernel size of 1×1 are superimposed on heterogeneous channels of the same scale for convolution, and then the CBAM attention mechanism is added to increase the weight of important channel features. Finally, multi-scale features are learned through upsampling and superposition convolution to restore the image size.

[0027] Furthermore, the FGAN network uses the PathGAN multi-block discriminator to expand and discriminate the image blocks; the SSIM loss function is used in the discriminator to calculate the similarity between images based on brightness, contrast and structural information.

[0028] Furthermore, the data in the sample data set are used to conduct supervised learning on the crops in the field mask area to obtain crop classification results, including: after sample cleaning of the crop sample set based on the field mask area, different band combinations of the input data are performed, and the Xgboost gradient boosting tree model is used to perform field-level crop classification training and prediction, the overall classification accuracy and Kappa coefficient are selected to evaluate the overall crop classification accuracy, and the band combination with the highest classification accuracy is selected to perform crop classification mapping in the field mask area.

[0029] Furthermore, let TN represent the number of correctly classified background object pixels, TP represent the number of correctly classified target object pixels, FN represent the number of target object pixels misclassified as background object pixels, FP represent the number of background object pixels misclassified as target object pixels, OA represent the overall classification accuracy, p e It represents the quotient of the sum of the products of the actual and predicted numbers of all objects and the square of the total number of samples, then:

[0030]

[0031]

[0032]

[0033] The beneficial effect of the present invention is that the present invention realizes efficient and high-precision mapping in the process of field crop extraction by fusing SAR remote sensing image data and optical image data and utilizing dense time series remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 The schematic diagram of the field-level crop extraction method integrating SAR and optical images provided by the present invention;

[0035] Figure 2 This is the structural principle diagram of the FGAN network;

[0036] Figure 3 This is the structural principle diagram of each module of the FGAN network;

[0037] Figure 4 This is the schematic diagram of the ASPP spatial pyramid pooling structure;

[0038] Figure 5 Schematic diagram of the CBAM attention structure;

[0039] Figure 6 This is the schematic diagram of the discriminator network of the FGAN network;

[0040] Figure 7 Schematic diagram of the structure of the semantic segmentation network for field extraction. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0042] As an example, Figure 1As shown, in order to solve the above technical problems, this embodiment provides a field-level crop extraction method integrating SAR and optical images, including:

[0043] Collect SAR remote sensing image data, optical remote sensing image data, ground crop sample data and field data;

[0044] The amplitude pixel information of SAR remote sensing image data is filtered and resampled after coordinate conversion;

[0045] Calculate remote sensing index data based on optical remote sensing image data, use remote sensing index data to enhance optical remote sensing image features, perform cloud removal on the optical remote sensing image after enhanced image features, and resample after coordinate conversion;

[0046] The resampled SAR remote sensing image data, optical remote sensing image data and field data are superimposed and registered, and then the clippings are processed to establish the sample data set required for image fusion network training;

[0047] Based on the generative adversarial network, the data in the sample data set is used to train the generative adversarial network, and a fusion learning generative network based on the fusion image of SAR remote sensing image data and optical data is obtained;

[0048] A fusion learning generative network is used to generate a fusion dataset of SAR remote sensing image data and optical remote sensing image data based on a sample dataset;

[0049] Based on the fusion data set of SAR remote sensing image data and optical remote sensing image data and field data, a field extraction semantic segmentation network model is constructed. Field feature information is learned through the field extraction semantic segmentation network model to obtain a field extraction semantic segmentation network.

[0050] The field information is extracted using the field extraction semantic segmentation network to obtain the field mask area;

[0051] Using the data in the sample dataset, supervised learning is performed on the crops in the masked area of ​​the field to obtain the crop classification results.

[0052] Optionally, the SAR remote sensing image data are Sentinel-1GRD satellite data during the crop growth period, the scanning mode is IW, and the polarization mode is VV polarization; the optical remote sensing image data are several spectral band data of Sentinel-2MS IL2A satellite data; the field data are field distribution label raster data; and the ground crop sample data are ground sample points containing longitude and latitude information and crop types.

[0053] Spectral band data include visible light (B2~B4 band), red edge (B5~B7 band), near infrared (B8 / B8A band) and shortwave infrared data (B11 / B12 band).

[0054] The amplitude pixel information of SAR remote sensing image data is filtered. In order to reduce the error influence of speckle noise and retain the edge details of the image, the amplitude pixel information of SAR remote sensing image data is filtered by using Lees igma algorithm through a 5×5 sliding window. In order to facilitate the image registration of SAR remote sensing image data and optical remote sensing image data, the image coordinate system of SAR remote sensing image data is converted to WGS84 coordinate system (World Geodetic System) and the image azimuth and range resolutions are resampled to 10m. SAR remote sensing image data is formed by radar backscattering, so it is easy to be disturbed and noisy. It is smoothed by SG (Savitzky-Golay, threshold filtering algorithm) smoothing filter to make it easier to reflect the change trend reflected by the data itself.

[0055] Optionally, remote sensing index data include normalized difference vegetation index NDVI, chlorophyll vegetation index GCVI, land surface moisture index LSWI, normalized difference building index NDBI, tassel cap brightness coefficient TCB and modified normalized difference water index MNDWI;

[0056] The remote sensing image after enhancing the image features is subjected to cloud removal processing, including: using a cloud removal algorithm to perform cloud projection through dark pixels to produce a cloud mask for cloud removal processing, and setting a fixed time domain and time step, performing median synthesis within the time domain and time step, and reducing the hole values ​​generated by the cloud removal algorithm.

[0057] In the actual application process, in order to align with the SAR remote sensing image data, the optical remote sensing image data is projected to the WGS84 coordinate system and then the B5 / 6 / 11 / 12 bands of the 20m resolution data are resampled to 10m resolution.

[0058] Optionally, the resampled SAR remote sensing image data, optical remote sensing image data and field data are superimposed and registered, and then clipped to establish a sample data set of fused images, including:

[0059] Set the slice image size, retain the set ratio of overlapping pixels for rectangular sliding window slicing, and set N W is the number of image slice rows, N H is the number of image slice columns, Width is the image length, Height is the image width, Clip x is the pixel length of the image slice, Clip y is the pixel width of the image slice, Stepx is the pixel size of the overlap between adjacent slices, Step y is the pixel size of adjacent slice overlap, the image slice formula is:

[0060]

[0061]

[0062] In actual application, since SAR remote sensing image data and optical remote sensing image data are rich in image spectral information and have a large width, if they are used directly as samples, extremely strong computing power is required, information reading and writing is slow, and it is easy to cause computer memory overflow. In order to facilitate efficient data access, the SAR remote sensing image data, optical remote sensing image data and the above-mentioned label raster data are superimposed and then cut and processed. Optionally, the slice image size is set to 512*512*(8+1+1), the first 8 layers are Sentinel-2 optical image B4 / 3 / 2 / 8 / 5 / 6 / 11 / 12 bands, the 9th layer is the processed S1 image, and the last layer is the sample label data. Generally, in order to ensure the integrity of the slice boundary information, 25% of the overlapping pixels are retained for rectangular sliding window slicing. In order to facilitate network training, the image slice pixel length Clip is generally taken. x Clip pixel width y The same value, optional, is set to 512; the overlapping pixel size of adjacent slices Step x Overlap pixel size with adjacent slices Step y are all set to 128. In practical applications, the irregularity of remote sensing images may result in the number of image slices N W and the number of image slice columns N H If it is not an integer, the first slice of the reverse sliding window is added to preserve data integrity.

[0063] Let the number of final slice rows be N Wend , the final number of slice columns is N Hend , by adding the first slice of the reverse sliding window to preserve data integrity, the extended slice calculation formula is:

[0064] N Wend =math.floor(N W )+1;

[0065] N Hend =math.floor(N H )+1.

[0066] Optionally, the fusion learning generation network is an FGAN network (Fusion Generative Adversar ialNetwork); the FGAN network includes two branch architectures, each of which is based on the Unet model and modifies the feature extraction network. After extracting low-dimensional semantic information with two layers of two-dimensional convolution with a kernel size of 5×5, three multi-layer convolutions with a kernel size of 3×3 are superimposed to obtain feature maps of different scales; after introducing the ASPP spatial pyramid pooling structure to obtain feature maps of several scales, four CatConv2d convolutions with a kernel size of 1×1 (Cat represents superimposed links of channels from different sources) are used to superimpose heterogeneous channels of the same scale for convolution, and then a CBAM attention mechanism is added to increase the weights of important channel features. Finally, multi-scale features are learned through upsampling and superposition convolution to restore the image size.

[0067] Optionally, the FGAN network uses a PathGAN (Pathways Generative Adversar ial Network) multi-block discriminator to expand and discriminate image blocks; the SSIM loss function is used in the discriminator to calculate the similarity between images based on brightness, contrast and structural information.

[0068] The FGAN network structure introduces the ASPP spatial pyramid pooling structure to obtain the influencing features of different scales, introduces CBAM as the attention mechanism, and adaptively learns the channel features of different scales. It introduces the conventional 1*1 two-dimensional convolution to perform channel dimension increase and decrease and superimpose convolution with the heterogeneous image feature map, and combines the upsampling structure to obtain image semantic information at multiple scales to ensure the integrity of the information.

[0069] As attached Figure 2 In the FGAN structure shown in the figure, the upper and lower groups of input units input SAR remote sensing image data and optical data respectively. Each input unit is composed of two low-level semantic extraction modules. The output end of each input unit is connected to a multidimensional semantic extraction unit composed of three multidimensional semantic extraction modules connected end to end in sequence. One output end of each multidimensional semantic extraction unit is connected to a two-dimensional convolution module with a kernel of 3, and the other output end is connected to a pooling unit composed of four hollow pyramid modules. The four hollow pyramid modules are connected end to end in sequence, and the output end of the two-dimensional convolution module with a kernel of 3 is connected to the input end of the hollow pyramid module. One output end of the pooling unit is connected to an input end of the upsampling and superposition module through the superposition convolution module and the attention module, and the other output end of the pooling unit is connected to the other input end of the upsampling and superposition module. The output end of the upsampling and superposition module is connected to the activation layer output through a two-dimensional convolution with a kernel of 1 to obtain the field image extraction result.

[0070] As attached Figure 3As shown, the low-level semantic extraction module includes a conventional two-dimensional convolution with a kernel of 4, a batch normalization layer, and a LeakyRe lu activation layer connected in sequence; the multi-dimensional semantic extraction module includes a low-level semantic extraction module arranged in sequence, two groups of conventional two-dimensional convolutions with a kernel of 1 arranged in parallel and respectively connected to the output end of the low-level semantic extraction module, each conventional two-dimensional convolution output end is connected to a batch normalization layer, and the output ends of the two batch normalization layers are connected to the LeakyRe lu activation layer. The two-dimensional convolution module includes a conventional two-dimensional convolution with a kernel of 3, a batch normalization layer, and a Re lu activation function connected in sequence. In the upsampling and superposition module, the upsampling module is connected to the input end of the semantic extraction module composed of a conventional two-dimensional convolution with a kernel of 3, a batch normalization layer, and a LeakyRe lu activation layer; the output end of the semantic extraction module is connected to the input end of the semantic extraction module composed of a conventional two-dimensional convolution with a kernel of 1, a batch normalization layer, and a LeakyRe lu activation layer.

[0071] The ASPP spatial pyramid pooling structure is shown in the attached Figure 4 As shown in the figure, after the feature map is input, the convolution layer with a conventional convolution step size of 1, the dilated convolution layer with three different dilation rates of 6, 12 and 18, and the global average pooling layer are connected respectively. The superposition of the three types of outputs is the output feature map. The dilated convolution in this structure can expand the receptive field of the convolution kernel, reduce the number of parameters and improve efficiency.

[0072] Let K Dilated is the convolution kernel size after dilation of the dilated convolution, k i is the regular convolution size of the i-th layer, d i is the hole filling size, RF i is the receptive field size, RF i+1 is the receptive field size of the i+1th convolution layer, S i is the step size, Stride i is the product of all layer steps, then:

[0073] K Dilated =k i +(k i -1)×(d i -1);

[0074] RF i+1 =RF i +(K Dilated -1)×Stride i ;

[0075]

[0076] After the two branches obtain feature maps of different scales, they are convolved by superimposing heterogeneous channels of the same scale through four two-dimensional convolutions with a kernel size of 1×1, and then the CBAM (Convolut ional Block Attention Module, used to enhance the attention mechanism of convolutional neural networks) attention structure is added to increase the weight of important channel features. Finally, multi-scale features are learned by upsampling and superimposing convolutions to restore the image size.

[0077] As attached Figure 5 As shown, the CBAM attention structure includes a channel attention module and a spatial attention module.

[0078] After the feature map is input, it enters the multilayer perceptron through the maximum pooling layer and average pooling layer of the channel attention module, and obtains the channel attention weight after the activation function; the input feature map is input by the input feature map module, and enters the two-dimensional convolution and activation function through the maximum pooling layer and average pooling layer of the channel attention module to obtain the spatial attention weight.

[0079] The discriminator network of the FGAN network is shown in the attached Figure 6 As shown, a PathGAN multi-block discriminator is used, including an input image unit, a downsampling layer, an activation layer, and an output feature map unit. Specifically, the downsampling layer includes four groups of downsampling units. The downsampling unit includes a conventional convolution with a kernel of 4, a batch normalization layer (Batch Normalization, i.e., BN layer) and an activation layer, and the activation layer is a Leakyrelu activation function. The activation layer includes an edge mirror filling unit, a conventional convolution with a kernel of 4, and an activation layer, and the activation layer is a Sigmoid activation layer. Conventional convolution uses two-dimensional convolution, and a conventional convolution with a kernel of 4 is represented as 4*4Conv2d.

[0080] Compared with the traditional GAN ​​overall discrimination, the PathGAN multi-block discriminator is used to expand and discriminate the image blocks to evaluate the authenticity of each image block and obtain more accurate generated images. The SSIM loss function is used in the discriminator to more accurately reflect the similarity between images and reduce the difference between adjacent pixels by considering brightness, contrast and structural information. The mean of the generated image is expressed as μ X , the mean of the original image is represented by μ Y , the variance of the generated image is expressed as σ X , the variance of the original image is expressed as σ Y , C1 and C2 are constants, then the simplified SSIM index calculation formula is:

[0081]

[0082] In the process of constructing the semantic segmentation network for field extraction, the SepConv2d deep separable convolution module is introduced as the backbone network with the improved Xcept i on structure to improve the network fitting performance, and the conventional 1*1Conv2d convolution is introduced to retain richer feature map extraction information, and the ASPP porous pyramid pooling architecture is added to obtain high-dimensional semantic features. The up-sampling and down-sampling feature maps of the same scale are combined with 1*1Conv2d+3*3Conv2d multi-layer convolution to add more weight parameters to accurately distinguish field information.

[0083] As an optional implementation, as shown in the attached Figure 7 As shown in the figure, the structural principle diagram of the field extraction semantic segmentation network, the fused image is input through the image input unit, and the first downsampling is completed through the superposition module of the low-level semantic convolution module and the depthwise separable convolution module, and the second downsampling is completed after iteration; then it enters the depthwise separable convolution module for multiple iterations, and enters the superposition module for iteration again to complete the third downsampling; after the third downsampling, it passes through the void space pooling pyramid layer and then enters the upsampling layer, and enters the two groups of low-level semantic convolution modules together with the second downsampling data to reach the random dropout layer, and after passing through another sampling layer, it enters the two groups of low-level semantic convolution modules together with the first downsampling data to reach the random dropout layer, and passes through the upsampling layer, the conventional two-dimensional convolution with a kernel of 1, and the upsampling layer in turn to obtain the predicted image as the output.

[0084] The low-level semantic convolution module includes two sets of conventional two-dimensional convolution, batch normalization layer and Relu activation layer. The depthwise separable convolution module includes depthwise separable convolution, batch normalization layer and Relu activation layer. The multidimensional semantic extraction module includes conventional two-dimensional convolution, batch normalization layer and Relu activation layer. Upsampling is used to restore the image size, and the image output unit obtains the field distribution map.

[0085] Optionally, supervised learning is performed on crops in field mask areas using data in the sample data set to obtain crop classification results, including: cleaning the crop sample set based on the field mask area, combining the input data with different bands, using the Xgboost gradient boosting tree model to perform field-level crop classification training and prediction, selecting the overall classification accuracy and Kappa coefficient to evaluate the overall crop classification accuracy, and selecting the band combination with the highest classification accuracy to perform crop classification mapping on the field mask area.

[0086] Optionally, let TN represent the number of correctly classified background object pixels, TP represent the number of correctly classified target object pixels, FN represent the number of target object pixels misclassified as background object pixels, FP represent the number of background object pixels misclassified as target object pixels, OA represent the overall classification accuracy, p eIt represents the quotient of the sum of the products of the actual and predicted numbers of all objects and the square of the total number of samples, then:

[0087]

[0088]

[0089]

[0090] Compared with optical images, SAR remote sensing image data has the advantages of full-time and cloud-penetrating, can provide large-area coverage with high spatial resolution and high frequency, and can supplement invalid pixels generated by interference of cloud and fog in optical images. Therefore, the present invention can make full use of multivariate remote sensing sequence image information, and realize efficient and high-precision mapping in the process of field crop extraction by fusing SAR remote sensing image data with optical image data and using dense time series remote sensing images.

[0091] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A field-level crop extraction method integrating SAR and optical images, characterized in that: include: Collect SAR remote sensing image data, optical remote sensing image data, ground crop sample data and field data; The amplitude pixel information of SAR remote sensing image data is filtered and resampled after coordinate conversion; Calculate remote sensing index data based on optical remote sensing image data, use remote sensing index data to enhance optical remote sensing image features, perform cloud removal on the optical remote sensing image after enhanced image features, and resample after coordinate conversion; The resampled SAR remote sensing image data, resampled optical remote sensing image data and field data are superimposed and registered, and then the clipping is processed to establish the sample data set required for image fusion network training; Based on the generative adversarial network, the data in the sample data set is used to train the generative adversarial network, and a fusion learning generative network based on the fusion image of SAR remote sensing image data and optical remote sensing image data is obtained; A fusion learning generative network is used to generate a fusion dataset of SAR remote sensing image data and optical remote sensing image data based on a sample dataset; Based on the fusion data set of SAR remote sensing image data and optical remote sensing image data and field data, a field extraction semantic segmentation network model is constructed. Field feature information is learned through the field extraction semantic segmentation network model to obtain a field extraction semantic segmentation network. The field information is extracted using the field extraction semantic segmentation network to obtain the field mask area; Using the data in the sample dataset, supervised learning is performed on the crops in the masked area of ​​the field to obtain the crop classification results.

2. The field-level crop extraction method by fusing SAR and optical images according to claim 1, characterized in that: The SAR remote sensing image data are Sentinel-1GRD satellite data during the crop growth period, with an IW scanning mode and a VV polarization mode; the optical remote sensing image data are several spectral band data of Sentinel-2MSI L2A satellite data; the field data are field distribution label raster data; and the ground crop sample data are ground sample points containing latitude and longitude information and crop types.

3. The field-level crop extraction method of fusing SAR and optical images according to claim 1, characterized in that: Remote sensing index data include normalized difference vegetation index NDVI, chlorophyll vegetation index GCVI, land surface moisture index LSWI, normalized difference building index NDBI, tassel cap brightness coefficient TCB and modified normalized difference water index MNDWI; The optical remote sensing image after enhancing the image features is subjected to cloud removal processing, including: using a cloud removal algorithm to perform cloud projection through dark pixels to produce a cloud mask for cloud removal processing, and setting a fixed time domain and time step, and performing median synthesis within the time domain and time step.

4. The field-level crop extraction method of fusing SAR and optical images according to claim 1, characterized in that: The resampled SAR remote sensing image data, the resampled optical remote sensing image data and the field data are superimposed and registered, and then the clipping is processed to establish a sample data set of fused images, including: Set the slice image size, retain the set ratio of overlapping pixels for rectangular sliding window slicing, and set is the number of image slice rows, is the number of image slice columns, is the image length, is the image width, is the pixel length of the image slice, is the pixel width of the image slice, is the pixel size of the overlap between adjacent slices, is the pixel size of adjacent slice overlap, the image slice formula is: ; ; Let the final number of slice rows be , the final number of slice columns is , by adding the first slice of the reverse sliding window to preserve data integrity, the extended slice calculation formula is: ; 。 5. The field-level crop extraction method of fusing SAR and optical images according to claim 1, characterized in that: The fusion learning generation network is the FGAN network; the FGAN network includes two branch architectures, each of which is based on the Unet model and modifies the feature extraction network. After extracting low-dimensional semantic information with two layers of two-dimensional convolution with a kernel size of 5×5, three multi-layer convolutions with a kernel size of 3×3 are superimposed to obtain feature maps of different scales; the ASPP spatial pyramid pooling structure is introduced to obtain feature maps of several scales, and then four CatConv2d convolutions with a kernel size of 1×1 are superimposed on heterogeneous channels of the same scale for convolution, and then the CBAM attention mechanism is added to increase the weight of important channel features. Finally, multi-scale features are learned through upsampling and superposition convolution to restore the image size.

6. The field-level crop extraction method by fusing SAR and optical images according to claim 5, characterized in that: The FGAN network uses the PathGAN multi-block discriminator to expand and discriminate image blocks; the SSIM loss function is used in the discriminator training process to calculate the similarity between images based on brightness, contrast and structural information.

7. The field-level crop extraction method by fusing SAR and optical images according to claim 1, characterized in that: Using the data in the sample data set, supervised learning is performed on the crops in the field mask area to obtain the crop classification results, including: after sample cleaning of the crop sample set based on the field mask area, different band combinations are performed on the input data, and the Xgboost gradient boosting tree model is used to perform field-level crop classification training and prediction. The overall classification accuracy and Kappa coefficient are selected to evaluate the overall crop classification accuracy, and the band combination with the highest classification accuracy is selected to perform crop classification mapping in the field mask area.

8. The field-level crop extraction method by fusing SAR and optical images according to claim 7, characterized in that: set up Indicates the number of correctly classified background object pixels, Indicates the number of correctly classified target object pixels, Indicates the number of pixels that are misclassified as target objects as background objects. Indicates the number of pixels that are misclassified as background objects as target objects. represents the overall classification accuracy, It represents the quotient of the sum of the products of the actual and predicted numbers of all objects and the square of the total number of samples, then: ; ; 。

Citation Information

Patent Citations

  • Method for crop mapping by using Gaofen-2 and Gaofen-3 based on field combination

    CN110189616A

  • SAR image target detection method based on adversarial domain adaptation

    CN112132042A

Cited By

  • A method for identifying cultivated land parcels based on optical and SAR image fusion

    CN122676357A