Wetland detailed mapping method based on remote sensing image and deep learning super-resolution algorithm
Through deep learning super-score algorithm and U-Net architecture, the resolution of low-resolution remote sensing images is improved and semantic segmentation is performed, which solves the problem of the existing wetland recognition methods having long acquisition periods at complex boundaries and high-resolution images, and achieves high-precision and space-time adaptability of wetland fine mapping.
Patent Information
- Application Number
- CN202411977076.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing wetland identification methods have low accuracy when wetlands and other lands are mixed with or complex boundaries, and cannot effectively utilize spatial context information. The high-resolution remote sensing image is costly and has a long acquisition period, making it difficult to support large-scale and long-time series wetland monitoring.
The fine wetland mapping method based on remote sensing images and deep learning super-segment algorithm is adopted. Through the convolution transform super-segment network of the convolution block attention module and the U-Net architecture based on Scharr convolution and fast Fourier convolution, the spatial resolution of low-resolution images is improved, and semantic segmentation is performed to generate high-precision wetland prediction confidence maps.
It significantly improves the accuracy of wetland mapping, reduces the impact of false positives and false negatives, and can obtain accurate wetland segmentation and mapping results with low spatial resolution. It has strong spatial and temporal adaptability and is suitable for a variety of remote sensing image data and environmental conditions.
Smart Images

Figure CN119399314B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular to a wetland fine-grained mapping method based on remote sensing images and a deep learning super-resolution algorithm. Background Art
[0002] The core task of wetland fine-grained mapping is to extract the spatial distribution and type information of wetlands from remote sensing images and analyze their dynamic changes. This process requires accurate determination of wetland boundaries and differentiation of different wetland types. In wetland monitoring practice, remote sensing technology has become a key means due to its wide coverage, strong timeliness, and large amount of information. In particular, the widespread use of high-resolution satellite images has greatly improved the accuracy of wetland distribution and type identification, and can capture the fine structure and characteristics of wetlands, accurately identify different types of wetlands and demarcate their boundaries, providing strong support for wetland management and protection.
[0003] However, when using remote sensing images for wetland mapping, the following problems are still encountered: 1) Although traditional wetland identification methods such as maximum likelihood method and support vector machine can extract wetland information, they have low accuracy when wetlands are mixed with other landforms or have complex boundaries, and they cannot effectively utilize spatial context information; 2) Although low-resolution images have a wide coverage area, their resolution is low and it is difficult to capture the detailed changes in wetlands. In addition, wetlands have similar spectral characteristics to other landforms, and identification often requires combination with high-resolution data or ground surveys, which limits the accuracy of wetland monitoring and thus affects the scientificity and accuracy of protection and management decisions; 3) High-resolution remote sensing images are expensive and have a long acquisition cycle, resulting in a small amount of data, which makes it difficult to support large-scale and long-term wetland monitoring, limiting the ability to monitor wetland changes in real time and continuously. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, comprising the following steps:
[0005] S1. Obtain high-resolution remote sensing images with a spatial resolution of 1m or less, low-resolution remote sensing images with a spatial resolution of more than 1m, and land cover datasets in the study area, perform radiation correction on the high-resolution remote sensing images, and perform atmospheric correction on the low-resolution remote sensing images to obtain high-resolution remote sensing image data and low-resolution remote sensing image data of the study area; extract wetland information from the land cover dataset, align it with the low-resolution remote sensing images, and generate a low-resolution wetland dataset LR_LR;
[0006] S2. Based on the high-resolution remote sensing image data, visual interpretation is performed in combination with the geographical environment characteristics of the study area, the wetland area is marked, and high-resolution wetland samples are generated;
[0007] S3, pairing high-resolution wetland samples with low-resolution remote sensing images to generate a high-resolution wetland dataset LR_HR, and dividing the high-resolution wetland dataset LR_HR into an LR_HR training set and an LR_HR validation set;
[0008] S4. Construct a wetland fine-grained mapping model, which includes a super-resolution module and a semantic segmentation module. The super-resolution module is set as a convolutional transformation super-resolution network based on a convolutional block attention module, and the semantic segmentation module is set as a U-Net architecture based on Scharr convolution and fast Fourier convolution. The low-resolution wetland dataset LR_LR and LR_HR training set are input into the convolutional transformation super-resolution network based on the convolutional block attention module, and shallow feature extraction, deep feature extraction and image feature reconstruction are performed through the convolution layer to achieve image feature dimension expansion and upgrade the low-resolution remote sensing image to a high-resolution remote sensing image.
[0009] S5. Use the U-Net architecture based on Scharr convolution and fast Fourier convolution to perform semantic segmentation on the super-resolution features obtained in the previous step to generate a high-resolution wetland prediction confidence map; and add an average pooling layer to the segmentation result of the low-resolution wetland dataset LR_LR to generate a low-resolution wetland prediction confidence map;
[0010] S6, calculating the high-resolution loss for the high-resolution wetland prediction confidence map, calculating the spatial generalization loss using the low-resolution wetland prediction confidence map, and calculating the temporal contrast loss by the similarity of temporal representations at the same location;
[0011] S7, the three losses in the previous step are weighted to obtain the total loss function, the parameters of the wetland fine-grained mapping model are updated, spatiotemporal perception learning is formed, and the wetland fine-grained mapping model is evaluated through the LR_HR validation set, and the wetland fine-grained mapping model corresponding to the best evaluation index is selected as the optimal model;
[0012] S8. Apply the optimal model trained in the previous step to the new low-resolution remote sensing image to perform detailed wetland mapping in an extended manner.
[0013] The technical solution further defined in the present invention is:
[0014] Further, in step S4, the convolutional transformation super-resolution network based on the convolutional block attention module includes a shallow feature extraction module, a deep feature extraction module and an image feature reconstruction module, and the convolutional block attention module is embedded in the deep feature extraction module, that is, the convolutional block attention module is embedded after each basic residual block structure, and the convolutional block attention module includes a channel attention module and a spatial attention module, both of which weight the feature map in the channel dimension and the spatial dimension respectively;
[0015] For the input low-resolution remote sensing image X∈R C×H×W , R C×H×W Represents a tensor with C channels and the size of each channel is a height H and a width W, where each element is a real value. The shallow feature extraction module uses a 3×3 convolution layer to map the low-resolution remote sensing image to the potential feature space to obtain shallow features; the shallow features pass through two basic residual block structures and the convolution block attention module, and then add a 3×3 convolution layer to obtain deep features; the deep features are combined with the shallow features by jump addition, and the features are reconstructed through an upsampling layer including a 3×3 convolution layer and a Pixel Shuffle layer. The Pixel Shuffle layer rearranges multiple channels of each pixel into a sub-block so that the feature map is rearranged as X∈R C×rH×rW , the H and W of the image are magnified r times to obtain the corresponding high-resolution remote sensing image output.
[0016] As mentioned above, the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, in the channel attention module, the input feature map F∈R C×H×W When , the feature map F is averaged pooled in the spatial dimension to obtain the feature vector F1. The formula is as follows:
[0017]
[0018] In the formula, the feature vector F1∈R C×1×1 ,c∈{1,...,C}; F(c,i,j) represents the value of the cth channel of the feature map F at position (i,j); the feature vector F1 is sparsely convolved and mapped to [0,1] by the Sigmoid function, and the formula is as follows:
[0019]
[0020] In the formula, represents the feature vector obtained by sparse convolution of feature vector F1; f(x) represents the Sigmoid activation function; x represents Each element in W c Indicates that the channel attention weight vector is obtained, and the obtained weight vector W c Multiply it with the input feature F to get the feature vector F2 weighted by channel attention. The formula is as follows:
[0021]
[0022] In the formula, the eigenvector F2∈R C×H×W , Represents the pixel-by-pixel multiplication operator.
[0023] As mentioned above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, in the spatial attention module, the feature map F2 obtained by the channel attention module is subjected to maximum pooling and average pooling respectively, so that the feature map is compressed in the channel dimension, and two feature maps F3 and F4 with a size of 1×H×W are obtained. The feature map F3 obtained by maximum pooling retains the maximum eigenvalue at each spatial position, and the feature map F4 obtained by average pooling retains the average eigenvalue at each spatial position. The formula is as follows:
[0024]
[0025] Where i∈{1,...,H}, j∈{1,...,W}, each pixel of feature map F3 is the average value of each channel at the corresponding spatial position, and each pixel of feature map F4 is the maximum value of each channel at the corresponding spatial position; feature map F3 and feature map F4 are concatenated in the channel dimension to obtain feature map F5, and the formula is as follows:
[0026]
[0027] In the formula, the feature map F5∈R 2×H×W , Represents an operator for concatenation and merging in a specific dimension; then a 7×7 convolution operation is performed on the feature map F5 to learn the relationship between spatial positions and obtain the feature map F6∈R 1×H×W , that is, F6=Conv 7×7 (F5); Finally, the sigmoid function is used to map it to [0,1] to obtain the spatial attention weight map W s , and the obtained weight W s Multiply it with the feature map F2 obtained by the channel attention module to obtain the feature map F7∈R after spatial attention weighting C×H×W .
[0028] As described above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, in step S5, each encoder and decoder in the U-Net architecture based on Scharr convolution and fast Fourier convolution is composed of two convolution modules; the first layer of the encoder is a CBR layer composed of a convolution layer, a batch normalization layer, and a ReLU activation function layer, and the second layer is composed of a parallel structure based on Scharr convolution and fast Fourier convolution. The Scharr convolution captures local edge details, and then the edge features are integrated and optimized through the CBR layer; the global context information is mined in the frequency domain through two fast Fourier convolution-CBR layers, and the results of the two parallel structures are added and optimized through the CBR layer, and then the global average pooling layer and the multi-layer perceptron are connected; after the global average pooling, an additional fully connected layer is added to map the extracted features into a vector z∈R dAs a time representation, it represents the global information related to the current image and time, where d represents the dimension; each encoder ends with a maximum pooling layer to achieve downsampling.
[0029] As mentioned above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, when the decoder decodes, it first gradually restores the spatial resolution of the image through upsampling, and uses jump connections to splice the features of the encoder with the features of the decoder; then, Scharr convolution is used to enhance edge information, and global context information is extracted through fast Fourier convolution; each layer is processed by batch normalization and ReLU activation function; the final upsampling layer is replaced by a deconvolution layer, and a high-resolution wetland prediction confidence map is generated through the Sigmoid activation function. ; High-resolution wetland prediction confidence map Input the average pooling layer for downsampling to obtain a low-resolution wetland prediction confidence map .
[0030] As described above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, in step S6, calculating the high-resolution loss specifically includes the following steps:
[0031] S6.1. Calculate the cross entropy loss. The formula is as follows:
[0032]
[0033] In the formula, Represents a high-resolution wetland prediction confidence map generated by a U-Net architecture based on Scharr convolution and fast Fourier convolution, P h Represents a realistic high-resolution wetland label map;
[0034] S6.2. Calculate the Focal Tversky loss. The formula is as follows:
[0035]
[0036] In the formula, α and β Represents the weight used to control the false positive and false negative; l represents the focus parameter, which is used to adjust the weight of difficult samples; e is a constant;
[0037] S6.3, high-resolution loss is expressed as the weighted sum of cross entropy loss and Focal Tversky loss:
[0038]
[0039] In the formula, cRepresents the weight factor used to control the relative weight of cross entropy loss and Focal Tversky loss, L ce represents the cross entropy loss value, L ft Represents the Focal Tversky loss value, L HR Represents the high-resolution loss value.
[0040] As described above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, in step S6, the time contrast loss is the z and The temperature parameter gradually decays with the training step number t, and the decay formula is as follows:
[0041]
[0042] In the formula, t 0 represents the initial temperature, k represents the attenuation factor; the information noise contrast estimation is used as the similarity measure, and its formula is as follows:
[0043]
[0044] Where, L TC Represents the time contrast loss value; z represents the feature vector of the image at a certain moment; is the eigenvector at the same position as z at different times; t represents the dynamic temperature hyperparameter, which is used to control the sensitivity of similarity and decays as training progresses; m j The representation vector representing the negative sample represents the feature representation of other images that are irrelevant to the image and the corresponding image; N represents the number of negative samples.
[0045] As mentioned above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, in step S6, the spatial generalization loss is explained by the Kullback-Leibler divergence, and the formula is:
[0046]
[0047] Where, L SG represents the spatial generalization loss, D KL (P‖Q) represents the Kullback-Leibler divergence, P(x) represents the distribution obtained from the LR_HR training set, and Q(x) represents the wetland fine-grained mapping model based on the prediction results. The calculated distribution, x represents the value of the random variable, m and s They represent the mean and standard deviation obtained from the LR_HR training set. and Respectively represent the wetland fine-grained mapping model according to The calculated mean and standard deviation are It represents the low-resolution feature map P constructed by the wetland fine-grained mapping model based on the low-resolution wetland dataset LR_LR. l Prediction results with low-resolution feature maps Calculated cross entropy loss.
[0048] As described above, in the wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithms, in step S8, the extended fine-grained wetland mapping expands the low-resolution remote sensing image into an integer number of overlapping sliding windows, and extracts image blocks by creating H×W sliding windows; during the movement process, each movement overlaps the previous movement area by t pixels, and the obtained H×W image blocks are input into the wetland fine-grained mapping model to obtain the Sigmoid confidence of the wetland, and the confidence map is spliced back into a complete image by calculating the maximum value of the overlapping area of each pixel, and a threshold is used to distinguish wetland pixels from background pixels.
[0049] The beneficial effects of the present invention are:
[0050] (1) In the present invention, the spatial resolution of low-resolution images is improved by using super-resolution technology based on the convolutional block attention module, and the wetland is segmented with high precision in combination with the semantic segmentation model, which makes up for the lack of spatial details of low-resolution images. This not only ensures the timeliness of wetland mapping, but also optimizes the wetland fine-grained mapping model through multiple loss functions, significantly improves the accuracy of wetland mapping, reduces the impact of false positives and false negatives, and thus achieves more accurate and reliable wetland mapping;
[0051] (2) The present invention has strong spatiotemporal adaptability in wetland monitoring and mapping, and can obtain accurate wetland segmentation and mapping results even when the spatial resolution is low, effectively solving the contradiction between time and spatial resolution in wetland mapping, so that it can be applied to a variety of remote sensing image data and environmental conditions, providing strong technical support for wetland protection, resource management and ecological restoration;
[0052] (3) In the present invention, Scharr convolution and fast Fourier convolution modules are introduced into the U-Net network architecture to optimize wetland image processing. The combination of Scharr convolution and fast Fourier convolution can not only accurately extract wetland boundaries and water body details and sharpen boundary contours, but also enhance the network's perception of global information through frequency domain processing, thereby improving the semantic segmentation accuracy of wetland images. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0054] Figure 2 Schematic diagram of low-resolution remote sensing images and corresponding high-resolution wetland labels in an embodiment of the present invention, (a), (b), (c), and (d) in the figure correspond to a label example, and each row from left to right is a low-resolution remote sensing image, a high-resolution remote sensing image, and high-resolution annotation data;
[0055] Figure 3 It is a structural schematic diagram of a wetland refined mapping model in an embodiment of the present invention;
[0056] Figure 4 Schematic diagram of the structure of a convolutional transform super-resolution network based on a convolutional block attention module in an embodiment of the present invention;
[0057] Figure 5 Schematic diagram of the structure of the convolutional block attention module in an embodiment of the present invention;
[0058] Figure 6 Schematic diagram of the encoder structure of the U-Net architecture based on Scharr convolution and fast Fourier convolution in an embodiment of the present invention;
[0059] Figure 7 Schematic diagram of the decoder structure of the U-Net architecture based on Scharr convolution and fast Fourier convolution in an embodiment of the present invention;
[0060] Figure 8 Schematic diagram of the result of wetland fine-grained mapping in an embodiment of the present invention, wherein (a) is a schematic diagram of the Sentinel-2 image input into the wetland fine-grained mapping model, and (b) is a schematic diagram of the high-resolution wetland mapping result output by the wetland fine-grained mapping model. DETAILED DESCRIPTION
[0061] This embodiment provides a wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithms, such as Figure 1 As shown, the following steps are included:
[0062] S1. Obtain high-resolution remote sensing images with a spatial resolution of 1m or less, low-resolution remote sensing images with a spatial resolution of more than 1m, and land cover datasets of the study area, perform radiation correction on the high-resolution remote sensing images, and perform atmospheric correction on the low-resolution remote sensing images to obtain high-resolution remote sensing image data and low-resolution remote sensing image data of the study area; extract wetland information from the land cover dataset, align it with the low-resolution remote sensing images, and generate a low-resolution wetland dataset LR_LR.
[0063] In this embodiment, the high-resolution remote sensing image is the Gaofen-2 satellite image, the low-resolution remote sensing image is the Sentinel-2 satellite image, and the land cover data is the GWL_FCS30D data. The radiation calibration is completed by applying the sensor calibration coefficient, and the Gaofen-2 satellite image is calibrated to obtain the high and low resolution wetland datasets of the study area. The atmospheric radiation transfer model is used to perform atmospheric correction on the Sentinel-2 satellite image to obtain the low resolution wetland dataset of the study area. The GWL_FCS30D data is the wetland dataset of the study area.
[0064] S2. Based on the high-resolution remote sensing image data, visual interpretation is performed in combination with the geographical environment characteristics of the study area, and the wetland area is marked to generate a high-resolution binary wetland label; in this embodiment, the wetland area is a natural wetland that does not include rivers, lakes and seawater.
[0065] S3. Pair high-resolution wetland samples with low-resolution remote sensing images to generate a high-resolution wetland dataset LR_HR, and divide the high-resolution wetland dataset LR_HR into an LR_HR training set and an LR_HR validation set in a ratio of 8:2.
[0066] In this embodiment, the size of the low-resolution remote sensing image is 64×64, and the size of the high-resolution remote sensing image is 640×640. The generated low-resolution remote sensing image and the corresponding high-resolution wetland label are as follows: Figure 2 As shown in the figure, (a), (b), (c) and (d) correspond to a label example respectively. Each row from left to right is a low-resolution remote sensing image, a high-resolution remote sensing image and a high-resolution annotation data. The white part in the high-resolution annotation data represents the wetland area, and the black part represents the non-wetland area.
[0067] S4, such as Figure 3 As shown in the figure, a wetland refined mapping model is constructed, which includes a super-resolution module and a semantic segmentation module. The super-resolution module is set to a convolutional transformation super-resolution network based on a convolutional block attention module, and the semantic segmentation module is set to a U-Net architecture based on Scharr convolution and fast Fourier convolution.
[0068] The low-resolution wetland dataset LR_LR and LR_HR training set are input into the convolutional transform super-resolution network based on the convolutional block attention module. Shallow feature extraction, deep feature extraction and image feature reconstruction are performed through the convolutional layer to achieve image feature dimension expansion and upgrade the low-resolution remote sensing images to high-resolution remote sensing images. The low-resolution wetland dataset LR_LR and LR_HR training set are collaboratively learned. Through the weakly supervised learning strategy, the low-resolution land cover data is used as auxiliary supervision information to provide guidance for low-resolution data that lack high-resolution references.
[0069] like Figure 4As shown in the figure, the convolutional transform super-resolution network based on the convolutional block attention module includes a shallow feature extraction module, a deep feature extraction module and an image feature reconstruction module. The convolutional block attention module is embedded in the deep feature extraction module, that is, the convolutional block attention module is embedded after each basic residual block structure (Basic Residual Blocks, BRB).
[0070] For the input low-resolution remote sensing image X∈R C×H×W , R C×H×W Represents a tensor with C channels and the size of each channel is a height H and a width W, where each element is a real value. The shallow feature extraction module uses a 3×3 convolution layer to map the low-resolution remote sensing image to the potential feature space to obtain shallow features; the shallow features pass through two basic residual block structures and the convolution block attention module, and then add a 3×3 convolution layer to obtain deep features; the deep features are combined with the shallow features by jump addition, and the features are reconstructed through an upsampling layer including a 3×3 convolution layer and a Pixel Shuffle layer. The Pixel Shuffle layer rearranges multiple channels of each pixel into a sub-block so that the feature map is rearranged as X∈R C×rH×rW , the H and W of the image are magnified r times to obtain the corresponding high-resolution remote sensing image output. In this embodiment, r=10.
[0071] like Figure 5 As shown in the figure, the convolutional block attention module includes a channel attention module and a spatial attention module, which weight the feature map in the channel dimension and the spatial dimension respectively; in the channel attention module, the input feature map F∈R 32 ×64×64 When , the feature map F is averaged pooled in the spatial dimension to obtain the feature vector F1. The formula is as follows:
[0072]
[0073] In the formula, the feature vector F1∈R 32×1×1 ,c∈{1,...,32}.
[0074] The feature vector F1 is sparsely convolved and mapped to [0,1] by the Sigmoid function. The formula is as follows:
[0075]
[0076] In the formula, represents the feature vector obtained by sparse convolution of feature vector F1; f(x) represents the Sigmoid activation function; x represents Each element in W c Indicates that the channel attention weight vector is obtained, and the obtained weight vector Wc Multiply it with the input feature F to get the feature vector F2 weighted by channel attention. The formula is as follows:
[0077]
[0078] In the formula, the eigenvector F2∈R 32×64×64 , Represents the pixel-by-pixel multiplication operator.
[0079] In the spatial attention module, the feature map F2 obtained by the channel attention module is subjected to maximum pooling and average pooling respectively, so that the feature map is compressed in the channel dimension, and two feature maps F3 and F4 with a size of 1×64×64 are obtained. The feature map F3 obtained by maximum pooling retains the maximum eigenvalue at each spatial position, and the feature map F4 obtained by average pooling retains the average eigenvalue at each spatial position. The formula is as follows:
[0080]
[0081] Where, i∈{1,...,64}, j∈{1,...,64}, each pixel of the feature map F3 is the average value of each channel at the corresponding spatial position, and each pixel of the feature map F4 is the maximum value of each channel at the corresponding spatial position.
[0082] The feature map F3 and the feature map F4 are concatenated in the channel dimension to obtain the feature map F5. The formula is as follows:
[0083]
[0084] In the formula, the feature map F5∈R 2×64×64 , Represents an operator for concatenation and merging in a specific dimension; then a 7×7 convolution operation is performed on the feature map F5 to learn the relationship between spatial positions and obtain the feature map F6∈R 1×64×64 , that is, F6=Conv 7×7 (F5); Finally, the sigmoid function is used to map it to [0,1] to obtain the spatial attention weight map W s , and the obtained weight W s Multiply it with the feature map F2 obtained by the channel attention module to obtain the feature map F7∈R after spatial attention weighting 32 ×64×64 .
[0085] S5. Use the U-Net architecture based on Scharr convolution and fast Fourier convolution to perform semantic segmentation on the super-resolution features obtained in the previous step to generate a high-resolution wetland prediction confidence map; and add an average pooling layer to the segmentation result of the low-resolution wetland dataset LR_LR to generate a low-resolution wetland prediction confidence map.
[0086] Each encoder and decoder in the U-Net architecture based on Scharr convolution and fast Fourier convolution consists of two convolution modules. In the encoder part, Scharr convolution and fast Fourier convolution are introduced to enhance the ability to extract local edge details and the ability to perceive global context information respectively; the output feature map size of the encoder is gradually reduced to 32×320×320, 64×160×160 and 128×80×80, and is gradually restored to the original resolution through a symmetrical decoder, while generating a high-resolution wetland prediction confidence map. , and add an average pooling layer to the segmentation results of the low-resolution wetland dataset LR_LR to generate a low-resolution wetland prediction confidence map .
[0087] like Figure 6 As shown in Figure 1, the first layer of the encoder is a convolution-batch normalization-ReLU activation composite layer (Convolution-Batch Normalization-ReLU, CBR) composed of a convolution layer, a batch normalization layer, and a ReLU activation function layer. The second layer is composed of a parallel structure based on Scharr convolution and fast Fourier convolution. Scharr convolution captures local edge details, and then the edge features are integrated and optimized through the CBR layer; the global context information is mined in the frequency domain through two fast Fourier convolution-CBR layers, and the results of the two parallel structures are added and optimized through the CBR layer, followed by a global average pooling layer (Global Average Pooling, GAP) and a multilayer perceptron (Multilayer Perceptron, MLP); after global average pooling, an additional fully connected layer is added to map the extracted features into a vector z∈R d As a time representation, it represents the global information related to the current image and time, where d represents the dimension; each encoder ends with a maximum pooling layer to achieve downsampling.
[0088] like Figure 7As shown in the figure, when the decoder decodes, it first gradually restores the spatial resolution of the image through upsampling, and uses jump connections to splice the low-level features of the encoder with the high-level features of the decoder to ensure detail recovery; then, Scharr convolution is used to enhance edge information, and global context information is extracted through fast Fourier convolution to improve the recognition ability of complex image structures; each layer is processed by batch normalization and ReLU activation function to ensure training stability and nonlinear feature learning; the final upsampling layer is replaced by a deconvolution layer, and a high-resolution wetland prediction confidence map is generated through the Sigmoid activation function ; High-resolution wetland prediction confidence maps where no high-resolution reference exists Input the average pooling layer for downsampling to obtain a low-resolution wetland prediction confidence map .
[0089] S6. Calculate the high-resolution loss for the high-resolution wetland prediction confidence map, calculate the spatial generalization loss using the low-resolution wetland prediction confidence map, and calculate the temporal comparison loss by the similarity of the temporal representation of the same location.
[0090] Calculating high-resolution loss specifically includes the following steps:
[0091] S6.1. Calculate the cross entropy loss. The formula is as follows:
[0092]
[0093] In the formula, Represents a high-resolution wetland prediction confidence map generated by a U-Net architecture based on Scharr convolution and fast Fourier convolution, Represents a realistic high-resolution wetland label map.
[0094] S6.2. Calculate the Focal Tversky loss. The formula is as follows:
[0095]
[0096] In the formula, α and β Represents the weight used to control the false positive and false negative. In this embodiment α =0.7, β =0.3; l Represents the focus parameter, which is used to adjust the weight of difficult samples. In this embodiment l =2; e is a constant. In this embodiment e =10 -6 .
[0097] S6.3, high-resolution loss is expressed as the weighted sum of cross entropy loss and Focal Tversky loss:
[0098]
[0099] In the formula, c Represents the weight factor used to control the relative weight of the cross entropy loss and the Focal Tversky loss. In this embodiment c =0.5; L ce represents the cross entropy loss value; L ft Represents the Focal Tversky loss value; L HR Represents the high-resolution loss value.
[0100] The time contrast loss is the z and In order to improve the performance of the model at different training stages, the temperature parameter gradually decays with the number of training steps t. The decay formula is as follows:
[0101]
[0102] In the formula, t 0 means initial temperature t 0=1, in this embodiment; k represents the attenuation factor, in this embodiment k=0.001; the temperature gradually decreases, prompting the model to pay more attention to fine-grained feature differences in the later stage of training, thereby improving the model's learning ability and detail capture ability.
[0103] On this basis, information noise contrast estimation is used as the similarity measure, and its formula is as follows:
[0104]
[0105] Where, L TC Represents the time contrast loss value; z represents the feature vector of the image at a certain moment; is the eigenvector at the same position as z at different times; t represents the dynamic temperature hyperparameter, which is used to control the sensitivity of similarity and decays as training progresses; m j The representation vector representing the negative sample represents the feature representation of other images that are irrelevant to the image and the corresponding image; N represents the number of negative samples.
[0106] The spatial generalization loss is explained by the Kullback-Leibler divergence, which is formulated as:
[0107]
[0108] Where, L SG represents the spatial generalization loss, DKL (P‖Q) represents the Kullback-Leibler divergence, P(x) represents the distribution obtained from the LR_HR training set, and Q(x) represents the wetland fine-grained mapping model based on the prediction results. The calculated distribution, x represents the value of the random variable, m and s They represent the mean and standard deviation obtained from the LR_HR training set. and Respectively represent the wetland fine-grained mapping model according to The calculated mean and standard deviation are It represents the low-resolution feature map P constructed by the wetland fine-grained mapping model based on the low-resolution wetland dataset LR_LR. l Prediction results with low-resolution feature maps Calculated cross entropy loss.
[0109] S7. The three losses in step S6 are weightedly combined to obtain the total loss function, the parameters of the wetland refined mapping model are updated, the spatiotemporal perception learning of the wetland refined mapping model is realized, the spatiotemporal perception learning is formed, and the wetland refined mapping model is evaluated through the LR_HR validation set, and the wetland refined mapping model corresponding to the best evaluation index is selected as the optimal model.
[0110] The three loss functions are weighted and the model parameters are updated simultaneously. The formula is:
[0111]
[0112] In the formula, f , oh and t They are respectively the high-resolution loss L HR , time contrast loss L TC and the spatial generalization loss L SG The weight coefficient, in this embodiment, f : oh : t =200:1:5.
[0113] S8. Apply the optimal model trained in the previous step to a new low-resolution remote sensing image. The low-resolution remote sensing image should be expanded to contain an integer number of sliding windows of size 64×64 that overlap each other. Extract 64×64 image blocks in sequence, and ensure that each extracted image block overlaps with the previous extracted area by 6 pixels. Input the obtained 64×64 image blocks into the wetland fine mapping model to obtain the Sigmoid confidence of the wetland. By calculating the maximum value of the overlapping area of each pixel, the confidence map is spliced into a complete image, and a fixed threshold of 0.5 is used to distinguish wetland pixels from background pixels. The wetland fine mapping result of this embodiment is as follows: Figure 8 As shown, Figure (a) is the Sentinel-2 image input to the wetland refined mapping model, and Figure (b) is the high-resolution wetland mapping result output by the wetland refined mapping model. The white part in the figure represents the wetland area, and the black part is the non-wetland area.
[0114] In recent years, deep learning methods have brought significant advantages to the refined mapping of wetlands. They can automatically learn complex feature patterns from remote sensing images. Through large amounts of data training, deep learning models can highly accurately distinguish wetlands from other landforms, overcome the similarities between wetlands and their surroundings in terms of spectrum, texture, etc., and thus solve the problem of identification.
[0115] Convolutional neural networks can automatically mine the complex mapping relationship between low-resolution and high-resolution images from massive remote sensing image data, and accurately reconstruct high-resolution images; in addition, deep learning can adapt to remote sensing image data with a wide range of sources, resolutions, spectral ranges and imaging times, and achieve high-quality super-resolution processing through continuous adjustment and training; with the continuous optimization of network structure and parameters, deep learning has great potential for performance improvement. For example, residual networks can restore finer details and provide better image data for wetland monitoring and analysis; in view of the above, this embodiment proposes a wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithms to provide more accurate and efficient technical support for wetland protection and management.
[0116] The method of this embodiment has the following main innovations:
[0117] 1) The method of this embodiment innovatively combines super-resolution technology and semantic segmentation technology in the wetland identification task, improves the accuracy of low-resolution remote sensing images in wetland feature extraction, and provides technical support for the refined mapping of wetlands.
[0118] 2) In the convolutional transform super-resolution network, a convolutional block attention module is introduced to enable the model to focus on the key areas of wetland images, enhance the recovery of key channels and spatial features, and improve the reconstruction accuracy of wetland details and boundaries.
[0119] 3) Introduce Scharr convolution and fast Fourier convolution modules into the U-Net network architecture to optimize wetland image processing. The combination of Scharr convolution and fast Fourier convolution can not only accurately extract wetland boundaries and water details and sharpen boundary contours, but also enhance the network's perception of global information through frequency domain processing, thereby improving the semantic segmentation accuracy of wetland images.
[0120] 4) A comprehensive loss function is designed for the wetland fine-grained mapping task, integrating high-resolution loss, temporal contrast loss and spatial generalization loss to optimize image details, temporal consistency and low-resolution image consistency, thereby significantly improving detail recovery and boundary accuracy, and enhancing the overall image reconstruction and segmentation effects.
[0121] In addition to the above embodiments, the present invention may also have other implementation modes. Any technical solution formed by equivalent replacement or equivalent transformation falls within the protection scope required by the present invention.
Claims
1. A wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm, characterized by: The following steps are involved: S1. Obtain high-resolution remote sensing images with a spatial resolution of 1m or less, low-resolution remote sensing images with a spatial resolution of more than 1m, and land cover datasets in the study area, perform radiation correction on the high-resolution remote sensing images, and perform atmospheric correction on the low-resolution remote sensing images to obtain high-resolution remote sensing image data and low-resolution remote sensing image data of the study area; extract wetland information from the land cover dataset, align it with the low-resolution remote sensing images, and generate a low-resolution wetland dataset LR_LR; S2. Based on the high-resolution remote sensing image data, visual interpretation is performed in combination with the geographical environment characteristics of the study area, the wetland area is marked, and high-resolution wetland samples are generated; S3, pairing high-resolution wetland samples with low-resolution remote sensing images to generate a high-resolution wetland dataset LR_HR, and dividing the high-resolution wetland dataset LR_HR into an LR_HR training set and an LR_HR validation set; S4. Construct a wetland fine-grained mapping model, which includes a super-resolution module and a semantic segmentation module. The super-resolution module is set as a convolutional transformation super-resolution network based on a convolutional block attention module, and the semantic segmentation module is set as a U-Net architecture based on Scharr convolution and fast Fourier convolution. The low-resolution wetland dataset LR_LR and LR_HR training set are input into the convolutional transformation super-resolution network based on the convolutional block attention module, and shallow feature extraction, deep feature extraction and image feature reconstruction are performed through the convolution layer to achieve image feature dimension expansion and upgrade the low-resolution remote sensing image to a high-resolution remote sensing image. S5. Use the U-Net architecture based on Scharr convolution and fast Fourier convolution to perform semantic segmentation on the super-resolution features obtained in the previous step to generate a high-resolution wetland prediction confidence map; and add an average pooling layer to the segmentation result of the low-resolution wetland dataset LR_LR to generate a low-resolution wetland prediction confidence map; S6, calculating the high-resolution loss for the high-resolution wetland prediction confidence map, calculating the spatial generalization loss using the low-resolution wetland prediction confidence map, and calculating the temporal contrast loss by the similarity of temporal representations at the same location; S7, the three losses in the previous step are weighted to obtain the total loss function, the parameters of the wetland fine-grained mapping model are updated, spatiotemporal perception learning is formed, and the wetland fine-grained mapping model is evaluated through the LR_HR validation set, and the wetland fine-grained mapping model corresponding to the best evaluation index is selected as the optimal model; S8, applying the optimal model trained in the previous step to the new low-resolution remote sensing image to perform wetland fine-grained mapping in an extended manner; In step S6, calculating the high-resolution loss specifically includes the following steps: S6.
1. Calculate the cross entropy loss. The formula is as follows: In the formula, Represents a high-resolution wetland prediction confidence map generated by a U-Net architecture based on Scharr convolution and fast Fourier convolution, P h Represents a realistic high-resolution wetland label map; S6.
2. Calculate the Focal Tversky loss. The formula is as follows: In the formula, α and β Represents the weight used to control the false positive and false negative; λ represents the focus parameter, which is used to adjust the weight of difficult samples; ε is a constant; S6.3, high-resolution loss is expressed as the weighted sum of cross entropy loss and Focal Tversky loss: In the formula, γ Represents the weight factor used to control the relative weight of cross entropy loss and Focal Tversky loss, L ce represents the cross entropy loss value, L ft Represents the Focal Tversky loss value, L HR Represents the high-resolution loss value; In step S6, the spatial generalization loss is explained by the Kullback-Leibler divergence, which is: Where, L SG represents the spatial generalization loss, D KL (P‖Q) represents the Kullback-Leibler divergence, P(x) represents the distribution obtained from the LR_HR training set, and Q(x) represents the wetland fine-grained mapping model based on the prediction results. The calculated distribution, x represents the value of the random variable, μ and σ They represent the mean and standard deviation obtained from the LR_HR training set. and Respectively represent the wetland fine-grained mapping model according to The calculated mean and standard deviation are It represents the low-resolution feature map P constructed by the wetland fine-grained mapping model based on the low-resolution wetland dataset LR_LR. l Prediction results with low-resolution feature maps Calculated cross entropy loss.
2. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 1 is characterized by: In the step S4, the convolutional transformation super-resolution network based on the convolutional block attention module includes a shallow feature extraction module, a deep feature extraction module and an image feature reconstruction module, and the convolutional block attention module is embedded in the deep feature extraction module, that is, the convolutional block attention module is embedded after each basic residual block structure, and the convolutional block attention module includes a channel attention module and a spatial attention module, both of which weight the feature map in the channel dimension and the spatial dimension respectively; For the input low-resolution remote sensing image X∈R C×H×W , R C×H×W Represents a tensor with C channels and the size of each channel is a height H and a width W, where each element is a real value. The shallow feature extraction module uses a 3×3 convolution layer to map the low-resolution remote sensing image to the potential feature space to obtain shallow features; the shallow features pass through two basic residual block structures and the convolution block attention module, and then add a 3×3 convolution layer to obtain deep features; the deep features are combined with the shallow features by jump addition, and the features are reconstructed through an upsampling layer including a 3×3 convolution layer and a Pixel Shuffle layer. The Pixel Shuffle layer rearranges multiple channels of each pixel into a sub-block so that the feature map is rearranged as X∈R C×rH×rW , the H and W of the image are magnified r times to obtain the corresponding high-resolution remote sensing image output.
3. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 2 is characterized by: In the channel attention module, the input feature map F∈R C×H×W When , the feature map F is averaged pooled in the spatial dimension to obtain the feature vector F1. The formula is as follows: In the formula, the feature vector F1∈R C×1×1 ,c∈{1,...,C}; F(c,i,j) represents the value of the cth channel of the feature map F at position (i,j); the feature vector F1 is sparsely convolved and mapped to [0,1] by the Sigmoid function, and the formula is as follows: In the formula, represents the feature vector obtained by sparse convolution of feature vector F1; f(x) represents the Sigmoid activation function; x represents Each element in W c Indicates that the channel attention weight vector is obtained, and the obtained weight vector W c Multiply it with the input feature F to get the feature vector F2 weighted by channel attention. The formula is as follows: In the formula, the feature vector F2∈R C×H×W , Represents the pixel-by-pixel multiplication operator.
4. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 3 is characterized by: In the spatial attention module, the feature map F2 obtained by the channel attention module is subjected to maximum pooling and average pooling respectively, so that the feature map is compressed in the channel dimension to obtain two feature maps F3 and F4 of size 1×H×W, wherein the feature map F3 obtained by maximum pooling retains the maximum eigenvalue at each spatial position, and the feature map F4 obtained by average pooling retains the average eigenvalue at each spatial position, and the formula is as follows: Where i∈{1,...,H}, j∈{1,...,W}, each pixel of feature map F3 is the average value of each channel at the corresponding spatial position, and each pixel of feature map F4 is the maximum value of each channel at the corresponding spatial position; feature map F3 and feature map F4 are concatenated in the channel dimension to obtain feature map F5, and the formula is as follows: In the formula, the feature map F5∈R 2×H×W , Represents an operator for concatenation and merging in a specific dimension; then a 7×7 convolution operation is performed on the feature map F5 to learn the relationship between spatial positions and obtain the feature map F6∈R 1×H×W , that is, F6=Conv 7×7 (F5); Finally, the sigmoid function is used to map it to [0,1] to obtain the spatial attention weight map W s , and the obtained weight W s Multiply it with the feature map F2 obtained by the channel attention module to obtain the feature map F7∈R after spatial attention weighting C×H×W .
5. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 1 is characterized by: In the step S5, each encoder and decoder in the U-Net architecture based on Scharr convolution and fast Fourier convolution is composed of two convolution modules; the first layer of the encoder is a CBR layer composed of a convolution layer, a batch normalization layer and a ReLU activation function layer, and the second layer is composed of a parallel structure based on Scharr convolution and fast Fourier convolution, the Scharr convolution captures local edge details, and then the edge features are integrated and optimized through the CBR layer; the global context information is mined in the frequency domain through two fast Fourier convolution-CBR layers, and the results of the two parallel structures are added and optimized through the CBR layer, and then the global average pooling layer and the multi-layer perceptron are connected; after the global average pooling, an additional fully connected layer is added to map the extracted features into a vector z∈R d As a time representation, it represents the global information related to the current image and time, where d represents the dimension; each encoder ends with a maximum pooling layer to achieve downsampling.
6. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 5 is characterized by: When the decoder performs decoding, the spatial resolution of the image is gradually restored by upsampling, and the features of the encoder and the decoder are spliced by using jump connections; then, the edge information is enhanced by using Scharr convolution, and the global context information is extracted by fast Fourier convolution; Each layer is processed by batch normalization and ReLU activation function; the final upsampling layer is replaced by a deconvolution layer, and a high-resolution wetland prediction confidence map is generated by the Sigmoid activation function. ; High-resolution wetland prediction confidence map Input the average pooling layer for downsampling to obtain a low-resolution wetland prediction confidence map .
7. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 1 is characterized by: In step S6, the time contrast loss is the z and The temperature parameter gradually decays with the training step number t, and the decay formula is as follows: In the formula, τ 0 represents the initial temperature, k represents the attenuation factor; the information noise contrast estimation is used as the similarity measure, and its formula is as follows: Where, L TC Represents the time contrast loss value; z represents the feature vector of the image at a certain moment; is the eigenvector at the same position as z at different times; τ represents the dynamic temperature hyperparameter, which is used to control the sensitivity of similarity and decays as training progresses; m j The representation vector representing the negative sample represents the feature representation of other images that are irrelevant to the image and the corresponding image; N represents the number of negative samples.
8. The wetland fine-grained mapping method based on remote sensing images and deep learning super-resolution algorithm according to claim 1 is characterized by: In step S8, the extended refined wetland mapping expands the low-resolution remote sensing image into an integer number of overlapping sliding windows, and extracts image blocks by creating H×W sliding windows; during the movement process, each movement overlaps the previous movement area by t pixels, and the obtained H×W image blocks are input into the wetland refined mapping model to obtain the Sigmoid confidence of the wetland, and the confidence map is spliced back into a complete image by calculating the maximum value of the overlapping area of each pixel, and a threshold is used to distinguish wetland pixels from background pixels.
Citation Information
Patent Citations
Multi-scale feature integrated high-resolution image tea garden automatic identification method
CN115909065A
Method, apparatus, and system for task driven approaches to super resolution
US20200364830A1