SAR image water body detection method based on double encoders and attention mechanism
By constructing a SAR image water body detection method based on dual encoders and attention mechanism, the problem of insufficient water body detection accuracy in complex backgrounds is solved, the dynamic interaction and fusion of features are realized, and the accuracy and stability of water body detection are improved, which is suitable for ecological environment monitoring and disaster prevention.
Patent Information
- Application Number
- CN202510823069.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing SAR image water body detection methods have problems such as insufficient feature extraction, low feature fusion efficiency, weak detail restoration ability and single loss function under complex backgrounds, resulting in insufficient water body detection accuracy.
A method based on dual encoders and attention mechanism is adopted to construct a dual encoder including ResNet branch and Swin Transformer branch. The adaptive feature fusion module is used to dynamically interact and fuse local texture information with global context information. The multi-scale convolution pooling module and the content-aware upsampling operator CARAFE are used for decoding. The training is carried out with the composite loss function of Dice loss, Focal loss and active contour loss.
It significantly enhances the completeness of feature representation, improves the accuracy and stability of water body detection, and can accurately identify water areas under complex backgrounds, providing reliable technical support for ecological environment monitoring and disaster prevention and control.
Smart Images

Figure CN120747734A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and in particular relates to a SAR image water body detection method based on a dual encoder and an attention mechanism. Background Art
[0002] With the rapid development of remote sensing technology, synthetic aperture radar (SAR) has become an important tool for water body detection under complex meteorological conditions due to its all-weather and all-day observation capabilities. The significant differences in the electromagnetic wave reflection characteristics of water and land in SAR images make water body detection based on SAR images have broad application prospects in ecological and environmental monitoring, water resources management, and disaster response. However, the interference of complex backgrounds (such as terrain undulation and vegetation cover) often leads to false detections and missed detections in traditional water body detection methods, seriously affecting the accuracy and reliability of detection. In recent years, deep learning technology has made significant progress in water body detection in SAR images, especially the application of convolutional neural networks (CNN) and Transformer architectures. However, existing methods still suffer from problems such as insufficient feature extraction, low feature fusion efficiency, weak detail restoration capabilities, and a single loss function. A single encoder structure is difficult to simultaneously capture the local texture details and global contextual information of water bodies, resulting in insufficient completeness of feature expression; traditional methods lack a dynamic interaction mechanism when fusing local and global features, making it difficult to effectively improve the representation ability of water body structure; upsampling operations in the decoding process are prone to introduce feature distortion, especially the insufficient accuracy of restoring water body edge details; therefore, the existing technology has the problem of insufficient water body detection accuracy under the interference of complex backgrounds. Summary of the Invention
[0003] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a SAR image water detection method based on dual encoders and attention mechanism, which solves the problem of insufficient water body detection accuracy under the interference of complex background in the existing technology.
[0004] The purpose of the present invention can be achieved through the following technical solutions:
[0005] A water body detection method for SAR images based on dual encoders and attention mechanism includes the following steps:
[0006] Obtain SAR image dataset;
[0007] A dual encoder consisting of a parallel ResNet branch and a Swin Transformer branch is constructed to receive SAR images and extract local texture information and global context information of the SAR images through the two branches respectively;
[0008] Build an adaptive feature fusion module based on the attention mechanism to dynamically interact and fuse local texture information with global context information to generate a fused feature map;
[0009] Construct a decoding module including multiple decoders to receive and decode the fused feature map to generate a water body detection result with the same size as the SAR image;
[0010] The dual encoder, adaptive feature fusion module and decoder module are jointly constructed as the detection model;
[0011] The SAR image dataset is divided into training set, validation set and test set;
[0012] Construct a composite loss function including Dice loss, Focal loss and active contour loss. Based on the composite loss function, the detection model is trained using the training set and validation set.
[0013] The SAR images in the test set are input into the trained detection model and the water body detection results are output.
[0014] The ResNet branch consists of a 7×7 convolutional layer, the first batch normalization layer, the first ReLU activation function, a maximum pooling layer, and multiple residual layers connected in sequence;
[0015] The 7×7 convolutional layer is used to perform the initial convolution operation on the received SAR image, extract the initial feature map and generate the initial feature map;
[0016] The first batch normalization layer is used to receive the initial feature map and perform normalization;
[0017] The first ReLU activation function is used to receive the normalized initial feature map for nonlinear activation, and the output end of the first ReLU activation function is independently connected to the maximum pooling layer and the decoder module respectively;
[0018] The maximum pooling layer is used to downsample the initial feature map after activation and input it into the residual layer;
[0019] Each residual layer is connected in sequence, and any residual layer includes multiple sequentially connected residual units;
[0020] Let the number of residual layers be N, where N is a positive integer greater than or equal to 2;
[0021] The input of the first residual layer is connected to the output of the maximum pooling layer;
[0022] The output ends of the first N-1 residual layers are independently connected to the next residual layer and the adaptive feature fusion module;
[0023] The output of the Nth residual layer is directly connected to the adaptive feature fusion module;
[0024] The output of the nth residual layer is used to output local texture information Y n , where n∈[1,N].
[0025] The residual units all include the first 3×3 convolutional layer, the second 3×3 convolutional layer and the sixth 1×1 convolutional layer;
[0026] The input ends of the first 3×3 convolutional layer and the sixth 1×1 convolutional layer are both used to receive the feature maps in the input residual unit;
[0027] The first 3×3 convolutional layer is connected to the third batch normalization layer and the fourth ReLU activation function in sequence to process the input feature map and extract local feature information;
[0028] The second 3×3 convolutional layer is used to receive the local feature information output by the fourth ReLU activation function. The output end of the second 3×3 convolutional layer is connected to the fourth batch normalization layer for further deepening the extraction of local feature information. The second 3×3 convolutional layer and the fourth batch normalization layer serve as the main path output;
[0029] The sixth 1×1 convolutional layer is used to adjust the number of channels of the received feature map so that the adjusted number of channels matches the number of output channels of the main path, and to construct a residual path. The residual path and the main path are connected to form a residual connection and then added for output.
[0030] The Swin Transformer branch includes:
[0031] The segmentation module is used to receive the SAR image and segment it into multiple non-overlapping 4×4 patches;
[0032] The linear embedding module is used to receive the patch block and perform linear mapping on each patch block to obtain the patch feature map;
[0033] Multiple Swin Transformer layers are arranged in sequence, where the number of Swin Transformer layers is equal to the number of residual layers, and the input of the first Swin Transformer layer is connected to the output of the linear embedding module;
[0034] Multiple Patch Merging modules: Any two adjacent Swin Transformer layers are connected via a Patch Merging module. The Patch Merging module is used to downsample, reduce the spatial dimension, and increase the channel dimension.
[0035] The outputs of the first N-1 Swin Transformer layers are independently connected to the corresponding Patch Merging module and the adaptive feature fusion module;
[0036] The output of the Nth Swin Transformer layer is directly and independently connected to the adaptive feature fusion module;
[0037] The output of the nth Swin Transformer layer is used to output the global context information X n , where n∈[1,N].
[0038] Construct an adaptive feature fusion module based on the attention mechanism to dynamically interact and fuse local texture information with global context information to generate a fused feature map. The specific steps include:
[0039] Constructing an adaptive feature fusion module including multiple adaptive feature fusion modules, wherein the multiple adaptive feature fusion modules are arranged in sequence, and the number of the adaptive feature fusion modules is equal to the number of the residual layers and corresponds one to one;
[0040] One output terminal of the nth residual layer and the nth Swin Transformer layer is connected to the input terminal of the corresponding nth adaptive feature fusion module;
[0041] In each adaptive feature fusion module, the input Y n and X n Perform channel compression and linear projection respectively, and convert Y n and X n Unify the characteristic dimensions of
[0042] Constructing bidirectional cross-attention paths:
[0043] In one path, the output of the ResNet branch is used as the query, and the output of the Swin Transformer branch is used as the key and value, forming a cross-attention from global to local, enhancing the semantic information of local features;
[0044] In the other path, the output of the Swin Transformer branch is used as the query, and the output of the ResNet branch is used as the key and value, forming a cross-attention from local to global, enhancing the spatial details of the global features;
[0045] Based on the attention mechanism, the attention weight is calculated for any path, and the calculated attention weight is matrix multiplied with the Value in the corresponding path to obtain the first intermediate feature map;
[0046] The residual connection is used to retain the two sets of original output features of local texture information and global context information output by the two branches of the dual encoder;
[0047] A lightweight channel attention mechanism is introduced into the four-way feature consisting of two original output features and two first intermediate feature maps, and the weighted coefficient corresponding to each feature is calculated respectively;
[0048] Multiply the weighted coefficient by the corresponding road feature to finally obtain four weighted features;
[0049] Concatenate the four weighted features according to the channel dimension to obtain the weighted concatenated features;
[0050] The weighted concatenated features are simultaneously input into the second 1×1 convolutional layer and the third 1×1 convolutional layer;
[0051] The weighted concatenated features are processed by the second 1×1 convolutional layer to obtain the second intermediate feature map;
[0052] After the weighted concatenated features pass through the third 1×1 convolution layer, they are sequentially input into the first 3×3 depthwise convolution, the second ReLU activation function, and then into the fourth 1×1 convolution layer for channel compression to obtain the third intermediate feature map;
[0053] Add the second intermediate feature map and the third intermediate feature map, perform normalization, and output the fused feature map Z n , n∈[1,N].
[0054] Based on the attention mechanism, the attention weight is calculated for each path, and the calculated attention weight is matrix multiplied with the Value in the corresponding path to obtain the first intermediate feature map. Specifically, the following steps are included:
[0055] In any path, the query vector of the path is dot-producted with the key vector of the other path element by element to obtain a dot product matrix. The expression in the dot product matrix is as follows:
[0056] S=QK T
[0057] Where QK T Indicates the similarity between the Query vector and the Key vector;
[0058] Scale each element in the dot product matrix to get S ij , the specific scaling calculation formula is as follows:
[0059]
[0060] Where, Represents the square root of the feature dimension of the query or key in the path;
[0061] Q i Indicates the Query value of position i, K j Indicates the key value of position j;
[0062] For the scaled element S ij Perform Softmax normalization to obtain the attention weight A ij , the specific calculation formula is as follows:
[0063]
[0064] Where N k Indicates the sequence length of the Key;
[0065] Perform matrix multiplication of the attention weight A and the Value in the corresponding path to obtain:
[0066]
[0067] Where O represents the first intermediate feature map; V represents the Value in the corresponding path.
[0068] The decoding module also includes a fifth 1×1 convolutional layer and a Sigmoid activation function;
[0069] The number of decoders is N+1;
[0070] The output of the first decoder is connected to the input of the fifth 1×1 convolutional layer, and the output of the fifth 1×1 convolutional layer is connected to the input of the Sigmoid activation function;
[0071] The fused feature map Z output by the Nth adaptive feature fusion module N Directly input into the N+1th decoder, decode to generate feature map M N+1 ;
[0072] The feature M N+1 The fused feature map Z output by the N-1th adaptive feature fusion module N-1 After splicing, they are input into the Nth decoder together and decoded to generate feature map M N ;
[0073] Repeat this process until the feature M3 generated by the third decoder is concatenated with the fused feature map Z1 output by the first adaptive feature fusion module. The concatenated features are then input into the second decoder and decoded to generate the feature map M2.
[0074] The output result of the first ReLU activation function in the ResNet branch is concatenated with the feature M2, and then input into the first decoder to generate the feature map M1.
[0075] After the feature M1 is processed by the fifth 1×1 convolution layer and the Sigmoid activation function, a water body detection result with the same size as the SAR image is generated.
[0076] The decoders all include a multi-scale convolutional pooling module, a content-aware upsampling operator CARAFE, two second 3×3 depth convolutions, a second batch normalization layer, and a third ReLU activation function;
[0077] The feature map input to the decoder is first input into the multi-scale convolution pooling module;
[0078] The multi-scale convolution pooling module performs convolution operations on the input feature map with dilation rates of 1, 2, 4, and 6 respectively, and integrates the convolution processing results of different dilation rates into the content-aware upsampling operator CARAFE;
[0079] The content-aware upsampling operator CARAFE generates a reorganization weight by analyzing the local content, and upsamples the input feature map based on the reorganization weight. The upsampling result is sequentially input into two second 3×3 depth convolutions;
[0080] After two second 3×3 depth convolution processes, the second batch normalization layer and the third ReLU activation function are input for processing, and the decoding result is output.
[0081] Construct a composite loss function including Dice loss, Focal loss and active contour loss. Based on the composite loss function, use the training set and validation set to train the detection model. The specific steps include:
[0082] The expressions for defining the Dice loss function and the Focal loss function are as follows:
[0083]
[0084] L F =α(1-p t ) γ ·L BCE (p,g)
[0085] Where, L D represents the Dice loss function; p represents the prediction result; g is the true label, and ε is the smoothing term;
[0086] L F represents the Focal loss function; p t =exp(-L BCE ), L BCE is the binary cross entropy loss; α and γ are both focusing factors;
[0087] The active contour loss function combines the regional energy term and the boundary energy term to guide the model to learn the target area and accurately locate the edge position. The active contour loss function L is defined as AC The expression is as follows:
[0088] L AC =Length+Region
[0089]
[0090] Where, Region represents the regional energy term; Length represents the boundary energy term;
[0091] P i,j Y is the probability map predicted by the detection model; i,j are the average intensities of the foreground and background regions of the true labels c1 and c2 respectively; x P i,j Represents the gradient of the probability graph on the x-axis; ▽ y P i,j represents the gradient of the probability map on the y-axis; H and W represent the height and width of the image respectively; μ, ν and λ are the weight parameters of the corresponding items;
[0092] Constructing the composite loss function L total , the specific expression is as follows:
[0093] L total =w D ·L D +w F ·L F +w AC ·L AC
[0094] Where w D 、w F 、w AC are the weight parameters of Dice loss, Focal loss and active contour loss respectively;
[0095] Set a convergence threshold or early stopping mechanism for validation loss;
[0096] With the composite loss function L total The training objective is to minimize the value of , and the training set is used to train the detection model;
[0097] After each training cycle, the validation set is used to calculate the validation loss. When the validation loss reaches the convergence threshold or the early stopping mechanism is triggered, the detection model training is completed.
[0098] The number of residual layers and Swin Transformer layers is four each;
[0099] The number of residual units in the four residual layers is 3, 4, 6, and 3 respectively;
[0100] The number of Swin Transformer blocks in the four Swin Transformer layers are 2, 2, 6, and 2, respectively.
[0101] Beneficial effects of the present invention:
[0102] The present invention constructs a dual encoder structure consisting of a ResNet branch and a Swin Transformer branch to collaboratively extract local texture information and global context information from SAR images, significantly enhancing the completeness of feature representation and accurately characterizing water morphological features and edge details in complex backgrounds. An adaptive feature fusion module designed based on a cross-attention mechanism achieves deep interaction and dynamic fusion of local and global features, effectively improving the model's ability to represent multi-scale water structures. In the decoding stage, the introduction of a multi-scale convolutional pooling module and a content-aware upsampling operator (CARAFE) not only enhances the ability to restore spatial details but also addresses the feature distortion problem in traditional upsampling. Furthermore, a composite loss function that integrates Dice loss, Focal loss, and active contour loss takes into account boundary detail constraints and class imbalance optimization, comprehensively improving detection accuracy and stability. This effectively addresses the problem of insufficient water detection accuracy under complex background interference in existing technologies. This technical solution can accurately identify water areas in complex scene SAR images, providing reliable technical support for application fields such as ecological environment monitoring, water resources management, and disaster prevention. It has important practical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0104] Figure 1 It is a schematic diagram of the overall structure of the detection model of the present invention;
[0105] Figure 2 Schematic diagram of the structure of the adaptive feature fusion module of the present invention;
[0106] Figure 3 It is a schematic diagram of the decoder structure of the present invention;
[0107] Figure 4 It is a detection effect diagram in an embodiment of the present invention. DETAILED DESCRIPTION
[0108] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0109] like Figures 1 to 4 As shown in FIG, a SAR image water body detection method based on dual encoder and attention mechanism includes the following steps:
[0110] Obtain SAR image dataset;
[0111] A dual encoder consisting of a parallel ResNet branch and a Swin Transformer branch is constructed to receive SAR images and extract local texture information and global context information of the SAR images through the two branches respectively;
[0112] Build an adaptive feature fusion module based on the attention mechanism to dynamically interact and fuse local texture information with global context information to generate a fused feature map;
[0113] Construct a decoding module including multiple decoders to receive and decode the fused feature map to generate a water body detection result with the same size as the SAR image;
[0114] The dual encoder, adaptive feature fusion module and decoder module are jointly constructed as the detection model;
[0115] The SAR image dataset is divided into training set, validation set and test set;
[0116] Construct a composite loss function including Dice loss, Focal loss and active contour loss. Based on the composite loss function, the detection model is trained using the training set and validation set.
[0117] The SAR images in the test set are input into the trained detection model and the water body detection results are output.
[0118] The ResNet branch consists of a 7×7 convolutional layer, the first batch normalization layer, the first ReLU activation function, a maximum pooling layer, and multiple residual layers connected in sequence;
[0119] The 7×7 convolutional layer is used to perform the initial convolution operation on the received SAR image, extract the initial feature map and generate the initial feature map;
[0120] The first batch normalization layer is used to receive the initial feature map and perform normalization;
[0121] The first ReLU activation function is used to receive the normalized initial feature map for nonlinear activation, and the output end of the first ReLU activation function is independently connected to the maximum pooling layer and the decoder module respectively;
[0122] The maximum pooling layer is used to downsample the initial feature map after activation and input it into the residual layer;
[0123] Each residual layer is connected in sequence, and any residual layer includes multiple sequentially connected residual units;
[0124] Let the number of residual layers be N, where N is a positive integer greater than or equal to 2;
[0125] The input of the first residual layer is connected to the output of the maximum pooling layer;
[0126] The output ends of the first N-1 residual layers are independently connected to the next residual layer and the adaptive feature fusion module;
[0127] The output of the Nth residual layer is directly connected to the adaptive feature fusion module;
[0128] The output of the nth residual layer is used to output local texture information Y n , where n∈[1,N].
[0129] The residual units all include the first 3×3 convolutional layer, the second 3×3 convolutional layer and the sixth 1×1 convolutional layer;
[0130] The input ends of the first 3×3 convolutional layer and the sixth 1×1 convolutional layer are both used to receive the feature maps in the input residual unit;
[0131] The first 3×3 convolutional layer is connected to the third batch normalization layer and the fourth ReLU activation function in sequence to process the input feature map and extract local feature information;
[0132] The second 3×3 convolutional layer is used to receive the local feature information output by the fourth ReLU activation function. The output end of the second 3×3 convolutional layer is connected to the fourth batch normalization layer for further deepening the extraction of local feature information. The second 3×3 convolutional layer and the fourth batch normalization layer serve as the main path output;
[0133] The sixth 1×1 convolutional layer is used to adjust the number of channels of the received feature map so that the adjusted number of channels matches the number of output channels of the main path, and to construct a residual path. The residual path and the main path are connected to form a residual connection and then added for output.
[0134] The Swin Transformer branch includes:
[0135] The Patch Partition module is used to receive the SAR image and split it into multiple non-overlapping 4×4 patches;
[0136] The linear embedding module is used to receive the patch block and perform linear mapping on each patch block to obtain the patch feature map;
[0137] Multiple Swin Transformer layers are arranged in sequence, where the number of Swin Transformer layers is equal to the number of residual layers, and the input of the first Swin Transformer layer is connected to the output of the linear embedding module;
[0138] Multiple Patch Merging modules: Any two adjacent Swin Transformer layers are connected via a Patch Merging module. The Patch Merging module is used to downsample, reduce the spatial dimension, and increase the channel dimension.
[0139] The outputs of the first N-1 Swin Transformer layers are independently connected to the corresponding Patch Merging module and the adaptive feature fusion module;
[0140] The output of the Nth Swin Transformer layer is directly and independently connected to the adaptive feature fusion module;
[0141] The output of the nth Swin Transformer layer is used to output the global context information X n , where n∈[1,N].
[0142] Construct an adaptive feature fusion module based on the attention mechanism to dynamically interact and fuse local texture information with global context information to generate a fused feature map. The specific steps include:
[0143] Constructing an adaptive feature fusion module including multiple adaptive feature fusion modules, wherein the multiple adaptive feature fusion modules are arranged in sequence, and the number of the adaptive feature fusion modules is equal to the number of the residual layers and corresponds one to one;
[0144] One output terminal of the nth residual layer and the nth Swin Transformer layer is connected to the input terminal of the corresponding nth adaptive feature fusion module;
[0145] In each adaptive feature fusion module, the input Y n and X n Perform channel compression and linear projection respectively, and convert Y n and Xn Unify the characteristic dimensions of
[0146] Constructing bidirectional cross-attention paths:
[0147] In one path, the output of the ResNet branch is used as the query, and the output of the Swin Transformer branch is used as the key and value, forming a cross-attention from global to local, enhancing the semantic information of local features;
[0148] In the other path, the output of the Swin Transformer branch is used as the query, and the output of the ResNet branch is used as the key and value, forming a cross-attention from local to global, enhancing the spatial details of the global features;
[0149] Based on the attention mechanism, the attention weight is calculated for any path, and the calculated attention weight is matrix multiplied with the Value in the corresponding path to obtain the first intermediate feature map;
[0150] The residual connection is used to retain the two sets of original output features of local texture information and global context information output by the two branches of the dual encoder;
[0151] A lightweight channel attention mechanism is introduced into the four-way feature consisting of two original output features and two first intermediate feature maps, and the weighted coefficient corresponding to each feature is calculated respectively;
[0152] Multiply the weighted coefficient by the corresponding road feature to finally obtain four weighted features;
[0153] Concatenate the four weighted features according to the channel dimension to obtain the weighted concatenated features;
[0154] The weighted concatenated features are simultaneously input into the second 1×1 convolutional layer and the third 1×1 convolutional layer;
[0155] The weighted concatenated features are processed by the second 1×1 convolutional layer to obtain the second intermediate feature map;
[0156] After the weighted concatenated features pass through the third 1×1 convolution layer, they are sequentially input into the first 3×3 depthwise convolution, the second ReLU activation function, and then into the fourth 1×1 convolution layer for channel compression to obtain the third intermediate feature map;
[0157] Add the second intermediate feature map and the third intermediate feature map, perform normalization, and output the fused feature map Z n , n∈[1,N].
[0158] Based on the attention mechanism, the attention weight is calculated for each path, and the calculated attention weight is matrix multiplied with the Value in the corresponding path to obtain the first intermediate feature map. Specifically, the following steps are included:
[0159] In any path, the query vector of the path is dot-producted with the key vector of the other path element by element to obtain a dot product matrix. The expression in the dot product matrix is as follows:
[0160] S=QK T
[0161] Where QK T Indicates the similarity between the Query vector and the Key vector;
[0162] Scale each element in the dot product matrix to get S ij , the specific scaling calculation formula is as follows:
[0163]
[0164] Where, Represents the square root of the feature dimension of the query or key in the path;
[0165] Q i Indicates the Query value of position i, K j Indicates the key value of position j;
[0166] For the scaled element S ij Perform Softmax normalization to obtain the attention weight A ij , the specific calculation formula is as follows:
[0167]
[0168] Where N k Indicates the sequence length of the Key;
[0169] Perform matrix multiplication of the attention weight A and the Value in the corresponding path to obtain:
[0170]
[0171] Where O represents the first intermediate feature map; V represents the Value in the corresponding path.
[0172] The decoding module also includes a fifth 1×1 convolutional layer and a Sigmoid activation function;
[0173] The number of decoders is N+1;
[0174] The output of the first decoder is connected to the input of the fifth 1×1 convolutional layer, and the output of the fifth 1×1 convolutional layer is connected to the input of the Sigmoid activation function;
[0175] The fused feature map Z output by the Nth adaptive feature fusion module N Directly input into the N+1th decoder, decode to generate feature map M N+1 ;
[0176] The feature M N+1 The fused feature map Z output by the N-1th adaptive feature fusion module N-1 After splicing, they are input into the Nth decoder together and decoded to generate feature map M N ;
[0177] Repeat this process until the feature M3 generated by the third decoder is concatenated with the fused feature map Z1 output by the first adaptive feature fusion module. The concatenated features are then input into the second decoder and decoded to generate the feature map M2.
[0178] The output result of the first ReLU activation function in the ResNet branch is concatenated with the feature M2, and then input into the first decoder to generate the feature map M1.
[0179] After the feature M1 is processed by the fifth 1×1 convolution layer and the Sigmoid activation function, a water body detection result with the same size as the SAR image is generated.
[0180] The decoders all include a multi-scale convolutional pooling module, a content-aware upsampling operator CARAFE, two second 3×3 depth convolutions, a second batch normalization layer, and a third ReLU activation function;
[0181] The feature map input to the decoder is first input into the multi-scale convolution pooling module;
[0182] The multi-scale convolution pooling module performs convolution operations on the input feature map with dilation rates of 1, 2, 4, and 6 respectively, and integrates the convolution processing results of different dilation rates into the content-aware upsampling operator CARAFE;
[0183] Preferably, when integrating the convolution processing results of different expansion rates, the convolution results processed with different expansion rates can be spliced first, and then integrated through the seventh 1×1 convolution layer;
[0184] The content-aware upsampling operator CARAFE generates a reorganization weight by analyzing the local content, and upsamples the input feature map based on the reorganization weight. The upsampling result is sequentially input into two second 3×3 depth convolutions;
[0185] After two second 3×3 depth convolution processes in sequence, the amount of calculation can be effectively reduced. The processed results are input into the second batch normalization layer and the third ReLU activation function for processing, and the decoding results are output.
[0186] Construct a composite loss function including Dice loss, Focal loss and active contour loss. Based on the composite loss function, use the training set and validation set to train the detection model. The specific steps include:
[0187] The expressions for defining the Dice loss function and the Focal loss function are as follows:
[0188]
[0189] L F =α(1-p t ) γ ·L BCE (p,g)
[0190] Where, L D represents the Dice loss function; p represents the prediction result; g is the true label, and ε is the smoothing term;
[0191] L F represents the Focal loss function; p t =exp(-L BCE ), L BCE is the binary cross entropy loss; α and γ are both focusing factors;
[0192] Preferably, the values of α and γ in this application are 0.8 and 0.2 respectively;
[0193] The active contour loss function combines the regional energy term and the boundary energy term to guide the model to learn the target area and accurately locate the edge position. The active contour loss function L is defined as AC The expression is as follows:
[0194] L AC =Length+Region
[0195]
[0196] Where, Region represents the regional energy term; Length represents the boundary energy term;
[0197] P i,j Y is the probability map predicted by the detection model; i,j are the average intensities of the foreground and background regions of the true labels c1 and c2 respectively; x P i,j Represents the gradient of the probability graph on the x-axis; ▽ y Pi,j represents the gradient of the probability map on the y-axis; H and W represent the height and width of the image respectively; μ, ν and λ are the weight parameters of the corresponding items;
[0198] Preferably, in this application, μ, ν and λ are weight parameters of the corresponding items, which are 0.8, 0.4 and 0.01 respectively;
[0199] Constructing the composite loss function L total , the specific expression is as follows:
[0200] L total =w D ·L D +w F ·L F +w AC ·L AC
[0201] Where w D 、w F 、w AC are the weight parameters of Dice loss, Focal loss and active contour loss respectively;
[0202] Preferably, in this application D 、w F 、w AC The values of are 0.45, 0.35 and 0.2 respectively;
[0203] Set a convergence threshold or early stopping mechanism for validation loss;
[0204] With the composite loss function L total The training objective is to minimize the value of , and the training set is used to train the detection model;
[0205] After each training cycle, the validation set is used to calculate the validation loss. When the validation loss reaches the convergence threshold or the early stopping mechanism is triggered, the detection model training is completed.
[0206] Preferably, the convergence threshold in this application can be set to 0.00001. When the verification loss changes little in several consecutive training cycles and the verification loss is less than the threshold, the training is determined to be finished.
[0207] The early stopping mechanism means stopping training when the training loss does not improve within the patience value consecutive training cycles. Preferably, the patience value is set to 10 in this application;
[0208] The number of residual layers and Swin Transformer layers is four each;
[0209] The number of residual units in the four residual layers is 3, 4, 6, and 3 respectively;
[0210] The number of Swin Transformer blocks in the four Swin Transformer layers is 2, 2, 6, and 2 respectively;
[0211] In this application, SAR images are input in the dimensions of C × H × W;
[0212] Where C represents the number of channels; H represents the height of the image; W represents the width of the image;
[0213] In the embodiment of the present application, the SAR image dimension is 3×256×256;
[0214] In the Swin Transformer branch, the SAR image is first divided into multiple non-overlapping 4×4 patches by the patch partition module. Each patch is linearly mapped to obtain a patch feature map with a dimension of 48×64×64. The global context information output by the subsequent four Swin Transformer layers has dimensions of 96×64×64, 192×32×32, 384×16×16, and 768×8×8, respectively.
[0215] In the ResNet branch, the initial feature map with a dimension of 64×128×128 is obtained after the 7×7 convolution layer, the first batch normalization layer, and the first ReLU activation function. Then, the dimensions of the local texture information output by the first to fourth residual layers are 64×64×64, 128×32×32, 256×16×16, and 512×8×8, respectively.
[0216] In the decoding module, the dimensions of the feature maps output by the 5th decoder to the 1st decoder after decoding are 512×16×16, 256×32×32, 128×64×64, 64×128×128 and 64×256×256 respectively; the feature M1 (dimension is 64×256×256) is processed by the fifth 1×1 convolution layer and the Sigmoid activation function in sequence to generate a water body detection result (i.e., prediction map) with a dimension of 3×256×256.
[0217] The present invention constructs a dual encoder structure consisting of a ResNet branch and a Swin Transformer branch to collaboratively extract local texture information and global context information from SAR images, significantly enhancing the completeness of feature representation and accurately characterizing water morphological features and edge details in complex backgrounds. An adaptive feature fusion module designed based on a cross-attention mechanism achieves deep interaction and dynamic fusion of local and global features, effectively improving the model's ability to represent multi-scale water structures. In the decoding stage, the introduction of a multi-scale convolutional pooling module and a content-aware upsampling operator (CARAFE) not only enhances the ability to restore spatial details but also addresses the feature distortion problem in traditional upsampling. Furthermore, a composite loss function that integrates Dice loss, Focal loss, and active contour loss takes into account boundary detail constraints and class imbalance optimization, comprehensively improving detection accuracy and stability. This effectively addresses the problem of insufficient water detection accuracy under complex background interference in existing technologies. This technical solution can accurately identify water areas in complex scene SAR images, providing reliable technical support for application fields such as ecological environment monitoring, water resources management, and disaster prevention. It has important practical value and broad application prospects.
[0218] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0219] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.
Claims
1. A water body detection method for SAR images based on dual encoders and attention mechanism, characterized in that: The specific steps include: Obtain SAR image dataset; A dual encoder consisting of a parallel ResNet branch and a Swin Transformer branch is constructed to receive SAR images and extract local texture information and global context information of the SAR images through the two branches respectively; Build an adaptive feature fusion module based on the attention mechanism to dynamically interact and fuse local texture information with global context information to generate a fused feature map; Construct a decoding module including multiple decoders to receive and decode the fused feature map to generate a water body detection result with the same size as the SAR image; The dual encoder, adaptive feature fusion module and decoder module are jointly constructed as the detection model; The SAR image dataset is divided into training set, validation set and test set; Construct a composite loss function including Dice loss, Focal loss and active contour loss. Based on the composite loss function, the detection model is trained using the training set and validation set. The SAR images in the test set are input into the trained detection model and the water body detection results are output.
2. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 1 is characterized in that: The ResNet branch consists of a 7×7 convolutional layer, the first batch normalization layer, the first ReLU activation function, a maximum pooling layer, and multiple residual layers connected in sequence; The 7×7 convolutional layer is used to perform the initial convolution operation on the received SAR image, extract the initial feature map and generate the initial feature map; The first batch normalization layer is used to receive the initial feature map and perform normalization; The first ReLU activation function is used to receive the normalized initial feature map for nonlinear activation, and the output end of the first ReLU activation function is independently connected to the maximum pooling layer and the decoder module respectively; The maximum pooling layer is used to downsample the initial feature map after activation and input it into the residual layer; Each residual layer is connected in sequence, and any residual layer includes multiple sequentially connected residual units; Let the number of residual layers be N, where N is a positive integer greater than or equal to 2; The input of the first residual layer is connected to the output of the maximum pooling layer; The output ends of the first N-1 residual layers are independently connected to the next residual layer and the adaptive feature fusion module; The output of the Nth residual layer is directly connected to the adaptive feature fusion module; The output of the nth residual layer is used to output local texture information Y n , where n∈[1,N].
3. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 2 is characterized in that: The residual units all include the first 3×3 convolutional layer, the second 3×3 convolutional layer and the sixth 1×1 convolutional layer; The input ends of the first 3×3 convolutional layer and the sixth 1×1 convolutional layer are both used to receive the feature maps in the input residual unit; The first 3×3 convolutional layer is connected to the third batch normalization layer and the fourth ReLU activation function in sequence to process the input feature map and extract local feature information; The second 3×3 convolutional layer is used to receive the local feature information output by the fourth ReLU activation function. The output end of the second 3×3 convolutional layer is connected to the fourth batch normalization layer for further deepening the extraction of local feature information. The second 3×3 convolutional layer and the fourth batch normalization layer serve as the main path output; The sixth 1×1 convolutional layer is used to adjust the number of channels of the received feature map so that the adjusted number of channels matches the number of output channels of the main path, and to construct a residual path. The residual path and the main path are connected to form a residual connection and then added for output.
4. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 3 is characterized in that: The Swin Transformer branch includes: The segmentation module is used to receive the SAR image and segment it into multiple non-overlapping 4×4 patches; The linear embedding module is used to receive the patch block and perform linear mapping on each patch block to obtain the patch feature map; Multiple Swin Transformer layers are arranged in sequence, where the number of Swin Transformer layers is equal to the number of residual layers, and the input of the first Swin Transformer layer is connected to the output of the linear embedding module; Multiple Patch Merging modules: Any two adjacent Swin Transformer layers are connected via a Patch Merging module. The Patch Merging module is used to downsample, reduce the spatial dimension, and increase the channel dimension. The outputs of the first N-1 Swin Transformer layers are independently connected to the corresponding Patch Merging module and the adaptive feature fusion module; The output of the Nth Swin Transformer layer is directly and independently connected to the adaptive feature fusion module; The output of the nth Swin Transformer layer is used to output the global context information X n , where n∈[1,N].
5. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 4 is characterized in that: Construct an adaptive feature fusion module based on the attention mechanism to dynamically interact and fuse local texture information with global context information to generate a fused feature map. The specific steps include: Constructing an adaptive feature fusion module including multiple adaptive feature fusion modules, wherein the multiple adaptive feature fusion modules are arranged in sequence, and the number of the adaptive feature fusion modules is equal to the number of the residual layers and corresponds one to one; One output terminal of the nth residual layer and the nth Swin Transformer layer is connected to the input terminal of the corresponding nth adaptive feature fusion module; In each adaptive feature fusion module, the input Y n and X n Perform channel compression and linear projection respectively, and convert Y n and X n Unify the characteristic dimensions of Constructing bidirectional cross-attention paths: In one path, the output of the ResNet branch is used as the query, and the output of the Swin Transformer branch is used as the key and value, forming a cross-attention from global to local, enhancing the semantic information of local features; In the other path, the output of the Swin Transformer branch is used as the query, and the output of the ResNet branch is used as the key and value, forming a cross-attention from local to global, enhancing the spatial details of the global features; Based on the attention mechanism, the attention weight is calculated for any path, and the calculated attention weight is matrix multiplied with the Value in the corresponding path to obtain the first intermediate feature map; The residual connection is used to retain the two sets of original output features of local texture information and global context information output by the two branches of the dual encoder; A lightweight channel attention mechanism is introduced into the four-way feature consisting of two original output features and two first intermediate feature maps, and the weighted coefficient corresponding to each feature is calculated respectively; Multiply the weighted coefficient by the corresponding road feature to finally obtain four weighted features; Concatenate the four weighted features according to the channel dimension to obtain the weighted concatenated features; The weighted concatenated features are simultaneously input into the second 1×1 convolutional layer and the third 1×1 convolutional layer; The weighted concatenated features are processed by the second 1×1 convolutional layer to obtain the second intermediate feature map; After the weighted concatenated features pass through the third 1×1 convolution layer, they are sequentially input into the first 3×3 depthwise convolution, the second ReLU activation function, and then into the fourth 1×1 convolution layer for channel compression to obtain the third intermediate feature map; Add the second intermediate feature map and the third intermediate feature map, perform normalization, and output the fused feature map Z n , n∈[1,N].
6. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 5 is characterized in that: Based on the attention mechanism, the attention weight is calculated for each path, and the calculated attention weight is matrix multiplied with the Value in the corresponding path to obtain the first intermediate feature map. Specifically, the following steps are included: In any path, the query vector of the path is dot-producted with the key vector of the other path element by element to obtain a dot product matrix. The expression in the dot product matrix is as follows: S=QK T Where QK T Indicates the similarity between the Query vector and the Key vector; Scale each element in the dot product matrix to get S ij , the specific scaling calculation formula is as follows: Where, Represents the square root of the feature dimension of the query or key in the path; Q i Indicates the Query value of position i, K j Indicates the key value of position j; For the scaled element S ij Perform Softmax normalization to obtain the attention weight A ij , the specific calculation formula is as follows: Where N k Indicates the sequence length of the Key; Perform matrix multiplication of the attention weight A and the Value in the corresponding path to obtain: Where O represents the first intermediate feature map; V represents the Value in the corresponding path.
7. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 6 is characterized in that: The decoding module also includes a fifth 1×1 convolutional layer and a Sigmoid activation function; The number of decoders is N+1; The output of the first decoder is connected to the input of the fifth 1×1 convolutional layer, and the output of the fifth 1×1 convolutional layer is connected to the input of the Sigmoid activation function; The fused feature map Z output by the Nth adaptive feature fusion module N Directly input into the N+1th decoder, decode to generate feature map M N+1 ; The feature M N+1 The fused feature map Z output by the N-1th adaptive feature fusion module N-1 After splicing, they are input into the Nth decoder together and decoded to generate feature map M N ; Repeat this process until the feature M3 generated by the third decoder is concatenated with the fused feature map Z1 output by the first adaptive feature fusion module. The concatenated features are then input into the second decoder and decoded to generate the feature map M2. The output result of the first ReLU activation function in the ResNet branch is concatenated with the feature M2, and then input into the first decoder to generate the feature map M1. After the feature M1 is processed by the fifth 1×1 convolution layer and the Sigmoid activation function, a water body detection result with the same size as the SAR image is generated.
8. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 7 is characterized in that: The decoders all include a multi-scale convolutional pooling module, a content-aware upsampling operator CARAFE, two second 3×3 depth convolutions, a second batch normalization layer, and a third ReLU activation function; The feature map input to the decoder is first input into the multi-scale convolution pooling module; The multi-scale convolution pooling module performs convolution operations on the input feature map with dilation rates of 1, 2, 4, and 6 respectively, and integrates the convolution processing results of different dilation rates into the content-aware upsampling operator CARAFE; The content-aware upsampling operator CARAFE generates a reorganization weight by analyzing the local content, and upsamples the input feature map based on the reorganization weight. The upsampling result is sequentially input into two second 3×3 depth convolutions; After two second 3×3 depth convolution processes, the second batch normalization layer and the third ReLU activation function are input for processing, and the decoding result is output.
9. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 8 is characterized in that: Construct a composite loss function including Dice loss, Focal loss and active contour loss. Based on the composite loss function, use the training set and validation set to train the detection model. The specific steps include: The expressions for defining the Dice loss function and the Focal loss function are as follows: L F =α(1-p t ) γ ·L BCE (p,g) Where, L D represents the Dice loss function; p represents the prediction result; g is the true label, and ε is the smoothing term; L F represents the Focal loss function; p t =exp(-L BCE ), L BCE is the binary cross entropy loss; α and γ are both focusing factors; The active contour loss function combines the regional energy term and the boundary energy term to guide the model to learn the target area and accurately locate the edge position. The active contour loss function L is defined as AC The expression is as follows: L AC =Length+Region Where, Region represents the regional energy term; Length represents the boundary energy term; P i,j Y is the probability map predicted by the detection model; i,j The true labels c1 and c2 are the average intensities of the foreground and background regions respectively; Represents the gradient of the probability map on the x-axis; represents the gradient of the probability map on the y-axis; H and W represent the height and width of the image respectively; μ, ν and λ are the weight parameters of the corresponding items; Constructing the composite loss function L total , the specific expression is as follows: L total =w D ·L D +w F ·L F +w AC ·L AC Where w D 、w F 、w AC are the weight parameters of Dice loss, Focal loss and active contour loss respectively; Set a convergence threshold or early stopping mechanism for validation loss; With the composite loss function L total The training objective is to minimize the value of , and the training set is used to train the detection model; After each training cycle, the validation set is used to calculate the validation loss. When the validation loss reaches the convergence threshold or the early stopping mechanism is triggered, the detection model training is completed.
10. The SAR image water body detection method based on dual encoder and attention mechanism according to claim 9 is characterized in that: The number of residual layers and Swin Transformer layers is four each; The number of residual units in the four residual layers is 3, 4, 6, and 3 respectively; The number of Swin Transformer blocks in the four Swin Transformer layers are 2, 2, 6, and 2, respectively.