Water segmentation method for remote sensing images based on attention-based Water-Lite-HRNet
By constructing an attention-based Water-Lite-HRNet network, combining structured pruning and attention modules, the problem of inaccurate segmentation results in water body segmentation in remote sensing images is solved, and high-precision and efficient water body segmentation effect is achieved.
Patent Information
- Application Number
- CN202310992432.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-08-08
AI Technical Summary
In the prior art, the remote sensing image water body segmentation method is poor, the segmentation results are inaccurate or discontinuous, and cannot meet the actual application needs.
The attention-based Water-Lite-HRNet network is adopted to construct four branches with different resolutions through structured pruning and embedded attention modules, and train them with adjacent branches' attention modules and mixed loss functions. Full-connection conditional random field is used for post-processing to improve segmentation accuracy.
While reducing the amount of model parameters, the accuracy of water segmentation is significantly improved, the average cross-border ratio reaches 97.96%, and the model training speed is improved.
Smart Images

Figure CN117058674B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent interpretation of remote sensing images, and specifically relates to an attention-based Water-Lite-HRNet remote sensing image water body segmentation method. Background Art
[0002] Water bodies are water that exists in nature in a specific form. They are an important component of the Earth's surface hydrosphere and are natural bodies of water bounded by relatively stable land. They include groundwater, atmospheric water vapor, rivers, lakes, seas, glaciers, snow, reservoirs, and ponds. Water bodies in remote sensing images refer to surface water, which is defined by clear topographic boundaries and water. Within a geographic area, it is considered a water body if the water is concentrated in a specific location, such as an ocean, lake, river, or reservoir. Within the topographic and water body boundaries, there is usually soil, vegetation, or a combination of soil and vegetation.
[0003] Water segmentation is a technique for accurately segmenting and characterizing water bodies against diverse background features. It analyzes acquired remote sensing images and performs rigorous, pixel-accurate water-land separation. Its results are crucial for applications such as water quality monitoring, military reconnaissance, environmental protection, land planning, and geographic mapping. However, the boundaries of water bodies in remote sensing images are often complex and diverse, encompassing lakes, rivers, and coastlines. These boundaries can be affected by factors such as topography, shape variations, shadows, and reflected light, resulting in poor results in these applications. Conventional methods cannot meet the requirements of water segmentation, leading to inaccurate or discontinuous segmentation results. Summary of the Invention
[0004] Purpose of the invention: To address the problems of poor field application effects, inaccurate or discontinuous segmentation results in the existing technology, the present invention provides an attention-based Water-Lite-HRNet remote sensing image water body segmentation method, which greatly reduces the number of HRNet parameters and improves the model training speed while improving the model segmentation accuracy.
[0005] Technical solution: The present invention provides an attention-based Water-Lite-HRNet remote sensing image water body segmentation method, comprising the following steps:
[0006] Step 1: Data preparation: Use LandCove.ai's original high-resolution remote sensing images and preprocess them to obtain a water body segmentation dataset. The dataset includes a training set, a test set, and a validation set.
[0007] Step 2: Build a network and perform structured pruning on the original HRNet network to obtain a Water-Lite-HRNet network, while embedding the attention module and the adjacent branch attention module. The Water-Lite-HRNet network is composed of four branches stacked together: branch 1, branch 2, branch 3, and branch 4. The attention module is embedded in the four branches, and the four branches only fuse features with adjacent branches of different resolutions. The adjacent branch attention module is introduced before the cross-branch fusion of high-level and low-level features to extract adjacent branch features.
[0008] Step 3: Train the model. Fine-tune and train the Water-Lite-HRNet network using the public dataset and the training and validation sets of the water segmentation dataset in step 1 to obtain the segmentation model Y.
[0009] Step 4: Segmentation and post-processing: Input the test set of the water body segmentation dataset in step 1 into the segmentation model Y to obtain the initial water body segmentation result, and then post-process it to obtain the final water body segmentation result.
[0010] Furthermore, in step 1, the LandCove.ai original dataset is labeled into four categories, including buildings, woodlands, water, and roads, and the sliding window overlapping cropping method is used for the original image preprocessing.
[0011] Furthermore, the specific implementation process of the attention-based Water-Lite-HRNet model in step 2 is as follows:
[0012] S2.1: Perform structured pruning on the HRNet network. The HRNet network is divided into five stages: the starting module, the first stage, the second stage, the third stage, and the fourth stage. The starting module performs preliminary processing and feature extraction on the input image, and the remaining stages are used for further feature extraction, information transmission, and prediction of the final segmentation results. The function of generating low-resolution branches is at the beginning of the first, second, and third stages, the multi-scale fusion function is between the first, second, third, and fourth stages, and the output of the network ends at the fourth stage.
[0013] S2.2: Starting module; the original resolution of the corresponding network is used as branch 1. Through the convolution operation, the input image is convolved and processed using Basicblock or Bottleblock to extract features. The attention module is used on the last convolution layer and the result is input into the first stage.
[0014] S2.3: The first stage: downsample branch 1 from the previous stage to generate low-resolution branch 2. Each branch uses Basicblock or Bottleblock for feature extraction, and uses the attention module on the last convolutional layer to input the result into the second stage.
[0015] S2.4: In the second stage, branch 2 of the previous stage is downsampled to generate a low-resolution branch 3. Branch 1 is only fused with the features of the adjacent branch 2, branch 2 is only fused with the features of the adjacent branch 1, and branch 3 is only fused with the features of the adjacent branch 2, cutting out the cross-branch feature fusion function; each branch uses Basicblock or Bottleblock for feature extraction, and an attention module is used on the last convolutional layer, and the result is input into the third stage; in the original HRNet network, all branches in the current stage are multi-scale fused with all branches in the previous stage, requiring 4 fusions. In Water-Lite-HRNet, 3 fusions are required;
[0016] S2.5: The third stage; downsample branch 3 of the previous stage to generate a low-resolution branch 4. Branch 1 of the current stage is only fused with the features of the adjacent branch 2, branch 2 is fused with the features of the adjacent branches 1 and 3, branch 3 is fused with the features of the adjacent branch 2, and branch 4 is only fused with the adjacent branch 3. The function of cross-branch feature fusion is cut off; each branch uses Basicblock or Bottleblock for feature extraction, uses the attention module on the last convolutional layer, and inputs the result into the fourth stage; in the original HRNet network, all branches in the current stage are multi-scale fused with all branches in the previous stage, which requires 9 fusions. In Water-Lite-HRNet, 5 fusions are required;
[0017] S2.6: The fourth stage: Branch 1 is only fused with the features of the adjacent branch 2, branch 2 is only fused with the features of the adjacent branches 1 and 3, branch 3 is only fused with the features of the adjacent branches 2 and 4, and branch 4 is only fused with the features of the adjacent branch 3. The cross-branch feature fusion function is cut off. After feature fusion, the result is input into the attention module of the adjacent branch. In the original HRNet network, all branches in the current stage are multi-scale fused with all branches in the previous stage, requiring 12 fusions. In Water-Lite-HRNet, 6 fusions are required.
[0018] S2.7: At the end of the fourth stage, that is, the results of the adjacent branch attention modules, it is necessary to retain the function of cross-branch feature fusion, upsample the three parallel low-resolution outputs to the size of branch 1, splice the results of the four branches, and then obtain the final segmentation result through 1×1 convolution.
[0019] Furthermore, the attention module includes a channel attention branch and a spatial attention branch, and the specific implementation process is as follows:
[0020] The channel attention branch transforms the input feature map F of dimension H×W×C into a feature representation F1∈R through two separate 1×1 convolutions H×W×C and Q1∈R H×W×C ; Where F1 is the feature response map, Q1 is the spatial weight map, H represents the height, W represents the width, and C represents the number of channels; the feature representations F1 and Q1 are respectively transformed into feature matrices V1∈R C×N and eigenvector S W1 ∈R N ×1 , where N = H × W; in S W1 Then add a fully connected layer S W2 ∈R N×1 , V1 and S W2 Multiply to get C1∈R C×1 , add a fully connected layer C after C1 A ∈R C×1 As the output channel attention vector; use the Sigmoid function to A Normalize and finally C A Multiply it by the input feature map to get F C ∈R H×W×C ;
[0021] The spatial attention branch transforms the input feature map F into a feature representation F2∈R via two separate 1×1 convolutions H ×W×C and Q2∈R H×W×C ; F2 is the feature response map, and then F2 is transformed into the feature matrix V2∈R C×N , where N = H × W, Q2 compresses each channel into a single element through global average pooling, thus obtaining the feature vector C W1 ∈R C×1 ; in C W1 Then add a fully connected layer C W2 ∈R C×1 , use the Softmax function to C W2 Perform the operation and multiply it with V2 to get S1∈R 1×N , add a fully connected layer S2∈R after S1 1×N As the output spatial attention vector, S2 is transformed into S A ∈R H×W×1 ; and use the Sigmoid function to A Normalize and finally convert S A Multiply it by the input feature map to get F S ∈RH×W×C ;
[0022] Then for F C and F S Using concatenation and 1×1 convolution operations, we get the output feature map of the attention module.
[0023] Furthermore, the adjacent branch attention module includes an upsampling two-input module, a downsampling two-input module and an up- and downsampling three-input module. The upsampling two-input module inputs the feature maps of branch 2 and branch 3; the downsampling two-input module inputs the feature maps of branch 3 and branch 4; the up- and downsampling three-input module inputs the feature maps of branch 2, branch 3 and branch 4, and the outputs of the upsampling two-input module, the downsampling two-input module and the up- and downsampling three-input module are B2, B3 and B4 respectively; B2, B3 and B4 are restored to the size of branch 1 after step-by-step upsampling, and information is transferred from the low-resolution branch to the high-resolution branch. The four are spliced together and then a 1×1 convolution is used as the final output of the model.
[0024] Furthermore, the upsampling two-input module upsamples the output feature map size in the branch 3 to the feature map size of the branch 2, and inputs it into the spatial attention module to obtain B21. The output feature map of the branch 2 is subjected to the void spatial pyramid pooling module and the 3×3 convolution to obtain B22; B22 is subjected to the channel attention module to obtain B23, B22 and B23 are multiplied and passed through the spatial attention module to obtain B24, and B22 and B24 are multiplied to obtain B25; B21 and B22 are multiplied to obtain B26, and B25 and B26 are added to obtain the final output B2 of the downsampling two-input module.
[0025] Furthermore, the downsampling two-input module downsamples the output feature map size in the branch 3 to the feature map size of the branch 4, and inputs it into the spatial attention module to obtain B41. The output feature map of the branch 4 is passed through the void spatial pyramid pooling module and the 3×3 convolution to obtain B42; B42 is passed through the channel attention module to obtain B43, B42 and B43 are multiplied and passed through the spatial attention module to obtain B44, and B42 and B44 are multiplied to obtain B45; B41 and B42 are multiplied to obtain B46, and B45 and B46 are added to obtain the final output B4 of the downsampling two-input module.
[0026] Furthermore, the up-down sampling three-input module downsamples the output feature map size in the branch 2 to the feature map size of the branch 3, and inputs it into the spatial attention module to obtain B31. The output feature map of the branch 3 is passed through the void spatial pyramid pooling module and the 3×3 convolution to obtain B32; B32 is passed through the channel attention module to obtain B33, B32 and B33 are multiplied and passed through the spatial attention module to obtain B34, B32 and B34 are multiplied to obtain B35, and B31 and B32 are multiplied to obtain B36; the output feature map size in the branch 4 is upsampled to the feature map size of the branch 3, and inputted into the spatial attention module to obtain B37, and B32 and B37 are multiplied to obtain B38; B35, B36 and B38 are added to obtain the final output B3 of the up-down sampling three-input module.
[0027] Furthermore, the loss function during network training in step 3 adopts a hybrid loss function, and the specific calculation formula is as follows:
[0028] L total =λL dice +(1-λ)L ce
[0029] Among them, the Dice loss function is recorded as L dice , the cross entropy loss function is recorded as L ce , λ is a hyperparameter used to balance the two loss functions, and the parameter range is [0,1];
[0030] For the Dice loss function, the specific calculation formula is as follows:
[0031]
[0032] Among them, p i Represents the probability value predicted by the model, t i Represents the corresponding target value, and M represents the number of samples;
[0033] For the cross entropy loss function, the specific calculation formula is as follows:
[0034]
[0035] Among them, p i Represents the probability value predicted by the model, t i Represents the corresponding target value, and M represents the number of samples.
[0036] Furthermore, the step 4 uses a fully connected conditional random field method to post-process the initial water body segmentation results that contain errors, uncertainties, and blurred classification target boundaries, to obtain an accurate and smooth final classification result.
[0037] Beneficial effects:
[0038] 1. The main feature of the Water-Lite-HRNet of the present invention is that the network is composed of four branches of different resolutions, namely branch 1, branch 2, branch 3 and branch 4, which are stacked together. Through these branches of different resolutions, low-level details and high-level semantic information can be retained at the same time. The original HRNet network is subjected to structured pruning to obtain the Water-Lite-HRNet network. The structured pruning operation refers to retaining only the function of outputting the segmentation result of the cross-branch fusion features at the end of the fourth stage, and removing the function of the cross-branch fusion features of the remaining stages in the network, which can reduce the number of network parameters and increase the training speed. An attention module is embedded in the four branches of Water-Lite-HRNet so that the network focuses on features that are more meaningful for water body segmentation, thereby improving the performance of the network in water body segmentation. The channel attention branch can increase the weight of the feature channel related to water body segmentation, and the spatial attention branch can focus on the semantics related to the foreground position. The adjacent branch attention module is introduced before the cross-branch fusion of high and low-level features to reduce the impact of noise on the feature map during the cross-layer fusion of high and low-level features.
[0039] 2. The present invention adopts the fully connected conditional random field method in the post-processing process to effectively reduce the noise and incoherence in the segmentation results and provide more accurate and smooth semantic segmentation results.
[0040] 3. The attention-based Water-Lite-HRNet remote sensing image water segmentation algorithm adopted in this paper improves the accuracy of water segmentation while reducing the number of model parameters, with a mean intersection over Union (mIoU) of 97.96%. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flowchart of the attention-based Water-Lite-HRNet remote sensing image water segmentation method;
[0042] Figure 2 This is the original network structure diagram of HRNet;
[0043] Figure 3 This is the network structure diagram of the attention-based Water-Lite-HRNet proposed in this paper;
[0044] Figure 4 This is the structural diagram of the attention module proposed in the present invention;
[0045] Figure 5 It is the upsampling two-input module in the adjacent branch attention module proposed in this invention;
[0046] Figure 6It is the downsampling two-input module in the adjacent branch attention module proposed in this invention;
[0047] Figure 7 It is the up- and down-sampling three-input module in the adjacent branch attention module proposed in this invention;
[0048] Figure 8 This is a comparison chart of the water body segmentation effects using the Water-Lite-HRNet, U-Net and DeepLabv3 models in this paper. DETAILED DESCRIPTION
[0049] The following combination Figures 1 to 8 The present invention is further described in the following examples, which are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of the present invention.
[0050] The present invention provides a water body segmentation method for remote sensing images based on attention Water-Lite-HRNet. The whole process is as follows: Figure 1 The specific implementation operations are as follows:
[0051] Step 1: First, the LandCover.ai raw dataset is annotated into four categories: buildings, woodlands, water, and roads. It includes 33 high-resolution remote sensing images of approximately 9000×9500 pixels and 8 high-resolution remote sensing images of approximately 4200×4700 pixels. Next, the raw images are preprocessed using a sliding window overlapping cropping method. Based on the raw high-resolution remote sensing images, a window size of 800×800 and a stride of 600 are defined. Starting from the top left corner of each image, the window is slid in units of the stride until the entire image is covered. Each time the window is slid, the image patch at the current window position is cropped. The final dataset contains 8368 optical images with red, green, and blue channels. For this invention, each image is simply labeled as water or non-water. 60% of the images are randomly selected as the training set, 5% of the images are randomly selected as the validation set, and the remaining 35% of the images are used as the test set.
[0052] Step 2: Build the network and perform structured pruning on the original HRNet network to obtain the Water-Lite-HRNet network, while embedding the designed attention module and adjacent branch attention module. The original HRNet network structure is as follows Figure 2 As shown, the Water-Lite-HRNet network structure is as follows Figure 3 The specific steps are as follows:
[0053] 2.1: Structured pruning of the HRNet network. The HRNet network maintains a high-resolution representation by concatenating parallel representations of different resolutions and repeatedly performing multi-scale fusion. The entire network is divided into five stages: the initiation module, stage 1, stage 2, stage 3, and stage 4. The initiation module performs preliminary processing of the input image and extracts features. The remaining stages are responsible for further feature extraction, information transfer, and prediction of the final segmentation result. The functions for generating low-resolution branches begin in stages 1, 2, and 3, while the multi-scale fusion function occurs between stages 1, 2, 3, and 4. The network output ends at stage 4.
[0054] 2.2: Starting Module. This corresponds to the network's original resolution as branch 1. The input image is convolved using a convolution operation, processed using a Basicblock or Bottleblock to extract features. An attention module is applied to the final convolutional layer, and the result is fed into the first stage. The resulting image is of size H × W × 512.
[0055] 2.3: Stage 1. Branch 1 from the previous stage is downsampled to produce a low-resolution branch 2. Each branch uses Basicblock or Bottleblock for feature extraction. An attention module is used on the last convolutional layer, and the results are fed into the second stage. The resulting sizes are H × W × 512 and (H / 2) × (W / 2) × 512, respectively.
[0056] 2.4: The second stage. The branch 2 of the previous stage is downsampled to produce a low-resolution branch 3. Branch 1 is only fused with the features of the adjacent branch 2, branch 2 is only fused with the features of the adjacent branch 1, and branch 3 is only fused with the features of the adjacent branch 2. The function of cross-branch feature fusion is cut off. Each branch uses Basicblock or Bottleblock for feature extraction, and an attention module is used on the last convolutional layer to input the result into the third stage. The result sizes are H×W×512, (H / 2)×(W / 2)×512, and (H / 4)×(W / 4)×512, respectively. In the original HRNet network, all branches in the current stage are multi-scale fused with all branches in the previous stage, which requires 4 fusions. In Water-Lite-HRNet, 3 fusions are required.
[0057] 2.5: The third stage. Branch 3 from the previous stage is downsampled to produce a low-resolution branch 4. In the current stage, branch 1 is fused only with the features of the adjacent branch 2. Branch 2 is fused with the features of the adjacent branches 1 and 3. Branch 3 is fused with the features of the adjacent branches 2. Branch 4 is fused only with the adjacent branch 3. Cross-branch feature fusion is removed. Each branch uses Basicblock or Bottleblock for feature extraction. An attention module is applied to the last convolutional layer, and the results are input to the fourth stage. The resulting sizes are H×W×512, (H / 2)×(W / 2)×512, (H / 4)×(W / 4)×512, and (H / 8)×(W / 8)×512, respectively. In the original HRNet network, all branches in the current stage are multi-scale fused with all branches in the previous stage, requiring nine fusions. In Water-Lite-HRNet, five fusions are required.
[0058] 2.6: Fourth stage. Branch 1 is only fused with features from the adjacent branch 2, branch 2 is only fused with features from the adjacent branches 1 and 3, branch 3 is only fused with features from the adjacent branches 2 and 4, and branch 4 is only fused with features from the adjacent branch 3. Cross-branch feature fusion is removed. After feature fusion, the results are fed into the adjacent branch attention modules. The resulting sizes are H×W×512, (H / 2)×(W / 2)×512, (H / 4)×(W / 4)×512, and (H / 8)×(W / 8)×512, respectively. In the original HRNet network, all branches in the current stage are multi-scale fused with all branches in the previous stage, requiring 12 fusions. In Water-Lite-HRNet, 6 fusions are required.
[0059] 2.7: At the end of the fourth stage, that is, the results of the adjacent branch attention modules, it is necessary to retain the function of cross-branch feature fusion, upsample the three parallel low-resolution outputs to the size of branch 1, splice the results of the four branches, and then obtain the final segmentation result through 1×1 convolution.
[0060] The structure of the attention module is as follows Figure 4 As shown in Figure 3, it allows the network to focus on features that are more meaningful for water body segmentation, thereby improving the network's performance in water body segmentation. The attention module consists of two parts: the channel attention branch and the spatial attention branch.
[0061] The channel attention branch in the attention module is used to increase the weight of the feature channels related to water body segmentation. First, the input feature map F with dimension H×W×C is converted into a feature representation F1∈R through two separate 1×1 convolutions. H×W×C and Q1∈R H×W×C; Where F1 is the feature response map, Q1 is the spatial weight map, H represents the height, W represents the width, and C represents the number of channels; Secondly, the feature representations F1 and Q1 are transformed into feature matrices V1∈R C×N and eigenvector S W1 ∈R N×1 , where N = H × W; again in S W1 Then add a fully connected layer S W2 ∈R N×1 , V1 and S W2 Multiply to get C1∈R C×1 , add a fully connected layer C after C1 A ∈R C×1 As the output channel attention vector; then use the Sigmoid function to A Normalize and then C A Multiply it by the input feature map to get F C ∈R H×W×C .
[0062] The spatial attention branch in the attention module is used to focus more on the semantics related to the foreground position. First, the input feature map F is converted into a feature representation F2∈R by two separate 1×1 convolutions. H×W×C and Q2∈R H×W×C ; F2 is the feature response map, and then F2 is transformed into the feature matrix V2∈R C×N , where N = H × W, Q2 compresses each channel into a single element through global average pooling, thus obtaining the feature vector C W1 ∈R C×1 ; Again in C W1 Then add a fully connected layer C W2 ∈R C×1 , use the Softmax function to C W2 Perform the operation and multiply it with V2 to get S1∈R 1×N , add a fully connected layer S2∈R after S1 1×N As the output spatial attention vector, S2 is transformed into S A ∈R H×W×1 ; Then use the Sigmoid function to A Normalize and then convert S A Multiply it by the input feature map to get F S ∈R H×W×C .
[0063] Finally, F C and F S Using concatenation and 1×1 convolution operations, we get the output feature map of the attention module.
[0064] Before cross-branch fusion of high-level and low-level features, an adjacent branch attention module is introduced. This module can reduce the impact of noise on feature maps during the cross-layer fusion of high-level and low-level features, extract adjacent features, and ultimately obtain the required attention-based algorithm. This module includes an upsampling two-input module, a downsampling two-input module, and an upsampling and downsampling three-input module.
[0065] The structure of the upsampling two-input module is as follows Figure 5 As shown, first, the upsampling two-input module inputs the feature maps of branches 2 and 3, and the output feature map size in branch 3 is upsampled to the feature map size of branch 2, and input into the spatial attention module to obtain B21. Secondly, the output feature map of branch 2 passes through the void spatial pyramid pooling module and convolution to obtain B22; then B22 passes through the channel attention module to obtain B23, B22 and B23 are multiplied and passed through the spatial attention module to obtain B24, and B22 and B24 are multiplied to obtain B25; then B21 and B22 are multiplied to obtain B26, and B25 and B26 are added to obtain the final output B2 of the downsampling two-input module.
[0066] The structure of the downsampling two-input module is as follows Figure 6 As shown, first, the feature maps of branches 3 and 4 are input to the downsampling two-input module, the output feature map size in branch 3 is downsampled to the feature map size of branch 4, and input into the spatial attention module to obtain B41. Secondly, the output feature map of branch 4 is passed through the void spatial pyramid pooling module and convolution to obtain B42; then B42 passes through the channel attention module to obtain B43, B42 and B43 are multiplied and passed through the spatial attention module to obtain B44, B42 and B44 are multiplied to obtain B45; then B41 and B42 are multiplied to obtain B46, and B45 and B46 are added to obtain the final output B4 of the downsampling two-input module.
[0067] The structure of the three-input module for up and down sampling is as follows Figure 7 As shown, first, the feature maps of branches 2, 3, and 4 are input to the up-and-down sampling three-input module, the output feature map size in branch 2 is downsampled to the feature map size of branch 3, and input into the spatial attention module to obtain B31. Secondly, the output feature map of branch 3 is passed through the void spatial pyramid pooling module and convolution to obtain B32; again, B32 passes through the channel attention module to obtain B33, B32 and B33 are multiplied and passed through the spatial attention module to obtain B34, B32 and B34 are multiplied to obtain B35, and B31 and B32 are multiplied to obtain B36; then the output feature map size in branch 4 is upsampled to the feature map size of the branch 3, input into the spatial attention module to obtain B37, B32 and B37 are multiplied to obtain B38; then B35, B36, and B38 are added to obtain the final output B3 of the up-and-down sampling three-input module.
[0068] Finally, B2, B3, and B4 are restored to the size of branch 1 through step-by-step upsampling. The low-resolution branch transfers information to the high-resolution branch. The four are concatenated and convolved once more as the final output of the model.
[0069] Step 3: Train the model. Input the training set and validation set obtained in step 1 into the model in step 2 for training to obtain the segmentation model Y. The loss function during network training adopts a hybrid loss function. The specific calculation formula is as follows:
[0070] L total =λ Ldice +(1-λ)L ce
[0071] Among them, the Dice loss function is recorded as L dice , the cross entropy loss function is recorded as L ce , λ is a hyperparameter used to balance the two loss functions, and the parameter range is [0,1].
[0072] By adjusting the value of λ, you can control the weight ratio of the Dice loss function and the cross entropy loss function in the hybrid loss function. A larger λ value will place more emphasis on the Dice loss function, focusing on pixel-level accuracy; while a smaller λ value will place more emphasis on the cross entropy loss function, focusing on the consistency of the overall category distribution.
[0073] For the Dice loss function, the specific calculation formula is as follows:
[0074]
[0075] Among them, p i Represents the probability value predicted by the model, t i Represents the corresponding target value, and M represents the number of samples.
[0076] The Dice loss function measures the similarity of segmentation by calculating the overlap between the predicted results and the true labels. It calculates the loss by calculating the ratio of the intersection of the prediction and the label to their union, which is more effective in dealing with class imbalance problems.
[0077] For the cross entropy loss function, the specific calculation formula is as follows:
[0078]
[0079] Among them, p i Represents the probability value predicted by the model, t i Represents the corresponding target value, and M represents the number of samples.
[0080] The cross entropy loss function is a commonly used classification loss function that can be applied to the classification prediction of each pixel in the semantic segmentation task. It calculates the loss by comparing the category distribution of the model with the category distribution of the true label.
[0081] Step 4: Since the present invention performs a pruning operation on HRNet, there will be a certain loss of segmentation accuracy, and there will be errors, uncertainties and blurred classification target boundaries in the segmentation results. At the same time, in the traditional semantic segmentation model, the category probability map of each pixel is obtained through forward propagation. However, these probability maps usually have some undesirable characteristics, such as small breaks, incoherent boundaries, etc. In order to obtain a more accurate final classification result, the present invention adopts a fully connected conditional random field method, which combines the fully connected layer and the conditional random field to improve the output results of the semantic segmentation model, effectively reducing the noise and incoherence in the segmentation results, and providing more accurate and smooth semantic segmentation results.
[0082] The self-built remote sensing image water body dataset is trained through the Water-Lite-HRNet network to obtain a model that can segment water bodies in complex scenes. The model performance is verified through the test set in the dataset, such as Figure 8 As shown in the figure, the water segmentation performance of the present invention using Water-Lite-HRNet, U-Net, and DeepLabv3 models is compared. The present invention achieves a mean Intersection Over Union (MIoU) of 97.96% on a self-built remote sensing image water dataset, and the model segmentation speed reaches 20 frames per second, improving the training speed and segmentation accuracy of the semantic segmentation model.
[0083] The evaluation indicators used in this paper are mIoU, which is the ratio of the intersection and union of objects in the predicted segmentation image and the true segmentation image; the number of frames per second (FPS) is the number of images processed by the model per second.
[0084]
[0085] mIoU=(IoU1+IoU2+…+IoU n / n),
[0086]
[0087] Among them, IoU is the intersection over union ratio, mIoU is the mean intersection over union ratio, FPS is the number of frames, t is the time to segment a single image, there are water bodies and background in the dataset, n is the number of samples, TP (True Positive) indicates the number of pixels correctly predicted by the model as the positive category (that is, the number of pixels whose samples are water are identified as water); FP (False Positive) indicates the number of pixels incorrectly predicted by the model as the positive category (that is, the number of pixels whose samples are background and the model identifies them as water); FN (False Negative) indicates the number of pixels incorrectly predicted by the model as the negative category (that is, the number of pixels whose samples are water and the model identifies them as background).
[0088] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. A water segmentation method for remote sensing images based on attention-based Water-Lite-HRNet, characterized by: The following steps are involved: Step 1: Data preparation: Use LandCove.ai's original high-resolution remote sensing images and preprocess them to obtain a water body segmentation dataset. The dataset includes a training set, a test set, and a validation set. Step 2: Build a network and perform structured pruning on the original HRNet network to obtain a Water-Lite-HRNet network, while embedding the attention module and the adjacent branch attention module. The Water-Lite-HRNet network is composed of four branches stacked together: branch 1, branch 2, branch 3, and branch 4. The attention module is embedded in the four branches, and the four branches only fuse features with adjacent branches of different resolutions. The adjacent branch attention module is introduced before the cross-branch fusion of high-level and low-level features to extract adjacent branch features. S2.1: Perform structured pruning on the HRNet network. The HRNet network is divided into five stages: the starting module, the first stage, the second stage, the third stage, and the fourth stage. The starting module performs preliminary processing and feature extraction on the input image, and the remaining stages are used for further feature extraction, information transmission, and prediction of the final segmentation results. The function of generating low-resolution branches is at the beginning of the first, second, and third stages, the multi-scale fusion function is between the first, second, third, and fourth stages, and the output of the network ends at the fourth stage. S2.2: Starting module; the original resolution of the corresponding network is used as branch 1. Through the convolution operation, the input image is convolved and processed using Basicblock or Bottleblock to extract features. The attention module is used on the last convolution layer and the result is input into the first stage. S2.3: The first stage: downsample branch 1 from the previous stage to generate low-resolution branch 2. Each branch uses Basicblock or Bottleblock for feature extraction, and uses the attention module on the last convolutional layer to input the result into the second stage. S2.4: In the second stage, branch 2 in the previous stage is downsampled to generate a low-resolution branch 3. Branch 1 is only fused with the features of the adjacent branch 2, branch 2 is only fused with the features of the adjacent branch 1, and branch 3 is only fused with the features of the adjacent branch 2. The cross-branch feature fusion function is cut off. Each branch uses Basicblock or Bottleblock for feature extraction, and an attention module is used on the last convolutional layer. The result is input into the third stage. S2.5: The third stage: downsample branch 3 from the previous stage to generate a low-resolution branch 4. In the current stage, branch 1 is only fused with the features of the adjacent branch 2, branch 2 is fused with the features of the adjacent branches 1 and 3, branch 3 is fused with the features of the adjacent branch 2, and branch 4 is only fused with the adjacent branch 3. The cross-branch feature fusion function is cut off. Each branch uses Basicblock or Bottleblock for feature extraction, and an attention module is used on the last convolutional layer. The results are input into the fourth stage. S2.6: The fourth stage: Branch 1 is only fused with the features of the adjacent branch 2, branch 2 is only fused with the features of the adjacent branches 1 and 3, branch 3 is only fused with the features of the adjacent branches 2 and 4, and branch 4 is only fused with the features of the adjacent branch 3. The cross-branch feature fusion function is pruned. feature After fusion, the result is input into the adjacent branch attention module; S2.7: At the end of the fourth stage, i.e., the results of the adjacent branch attention modules, it is necessary to retain the cross-branch feature fusion function. The three parallel low-resolution outputs are upsampled to the size of branch 1, the results of the four branches are concatenated, and the final segmentation result is obtained through 1×1 convolution. Step 3: Train the model. Fine-tune and train the Water-Lite-HRNet network using the public dataset and the training and validation sets of the water segmentation dataset in step 1 to obtain the segmentation model Y. Step 4: Segmentation and post-processing: Input the test set of the water body segmentation dataset in step 1 into the segmentation model Y to obtain the initial water body segmentation result, and then post-process it to obtain the final water body segmentation result.
2. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 1 is characterized in that In step 1, the LandCove.ai original dataset is labeled into four categories, including buildings, woodlands, water, and roads, and the sliding window overlapping cropping method is used to preprocess the original images.
3. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 1, characterized in that The attention module includes a channel attention branch and a spatial attention branch. The specific implementation process is as follows: The channel attention branch transforms the input feature map F of dimension H×W×C into a feature representation F1∈R through two separate 1×1 convolutions H×W×C and Q1∈R H×W×C ; Where F1 is the feature response map, Q1 is the spatial weight map, H represents the height, W represents the width, and C represents the number of channels; the feature representations F1 and Q1 are respectively transformed into feature matrices V1∈R C×N and eigenvector S W1 ∈R N×1 , where N = H × W; in S W1 Then add a fully connected layer S W2 ∈R N×1 , V1 and S W2 Multiply to get C1∈R C×1 , add a fully connected layer C after C1 A ∈R C×1 As the output channel attention vector; use the Sigmoid function to A Normalize and finally C A Multiply it by the input feature map to get F C ∈R H×W×C ; The spatial attention branch transforms the input feature map F into a feature representation F2∈R via two separate 1×1 convolutions H×W×C and Q2∈R H×W×C ; F2 is the feature response map, and then F2 is transformed into the feature matrix V2∈R C×N , where N = H × W, Q2 compresses each channel into a single element through global average pooling, thus obtaining the feature vector C W1 ∈R C×1 ; in C W1 Then add a fully connected layer C W2 ∈R C×1 , use the Softmax function to C W2 Perform the operation and multiply it with V2 to get S1∈R 1×N , add a fully connected layer S2∈R after S1 1×N As the output spatial attention vector, S2 is transformed into S A ∈R H×W×1 ; And use the Sigmoid function to A Normalize and finally convert S A Multiply it by the input feature map to get F S ∈R H×W×C ; Then for F C and F S Using concatenation and 1×1 convolution operations, we get the output feature map of the attention module.
4. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 1, characterized in that The adjacent branch attention module includes an upsampling two-input module, a downsampling two-input module and an up- and downsampling three-input module. The upsampling two-input module inputs the feature maps of branch 2 and branch 3; the downsampling two-input module inputs the feature maps of branch 3 and branch 4; the up- and downsampling three-input module inputs the feature maps of branch 2, branch 3 and branch 4, and the outputs of the upsampling two-input module, the downsampling two-input module and the up- and downsampling three-input module are B2, B3 and B4 respectively; B2, B3 and B4 are restored to the size of branch 1 after step-by-step upsampling, and information is transferred from the low-resolution branch to the high-resolution branch. The four are spliced together and a 1×1 convolution is used as the final output of the model.
5. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 4 is characterized in that The upsampling two-input module upsamples the output feature map size in the branch 3 to the feature map size of the branch 2, and inputs it into the spatial attention module to obtain B21. The output feature map of the branch 2 is subjected to the void spatial pyramid pooling module and the 3×3 convolution to obtain B22; B22 is subjected to the channel attention module to obtain B23, B22 and B23 are multiplied and passed through the spatial attention module to obtain B24, and B22 and B24 are multiplied to obtain B25; B21 and B22 are multiplied to obtain B26, and B25 and B26 are added to obtain the final output B2 of the downsampling two-input module.
6. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 4, characterized in that The downsampling two-input module downsamples the output feature map size in the branch 3 to the feature map size of the branch 4, and inputs it into the spatial attention module to obtain B41. The output feature map of the branch 4 is subjected to the void spatial pyramid pooling module and the 3×3 convolution to obtain B42; B42 is subjected to the channel attention module to obtain B43, B42 and B43 are multiplied and passed through the spatial attention module to obtain B44, and B42 and B44 are multiplied to obtain B45; B41 and B42 are multiplied to obtain B46, and B45 and B46 are added to obtain the final output B4 of the downsampling two-input module.
7. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 4, characterized in that The up-down sampling three-input module downsamples the output feature map size in the branch 2 to the feature map size of the branch 3, and inputs it into the spatial attention module to obtain B31. The output feature map of the branch 3 is passed through the void spatial pyramid pooling module and the 3×3 convolution to obtain B32; B32 is passed through the channel attention module to obtain B33, B32 and B33 are multiplied and passed through the spatial attention module to obtain B34, B32 and B34 are multiplied to obtain B35, and B31 and B32 are multiplied to obtain B36; the output feature map size in the branch 4 is upsampled to the feature map size of the branch 3, and inputted into the spatial attention module to obtain B37, B32 and B37 are multiplied to obtain B38; B35, B36 and B38 are added to obtain the final output B3 of the up-down sampling three-input module.
8. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to claim 1, characterized in that The loss function during network training in step 3 adopts a hybrid loss function, and the specific calculation formula is as follows: THE total =λL dice +(1-λ)L ce Among them, the Dice loss function is recorded as L dice , the cross entropy loss function is recorded as L ce , λ is a hyperparameter used to balance the two loss functions, and the parameter range is [0,1]; For the Dice loss function, the specific calculation formula is as follows: Among them, p i Represents the probability value predicted by the model, t i Represents the corresponding target value, and M represents the number of samples; For the cross entropy loss function, the specific calculation formula is as follows: Among them, p i Represents the probability value predicted by the model, t i Represents the corresponding target value, and M represents the number of samples.
9. The attention-based Water-Lite-HRNet remote sensing image water segmentation method according to any one of claims 1 to 8, characterized in that The step 4 uses a fully connected conditional random field method to post-process the initial water body segmentation results that contain errors, uncertainties, and blurred classification target boundaries, to obtain an accurate and smooth final classification result.
Citation Information
Patent Citations
Garage pedestrian detection method and system adopting branch fusion network lightweight
CN114187606A
Substation safety detection method and device under shielding condition
CN115690659A