Remote sensing image semantic segmentation network and segmentation method

By introducing context anchor residual blocks, multi-scale attention aggregation modules and mixed residual paths into the remote sensing image semantic segmentation network, the problem of feature information loss and segmentation area discontinuity in the prior art is solved, and richer feature extraction and better segmentation effects are achieved.

CN119942103AInactive Publication Date: 2025-05-06ANHUI NORMAL UNIV

Patent Information

Application Number
CN202411787777.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing remote sensing image semantic segmentation technology is prone to loss of feature information when processing complex scenarios, fails to make full use of shallow spatial information, and does not pay enough attention to long-distance dependence information in space, resulting in discontinuity of segmented areas.

Method used

A remote sensing image semantic segmentation network based on context anchor residual and multi-scale aggregation is proposed, including context anchor residual block (CARB), multi-scale attention aggregation module (MSAA) and hybrid residual path (MRP), to enhance the network's representation ability, capture a larger range of context information, and improve the ability to identify targets at different scales.

Benefits of technology

Through the proposed network structure, it can effectively capture the detailed spatial features and deep semantic information of the image, enhance the target recognition ability, achieve better segmentation effect, and show better segmentation performance on UMWD and WHDLD datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942103A_ABST
    Figure CN119942103A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of networks, in particular to a remote sensing image semantic segmentation network and method, and the method comprises the steps: constructing a remote sensing image semantic segmentation model based on context anchor point residual errors and multi-scale aggregation, the network structure of the model is mainly composed of a context anchor point residual block CARB, a multi-scale attention aggregation module MSAA and a mixed residual path MRP. In the model training stage, a composite loss function of a cross entropy loss function and a Dice coefficient loss function is used for assisting model training; original input images of different sizes are input into the network, through a series of feature extraction and classification of the network, a segmentation task with clear boundaries is finally completed on a target, a complete segmentation effect picture is output, the network can extract richer feature information, fine spatial features and deep semantic information of the images are effectively captured, and the image segmentation efficiency is improved. And the target identification capability is further enhanced, so that a better segmentation effect is realized in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network technology, and in particular to a remote sensing image semantic segmentation network and a segmentation method. Background Art

[0002] As a kind of image data containing rich semantic information, high resolution and multiple time periods, the semantic segmentation task of remote sensing images is of great significance in mining potential information. In recent years, remote sensing image semantic segmentation has achieved remarkable results in many fields, including water area change monitoring, natural resource investigation and research, and wetland vegetation management. These applications not only improve the understanding of environmental changes, but also provide a scientific basis for decision support.

[0003] High-resolution remote sensing images are significantly different from natural scene images and usually face challenges such as complex background, occlusion between objects, large scale variations between classes, and small differences in boundary features.

[0004] These problems pose great challenges to accurately segmenting target information in complex scenes. Considering the correlation between local and global features, Li et al. proposed a new model for remote sensing image segmentation tasks by combining local and global features of the target, emphasizing the complementary effect between local features such as edges and textures and the overall structure of the image. The model shows good segmentation accuracy and stability. Zhao et al. proposed a semantic segmentation network based on end-to-end attention, integrating the attention mechanism into the multi-scale module to enhance important features, making the model more sensitive to the boundary area of ​​small targets in remote sensing images. Geng et al. proposed a dual-path feature perception network that can effectively process spatial information and contextual information, ensuring that the network can simultaneously process information of different scales and levels. The two paths interact through the feature fusion module to enhance the segmentation effect. In order to improve the model's ability to perceive details, Li Linjuan et al. proposed a segmentation network guided by cross-layer detail perception and group attention, which strengthens important regional information through attention synergy and ensures the integrity of the segmented area of ​​remote sensing images.

[0005] Although the above methods have made some progress in the target segmentation of remote sensing images, there are still some limitations in their application in practical scenarios. First, affected by the inherent properties of the convolution kernel, it is easy to cause feature information loss; and it fails to fully utilize shallow spatial information, and the target segmentation of remote sensing images is not complete. Secondly, the above methods do not pay enough attention to the long-distance spatial dependency information, and the context information obtained is not rich enough, resulting in discontinuous segmentation areas.

[0006] To sum up the above problems, we propose a remote sensing image semantic segmentation network and segmentation method. Summary of the invention

[0007] The purpose of the present invention is to provide a remote sensing image semantic segmentation network and a segmentation method to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A remote sensing image semantic segmentation network and segmentation method, comprising:

[0010] A remote sensing image semantic segmentation model based on contextual anchor residual and multi-scale aggregation is constructed. The network structure of the model is mainly composed of contextual anchor residual block CARB, multi-scale attention aggregation module MSAA and hybrid residual path MRP.

[0011] The context anchor residual block is used to enhance the representation ability of the network and capture a wider range of context information;

[0012] The hybrid residual path is used to alleviate the differences between features and combine shallow spatial features with deep semantic information to improve the segmentation effect of edges;

[0013] The multi-scale attention aggregation module is used to obtain and fuse information from different receptive fields and improve the network's ability to recognize objects of different scales;

[0014] The remote sensing image semantic segmentation model is used to segment the original input image into predefined label types, and the network needs to classify each pixel of the image;

[0015] In the model training stage, a composite loss function of the cross entropy loss function and the Dice coefficient loss function is used to assist in model training;

[0016] The network inputs original input images of different sizes, and after a series of feature extraction and classification by the network, it finally completes the task of segmenting the target with clear boundaries and outputs a complete segmentation effect map.

[0017] Preferably, the context anchor residual module CARB comprises three 3×3 convolutional layers and two 1×1 convolutional layers, and one context anchor attention module CAA;

[0018] In the CAA module, after the feature map is processed by the second 3×3 convolutional layer, the feature information of two different levels is concatenated as the input of the third 3×3 convolutional layer;

[0019] After the third convolutional layer is calculated, 1×1 convolution is used to adjust the number of channels of different features, and then the feature information of multiple levels is fused and sent to the CAA module to enhance the key features for learning context information;

[0020] In CARB, three 3×3 convolutional layers are connected in series. The 3×3 convolution of the second layer combines the features extracted by the 3×3 convolutional block of the first layer and is regarded as the output of a 5×5 convolutional block, while the output of the third convolutional layer is equivalent to the result of a 7×7 convolutional block. CARB also replaces the 5×5 and 7×7 convolutional blocks with a series of smaller 3×3 convolutional blocks to obtain three different receptive field information.

[0021] Preferably, the CAA module is used to capture long-range contextual information around the image target and enhance the central feature expression capability in the feature extraction process;

[0022] When processing different feature maps, CAA uses different convolution kernel sizes;

[0023] When the size of the input feature map is C×H×W, CAA first performs an average pooling operation and then calculates it through a 1×1 convolution. The intermediate features obtained are:

[0024] F avgpool,n =Conv 1×1 (P avgpool (F input,n )),n=0,...,3(1)

[0025] Where P avgpool Represents the global average pooling operation, Conv 1×1 represents a 1×1 convolution operation, F input,n represents the input feature, F avgpool Represents the intermediate features of the output;

[0026] Then two depth-wise separable convolutions DWConv are used to capture features in the horizontal and vertical directions respectively:

[0027] F w,n =DWConv 1×k (F avgpool,n )(2)

[0028] F h,n =DWConv k×1 (F w,n ) (3)

[0029] Among them, F w,n represents the features extracted in the horizontal direction, F h,n Represents the features extracted in the vertical direction;

[0030] Finally, the result is mapped to [0,1] through the Sigmoid activation function to generate the attention weight of the feature map, which is then weighted fused with the original feature map to adjust the importance of different regions and obtain the final output:

[0031] W n =Sigmoid(Conv 1×1 (F h,n )) (4)

[0032]

[0033] W n Mapping weight values, represents element-by-element multiplication, F output,n Indicates the features after CAA enhancement.

[0034] Preferably, the multi-scale attention aggregation module MSAA is used to refine the spatial and channel information of the feature map. In the spatial aggregation module, two convolutional layers are used to aggregate spatial information, and no pooling operation is used to further retain the feature map.

[0035] Preferably, the remote sensing image semantic segmentation network CMNet uses a hybrid residual path module MRP to connect the encoder and the decoder, and performs a convolution operation before splicing the encoder features with the corresponding features of the decoder, rather than directly splicing the feature information.

[0036] Preferably, the hybrid residual path module MRP uses DWConv to extract deeper feature information from the shallow features of the encoder, adjusts the number of channels through 1×1 convolution, and finally performs feature fusion as the input of the next layer.

[0037] Preferably, during the model training stage, a composite loss function of the cross entropy loss function CELoss and the Dice coefficient loss function DiceLoss is used to assist in the training of the model;

[0038] The composite loss function is shown in Formula 6:

[0039] L=L ce (y,p)+D dice (y,p)(6)

[0040] Among them, L represents the composite loss function, y represents the true value of the label, and p represents the predicted value of the network model.

[0041] Preferably, the cross entropy loss function CELoss is a loss function in semantic segmentation, and is used when softmax is used to classify pixels in a CMNet network;

[0042] Among them, L ce (y,p) is calculated as shown in formula 7:

[0043]

[0044] Among them, y iRepresents the true value of the pixel i-th class label, p i It represents the probability that the model predicts that the pixel belongs to the i-th category, and c represents the number of category labels.

[0045] Preferably, the Dice coefficient loss function Dice Loss is a set similarity measurement function, which is usually used to evaluate the similarity between two samples. The function value range is [0,1]. In CMNet, the larger the value, the greater the overlap between the predicted result and the true result.

[0046] Among them, L dice(y,p) The calculation is shown in formula 8:

[0047]

[0048] Here, ε is a very small value to avoid division by 0.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] The present invention proposes a semantic segmentation network model based on a coding-decoding structure, and fuses shallow information with deep information. When processing specific scenes of remote sensing images, the network can extract richer feature information, effectively capture the detailed spatial features and deep semantic information of the image, and further enhance the target recognition capability, thereby achieving better segmentation effects in complex environments.

[0051] The present invention proposes a contextual anchor residual block to enhance the ability to understand long-distance contextual information and extract richer image features; introduces a multi-scale attention aggregation module to fuse multi-scale features, and combines it with a hybrid residual path to improve the utilization of shallow spatial feature information and enhance the boundary segmentation effect; experimental verification is carried out on the self-built dataset UMWD and the public dataset WHDLD. Compared with the comparison network model, CMNet has better segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the overall structure of CMNe of the present invention;

[0053] Figure 2 This is a schematic diagram of the overall structure of the CARB of the present invention;

[0054] Figure 3 It is a schematic diagram of the overall structure of CAA of the present invention;

[0055] Figure 4 It is a schematic diagram of the overall structure of MSAA of the present invention;

[0056] Figure 5 It is a schematic diagram of the overall structure of MRP of the present invention;

[0057] Figure 6 Schematic diagram of visualization results of different models of the present invention on UMWD dataset;

[0058] Figure 7 Schematic diagram of the visualization results of different models of the present invention on the WHDLD dataset. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] Embodiment 1:

[0061] like Figure 1-7 As shown, a remote sensing image semantic segmentation network and segmentation method include:

[0062] A remote sensing image semantic segmentation model based on contextual anchor residual and multi-scale aggregation is constructed. The network structure of the model is mainly composed of contextual anchor residual block CARB, multi-scale attention aggregation module MSAA and hybrid residual path MRP.

[0063] The context anchor residual block is used to enhance the representation ability of the network and capture a wider range of context information;

[0064] The hybrid residual path is used to alleviate the differences between features and combine shallow spatial features with deep semantic information to improve the segmentation effect of edges;

[0065] The multi-scale attention aggregation module is used to obtain and fuse information from different receptive fields and improve the network's ability to recognize objects of different scales;

[0066] The remote sensing image semantic segmentation model is used to segment the original input image into predefined label types, and the network needs to classify each pixel of the image;

[0067] In the model training stage, a composite loss function of the cross entropy loss function and the Dice coefficient loss function is used to assist in model training;

[0068] The network inputs original input images of different sizes, and after a series of feature extraction and classification by the network, it finally completes the task of segmenting the target with clear boundaries and outputs a complete segmentation effect map.

[0069] Embodiment 2:

[0070] Based on Example 1, a context anchor residual module is further disclosed.

[0071] The residual network structure is proposed to solve the problems of low training efficiency and gradient vanishing and exploding in deep CNN networks. In addition, the simple double convolution layer in the original network cannot fully extract feature information, resulting in loss of semantic information during the convolution process.

[0072] Inspired by the residual idea, CARB is used to replace the double convolutional layer. The overall structure of CARB is as follows: Figure 2 As shown in Figure 2. This module contains three 3×3 convolutional layers and two 1×1 convolutional layers, as well as a contextual anchor attention module CAA. The structure of the CAA module is shown in Figure 2. Figure 3 After the feature map is processed by the second 3×3 convolutional layer, the feature information of two different levels is concatenated as the input of the third 3×3 convolutional layer.

[0073] After the third convolution layer is calculated, 1×1 convolution is used to adjust the number of channels of different features, and then the feature information of multiple levels is fused and sent to the CAA module to enhance key features and better learn context information. Compared with the traditional double convolution layer, CARB not only deepens the network depth, but also enhances the feature extraction capability, and achieves a high degree of fusion of deep features and shallow features. In addition, the network introduces a residual structure, which effectively solves the problem of gradient disappearance and explosion.

[0074] In CARB, three 3×3 convolutional layers are connected in series. The 3×3 convolution of the second layer combines the features extracted by the 3×3 convolutional block of the first layer, which can be regarded as the output of a 5×5 convolutional block, while the output of the third convolutional layer is equivalent to the result of a 7×7 convolutional block. CARB uses a series of smaller 3×3 convolutional blocks to replace 5×5 and 7×7 convolutional blocks to obtain three different receptive field information, which not only improves the generalization ability of the network but also reduces memory usage.

[0075] Traditional attention mechanisms usually only focus on local areas of the image, while ignoring more distant contextual information related to the target object, which may lead to insufficient understanding of the target by the network. To solve this problem, the CAA module is introduced to capture the long-distance contextual information around the image target and enhance the central feature expression ability in the feature extraction process. CAA uses different convolution kernel sizes when processing different feature maps. When the size of the input feature map is C×H×W, CAA first performs an average pooling operation and then calculates it through a 1×1 convolution. The intermediate features obtained are:

[0076] F avgpool,n =Conv 1×1 (P avgpool (F input,n)),n=0,...,3(1)

[0077] Where P avgpool Represents the global average pooling operation, Conv 1×1 represents a 1×1 convolution operation, F input,n represents the input feature, F avgpool Indicates the intermediate features of the output. In CARB1 and CARB8, the input feature map size is 512×512, and n=0; in CARB2 and CARB7, the input feature map is n=1. And so on, the network adjusts the convolution kernel size according to different feature map sizes.

[0078] Then two depth-wise separable convolutions DWConv are used to capture features in the horizontal and vertical directions respectively:

[0079] F w,n =DWConv 1×k (F avgpool,n )(2)

[0080] F h,n =DWConv k×1 (F w,n )(3)

[0081] Among them, F w,n represents the features extracted in the horizontal direction, F h,n Represents the features extracted in the vertical direction.

[0082] CAA uses two DWConv1Ds to capture features in different directions, which can enhance the network's ability to recognize long-shaped objects (such as buildings). In order to enable CAA to better process feature information of different scales, the size of the k value is set to 11+2n, and the size of the receptive field is adaptively changed as the feature map changes to enrich the feature information.

[0083] Finally, the result is mapped to [0,1] through the Sigmoid activation function to generate the attention weight of the feature map, which is then weighted fused with the original feature map to adjust the importance of different regions and obtain the final output:

[0084] W n =Sigmoid(Conv 1×1 (F h,n ))(4)

[0085]

[0086] W n Mapping weight values, represents element-by-element multiplication, F output,n Indicates the features after CAA enhancement.

[0087] Embodiment 3:

[0088] Based on Example 1, a multi-scale attention aggregation module is further disclosed.

[0089] In order to further enrich the spatial and channel information of the feature map, CMNet uses MSAA to refine the feature information. In the spatial aggregation module, two convolutional layers are used to aggregate spatial information, and no pooling operation is used to further retain the feature map. Figure 4 The overall structure of the MSAA module is shown. Taking the input feature C1×H×W as an example, MSAA performs feature aggregation through spatial and channel dual-path parallel calculations. In the spatial refinement part, 1×1 convolution is first used to reduce the dimension and reduce the number of channels from C1 to C2, where the size of C2 is C1 / α, and α is set to 4. Subsequently, multiple convolutions with different kernel sizes are used for multi-scale feature fusion, and the output is sent to the next layer of the network. Next, a 7×7 convolution layer is used through spatial attention to reduce the number of channels from the input size C to C / r, and the value of r is set to 4, and then the spatial weight feature map is calculated by dimensionality increase and Sigmoid function. Finally, the output is element-wise multiplied with the feature map for spatial aggregation.

[0090] In the channel aggregation part, MSAA reduces the dimension of the feature map to C1×1×1 through global average pooling. Subsequently, after two 1×1 convolutions and ReLU activation functions, a channel attention weight map is generated and fused with the spatially aggregated feature map to enhance important channel features. Finally, this result is added to the input feature map to enhance the feature representation while retaining the original information. The MSAA module can effectively improve the spatial and channel information of subsequent network layers.

[0091] Embodiment 4:

[0092] Based on Example 1, a hybrid residual path module is further disclosed.

[0093] In order to reduce the information lost in the pooling operation, the traditional U-Net will splice the low-level features with the high-level features through jump connections to provide multi-level spatial information for subsequent image segmentation tasks. However, in the jump connection, the low-level information is obtained by shallow network calculations, and directly combining it with the high-level information may result in feature differences, affecting the segmentation effect.

[0094] CMNet uses MRP to connect the encoder and decoder, and performs some additional convolution operations before concatenating the encoder features with the corresponding features of the decoder, instead of directly concatenating the feature information. Figure 5As shown in the figure. MRP uses DWConv to extract deeper feature information from the shallow features of the encoder, adjusts the number of channels through 1×1 convolution, and finally performs feature fusion as the input of the next layer. By splicing the shallow feature map with the deep feature map through MRP, not only can the difference between feature information be reduced, but also the spatial information in the shallow layer can be fused with the deep semantic information, thus improving the network's ability to recognize boundary areas.

[0095] Embodiment 5:

[0096] Based on Example 1, a loss function is further disclosed.

[0097] The task of the semantic segmentation model is to segment the original input image into predefined label types. The network needs to classify each pixel of the image. Therefore, in the model training stage, the composite loss function of CELoss and Dice Loss is used to assist the training of the model. The composite loss function is shown in Formula 6.

[0098] L=L ce (y,p)+D dice (y,p)(6)

[0099] Among them, L represents the composite loss function, y represents the true value of the label, and p represents the predicted value of the network model.

[0100] CELoss is a loss function commonly used in semantic segmentation and is used when softmax is used to classify pixels in the CMNet network. ce The calculation of (y,p) is shown in Formula 7.

[0101]

[0102] Among them, y i Represents the true value of the pixel i-th class label, p i It represents the probability that the model predicts that the pixel belongs to the i-th category, and c represents the number of category labels.

[0103] Dice Loss is a set similarity measurement function, usually used to evaluate the similarity between two samples. The function value range is [0,1]. In CMNet, the larger the value, the greater the overlap between the predicted result and the true result. dice(y,p) See formula 8 for calculation.

[0104]

[0105] Here, ε is a very small value to avoid division by 0.

[0106] Embodiment 6:

[0107] Based on Example 1, a data set is further disclosed.

[0108] In order to evaluate the segmentation performance of the model, the experiment used the self-built urban micro-wetland dataset UMWD and the public dataset WHDLD for model testing. The selected datasets have the characteristics of diverse target types, complex backgrounds between classes, and large scale changes to meet the actual situation of real scenes.

[0109] UMWD self-built dataset:

[0110] Taking a small urban wetland as the research area, a professional satellite map software was used to download high-resolution remote sensing images of the research area, and the spatial resolution was set to 0.60 meters and the scale was set to 1:2260. Then the labeling tool labelme was used to annotate the data, and the data category labels were set to the main targets of the five urban wetlands: vegetation, water bodies, roads, buildings, and background. Since the resolution of the original remote sensing image was too large, in order to facilitate the network model to read the image for training and learning, a sliding window was used for cutting, the window size was set to 512, and the horizontal and vertical step sizes were both set to 256, and the output was a 512×512 image. Table 1 shows the number and proportion of pixels occupied by each category in the dataset. Through data enhancement technology, the data was rotated, blurred, cropped, and noise was added to expand the dataset to improve the robustness and generalization ability of the model, and finally 2135 images were generated. The dataset was divided into training set, validation set, and test set according to the ratio of 8:1:1.

[0111] Table 1 UMWD dataset category statistics

[0112]

[0113] WHDLD public dataset:

[0114] The WHDLD dataset was released by Wuhan University and is derived from 2-meter satellite images from the GaoFen-1 and ZiYuan-3 satellites. Table 2 shows its data distribution. This public dataset is one of the commonly used datasets in the field of semantic segmentation and contains 4940 images of size 256×256. These images are divided into 6 category labels: buildings, roads, sidewalks, vegetation, bare soil, and water bodies. During the training and testing of the network, the data source is divided into training set, validation set, and test set in a ratio of 8:1:1.

[0115] Table 2. Category statistics of WHDLD dataset

[0116]

[0117] Embodiment 7:

[0118] Based on Example 1, experimental results and analysis are further disclosed.

[0119] Evaluation Metrics:

[0120] In the experiment, four indicators, namely recall, intersection over union (IoU), precision and F1, are used to evaluate the segmentation performance of the model. Recall indicates the ratio of the number of pixels predicted correctly by the network to the total number of pixels in the category; IoU measures the ratio of the intersection and union of the predicted target area to the true target area, which is used to evaluate the correlation between the predicted value and the true value; Pre indicates the proportion of samples predicted by the model as positive that are truly positive; F1 is a comprehensive performance indicator that evaluates the performance of the semantic segmentation model by the harmonic mean of precision and recall. The value range of these four indicators is [0,1], and the larger the value, the better the segmentation performance of the model, and the lower the value, the poorer the segmentation performance. The following are the formula definitions of Recall, IoU, Pre and F1:

[0121]

[0122] Among them, GT represents the true label value; TP represents the number of targets correctly predicted by the model as the actual category; FP represents the number of pixels that the model incorrectly predicts as other categories, that is, not the true category of the pixel; FN represents the number of pixels that the model fails to predict as the actual category.

[0123] Experimental environment:

[0124] The experiment uses Python 3.8 to implement the algorithm under the Ubuntu 20.04.2LTS operating system, uses the deep learning framework Pytorch-1.11.0, Cuda11.3 to build the semantic segmentation network model, and uses the NVIDIA GeForceRTX4070 GPU with 32GB of memory as the image processor. Some network parameter settings are shown in Table 3.

[0125] Table 3 Experimental parameter settings

[0126]

[0127] Ablation experiment:

[0128] In order to verify the effectiveness and necessity of each module in the CMNet network model, an ablation experiment was conducted on the UMWD dataset. In the experiment, the output result of the U-Net model with VGG as the backbone network was selected as the Baseline, and its effectiveness was verified by adding the combination of the proposed modules respectively. From the data in Table 4, it can be seen that after adding the CARB module, the overall performance has been effectively improved compared with the network using the original double convolution. The network that introduces the CARB module alone has an mloU and mF1 that are 2.48% and 1.36% higher than the Baseline, respectively. The results show that the CARB module can extract richer feature information, emphasize key information through CAA, and suppress irrelevant content. After introducing the MSAA module alone, the mIoU and mFi of the segmentation network reached 93.76% and 96.73%, respectively, which are 1.49% and 0.82% higher than the original network, improving the network's ability to recognize multi-scale objects. The network with MRP also improved in performance. The mloU and mF1 of the segmentation network increased by 1.03% and 0.57% compared with the baseline, making better use of the shallow contour and edge space information. After introducing the three modules at the same time, the network achieved the best performance, with mloU and mF1 reaching 94.94% and 97.38% respectively, which were 2.67% and 1.47% higher than the baseline.

[0129] In summary, the three modules proposed in this paper have a good improvement on the network segmentation accuracy. The synergy between CARB, MSAA and MRP enables CMNet to significantly improve the performance of semantic segmentation of remote sensing images.

[0130] Table 4. Comparison of ablation experiment evaluation

[0131]

[0132] Note: The bold data are the optimal values, and √ means adding the module.

[0133] Model performance comparison:

[0134] This paper proposes a semantic segmentation network for high-resolution remote sensing images, aiming to effectively segment various types of targets in the image. Under the same experimental conditions, CMNet is compared with six representative deep learning semantic segmentation network models: U-Net, PSPNet, DeepI,abV3+1], SegFormer, Mask2Former, and CMTFNet. The backbones used by the six comparison models in the experiment are shown in Table 5.

[0135] Table 5 Backbone of different methods

[0136]

[0137] Among them, U-Net is a U-shaped network, which is divided into two parts: encoder and decoder. The high-resolution features in the encoder are transferred to the decoder through jump connections. PSPNet captures global context information at different scales to better understand complex image content. DeepLabV3+ combines dilated convolution with spatial pyramid pooling to build an encoder-decoder architecture, which performs well in multiple semantic segmentation tasks. SegFormer uses a layered Transformer encoder to capture the details and global information of the image, and combines a lightweight MLP decoder to fuse feature maps from different convolutional layers, thereby improving the segmentation accuracy of the model.

[0138] Mask2Former is an advanced visual segmentation model. Its core idea is to use the Transformer architecture to enhance the ability to understand contextual information in images, so that the model can capture long-distance dependencies in images. CMTFNet proposes a new segmentation network by combining the advantages of CNN and Transformer, further processes the extracted feature information, and achieves good performance in remote sensing image segmentation tasks. The experiment tests these five semantic segmentation models on the UMWD self-built dataset and the WHDLD public dataset, and performs statistical analysis on the resulting data.

[0139] Experimental results of UMWD dataset:

[0140] Through experiments, different models are quantitatively tested, and Table 6 lists the experimental results of 7 models on the UMWD dataset. According to the data in Table 6, the comprehensive segmentation performance of the segmentation model CMNet proposed in the present invention is better than that of other comparison models, and its network's mRecall, mloU, mFi and mPre reach 97.55%, 94.94%, 97.38% and 98.25% respectively. Among them, in terms of mIoU, it is 2.67%, 16.74%, 15.65%, 17.24%, 17.03% and 1.77% higher than UNet, PSPNet, DeepLabV3+, SegFormer, Mask2Former and CMTFNet respectively. Analysis of the experimental results of each category shows that the segmentation performance of the CMNet segmentation model is weaker than that of DeepLabV3+ in vegetation targets, but the network is better than the other 6 models in the recognition of other categories, especially in the recognition of road and building targets, and the segmentation effect is significantly better than other models.

[0141] Specifically, in terms of road segmentation indicators, CMNet's experimental indicators Recall, IoU, F1 and Pre are improved by 0.30%, 2.67%, 1.47% and 2.57% respectively compared with the suboptimal UMTFNet model. In terms of building target segmentation, Recall, IoU, F1 and Pre are improved by 1.31%, 1.42%, 0.74% and 0.18% respectively compared with the suboptimal model. The experimental results show that when identifying long and narrow objects, CARB improves the learning ability of long-distance contextual information, enabling the network to effectively identify long-shaped objects. The experimental results show that CMNet has achieved better segmentation results than other semantic segmentation network models in the semantic segmentation task of UMWD remote sensing images on the self-built dataset, showing obvious advantages.

[0142] Table 6 Test results of different models on UMWD dataset (%)

[0143]

[0144] Note: The bold data are the optimal values.

[0145] Experimental results of WHDLD dataset:

[0146] Through experimental tests on the public dataset WHDLD, the experimental results of 7 models are listed in Table 7. The analysis results show that the overall performance indicators of the CMNet model on the WHDLD dataset are mRecall of 76.83%, mIoU of 65.96%, mFl of 77.99%, and mPre of 79.72%.

[0147] Among them, in terms of mIoU index, CMNet is 2.61%, 7.68%, 4.61%, 2.92%, 6.81%, and 0.68% higher than UNet, PSPNet, DeepI abV3+, SegFormer, Mask2Former, and CMTFNet, respectively. In addition, in categories such as buildings, roads, pavements, and bare soil, the segmentation performance of the CMNet model is better than that of the other six segmentation models, showing a clear advantage. The experimental results show that CMNet's overall performance is better than other models in the multi-classification remote sensing image segmentation task of the WHDLD dataset, verifying its effectiveness and applicability.

[0148] Table 7 Test results of different models on the WHDLD dataset (%)

[0149]

[0150] Note: The bold data are the optimal values.

[0151] Visual display:

[0152] Figure 6 The segmentation results of different segmentation models on the UMWD dataset are shown, which helps to evaluate the performance of remote sensing image semantic segmentation models. Figure 7 It can be seen from the prediction results that the segmentation effect of the CMNet network model proposed in this invention on different test images is significantly better than that of the other five models, showing a more complete segmentation area and a more continuous and smooth segmentation boundary, and the prediction results are closer to the label image. Figure 7 In the test examples in the first row, within the white dotted box area, U_Net mistakenly identifies part of the background area as vegetation, PSPNet, DeepLabV3+ and CMTFNet have relatively rough boundaries for the identified buildings, and SegFormer and Mask2Former mistakenly identify part of the background content as vegetation when identifying vegetation. In contrast, CMNet performs well in identifying the marked area, can accurately identify different types, and the inter-class segmentation boundary is clearer. In the marked area in the second row, the boundaries between the road, vegetation and water bodies are relatively close, and the road targets are continuous. The contrast method mistakenly identifies the road as vegetation or water bodies, so that the segmentation results show the phenomenon of adhesion between the road, vegetation and water bodies, resulting in incomplete segmentation areas. CMNet can correctly identify the continuous road area, clearly segment the boundary, and there is almost no adhesion. In the marked area in the fourth row, the water body consists of a smaller part and a larger part, separated by a smaller background. Compared with the other five models, all of them failed to correctly identify the two targets, mistaking them for connected areas, and misclassifying the background as water. CMNet can accurately identify the water and background, and successfully segment the two target areas. In the test example in the 5th row, densely distributed vegetation appeared in the marked area, the vegetation areas were similar in shape, and there were continuous roads nearby, making the segmentation task difficult. PSPNet had the worst recognition effect, basically failing to identify the target. U-Net and DeepI abV3+ mistakenly identified the road as the background, and the vegetation recognition results were also relatively few; SegFormer, Mask2Former and CMTFNet had incomplete road recognition, unclear boundaries, and mistakenly classified vegetation as the background. In contrast, CMNet has the best segmentation results, which are closest to the label image, can accurately identify dense targets, and better complete the classification task of closely arranged objects. The visualization results show that the semantic segmentation network model proposed in the present invention can better focus on the segmented area, enhance the recognition ability of the target, and provide clearer inter-class boundary segmentation, thereby improving the overall segmentation performance.

[0153] Figure 7The visualization results of different models on the WHDLD dataset are shown. From the analysis of the experimental results, it can be seen that in the first row, the U-Net, PSPNet and DeepLabV3+ models have poor recognition capabilities for roads near water bodies and fail to fully recognize road targets. SegFormer and CMTFNet misclassify some roads as bare soil, while Mask2Former cannot fully recognize water areas, resulting in poor inter-class boundary segmentation. In contrast, CMNet can correctly distinguish various types of targets and clearly segment continuous inter-class boundaries. In the second row of data, there are dense buildings and road areas, which poses a considerable challenge to the segmentation task. PSPNet has the worst recognition effect and almost fails to distinguish targets. U-Net and DeepIabV3+ miss many small targets, SegFormer misclassifies a large number of road surfaces as vegetation, Mask2Former also misclassifies many vegetation as road surfaces, and CMTFNet classifies roads as road surfaces. CMNet performs well in recognizing densely arranged areas, can successfully complete the segmentation task of small targets, and ensure the integrity of the segmentation results. The visualization experimental results once again prove that the CMNet network can effectively handle targets with small regional differences, completely complete the segmentation task, and the fragmented areas are small, showing high segmentation performance.

[0154] MRP effectiveness analysis:

[0155] The connection path between the encoder and decoder in the network plays an important role in image segmentation and can make more full use of shallow information. Using MRP instead of the original direct connection to fuse shallow information with deep features can effectively improve the utilization of shallow spatial information. In order to test the effectiveness of selecting two MRPs in the CMNet proposed in this invention, a network containing different numbers of MRP paths was constructed experimentally, and the connection paths were numbered from the outer layer to the inner layer in sequence, and the segmentation performance was tested using a variety of combinations. The experimental results are shown in Table 8.

[0156] Table 8 Experimental comparison of different residual paths

[0157]

[0158] Note: The bold data are the optimal values.

[0159] From the experimental results in Table 8, we can see that there are one MRP and two MRPs in Network 2 and Network 3 respectively. After adding MRP, the segmentation performance of the network is improved, and the mIoU is increased by 0.32% and 0.42% respectively compared with the network without MRP. However, in Network 4 and Network 5, when there are three and four MRPs respectively, the increase in the number of MRPs leads to a decrease in the segmentation performance of the network, and the mloU is reduced by 0.03% and 0.02% compared with Network 3. The experimental results verify the effectiveness of the CMNet network, indicating that in the connection of 64 channels and 128 channels, that is, shallow features, the use of two MRPs can best improve the segmentation effect of the network.

[0160] in conclusion:

[0161] In view of the challenges of target segmentation in remote sensing images under complex backgrounds, this paper proposes a segmentation network model with significant effects, namely CMNet, which aims to improve the segmentation accuracy of multi-scale remote sensing image targets under complex backgrounds. The network uses CARB to enhance feature extraction capabilities, and introduces CAA attention in each CARB to enhance the ability to understand long-distance contextual information. In addition, MSAA is used to achieve multi-scale feature fusion, aggregate semantic information in space and channels, and improve the ability to identify multi-scale targets. Finally, MRP is used to splice image features to better utilize shallow spatial information and compensate for information loss in the pooling process. Experimental results show that the segmentation performance of CMNet on the UMWD self-built dataset is better than other segmentation models overall. And it has also achieved good results on the WHDLD public dataset, proving its effectiveness.

[0162] The network requires certain computing resources during inference. In future work, we should focus on optimizing the network structure, reducing the number of model parameters and computing energy consumption, so that it can be deployed on mobile devices with weaker computing power for inference, in order to build a remote sensing image semantic segmentation network model with better performance.

[0163] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A remote sensing image semantic segmentation network and segmentation method, characterized in that: include: A remote sensing image semantic segmentation model based on contextual anchor residual and multi-scale aggregation is constructed. The network structure of the model is mainly composed of contextual anchor residual block CARB, multi-scale attention aggregation module MSAA and hybrid residual path MRP. The context anchor residual block is used to enhance the representation ability of the network and capture a wider range of context information; The hybrid residual path is used to alleviate the differences between features and combine shallow spatial features with deep semantic information to improve the segmentation effect of edges; The multi-scale attention aggregation module is used to obtain and fuse information from different receptive fields and improve the network's ability to recognize objects of different scales; The remote sensing image semantic segmentation model is used to segment the original input image into predefined label types, and the network needs to classify each pixel of the image; In the model training stage, a composite loss function of the cross entropy loss function and the Dice coefficient loss function is used to assist in model training; The network inputs original input images of different sizes, and after a series of feature extraction and classification by the network, it finally completes the task of segmenting the target with clear boundaries and outputs a complete segmentation effect map.

2. A remote sensing image semantic segmentation network and segmentation method according to claim 1, characterized in that: The context anchor residual module CARB includes three 3×3 convolutional layers and two 1×1 convolutional layers, and a context anchor attention module CAA; In the CAA module, after the feature map is processed by the second 3×3 convolutional layer, the feature information of two different levels is concatenated as the input of the third 3×3 convolutional layer; After the third convolutional layer is calculated, 1×1 convolution is used to adjust the number of channels of different features, and then the feature information of multiple levels is fused and sent to the CAA module to enhance the key features for learning context information; In CARB, three 3×3 convolutional layers are connected in series. The 3×3 convolution of the second layer combines the features extracted by the 3×3 convolutional block of the first layer and is regarded as the output of a 5×5 convolutional block, while the output of the third convolutional layer is equivalent to the result of a 7×7 convolutional block. CARB also replaces the 5×5 and 7×7 convolutional blocks with a series of smaller 3×3 convolutional blocks to obtain three different receptive field information.

3. A remote sensing image semantic segmentation network and segmentation method according to claim 2, characterized in that: The CAA module is used to capture long-range context information around image targets and enhance the central feature expression capability in the feature extraction process; When processing different feature maps, CAA uses different convolution kernel sizes; When the size of the input feature map is C×H×W, CAA first performs an average pooling operation and then calculates it through a 1×1 convolution. The intermediate features obtained are: F avgpool,n =Conv 1×1 (P avgpool (F input,n )),n=0,...,3(1) Where P avgpool Represents the global average pooling operation, Conv 1×1 represents a 1×1 convolution operation, F input,n represents the input feature, F avgpool Represents the intermediate features of the output; Then two depth-wise separable convolutions DWConv are used to capture features in the horizontal and vertical directions respectively: F w,n =DWConv 1×k (F avgpool,n )(2) F h,n =DWConv k×1 (F w,n )(3) Among them, F w,n represents the features extracted in the horizontal direction, F h,n Represents the features extracted in the vertical direction; Finally, the result is mapped to [0,1] through the Sigmoid activation function to generate the attention weight of the feature map, which is then weighted fused with the original feature map to adjust the importance of different regions and obtain the final output: W n =Sigmoid(Conv 1×1 (F h,n ))(4) W n Mapping weight values, represents element-by-element multiplication, F output,n Indicates the features after CAA enhancement.

4. The remote sensing image semantic segmentation network and segmentation method according to claim 1, characterized in that: The multi-scale attention aggregation module MSAA is used to refine the spatial and channel information of the feature map. In the spatial aggregation module, two convolutional layers are used to aggregate spatial information, and no pooling operation is used to further retain the feature map.

5. The remote sensing image semantic segmentation network and segmentation method according to claim 1, characterized in that: The remote sensing image semantic segmentation network CMNet adopts a hybrid residual path module MRP to connect the encoder and the decoder, and performs a convolution operation before the encoder features are spliced ​​with the corresponding features of the decoder, instead of directly splicing the feature information.

6. A remote sensing image semantic segmentation network and segmentation method according to claim 1, characterized in that: The hybrid residual path module MRP uses DWConv to extract deeper feature information from the shallow features of the encoder, adjusts the number of channels through 1×1 convolution, and finally performs feature fusion as the input of the next layer.

7. The remote sensing image semantic segmentation network and segmentation method according to claim 1, characterized in that: In the model training stage, a composite loss function of the cross entropy loss function CELoss and the Dice coefficient loss function Dice Loss is used to assist in model training; The composite loss function is shown in Formula 6: L=L ce (y,p)+D dice (y,p)(6) Among them, L represents the composite loss function, y represents the true value of the label, and p represents the predicted value of the network model.

8. A remote sensing image semantic segmentation network and segmentation method according to claim 7, characterized in that: The cross entropy loss function CELoss is a loss function in semantic segmentation and is used when softmax is used to classify pixels in the CMNet network. Among them, L ce (y,p) is calculated as shown in formula 7: Among them, y i Represents the true value of the pixel i-th class label, p i It represents the probability that the model predicts that the pixel belongs to the i-th category, and c represents the number of category labels.

9. A remote sensing image semantic segmentation network and segmentation method according to claim 7, characterized in that: The Dice coefficient loss function Dice Loss is a set similarity measurement function, which is usually used to evaluate the similarity of two samples. The function value range is [0,1]. In CMNet, the larger the value, the greater the overlap between the predicted result and the true result. Among them, L dice(y,p) The calculation is shown in formula 8: Here, ε is a very small value to avoid division by 0.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on multi-scale feature fusion and attention mechanism

    CN117765409A

  • Remote sensing image semantic segmentation method based on semantic adaptive edge enhancement network

    CN118781596A

  • Landslide recognition method based on laplacian pyramid remote sensing image fusion

    US11521377B1

  • Boundary-optimized remote sensing image semantic segmentation method and apparatus, and device and medium

    WO2023077816A1

Cited By

  • Image semantic segmentation method and device based on multi-scale feature fusion and medium

    CN120182612A

  • Multi-modal ship target individual identification method and system

    CN120472250A

  • A multimodal ship target individual recognition method and system

    CN120472250B

  • Remote sensing image segmentation method based on integrated cross attention mechanism

    CN122336757A

  • A cloud segmentation method based on a two-stage attention residual fusion network

    CN122618249A