Edge enhanced remote sensing image segmentation method and system fusing attention and space state model

By introducing edge enhancement modules that integrate attention and spatial state models in remote sensing image segmentation and using jump connection structures, the difficulties of deep learning networks in edge prediction errors and discontinuities are solved, achieving higher edge accuracy and segmentation accuracy.

CN120088473AActive Publication Date: 2025-06-03UNIV OF JINAN

Patent Information

Application Number
CN202510034460.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-03
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Deep learning networks often face edge prediction errors and edge discontinuities in remote sensing image segmentation, especially when dealing with edge discontinuities and small targets.

Method used

The edge-enhanced remote sensing image segmentation method is adopted to integrate attention and spatial state model. By adding edge-enhanced modules of spatial state model combined with channel attention in the SegNext model, and using a jump connection structure, multi-stage feature maps are fused to improve the recognition ability of edge features.

Benefits of technology

The edge accuracy and network applicability of remote sensing image segmentation are significantly improved, especially when dealing with small targets and blurred boundaries, and segmentation accuracy and generalization capabilities are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088473A_ABST
    Figure CN120088473A_ABST
Patent Text Reader

Abstract

The invention provides an edge-enhanced remote sensing image segmentation method and system fusing attention and a space state model. The invention discloses an edge-enhanced remote sensing image segmentation method and system fusing attention and a space state model. The method comprises the following implementation steps: constructing an edge texture feature enhancement structure; the edge texture enhancement structure is introduced into a SegNext semantic segmentation model; dividing the remote sensing image segmentation data set to generate a training sample set, a verification sample set and a test sample set; preprocessing the data set; the method comprises the following steps: preliminarily extracting fine features of an optical remote sensing image by using a neural network, and enhancing a decoder training model by using channel attention and edge texture of a spatial state model; and finally, sending test sample data into the edge texture enhancement model of the trained attention and space state model to obtain a test result. According to the method, the constructed edge texture feature enhancement module and the SegNext semantic segmentation model are used for cooperative training, the edge texture features are enhanced while the ground feature features are ensured, and the segmentation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of remote sensing image semantic segmentation, and particularly to an edge-enhanced remote sensing image segmentation method and system that integrates attention and spatial state models; the present invention can be used to improve the accuracy of remote sensing image segmentation. Background Art

[0002] Semantic segmentation is to use a neural network to practice the end-to-end segmentation results. It involves assigning each pixel in the image to a specific class label to achieve fine object boundary recognition. In the past decade, the great success of deep learning in semantic segmentation has benefited from the research and improvement of deep learning networks by countless scholars: The Fully convolutional networks (FCN) laid the foundation for deep learning models in the field of semantic segmentation. SegNet proposed a unique encoder-decoder architecture. UNet effectively improved the accuracy of image segmentation by combining the encoder-decoder architecture with skip connections. The Pyramid Scene Parsing Network (PSPNet) and the DeepLab series noticed that the network needs to extract context information. Later, Transformer, which performs excellently in natural language processing, showed amazing performance in the field of computer vision using the self-attention mechanism. At the same time, various network frameworks based on Transformer are also developing rapidly, once defeating the neural network architecture using traditional convolution. SegNeXt proposed a convolutional architecture that is more efficient than the self-attention mechanism of Transformer in encoding context information. Applying remote sensing semantic segmentation technology to disaster warning can identify targets such as buildings, farmland, and disaster areas. By comparing with recent remote sensing image data, people can evaluate potential disasters in the area, such as floods and landslides, which helps in formulating agricultural safety, flood management, and corresponding disaster rescue plans. However, in remote sensing semantic segmentation, since satellite images contain a large number of targets, there are large differences in spatial and edge features between different targets, and there are differences in different scales, directions, etc. between the same targets, and it is difficult to identify the features and edges of small targets at different resolutions. As a result, deep learning networks often face edge prediction errors and edge discontinuities when dealing with discontinuous edges and small targets (such as water bodies, woods, and roads). Therefore, it is necessary to design a structure that can improve the accuracy of object edges in remote sensing image segmentation and enhance the applicability of the network. Summary of the Invention

[0003] An edge-enhanced remote sensing image segmentation method and system that integrates attention and spatial state models, characterized by changing the structure of the original model, using skip connections for splicing, using a convolutional neural network to extract fine feature maps of optical remote sensing images, using a multi-stage encoder-decoder structure to fuse feature maps, and using an edge enhancement module that combines a channel attention-based spatial state model in the decoder for edge enhancement. The specific steps of this method are as follows: Step 1, design an edge enhancement module that combines channel attention and spatial state models: Add a channel attention mechanism to the VSS block of VMamba. The structure of the VSS block is: after normalization, the input is divided into two branches. In the first branch, the input passes through a linear layer and calculates channel attention, and then passes through an activation function. In the second branch, the input is processed through a linear layer, a depthwise separable convolution, and an activation function, and then input into a 2D selective scan (SS2D) module for further feature extraction. Subsequently, the features are normalized using layer normalization, and then element-wise generation is performed using the output of the first branch to merge the two paths. Finally, the features are mixed using the product of a linear layer and channel attention, and this result is combined with a residual connection to form the output of the VSS block. Among them, the edge enhancement module is constructed by adding the channel attention mechanism after the linear layer of the first branch; Step 2, add the constructed edge enhancement module into the SegNext model: The encoder of the SegNext model has outputs in 4 stages. Place the constructed edge enhancement module after the last stage, and add the edge enhancement module simultaneously after the corresponding 4 upsampling stages in the decoder; Step 3, establish a skip connection structure: Add the outputs of each of the 4 stages of the SegNext model encoder and decoder to establish skip connections; Step 4, divide the ground object segmentation dataset to generate a training sample set, a validation sample set, and a test sample set: Select some labeled segmented remote sensing images, select at least 500 remote sensing image segmentation data of any size. Among them, the training set, validation set, and test set are divided according to the ratio of 8:1:1; Step 5, preprocess the remote sensing segmentation dataset: First, perform preprocessing such as random cropping and random flipping on the images in each batch, and crop the images to a size of 512*512; Step 6, use the modified edge enhancement structure SegNext network model to train the dataset: Step 1: Input the preprocessed training samples into the backbone network (MSCAN) of SegNext for feature extraction to generate 4 layers of different features; Step 2: Upsample these 4 feature maps at different levels. During the upsampling process, pass through the constructed edge enhancement module to obtain 4 different edge texture features respectively. Use skip connections to fuse these 4 layers of features to obtain 4 fused feature maps. Finally, send these 4 fused feature maps into the segmentation head for class prediction, and then calculate the cross-entropy loss with the ground truth labels; Step 7: Obtain the segmentation result: Send the test sample data into the trained SegNext edge feature enhancement model to obtain the test result; Step 8: Performance evaluation: Use the predicted result map in Step 7 and the label map of the remote sensing image to be predicted to calculate the class evaluation index and the overall evaluation index to evaluate the network performance; The categories of evaluation indices include: Intersection over Union (IoU): ; In the formula, i represents the positive example; j represents the negative example; IoU represents the Intersection over Union, that is, the ratio of the intersection to the union of the prediction result and the ground truth for each class; pii represents the total number of pixels with the true class i labeled as class i, that is, the true positive (TP); pij represents the total number of pixels with the true class j labeled as class i, that is, the false positive (FP); pji represents the total number of pixels with the true class i labeled as class j, that is, the false negative (FN); Mean Pixel Accuracy (MPA): ; In the formula, k represents the total number of classes, i represents the positive example; j represents the negative example; pii represents the total number of pixels with the true class i labeled as class i, that is, the true positive (TP); pij represents the total number of pixels with the true class j labeled as class i, that is, the false positive (FP). Description of the Drawings

[0004] Figure 1 is the flowchart of each module of the edge enhancement method.

[0005] Figure 2 is the test result graph of the present invention.

[0006] Figure 3 is the structural diagram of the module of the present invention. Detailed Embodiment

[0007] The present invention will be further described in detail below with reference to the drawings.

[0008] Refer to the appendixFigure 1 , a further detailed description of the steps of the present invention will be given.

[0009] 1. An edge-enhanced remote sensing image segmentation method and system integrating attention and spatial state model, characterized in that the structure of the original model is changed and skip connections are used for splicing. The fine feature map of the optical remote sensing image is extracted by using a convolutional neural network, and the feature maps are fused by using a multi-stage encoder-decoder structure. In the decoder, an edge enhancement module combining the spatial state model with channel attention is used for edge enhancement. The specific steps of this method are as follows: Step 1, design an edge enhancement module combining channel attention and spatial state model: Add a channel attention mechanism to the VSS block of VMamba. The structure of the VSS block is: after normalization, the input is divided into two branches. In the first branch, the input passes through a linear layer and calculates the channel attention, and then passes through an activation function. In the second branch, the input is processed through a linear layer, a depthwise separable convolution and an activation function, and then input into the 2D selective scan (SS2D) module for further feature extraction. Subsequently, the features are normalized using layer normalization, and then element-wise generation is performed using the output of the first branch to merge the two paths. Finally, the features are mixed using the product of the linear layer and the channel attention, and this result is combined with the residual connection to form the output of the VSS block. Among them, the edge enhancement module is constructed by adding the channel attention mechanism to the linear layer of the first branch; Step 2, add the constructed edge enhancement module into the SegNext model: The encoder of the SegNext model has outputs in 4 stages. The constructed edge enhancement module is placed after the last stage, and edge enhancement modules are added simultaneously after the corresponding 4 upsampling stages in the decoder; Step 3, establish a skip connection structure: Add the outputs of each of the 4 stages of the SegNext model encoder and decoder to establish skip connections; Step 4, divide the ground object segmentation dataset to generate a training sample set, a validation sample set and a test sample set: Select some labeled segmented remote sensing images, and select at least 500 remote sensing image segmentation data of any size. Among them, the training set, validation set and test set are divided according to the ratio of 8:1:1; Step 5, preprocess the remote sensing segmentation dataset: First, perform preprocessing such as random cropping and random flipping on the images in each batch, and crop the images to a size of 512*512; Step 6, use the modified edge enhancement structure SegNext network model to train the dataset: First step: Input the preprocessed training samples into the backbone network (MSCAN) of SegNext for feature extraction to generate 4 layers of different features. Second step: Upsample these 4 feature maps at different levels. During the upsampling process, through the constructed edge enhancement module, 4 different edge texture features are obtained respectively. Use skip connections to fuse these 4 layers of features to obtain 4 fused feature maps. Finally, send these 4 fused feature maps into the segmentation head for class prediction, and then calculate the cross-entropy loss with the ground truth labels. Step 7: Obtain the segmentation result: Send the test sample data into the trained SegNext edge feature enhancement model to obtain the test result. Step 8: Performance evaluation: Calculate the class evaluation index and the overall evaluation index using the predicted result map in Step 7 and the label map of the remote sensing image to be predicted to evaluate the network performance. The categories of evaluation indicators include: Intersection over Union (IoU): ; In the formula, i represents the positive example; j represents the negative example; IoU represents the Intersection over Union, that is, the ratio of the intersection to the union of the prediction result and the ground truth for each class; pii represents the total number of pixels with the true class i labeled as class i, that is, the true positive example TP; pij represents the total number of pixels with the true class j labeled as class i, that is, the false positive example FP; pji represents the total number of pixels with the true class i labeled as class j, the false negative example FN. Mean Pixel Accuracy (MPA): ; In the formula, k represents the total number of classes, i represents the positive example; j represents the negative example; pii represents the total number of pixels with the true class i labeled as class i, that is, the true positive example TP; pij represents the total number of pixels with the true class j labeled as class i, that is, the false positive example FP.

[0010] Simulation experiment conditions: 1. The simulation experiment conditions of the present invention: Server GPU: GeForce RTX3090 Ti, video memory 24G.

[0011] 2. The simulation experiment platform of this invention patent is: ubuntu 20.04.4 system, python 3.10.13, pytroch-gpu 2.1.1.

[0012] Simulation experiment content and analysis of its experimental results: The simulation experiment of the present invention is for using the present invention and three existing technologies (SegNext segmentation method, CBAM channel attention module, VMamba segmentation method, and spatial state model), designing an edge enhancement module by combining the channel attention module, introducing this module into SegNext to segment remote sensing images, and obtaining structural results. The datasets used in the simulation are: LRBS and WHDLD.

[0013] The existing technologies adopted in the simulation experiment are: The SegNext image segmentation method refers to the image segmentation method proposed by Guo Meng-Hao et al. in "Segnext: Rethinking convolutional attention design for semantic segmentation. Advances in Neural Information Processing Systems 35 (2022): 1140-1156.", abbreviated as the SegNext segmentation method.

[0014] The VMamba image segmentation method refers to the image segmentation method proposed by Zhu, Lianghui et al. in Vision mamba: Efficient visual representation learning with bidirectional state space model., abbreviated as the VMamba segmentation method; the spatial state model refers to the core module VSS block in VMamba and its structure.

[0015] The PSPNet image segmentation method refers to the image segmentation method proposed by Zhao, Hengshuang et al. in "Pyramid scene parsing network." Proceedings of the IEEE conference on computer vision and pattern recognition. 2017., abbreviated as the PSPNet segmentation method.

[0016] The existing CBAM attention mechanism refers to the simple and effective feed-forward convolutional neural network attention mechanism proposed by Woo, Sanghyun et al. in "Cbam: Convolutional block attention module. Proceedings of the European conference on computer vision (ECCV). 2018.", which is abbreviated as the CBAM attention mechanism.

[0017] The input images used in the simulation experiments of the present invention are the publicly available WHDLD remote sensing dataset. The WHDLD dataset is a dense label dataset released by Wuhan University, mainly used for semantic segmentation of remote sensing images. The WHDLD contains 4,940 RGB images of size 256 × 256 taken by the Gaofen-1 satellite and the Changyun-3 satellite over the urban area of Wuhan. Through image fusion and resampling, the image resolution reaches 2m / pixel. The images included in the WHDLD are labeled into 6 categories.

[0018] Simulation Experiment 1 is to introduce the edge feature enhancement structure into the experimental results of SegNext.

[0019] Simulation Experiment 2 is the experimental result under the above simulation conditions of the SegNext-l method in the prior art.

[0020] Simulation Experiment 3 is the experimental result under the above simulation conditions of the PSPNet method in the prior art.

[0021] Table 1. Comparison table of the simulation experiment results of the present invention

[0022] Combined with Table 1, it can be seen that compared with the existing two methods, PSPNet and SegNext, the average intersection over union of the present invention is 65.16, the average mAcc is 76.25, and most importantly, it is 82.53 in the vegetation class. All three indicators are higher than the two existing technical methods. Especially in the vegetation class, it has increased by 1.3 percentage points, which proves that the vegetation segmentation accuracy obtained by the present invention is higher.

[0023] Figure 2 This is the prediction result map of the LRBS remote sensing dataset obtained by the present invention under the above experimental conditions. From the prediction result map and the ground truth map, it can be seen that the boundary region in the prediction result map is close to the boundary region in the ground truth map, and the prediction result accurately shows the boundary.

[0024] The above simulation experiments show that the edge feature enhancement structure combining SegNext, spatial state model and attention mechanism adopted by the present invention can extract more accurate ground object edge details. To a certain extent, this structure solves the problem of unreasonable prediction for small targets and fuzzy boundaries. By improving the accuracy and generalization ability of the network for complex boundary segmentation, the present invention provides a method with high accuracy for the ground object segmentation task of remote sensing images.

Claims

1. A method and system for edge-enhanced remote sensing image segmentation integrating attention and spatial state models, characterized in that: The structure of the original model is changed, and skip connections are used for splicing. A convolutional neural network is used to extract fine feature maps of optical remote sensing images. A multi-stage encoder-decoder structure is used to fuse feature maps. An edge enhancement module combined with a spatial state model of channel attention is used in the decoder for edge enhancement. The specific steps of this method include the following: Step 1: Design an edge enhancement module that combines channel attention and spatial state model: Add a channel attention mechanism in the VSS block of VMamba. The structure of the VSS block is: after normalization, the input is divided into two branches. In the first branch, the input passes through a linear layer and calculates the channel attention, and then passes through an activation function. In the second branch, the input is processed by a linear layer, a depth-wise separable convolution, and an activation function, and then input into the 2D selective scan (SS2D) module for further feature extraction. Subsequently, the features are normalized using layer normalization, and then element-wise generation is performed using the output of the first branch to merge the two paths. Finally, the features are mixed using the product of the linear layer and the channel attention, and this result is combined with a residual connection to form the output of the VSS block. Among them, an edge enhancement module is constructed after the channel attention mechanism is added to the linear layer of the first branch; Step 2: Add the constructed edge enhancement module to the SegNext model: The encoder of the SegNext model has four stages of output, and the constructed edge enhancement module is placed after the last stage, and the edge enhancement module is added after the four upsampling stages corresponding to the decoder; Step 3: Establish a skip connection structure: Add the outputs of the 4 stages of the SegNext model encoder and decoder to establish a skip connection; Step 4: Divide the ground feature segmentation data set to generate training sample set, verification sample set and test sample set: Select some annotated segmented remote sensing images, and select at least 500 remote sensing image segmentation data of no size requirement, where the training set, validation set, and test set are divided in a ratio of 8:1:1; Step 5: Preprocess the remote sensing segmentation dataset: First, the images in each batch are randomly cropped, randomly flipped and other preprocessing, and the images are cropped to a size of 512*512; Step 6: Use the modified edge-enhanced structure SegNext network model to train the dataset: In the first step, the preprocessed training samples are input into the SegNext backbone network (MSCAN) for feature extraction to generate 4 layers of different features; In the second step, the feature maps of these four different levels are upsampled. During the upsampling process, four different edge texture features are obtained through the constructed edge enhancement module. The four layers of features are fused using skip connections to obtain four fused feature maps. Finally, the four layers of fused feature maps are sent to the segmentation head for category prediction, and then the cross entropy loss is calculated with the true label. Step 7, get the segmentation result: Send the test sample data into the trained SegNext edge feature enhancement model to obtain the test results; Step 8, performance evaluation: Use the predicted result map in step 7 and the label map of the remote sensing image to be predicted to calculate the category evaluation index and the overall evaluation index to evaluate the network performance; The categories of evaluation indicators include: Intersection over Union (IoU): ; In the formula, i represents a positive example; j represents a negative example; IoU represents the intersection-over-union ratio, that is, the ratio of the intersection and union of each type of prediction result and the true value; pii represents the total number of pixels whose true category is i and is marked as category i, that is, the true positive example TP; pij represents the total number of pixels whose true category is j and is marked as category i, that is, the false positive example FP; pji represents the total number of pixels whose true category is i and is marked as category j, that is, the false negative example FN; Average pixel accuracy MPA: ; In the formula, k represents the total number of categories, i represents positive examples, j represents negative examples, pii represents the total number of pixels whose true category is i and is marked as category i, i.e. true positive examples TP, and pij represents the total number of pixels whose true category is j and is marked as category i, i.e. false positive examples FP.

2. According to the method of claim 1, a method and system for edge-enhanced remote sensing image segmentation integrating attention and spatial state models, characterized in that: The construction described in steps 1, 2, 3 and 6 is based on the edge texture feature enhancement structure training of SegNext, and the corresponding edge enhancement structure is designed for edge refinement segmentation of large objects in remote sensing images for training, so as to enhance the accuracy of the network for object segmentation and improve the accuracy of segmentation.

Citation Information

Patent Citations

  • Semantic image segmentation method and system based on edge enhancement

    CN111462126A

  • Remote sensing image building segmentation method based on attention mechanism and multi-scale features

    CN113298818A

  • Remote sensing image semantic segmentation method based on multi-scale feature fusion and attention mechanism

    CN117765409A

  • Generating and visualizing planar surfaces in three-dimensional space

    CN118864759A

Cited By

  • Forest stand boundary extraction method and system

    CN120931948A