City remote sensing image semantic segmentation method based on global and local depth information fusion
Through the method of fusion of global and local deep information, the problems of large target scale differences, weak long-distance pixel correlation and vegetation shadow masking in urban scene remote sensing images are solved, and higher segmentation accuracy and continuity are achieved.
Patent Information
- Application Number
- CN202510647230.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-29
AI Technical Summary
In urban scene remote sensing images, the target scales vary greatly, the long-distance pixel correlation within the target is weak, the land objects are dense, and it is easily obscured by vegetation and shadows, resulting in a decrease in segmentation accuracy and discontinuity of results.
A semantic segmentation method for urban remote sensing images based on the fusion of global and local deep information is constructed, including multi-scale semantic enhancement module, global-local information fusion module and local high-frequency feature reconstruction module under boundary guidance. Through technical means such as hollow convolution, cross-scale interactive learning, and boundary information introduction, the accuracy of feature extraction and segmentation is enhanced.
It improves the segmentation accuracy of remote sensing images in urban scenes, reduces the loss of pixel details in blind spots, improves the occlusion problem between targets, and improves the continuity and accuracy of segmentation results.
Smart Images

Figure CN120564191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban remote sensing image semantic segmentation, and in particular to a method for urban remote sensing image semantic segmentation based on the fusion of global and local depth information. Background Art
[0002] In recent years, with the continued advancement of digital city construction, remote sensing imagery of urban scenes, characterized by their rich textures and dense features, has become an indispensable technical support for a variety of fields, including urban planning, geographic information system construction, smart city development, ecological and environmental monitoring, and population estimation. Accurately segmenting multiple types of objects in urban scenes has become a key technical challenge in this field. With the continuous innovation of remote sensing image processing technology, deep learning-based segmentation methods, with their powerful feature extraction capabilities, have become a research hotspot in remote sensing image processing. However, remote sensing images of urban scenes have complex backgrounds, and long-distance pixel correlation within objects is weak. Furthermore, the scales of various objects in urban scenes vary significantly. Furthermore, urban scenes are densely populated with features, and objects such as buildings are easily obscured by vegetation and shadows. Summary of the Invention
[0003] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a semantic segmentation method for urban remote sensing images based on the fusion of global and local depth information.
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0005] A semantic segmentation method for urban remote sensing images based on the fusion of global and local depth information includes the following steps:
[0006] Step 1: Construct a multi-scale semantic enhancement module based on context aggregation. Based on global context modeling, this module enhances the semantic association between multi-scale features through a cross-scale interactive learning mechanism, thereby improving the network's multi-scale feature extraction capability.
[0007] Step 2: Construct a feature decoding module based on multi-level global-local information fusion; introduce global feature extraction branches and local feature extraction branches to enhance the ability to extract complex spatial information in urban scenes;
[0008] Step three: construct a local high-frequency feature reconstruction module under boundary guidance; introduce boundary information into the process of feature fusion of the last layer decoder and encoder, further enhance the edge information through the local feature reconstruction module, and improve the network's ability to extract local information.
[0009] It should be noted that the multi-scale semantic enhancement module based on context aggregation includes void convolution technology and cross-scale interactive learning technology.
[0010] It should be noted that the cross-scale interactive learning technology achieves cross-scale semantic interaction by fusing features from adjacent branches through feature addition or concatenation. This design not only enhances the network's ability to extract multi-scale contextual features, but also effectively avoids the loss of feature information.
[0011] It should be noted that the global feature extraction branch uses a two-dimensional selective scanning mechanism to capture dependencies in the sequence.
[0012] It should be noted that the local feature extraction branch also pays attention to local detail information, thereby enhancing the ability to extract complex spatial information in urban scenes.
[0013] It should be noted that the boundary guidance module introduces explicit boundary information into the multi-level feature extraction process of the encoder.
[0014] It should be noted that the local feature reconstruction module realizes the extraction of boundary perception weights through a specially designed boundary autocorrelation operation.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] 1. To address the problem of large variations in object scales in remote sensing images, which can lead to reduced overall segmentation accuracy, we propose a multi-scale semantic enhancement module based on context aggregation. This module consists of a dilated convolution branch and a cross-scale interaction mechanism. The multi-scale semantic enhancement branch consists of dilated convolutions with different dilation rates, providing receptive fields of varying sizes. The cross-scale interaction mechanism fully exploits the semantic correlation between multi-scale features, addressing the problem of traditional multi-scale feature extraction methods that ignores semantic connections between branches. This module mitigates the loss of pixel details in blind areas and improves the network's segmentation accuracy for multi-scale objects.
[0017] 2. To address the problem of holes and discontinuities in segmentation results caused by complex scenes and weak correlation between distant pixels within objects in remote sensing images, a feature decoding module based on multi-level global-local information fusion is proposed. First, a global feature extraction branch is introduced, which uses a two-dimensional selective scanning mechanism to capture dependencies within the sequence. Subsequently, a local feature extraction branch is introduced, which also focuses on local details, enhancing the ability to extract complex spatial information in urban scenes. This module effectively resolves the problem of holes and discontinuities within objects during segmentation.
[0018] 3. To address the problem of objects in remote sensing images being easily obscured by vegetation and shadows, which can lead to decreased segmentation accuracy in boundary regions, we propose a boundary-guided local high-frequency feature reconstruction module. This algorithm first generates a boundary information map through the boundary guidance module and incorporates this boundary information map into the final decoder and encoder feature fusion process. The local feature reconstruction module further enhances edge information, improving the network's ability to extract local information. This module effectively mitigates the problem of inter-object occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the overall structure of the present invention;
[0020] Figure 2 Schematic diagram of the structure of the multi-scale semantic enhancement module based on context aggregation of the present invention;
[0021] Figure 3 Schematic diagram of the structure of the feature decoding module based on multi-level global-local information fusion of the present invention;
[0022] Figure 4 This is a schematic diagram of the global feature extraction branch structure of the present invention;
[0023] Figure 5 It is a local feature extraction branch of the present invention;
[0024] Figure 6 Schematic diagram of the structure of the local high-frequency feature reconstruction module under boundary guidance of the present invention;
[0025] Figure 7 Schematic diagram of the structure of the efficient local attention (ELA) method of the present invention. DETAILED DESCRIPTION
[0026] The present invention will be further described below in conjunction with the accompanying drawings. It should be noted that the following embodiments are based on the present technical solution and provide detailed implementation methods and specific operating processes, but the protection scope of the present invention is not limited to these embodiments.
[0027] like Figures 1 to 7 As shown, the present invention is a method for semantic segmentation of urban remote sensing images by fusing global and local depth information, comprising the following steps:
[0028] Step 1: A multi-scale semantic enhancement module based on context aggregation is constructed to address the problem of decreased segmentation accuracy caused by large object scale differences in remote sensing images of urban scenes. This module consists of two components: a dilated convolution branch and a cross-scale interaction mechanism. The dilated convolution branch first performs dilated convolution operations with different dilation rates, providing receptive fields of varying sizes. Subsequently, the cross-scale interaction mechanism strengthens the semantic correlation between multi-scale features, compensating for the shortcomings of traditional multi-scale feature extraction methods that ignore semantic connections between branches while also reducing the loss of pixel details in blind areas.
[0029] Specifically, the design of the multi-scale semantic enhancement module based on context aggregation mainly includes two key steps: initial feature aggregation and cross-scale interactive learning. First, the multi-scale receptive field information is obtained through four parallel branches. First, the input features are processed through a 1×1 convolution layer to generate the initial aggregated features f m The role of 1×1 convolution is to reduce the dimension of feature channels while retaining important semantic information, providing a basis for subsequent multi-scale feature interaction. Next, the initial aggregated feature f m Divide it into four sub-feature maps along the channel dimension, recorded as Through this division method, feature information can be distributed to different scale spaces, providing diverse feature representations for subsequent cross-scale interactions. In the cross-scale interactive learning stage, the module interactively fuses the features of adjacent branches through a series of dilated convolutions of different scales. The advantage of dilated convolution is that it can expand the receptive field while maintaining the resolution of the feature map, thereby capturing a wider range of contextual information. Specifically, dilated convolutions with different expansion rates are applied to each sub-feature map. For example, Use a smaller dilation rate to capture local details, and A large dilation rate is used to capture global contextual information. Subsequently, features from adjacent branches are fused through feature addition or concatenation, enabling cross-scale semantic interaction. This design not only enhances the network's ability to extract multi-scale contextual features but also effectively avoids feature loss.
[0030] Step 2: Construct a feature decoding module based on multi-level global-local information fusion to solve the problem of hollow segmentation results caused by complex backgrounds and long-distance dependencies between targets in urban scene remote sensing images. The specific structure is as follows: Figure 3First, a global feature extraction branch is introduced. This branch uses a two-dimensional selective scanning mechanism to obtain contextual knowledge by calculating the compressed hidden state of each image patch along its corresponding scanning path, capturing dependencies in the sequence and enhancing the network's ability to acquire global information. Subsequently, a local feature extraction branch is introduced to focus on local details and enhance the network's ability to extract complex spatial information, thereby improving the segmentation accuracy of objects with long-range dependencies.
[0031] Specifically, the initial architecture of the global feature extraction branch based on the 2D selective scanning mechanism is formulated by replacing the S6 module. The specific structure is as follows Figure 4 As shown. S6 is the core of Mamba, which implements the extraction of global features, selective mechanisms and linear complexity. The core part of this module is SSMs, which implements linear complexity modeling of long-range dependencies through structured state sequences and recursive calculation forms, providing a new technical path for visual tasks. The core idea of SSMs is to describe the dynamic behavior of the system through state equations and observation equations. The state equation defines the change of the system state over time, while the observation equation defines the relationship between the system state and the observed data. The state equation of SSMs is as follows:
[0032] h t =Ah t-1 +Bx t
[0033] Where ht is the state of the system at time step t, xt is the input, A and B are the state transition matrix and input matrix.
[0034] The observation equation of SSMs is as follows:
[0035] y t =Ch t +Dx t
[0036] Among them, yt is the output, C and D are the observation matrix and direct transfer matrix.
[0037] This global feature extraction branch is replaced by the newly proposed 2D Selective Sweep (SS2D) module. To further improve computational efficiency and while achieving the effect of a gating mechanism through SS2D's inherent selectivity, the entire multiplication branch is removed from the module. Consequently, the global feature extraction branch consists of a single network branch with two residual modules. The core component, SS2D, internally comprises three steps: cross-sweep, S6 block selective sweep, and cross-merge. Specifically, SS2D first flattens the input patch into sequences along four distinct traversal paths (i.e., cross-sweep). Each patch sequence is then processed in parallel using a separate S6 block, and the resulting sequences are reshaped and merged to form the output map (i.e., cross-merge). SS2D allows each pixel in the image to integrate information from all other pixels in different directions. This integration helps establish a global receptive field in two dimensions. By introducing the SS2D module, the global feature extraction branch, based on the 2D selective sweep mechanism, inherits the advantages of the Mamba model in sequence data processing and successfully extends it to the visual domain. This branched architecture achieves global context modeling in two-dimensional space through a cross-scanning strategy. Its unique sequence processing approach effectively captures long-range spatial dependencies in images while maintaining linear computational complexity, providing more expressive feature representations for subsequent visual tasks. This design retains the advantages of the original Mamba model in selective feature extraction while being specifically optimized for the two-dimensional nature of visual data, demonstrating dual advantages in computational efficiency and feature representation capabilities.
[0038] Subsequently, a local feature extraction branch based on depthwise separable convolution is constructed to increase the local feature extraction capability of the Mamba module without increasing the amount of computation, so as to achieve the purpose of enhancing local semantic information. The module structure is as follows Figure 5 As shown, the specific structure consists of two local feature extraction branches and a residual connection. First, the feature map is input into a 3*3 and 5*5 depthwise separable convolution to achieve preliminary extraction of multi-scale features from the input feature map. Subsequently, to enhance the feature extraction capability of local information, the Relu activation function is introduced to increase the nonlinearity of the model. Next, the model is input into a batch normalization layer to accelerate the training process and help stabilize gradient propagation, thereby improving the model's convergence speed and performance. Compared with traditional convolution-based local feature extraction algorithms, this algorithm reduces computational complexity. The residual connection branch is used to preserve the network's initial information and enhance the model's expressiveness. Finally, the final output is obtained through a 1*1 convolutional layer. The local feature extraction branch focuses on local details, improving the network's ability to extract complex spatial information, thereby improving the segmentation accuracy of objects with long-range dependencies.
[0039] Step 3: Construct a local high-frequency feature reconstruction module under boundary guidance to solve the problem that the objects in urban scene remote sensing images are easily obscured by vegetation and shadows, which leads to the problem that the extraction results are prone to adhesion between different categories. The specific structure is as follows: Figure 6 First, a boundary feature guidance module is constructed. Under the constraint of boundary information, a boundary feature map is extracted and fused through convolution operations. Then, the boundary feature map is fused with the output features of the last layer of the decoder. The boundary features are further enhanced through the local feature reconstruction module, which improves the network's ability to extract local information and thus enhances the segmentation accuracy of the occluded area.
[0040] Specifically, we first construct a boundary guidance module to introduce the explicit boundary information into the multi-level feature extraction process of the encoder, and finally input the fused boundary information map into the local feature reconstruction module. Before sending it to the final segmentation head, we use the determined boundary information to reduce the noise generated by fusion.
[0041] Under the constraints of boundary information, the boundary guidance module extracts multi-scale boundary information from the features extracted by the backbone network. Ultimately, through upsampling and concatenation, a boundary information map is generated, paving the way for the subsequent local feature reconstruction module. Specifically, the image is first input into the ResNet18 encoder, generating four multi-scale feature maps. To incorporate multi-scale edge information extraction into the backbone network learning process, four identical edge information extraction blocks are sequentially constructed using convolutional layers with a kernel size of 3*3 and ReLu activation functions. After passing through the edge information extraction blocks, the feature maps at the four scales are generated as boundary feature maps at the four scales. The resulting high-level boundary feature maps are then upsampled from the bottom up and summed with the lower-level features to produce a fused boundary prediction map. The boundary information map is then constructed by concatenating the four-scale feature maps by downsampling them separately. This process is supervised by the ground-truth boundary map.
[0042] Secondly, a local high-frequency feature reconstruction module was constructed. The specific process is as follows: First, the model enhances its attention to edge information through a series of reconstruction operations. Specifically, the initial input feature map of size C*H*W (where C represents the number of channels, and H and W represent height and width, respectively) is transformed into a form of size C*N, where N = H*W, and each channel corresponds to information about all pixels. This process aims to reorganize the spatial information for easier processing. Next, C*N is multiplied by its transposed matrix to obtain a boundary-aware matrix, which represents the correlation between each spatial location in M. The more similar the spatial features of two locations in M, the greater the correlation value. Therefore, this step further emphasizes edge-related features and correlates the dependencies between edge pixels. Subsequently, the features of the first layer of the encoder are also reconstructed. The reconstructed result is multiplied by the boundary-aware matrix to obtain a feature map that emphasizes edge information. This feature map is then reconstructed back to its original size of C*H*W, summed with the encoder features, and input into the efficient local attention (ELA) module to further associate and fuse the contextual spatial semantic information. The decoder's output feature map is summed with the edge-enhanced encoder features before the ELA module further enhances the context. This ultimately enhances high-frequency local information, effectively improving the segmentation of occluded areas.
[0043] The Efficient Local Attention (ELA) method uses strip pooling in the spatial dimension to obtain feature vectors in the horizontal and vertical directions, maintaining a narrow kernel shape to capture long-range dependencies and prevent irrelevant regions from affecting label predictions, thereby obtaining rich target position features in their respective directions. The specific structure is as follows: Figure 7 . ELA processes the above feature vectors independently for each direction to obtain attention predictions, and then combines them using a product operation to ensure accurate position information of the region of interest. Specifically, in the second step, 1D convolution is applied to locally interact with the two feature vectors respectively, and the kernel size can be optionally adjusted to indicate the coverage of the local interaction. The resulting feature vectors are subjected to group normalization (GN) and a nonlinear activation function to produce position attention predictions in two directions. The final position attention is obtained by multiplying the position attentions in the two directions. Compared with two-dimensional convolution, one-dimensional convolution is more suitable for processing sequential signals, and is lighter and faster. Compared with BN, GN has comparable performance and greater versatility.
[0044] Those skilled in the art can make various corresponding changes and modifications based on the above technical solutions and concepts, and all of these changes and modifications should be included in the scope of protection of the claims of the present invention.
Claims
1. A semantic segmentation method for urban remote sensing images based on the fusion of global and local depth information, characterized in that: The offense comprises the following steps: Step 1: Construct a multi-scale semantic enhancement module based on context aggregation. Based on global context modeling, this module enhances the semantic association between multi-scale features through a cross-scale interactive learning mechanism, thereby improving the network's multi-scale feature extraction capability. Step 2: Construct a feature decoding module based on multi-level global-local information fusion; introduce global feature extraction branches and local feature extraction branches to enhance the ability to extract complex spatial information in urban scenes; Step three: construct a local high-frequency feature reconstruction module under boundary guidance; introduce boundary information into the process of feature fusion of the last layer decoder and encoder, and further enhance the edge information through the local feature reconstruction module to improve the network's ability to extract local information.
2. The urban remote sensing image semantic segmentation method based on the fusion of global and local depth information according to claim 1 is characterized in that: In the step 1, the multi-scale semantic enhancement module based on context aggregation includes initial feature aggregation and cross-scale interactive learning; the input features are processed through a 1×1 convolution layer to generate the initial aggregated feature f m The role of 1×1 convolution is to reduce the dimension of feature channels while retaining important semantic information, providing a basis for subsequent multi-scale feature interaction, and then the initial aggregated feature f m Divide it into four sub-feature maps along the channel dimension, recorded as Through this division method, feature information can be allocated to different scale spaces, providing diverse feature representations for subsequent cross-scale interactions; In the cross-scale interactive learning stage, the module interactively fuses the features of adjacent branches through a series of dilated convolutions of different scales, applies dilated convolutions with different expansion rates to each sub-feature map, and then fuses the features of adjacent branches by adding or splicing features, thereby achieving cross-scale semantic interaction.
3. The urban remote sensing image semantic segmentation method based on the fusion of global and local depth information according to claim 1 is characterized in that: In the second step, a global feature extraction branch is introduced. This branch uses a two-dimensional selective scanning mechanism to obtain contextual knowledge by calculating the compressed hidden state of each image patch along its corresponding scanning path, capturing the dependencies in the sequence and enhancing the network's ability to acquire global information. Subsequently, a local feature extraction branch is introduced to focus on local detail information and enhance the network's ability to extract complex spatial information, thereby improving the segmentation accuracy of targets with long-range dependencies.
4. The urban remote sensing image semantic segmentation method based on the fusion of global and local depth information according to claim 1 is characterized in that: In the step three, a boundary feature guidance module is constructed. Under the constraint of boundary information, a boundary feature map is extracted and fused through convolution operation. Then, the boundary feature map is fused with the output features of the last layer of the decoder, and the boundary features are further enhanced through the local feature reconstruction module, thereby improving the network's ability to extract local information and thus improving the segmentation accuracy of the occluded area.
5. The urban remote sensing image semantic segmentation method based on the fusion of global and local depth information according to claim 4 is characterized in that: The local feature reconstruction module extracts the boundary perception weight through a specially designed boundary autocorrelation operation.
Citation Information
Cited By
Lightweight semantic segmentation method and system for remote sensing image of urban scene
CN121213934A
Remote sensing image semantic segmentation method and device, equipment and medium
CN121213935A