Degradation robust multi-scale context sensing network for remote sensing image segmentation
By using a degenerate robust multi-scale context-aware network, multi-scale sliding windows and four-way directional convolution kernels are employed to process remote sensing images, solving the problem of difficulty in capturing local context in traditional methods and achieving high consistency and fine-grained effect in remote sensing image segmentation.
Patent Information
- Application Number
- CN202511584291.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-06
AI Technical Summary
Traditional convolutional kernels struggle to capture non-axis-aligned and non-axis-symmetric local contexts in remote sensing images. This leads to the loss of details or semantic confusion during feature extraction when small-scale targets and large-scale structures are intertwined. Existing methods have failed to effectively address the problems of feature extraction and semantic understanding in complex scenes.
A degenerate robust multi-scale context-aware network is adopted, including a backbone feature extraction network module, a multi-granularity context aggregation module, and a robust four-directional feature fusion module. Through multi-scale sliding window parallel processing and four-way directional convolution kernel decoupling, it breaks through the traditional serial fusion paradigm and enhances the segmentation consistency of complex scenes and the discriminability of objects with special structures and directional characteristics.
It significantly improves the segmentation consistency of complex scenes in remote sensing image segmentation, solves the problems of difficult fine edge segmentation and directional small target recognition, reduces the model learning difficulty through structured priors, and enhances the discriminability of objects with special structures and directional characteristics.
Smart Images

Figure CN121482391A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the fields of deep learning and image processing technology, and in particular to a degradation-robust multi-scale context-aware network for remote sensing image segmentation. Background Technology
[0002] One of the core challenges in semantic segmentation of remote sensing images is the refined parsing of complex local structures. Land feature boundaries, road networks, and farmland furrows exhibit significant directional characteristics, making it difficult for traditional isotropic convolutional kernels to capture non-axis-aligned and non-axis-symmetric local contexts. The interweaving of small-scale targets and large-scale structures in local regions leads to a tendency for feature extraction from a single receptive field to result in lost details or semantic confusion.
[0003] Zeng Junying (“Multi-level branched cross-scale fusion remote sensing image semantic segmentation network”, Laser & Optoelectronics Progress, 2024, 1-20) constructed a multi-level branched network structure and designed a shallow Swing Transformer feature extraction module, spatial branch, semantic branch and boundary branch. Each branch focuses on extracting feature information at a specific level and introduces a multi-scale decoding module with a large receptive field to transmit feature information at different scales.
[0004] However, this method uses expanding the receptive field as the design principle for the multi-scale decoding module, without considering the challenges of feature extraction in specific directions and structures, which limits the model's feature extraction and semantic understanding of complex scenes. Summary of the Invention
[0005] To address the aforementioned technical issues, embodiments of this application propose a degradation-robust multi-scale context-aware network for remote sensing image segmentation. This network breaks through the serial fusion paradigm of traditional multi-scale methods, avoids upsampling noise, and significantly improves the segmentation consistency of complex scenes. Simultaneously, it reduces the learning difficulty of the model through structured priors, enhancing the discriminability of objects with special structures and directional characteristics.
[0006] To achieve the above objectives, embodiments of this application propose a degradation-robust multi-scale context-aware network for remote sensing image segmentation, the network comprising: a backbone feature extraction network module, a multi-granularity context aggregation module, and a robust four-way feature fusion module; The backbone feature extraction network module receives remote sensing images and performs multi-stage feature extraction operations on the remote sensing images to obtain an initial feature map, which is then input into the multi-granularity context aggregation module. The multi-granularity context aggregation module processes the initial feature map with multi-scale context through a multi-scale sliding window to obtain multi-granularity fused features, which are then input into the robust four-way feature fusion module. Different sliding windows are processed in parallel, and the granularity of the different sliding windows is different. The robust four-way feature fusion module decouples the directional components of multi-granularity fusion features through four-way directional convolution kernels, and fuses the directional features output by the four-way directional convolution kernels to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
[0007] To achieve the above objectives, embodiments of this application propose a degradation-robust multi-scale context-aware method for remote sensing image segmentation, the method comprising the following steps: Acquire remote sensing images; By using a backbone feature extraction network module, multi-stage feature extraction operations are performed on remote sensing images to obtain an initial feature map; The multi-granularity context aggregation module is used to process the initial feature map with multi-scale context through a multi-scale sliding window to obtain multi-granularity fused features. Different sliding windows are processed in parallel, and the granularity of the different sliding windows is different. A robust four-way feature fusion module is used to decouple the directional components of multi-granularity fusion features through four-way directional convolution kernels, and the directional features output by the four-way directional convolution kernels are fused to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
[0008] To achieve the above objectives, embodiments of this application also propose a degradation-robust multi-scale context-aware device for remote sensing image segmentation, the device comprising: The acquisition module is used to acquire remote sensing images; The extraction module is used to perform multi-stage feature extraction operations on remote sensing images using the backbone feature extraction network module to obtain an initial feature map; The processing module utilizes the multi-granularity context aggregation module to process the initial feature map using a multi-scale sliding window to obtain multi-granularity fused features. Different sliding windows are processed in parallel, and the granularity of the different sliding windows is different. The decoupling module utilizes the robust four-way feature fusion module to decouple the directional components of multi-granularity fusion features through four-way directional convolution kernels, and fuses the directional features output by the four-way directional convolution kernels to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
[0009] To achieve the above objectives, embodiments of this application also propose an electronic device, including: a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to execute the instructions such that the electronic device can implement a degradation-robust multi-scale context-aware method for remote sensing image segmentation as described above.
[0010] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables a degradation-robust multi-scale context-aware method for remote sensing image segmentation as described above.
[0011] This application proposes a degradation-robust multi-scale context-aware network for remote sensing image segmentation, comprising: a backbone feature extraction network module, a multi-granularity context aggregation module, and a robust four-way feature fusion module. The backbone feature extraction network module receives remote sensing images and performs multi-stage feature extraction operations to obtain an initial feature map with strong discriminative power, multi-scale, and hierarchical structure. This initial feature map is then input to the multi-granularity context aggregation module as the basis for further feature extraction, discrimination, and fusion to achieve final accurate segmentation. Subsequently, the multi-granularity context aggregation module processes the initial feature map using a multi-scale sliding window to obtain multi-granularity fused features, which are then input to the robust four-way feature fusion module. Because different sliding windows process in parallel... Furthermore, the different sliding windows have varying granularities, which breaks through the traditional serial fusion paradigm of multi-scale methods. It understands the dependencies within local regions from different spatial scales and then fuses these complementary information, avoiding upsampling noise and significantly improving the segmentation consistency of complex scenes. Finally, the robust four-directional feature fusion module decouples the directional components of the multi-granularity fusion features through four-way directional convolutional kernels and fuses the directional features output by the four-way directional convolutional kernels to obtain the target feature map for semantic segmentation of remote sensing images. Since the four-way directional convolutional kernels process in parallel, and the dimensions of the kernels in different paths are different, this allows for structured priors, reducing the learning difficulty of the model and enhancing the discriminability of objects with special structures and directional characteristics. This effectively solves the problems of difficult fine-grained edge segmentation and recognition of small directional targets within local context. Based on this, this scheme breaks through the traditional serial fusion paradigm of multi-scale methods, avoids upsampling noise, and significantly improves the segmentation consistency of complex scenes; at the same time, it reduces the learning difficulty of the model through structured priors and enhances the discriminability of objects with special structures and directional characteristics.
[0012] Optionally, the initial feature map is processed using a multi-scale sliding window to obtain multi-granularity fused features. This includes: obtaining multi-scale sliding windows with different granularities at different scales, and ensuring that the sliding windows at each scale do not overlap; constructing the initial feature map into local features at multiple scales using the multi-scale sliding window, performing self-attention computation within each scale sliding window to obtain the output features of each scale sliding window; wherein the output features of each scale sliding window are features whose spatial dimensions are restored after inverse transformation of the sliding window; and summing the output features of each scale sliding window element by element to obtain the multi-granularity fused features.
[0013] Optionally, the initial feature map is denoted as , The self-attention calculation performed within each scale sliding window yields the output features of each scale sliding window, which are expressed by the following formula (1): (1); in, , , For linear projection of features within the sliding window, Scaling factor This indicates the size of the corresponding sliding window. , H represents the height of the initial feature map, W represents the width of the initial feature map, and C represents the number of channels in the initial feature map. The output features of each scale sliding window are added element-wise to obtain multi-granularity fused features. It can be expressed by the following formula (2): (2); in, , , , These represent the features obtained after sliding window transformations of three different scales.
[0014] Optionally, the multi-scale sliding window includes: a fine-grained window, a medium-grained window, and a coarse-grained window; the fine-grained window is used to divide the initial feature map into... Non-overlapping Window; a medium-granularity window is used to divide the initial feature map into Non-overlapping Window; coarse-grained window is used to divide the initial feature map into Non-overlapping window.
[0015] Optionally, the directional components of the multi-granularity fusion features are decoupled using four-way directional convolution kernels, and the directional features output by the four-way directional convolution kernels are fused to obtain the target feature map, including: Obtain four-way directional convolutional kernels; wherein, the four-way directional convolutional kernels include a first-way convolutional kernel, a second-way convolutional kernel, a third-way convolutional kernel, and a fourth-way convolutional kernel, and the dimensions of the first-way convolutional kernel and the third-way convolutional kernel are respectively... The dimensions of the second and fourth convolutional kernels are respectively ; This is the default value; The first convolutional kernel scans the multi-granularity fusion features from top to bottom to obtain contextual information in the vertical direction. By using a second convolutional kernel, multi-granularity fusion features are scanned from left to right to obtain horizontal contextual information; The third convolutional kernel scans the multi-granularity fusion features from the top left to the bottom right to obtain contextual information in the first diagonal direction. The fourth convolutional kernel scans the multi-granularity fusion features from the top left to the bottom right to obtain contextual information along the second diagonal; the first diagonal is different from the second diagonal. The context information in the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction is fused to obtain the target feature map.
[0016] Optionally, the network also includes a decoding module; The multi-granularity context aggregation module inputs multi-granularity fused features into the decoding module; The decoding module receives multi-granularity fused features and performs upsampling on the multi-granularity fused features so that the size of the upsampled multi-granularity fused features is consistent with the size of the initial feature map; The decoding module fuses the upsampled multi-granularity fusion features with the initial feature map to obtain the target multi-granularity fusion features; The decoding module inputs the target multi-granularity fused features into the robust four-way feature fusion module. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0018] Figure 1 This is a schematic diagram of the structure of a degradation-robust multi-scale context-aware network for remote sensing image segmentation provided in one embodiment of this application; Figure 2 This is a schematic diagram of another degenerate robust multi-scale context-aware network provided in one embodiment of this application; Figure 3 This is a schematic diagram of the structure of a multi-granularity context aggregation module provided in one embodiment of this application; Figure 4 This is a schematic diagram of the structure of a robust four-way feature fusion module provided in one embodiment of this application; Figure 5 This is a schematic diagram illustrating an image processing effect provided in one embodiment of this application; Figure 6 This is a flowchart of a degradation-robust multi-scale context-aware method for remote sensing image segmentation provided in another embodiment of this application; Figure 7 This is a schematic diagram of the structure of a degradation-robust multi-scale context-aware device for remote sensing image segmentation provided in another embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0020] One of the core challenges in semantic segmentation of remote sensing images is the refined parsing of complex local structures. Land feature boundaries, road networks, and farmland furrows exhibit significant directional characteristics, making it difficult for traditional isotropic convolutional kernels to capture non-axis-aligned and non-axis-symmetric local contexts. The interweaving of small-scale targets and large-scale structures in local regions leads to a tendency for feature extraction from a single receptive field to result in lost details or semantic confusion.
[0021] Zeng Junying (“Multi-level branched cross-scale fusion remote sensing image semantic segmentation network”, Laser & Optoelectronics Progress, 2024, 1-20) constructed a multi-level branched network structure and designed a shallow Swing Transformer feature extraction module, spatial branch, semantic branch and boundary branch. Each branch focuses on extracting feature information at a specific level and introduces a multi-scale decoding module with a large receptive field to transmit feature information at different scales.
[0022] However, this method uses expanding the receptive field as the design principle for the multi-scale decoding module, without considering the challenges of feature extraction in specific directions and structures, which limits the model's feature extraction and semantic understanding of complex scenes.
[0023] Therefore, how to solve the problem of feature extraction in specific directions and structures, and realize the model's feature extraction and semantic understanding of complex scenes, is a technical problem that urgently needs to be solved.
[0024] To address the aforementioned challenges, this invention proposes a Degradation-Robust Multi-Scale Context Perception Network (RMCPNet) for remote sensing image semantic segmentation tasks. It comprises two innovative modules: a Multi-granularity Context Aggregation Module (MCAM) and a Robust Quadruple-Oriented Feature Fusion Module (RQFM). The role of MCAM is to execute these modules in parallel. , , Three sets of self-attention calculations at different window scales are used, and the three-way features are summed to achieve multi-scale local context aggregation. At the same time, it breaks through the serial fusion paradigm of traditional multi-scale methods, avoids upsampling noise, and significantly improves the segmentation consistency of complex scenes. RQFM designs four parallel directional convolution kernels to explicitly decouple the directional components that depend on local space. It reduces the learning difficulty of the model through structured priors and enhances the discriminability of objects with special structures and directional characteristics.
[0025] One embodiment of this application proposes a degradation-robust multi-scale context-aware network for remote sensing image segmentation. The implementation details of this degradation-robust multi-scale context-aware network for remote sensing image segmentation are described below. The schematic diagram of the degenerate robust multi-scale context-aware network for remote sensing image segmentation proposed in this embodiment can be seen as follows: Figure 1 and Figure 2As shown, it includes: a backbone feature extraction network module, a multi-granularity context aggregation module, and a robust four-way feature fusion module. Optionally, the perceptual network may also include a decoding module.
[0026] For the backbone feature extraction network module: The backbone feature extraction network module receives remote sensing images and performs multi-stage feature extraction operations on the remote sensing images to obtain an initial feature map, and then inputs the initial feature map into the multi-granularity context aggregation module.
[0027] For example, in this application embodiment, ConvNeXt can be used as the backbone feature extraction network module; after acquiring the remote sensing image, it gradually extracts and transforms the feature representation with strong discriminativeness, multi-scale and hierarchical structure from the remote sensing image, which serves as the basis for the subsequent multi-granularity context aggregation module and robust four-way feature fusion module to further extract, discriminate and fuse features and achieve the final accurate segmentation.
[0028] This application embodiment can use remote sensing images from the Vaihingen dataset as input images to illustrate the specific implementation of this application embodiment. The input image is... An RGB image of size , denoted as .
[0029] This application adopts an "encoding-decoding" technical framework. The backbone feature extraction network module used in the encoding stage is the ConvNeXt model. In specific implementation, to balance accuracy and computational efficiency, the ConvNeXt-tiny model is selected as the backbone feature extraction network module. That is, feature extraction is implemented in four stages during the encoding stage, and the feature map sizes corresponding to the four stages are as follows: , , , .
[0030] In the ConvNeXt-tiny encoding process, a multi-granularity context aggregation module designed in this application embodiment is added at the end of each of the four stages, which are denoted as follows according to the different stages: .
[0031] Taking the first phase as an example, The input size is To further reduce computational complexity, before dividing the window and calculating self-attention, the input is halved in width, height, and number of channels; for example, in height, 128 pixels are reduced to 64 pixels; in width, 128 pixels are reduced to 64 pixels; and in number of channels, 96 channels are reduced to 48 pixels. Updated to .
[0032] For the multi-granularity context aggregation module: This module processes the initial feature map using a multi-scale sliding window to obtain multi-granularity fused features, which are then input into the robust four-way feature fusion module. Different sliding windows process these features in parallel, and each window has a different granularity.
[0033] In one possible embodiment, the initial feature map is processed using a multi-scale sliding window to obtain multi-granularity fused features. This includes: acquiring multi-scale sliding windows, where the granularity of the sliding windows at different scales is different and the sliding windows at each scale do not overlap; when the initial feature map is constructed into local features at multiple scales using the multi-scale sliding windows, self-attention calculation is performed within each scale sliding window to obtain the output features of each scale sliding window; and the output features of each scale sliding window are added element-wise to obtain the multi-granularity fused features.
[0034] The output features of each scale sliding window are the features that restore the spatial dimension after the inverse transformation of the sliding window.
[0035] For example, the Multi-Granularity Context Aggregation Module (MCAM) can enhance local context modeling capabilities and avoid upsampling noise through a parallel multi-scale window self-attention mechanism. First, the initial feature map is divided into multiple non-overlapping sliding windows of different scales. Then, self-attention calculations are performed independently within each sliding window to generate multi-scale features. The three output features are then fused by adding them element-wise to output the enhanced multi-granularity fused features.
[0036] For example, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a multi-granularity context aggregation module provided in an embodiment of this application. The core idea of the multi-granularity context aggregation module is to simultaneously understand the dependencies within a local region from different spatial scales, and then fuse these complementary information to achieve a more comprehensive and robust representation of the local context. The local context modeling capability is enhanced through a parallel multi-scale window partitioning strategy and a feature fusion mechanism. The core operation consists of three stages: the multi-scale window partitioning stage, the parallel self-attention computation stage, and the feature fusion stage.
[0037] In one possible embodiment, the multi-scale sliding window includes: a fine-grained window, a medium-grained window, and a coarse-grained window; the fine-grained window is used to divide the initial feature map into... Non-overlapping Window; a medium-granularity window is used to divide the initial feature map into Non-overlapping Window; coarse-grained window is used to divide the initial feature map into Non-overlapping window.
[0038] For example, given the input feature map Construct local windows of three independent scales: 1) Fine-grained window: This involves viewing the feature map... Divided into Non-overlapping window; 2) Medium-grained window: This involves viewing the feature map... Divided into Non-overlapping window; 3) Coarse-grained window: This applies to the feature map. Divided into Non-overlapping window.
[0039] Specifically, if the input batch size is set to 4, meaning that each training iteration inputs 4 images to the network, then the dimension of the input tensor is 4. The tensor dimension for performing window self-attention computation is . .for The size of the window, the tensor dimension becomes That is, from the original 4 images of different sizes... The 48-channel tensor was transformed into 1024 frames of size. A 48-channel tensor. Window and The window process is exactly the same as described above, except for the window size.
[0040] It should be explained that, in order to reduce the complexity of subsequent self-attention calculations, the width, height, and number of channels of the input feature map can be halved before dividing the window for self-attention calculations. This reduces the complexity of self-attention calculations while retaining key feature information. For details, please refer to the above embodiment.
[0041] In one possible embodiment, the initial feature map is denoted as... , The self-attention calculation performed within each scale sliding window yields the output features of each scale sliding window, which are expressed by the following formula (1): (1); in, , , For linear projection of features within the sliding window, Scaling factor This indicates the size of the corresponding sliding window. , H represents the height of the initial feature map, W represents the width of the initial feature map, and C represents the number of channels in the initial feature map. For example, if s=2, it means the size of the sliding window is... If s=4, it means the size of the sliding window is... If s=8, it means the size of the sliding window is... .
[0042] The output features of each scale sliding window are added element-wise to obtain multi-granularity fused features. It can be expressed by the following formula (2): (2); in, , , , These represent the features obtained after sliding window transformations of three different scales.
[0043] For example, The output tensor is upsampled based on features, and its dimension is the same as the input tensor.
[0044] For the robust four-way feature fusion module: The robust four-way feature fusion module decouples the directional components of multi-granularity fusion features through four-way directional convolution kernels, and fuses the directional features output by the four-way directional convolution kernels to obtain the target feature map.
[0045] The four directional convolution kernels are processed in parallel, with different dimensions for each kernel. The target feature map is used for semantic segmentation of remote sensing images.
[0046] In one possible embodiment, the step of decoupling the directional components of the multi-granularity fusion features using four-way directional convolution kernels and fusing the directional features output by the four-way directional convolution kernels to obtain the target feature map includes: Obtain four-way directional convolution kernels; The first convolutional kernel scans the multi-granularity fusion features from top to bottom to obtain contextual information in the vertical direction. By using a second convolutional kernel, multi-granularity fusion features are scanned from left to right to obtain horizontal contextual information; The third convolutional kernel scans the multi-granularity fusion features from the top left to the bottom right to obtain contextual information in the first diagonal direction. The fourth convolutional kernel scans the multi-granularity fusion features from the top left to the bottom right to obtain contextual information along the second diagonal; the first diagonal is different from the second diagonal. The context information in the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction is fused to obtain the target feature map.
[0047] The four-way convolutional kernel includes a first-way convolutional kernel, a second-way convolutional kernel, a third-way convolutional kernel, and a fourth-way convolutional kernel. The dimensions of the first-way convolutional kernel and the third-way convolutional kernel are respectively... The dimensions of the second and fourth convolutional kernels are respectively ; This is the default value.
[0048] For example, vertical convolutional kernels (e.g., first-path convolutional kernels) can focus on vertical contextual information, such as building outlines, tree rows, etc. Specifically, after multi-granularity fused features pass through a cross-attention mechanism, each column of pixels is scanned by an asymmetric convolutional kernel to calculate the vertical dependencies of pixels within that column. This results in a strong response to vertical edges and vertically arranged objects, capturing intensity variations and consistency within columns.
[0049] For example, horizontally oriented convolutional kernels (e.g., second-path convolutional kernels) can capture intra-row dependencies and are sensitive to horizontal edges; for example, roads. Specifically, after multi-granularity fused features pass through a cross-attention mechanism, each row of pixels is scanned by an asymmetric convolutional kernel to calculate the horizontal dependencies of pixels within that row. This results in a strong response to horizontal edges and horizontally arranged objects, capturing intra-row intensity variations and consistency.
[0050] For example, a convolutional kernel along the first diagonal direction (e.g., a third-path convolutional kernel) can capture structural information from the top left to the bottom right; for example, the diagonal structure of farmland furrows. Specifically, this convolutional path scans a specific diagonal region along the top-left to bottom-right diagonal direction, calculating dependencies along that direction. It responds strongly to structures along the main diagonal edges and along this direction, effectively capturing continuity and corresponding patterns in that direction.
[0051] For example, the convolutional kernel in the second diagonal direction (e.g., the fourth convolutional kernel) can capture structural information from the top right to the bottom left; for example, specific feature arrangements. Specifically, it responds strongly to structural information in the anti-diagonal direction, working in conjunction with the third convolutional kernel to cover the local context and corresponding dependencies mainly in the diagonal direction.
[0052] For example, such as Figure 4 As shown, Figure 4This is a schematic diagram of the structure of a robust four-way feature fusion module provided in an embodiment of this application. The robust four-way feature fusion module includes four parallel branches. The features extracted by each branch are further processed by multi-scale convolution kernels, such as asymmetric convolution, padding, convolution, and restoration. Finally, they are fused into a single feature map through addition or splicing operations.
[0053] In one possible embodiment of this application, the network further includes: a decoding module; a multi-granularity context aggregation module inputting multi-granularity fused features to the decoding module; the decoding module receiving the multi-granularity fused features and performing an upsampling operation on the multi-granularity fused features so that the size of the upsampled multi-granularity fused features is consistent with the size of the initial feature map; the decoding module fusing the upsampled multi-granularity fused features with the initial feature map to obtain the target multi-granularity fused features; and the decoding module inputting the target multi-granularity fused features to a robust four-way feature fusion module.
[0054] For example, the output of each multi-granularity context aggregation module is denoted as follows: , , , .
[0055] in, The third-stage decoding features are obtained through decoding upsampling, denoted as... ; and With the same dimensions, element-wise summation followed by decoding upsampling yields the second-stage decoding features, denoted as... Similarly, and With the same dimensions, element-wise summation followed by decoding upsampling yields the second-stage decoding features, denoted as... .
[0056] The features decoded by the decoding module are denoted as: , which serves as the input feature of the robust four-way feature fusion module.
[0057] Specifically, when The input is fed into the robust four-way feature fusion module, and passes through four paths in parallel within the module. Specifically, the first path's convolutional kernel has a dimension of... The direction is from top to bottom; the dimension of the second convolutional kernel is... The direction is from left to right; the dimension of the third convolutional kernel is... The direction is from top left to bottom right; the dimension of the fourth convolutional kernel is... The direction is from the top left to the bottom right.
[0058] Each path contains three branches corresponding to convolutional kernels of different scales, enabling feature fusion across multiple scales and directions.
[0059] The output of the robust four-way feature fusion module and The feature tensors after convolution are summed and then fed into the robust four-way feature fusion module; this process is repeated to obtain... The output after the robust four-way feature fusion module is denoted as Subsequently, after upsampling, and with The final output tensor is obtained by adding the upsampled features element by element.
[0060] in, This indicates a 2x upsampling. This indicates a 4x upsampling.
[0061] Training is performed under the constraint of the cross-entropy loss function until the maximum number of iterations is reached.
[0062] Understandably, the robust four-way feature fusion module extracts and encodes the most representative spatial dependency design in a local area in an explicit and structured manner by performing parallel scanning convolutions in four directions: vertical, horizontal, main diagonal, and anti-diagonal. This significantly improves the ability to represent local features of remotely sensed objects with directional and linear structures, and effectively solves problems related to local context, such as the difficulty in fine-grained edge segmentation and the identification of small directional targets.
[0063] Based on the above experimental steps, experiments were conducted on remote sensing image datasets such as Vaihingen, Potsdam, and LoveDA. The average intersection-union ratio (mIoU) of each category was used as a measure of the network model's performance in semantic segmentation tasks on a given dataset.
[0064]
[0065] Numerical results and visualization results are as follows Figure 5 As shown, Figure 5 This is a schematic diagram illustrating an image processing effect provided by an embodiment of this application. Quantitative analysis shows that the mIoU value of this application on the Vaihingen dataset is significantly higher than that of the UNetFormer method and the PPMamba method; Figure 5 The image processing results shown demonstrate that this application achieves higher completeness and semantic consistency in the identification of elongated tree-type targets within the red box compared to the UNetFormer and PPMamba methods.
[0066] This application proposes a degradation-robust multi-scale context-aware network for remote sensing image segmentation, comprising: a backbone feature extraction network module, a multi-granularity context aggregation module, and a robust four-way feature fusion module. The backbone feature extraction network module receives remote sensing images and performs multi-stage feature extraction operations to obtain an initial feature map with strong discriminative power, multi-scale, and hierarchical structure. This initial feature map is then input to the multi-granularity context aggregation module as the basis for further feature extraction, discrimination, and fusion to achieve final accurate segmentation. Subsequently, the multi-granularity context aggregation module processes the initial feature map using a multi-scale sliding window to obtain multi-granularity fused features, which are then input to the robust four-way feature fusion module. Because different sliding windows process in parallel... Furthermore, the different sliding windows have varying granularities, which breaks through the traditional serial fusion paradigm of multi-scale methods. It understands the dependencies within local regions from different spatial scales and then fuses these complementary information, avoiding upsampling noise and significantly improving the segmentation consistency of complex scenes. Finally, the robust four-directional feature fusion module decouples the directional components of the multi-granularity fusion features through four-way directional convolutional kernels and fuses the directional features output by the four-way directional convolutional kernels to obtain the target feature map for semantic segmentation of remote sensing images. Since the four-way directional convolutional kernels process in parallel, and the dimensions of the kernels in different paths are different, this allows for structured priors, reducing the learning difficulty of the model and enhancing the discriminability of objects with special structures and directional characteristics. This effectively solves the problems of difficult fine-grained edge segmentation and recognition of small directional targets within local context. Based on this, this scheme breaks through the traditional serial fusion paradigm of multi-scale methods, avoids upsampling noise, and significantly improves the segmentation consistency of complex scenes; at the same time, it reduces the learning difficulty of the model through structured priors and enhances the discriminability of objects with special structures and directional characteristics.
[0067] Another embodiment of this application proposes a degradation-robust multi-scale context-aware method for remote sensing image segmentation, applied to an electronic device, wherein the electronic device can be a terminal or a server. This embodiment and the following embodiments will use a server as an example for description. The implementation details of the degradation-robust multi-scale context-aware method for remote sensing image segmentation proposed in this embodiment will be described in detail below. The following implementation details are provided for ease of understanding and are not necessary for implementing this solution.
[0068] The specific process of the degradation-robust multi-scale context-aware method for remote sensing image segmentation proposed in this embodiment can be described as follows: Figure 6 As shown, it includes: Step 601: Acquire remote sensing images.
[0069] Step 602: Using the backbone feature extraction network module, perform multi-stage feature extraction operations on the remote sensing image to obtain the initial feature map.
[0070] Step 603: Using the multi-granularity context aggregation module, the initial feature map is processed with multi-scale context through a multi-scale sliding window to obtain multi-granularity fused features.
[0071] Different sliding windows are processed in parallel, and the granularity of the different sliding windows is different.
[0072] Step 604: Using the robust four-way feature fusion module, the directional components of the multi-granularity fusion features are decoupled through four-way directional convolution kernels, and the directional features output by the four-way directional convolution kernels are fused to obtain the target feature map.
[0073] The four directional convolution kernels are processed in parallel, with different dimensions for each kernel. The target feature map is used for semantic segmentation of remote sensing images.
[0074] For the specific implementation of steps 601 to 604, please refer to the relevant description of a degradation robust multi-scale context-aware network for remote sensing image segmentation in the above embodiments, which will not be repeated here.
[0075] This embodiment proposes a degradation-robust multi-scale context-aware method for remote sensing image segmentation. It acquires remote sensing images and uses a backbone feature extraction network module to perform multi-stage feature extraction operations to obtain an initial feature map. Then, a multi-granularity context aggregation module processes the initial feature map using a multi-scale sliding window to obtain multi-granularity fused features. Because different sliding windows process in parallel with varying granularities, this approach breaks through the traditional serial fusion paradigm of multi-scale methods, understanding dependencies within local regions at different spatial scales before fusing these dependencies. Complementary information avoids upsampling noise, significantly improving segmentation consistency in complex scenes. Finally, a robust four-directional feature fusion module is used, employing four directional convolutional kernels to decouple the directional components of multi-granularity fusion features and fuse the directional features output by the four kernels to obtain the target feature map. Because the four directional convolutional kernels process in parallel, with different dimensions for each kernel, structured priors can reduce the model's learning difficulty, thereby enhancing the discriminability of objects with special structures and directional characteristics. This effectively solves the problem of local context, such as the difficulty in fine-grained edge segmentation and the recognition of small directional targets. Based on this, this scheme breaks through the traditional serial fusion paradigm of multi-scale methods, avoids upsampling noise, and significantly improves segmentation consistency in complex scenes; simultaneously, structured priors reduce the model's learning difficulty and enhance the discriminability of objects with special structures and directional characteristics.
[0076] The steps described above are for clarity only. In implementation, they can be combined into one step, or some steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0077] Another embodiment of this application proposes a degradation-robust multi-scale context-aware device for remote sensing image segmentation. The details of this device are described below for ease of understanding and are not essential for implementing this example. Figure 7 This is a schematic diagram of the structure of a degradation-robust multi-scale context-aware device for remote sensing image segmentation proposed in this embodiment, including: Acquisition module 701 is used to acquire remote sensing images; The extraction module 702 is used to perform multi-stage feature extraction operations on remote sensing images using the backbone feature extraction network module to obtain an initial feature map. The processing module 703 is used to process the initial feature map with multi-scale context through a multi-scale sliding window using a multi-granularity context aggregation module to obtain multi-granularity fused features; wherein, different sliding windows are processed in parallel, and the granularity of different sliding windows is different; The decoupling module 704 is used to utilize the robust four-way feature fusion module to decouple the directional components of multi-granularity fusion features through four-way directional convolution kernels, and to fuse the directional features output by the four-way directional convolution kernels to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
[0078] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0079] It is worth mentioning that all modules and units involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units do not exist in this embodiment.
[0080] Another embodiment of this application provides an electronic device, such as Figure 8 As shown, it includes a processor 81 and a memory 82. The memory 82 stores instructions that the processor 81 can execute. When the processor 81 is configured to execute the instructions, the electronic device can implement a degradation-robust multi-scale context-aware method for remote sensing image segmentation as described in the above method embodiment.
[0081] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0082] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0083] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a degradation-robust multi-scale context-aware method for remote sensing image segmentation as described in the above method embodiments.
[0084] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0085] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A degradation-robust multi-scale context-aware network for remote sensing image segmentation, characterized in that, The network includes: a backbone feature extraction network module, a multi-granularity context aggregation module, and a robust four-way feature fusion module; The backbone feature extraction network module receives remote sensing images and performs multi-stage feature extraction operations on the remote sensing images to obtain an initial feature map, which is then input into the multi-granularity context aggregation module. The multi-granularity context aggregation module performs multi-scale context processing on the initial feature map through a multi-scale sliding window to obtain multi-granularity fused features, and inputs the multi-granularity fused features into the robust four-way feature fusion module; different sliding windows are processed in parallel, and the granularity of different sliding windows is different; The robust four-way feature fusion module decouples the directional components of multi-granularity fusion features through four-way directional convolution kernels, and fuses the directional features output by the four-way directional convolution kernels to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
2. The degradation-robust multi-scale context-aware network for remote sensing image segmentation according to claim 1, characterized in that, The process of performing multi-scale contextual processing on the initial feature map through a multi-scale sliding window to obtain multi-granularity fused features includes: Obtain multi-scale sliding windows, where the granularity of sliding windows at different scales is different and the sliding windows at each scale do not overlap; When the initial feature map is constructed into local features at multiple scales through a multi-scale sliding window, self-attention computation is performed within each scale sliding window to obtain the output features of each scale sliding window; wherein, the output features of each scale sliding window are features whose spatial dimensions are restored after the inverse transformation of the sliding window. The output features of sliding windows at each scale are added element by element to obtain multi-granularity fused features.
3. The degradation-robust multi-scale context-aware network for remote sensing image segmentation according to claim 2, characterized in that, Let the initial feature map be... , The self-attention calculation performed within each scale sliding window yields the output features of each scale sliding window, which are expressed by the following formula (1): (1); in, , , For linear projection of features within the sliding window, Scaling factor This indicates the size of the corresponding sliding window. , H represents the height of the initial feature map, W represents the width of the initial feature map, and C represents the number of channels in the initial feature map. The output features of each scale sliding window are added element-wise to obtain multi-granularity fused features. It can be expressed by the following formula (2): (2); in, , , , These represent the features obtained after sliding window transformations of three different scales.
4. The degradation-robust multi-scale context-aware network for remote sensing image segmentation according to claim 3, characterized in that, Multi-scale sliding windows include: fine-grained windows, medium-grained windows, and coarse-grained windows; fine-grained windows are used to divide the initial feature map into... Non-overlapping Window; a medium-granularity window is used to divide the initial feature map into Non-overlapping Window; coarse-grained window is used to divide the initial feature map into Non-overlapping window.
5. The degradation-robust multi-scale context-aware network for remote sensing image segmentation according to claim 1, characterized in that, The process involves decoupling the directional components of multi-granularity fusion features using four-way directional convolution kernels, and fusing the directional features output by the four-way directional convolution kernels to obtain the target feature map, including: Obtain four-way directional convolutional kernels; wherein, the four-way directional convolutional kernels include a first-way convolutional kernel, a second-way convolutional kernel, a third-way convolutional kernel, and a fourth-way convolutional kernel, and the dimensions of the first-way convolutional kernel and the third-way convolutional kernel are respectively... The dimensions of the second and fourth convolutional kernels are respectively ; This is the default value; The first convolutional kernel scans the multi-granularity fusion features from top to bottom to obtain contextual information in the vertical direction. By using a second convolutional kernel, multi-granularity fusion features are scanned from left to right to obtain horizontal contextual information; The third convolutional kernel scans the multi-granularity fusion features from the top left to the bottom right to obtain contextual information in the first diagonal direction. The fourth convolutional kernel scans the multi-granularity fusion features from the top left to the bottom right to obtain contextual information along the second diagonal; the first diagonal is different from the second diagonal. The context information in the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction is fused to obtain the target feature map.
6. The degradation-robust multi-scale context-aware network for remote sensing image segmentation according to claim 1, characterized in that, The network also includes a decoding module; The multi-granularity context aggregation module inputs multi-granularity fused features into the decoding module; The decoding module receives multi-granularity fused features and performs upsampling on the multi-granularity fused features so that the size of the upsampled multi-granularity fused features is consistent with the size of the initial feature map; The decoding module fuses the upsampled multi-granularity fusion features with the initial feature map to obtain the target multi-granularity fusion features; The decoding module inputs the target multi-granularity fused features into the robust four-way feature fusion module.
7. A degradation-robust multi-scale context-aware method for remote sensing image segmentation, characterized in that, The method includes: Acquire remote sensing images; By using a backbone feature extraction network module, multi-stage feature extraction operations are performed on remote sensing images to obtain an initial feature map; The multi-granularity context aggregation module is used to process the initial feature map with multi-scale context through a multi-scale sliding window to obtain multi-granularity fused features. Different sliding windows are processed in parallel, and the granularity of the different sliding windows is different. A robust four-way feature fusion module is used to decouple the directional components of multi-granularity fusion features through four-way directional convolution kernels, and the directional features output by the four-way directional convolution kernels are fused to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
8. A degradation-robust multi-scale context-aware device for remote sensing image segmentation, characterized in that, The device includes: The acquisition module is used to acquire remote sensing images; The extraction module is used to perform multi-stage feature extraction operations on remote sensing images using the backbone feature extraction network module to obtain an initial feature map; The processing module utilizes the multi-granularity context aggregation module to process the initial feature map using a multi-scale sliding window to obtain multi-granularity fused features. Different sliding windows are processed in parallel, and the granularity of the different sliding windows is different. The decoupling module utilizes the robust four-way feature fusion module to decouple the directional components of multi-granularity fusion features through four-way directional convolution kernels, and fuses the directional features output by the four-way directional convolution kernels to obtain the target feature map. The four-way directional convolution kernels are processed in parallel, and the directional convolution kernels of different paths have different dimensions. The target feature map is used for semantic segmentation of remote sensing images.
9. An electronic device, characterized in that, include: The processor and memory, wherein the memory stores instructions that the processor can execute, and the processor is configured to, when executing the instructions, enable the electronic device to implement the degradation-robust multi-scale context-aware method for remote sensing image segmentation as described in claim 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the degradation-robust multi-scale context-aware method for remote sensing image segmentation as described in claim 7.