Remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception

Through a dual-path neural network architecture, combining high-resolution and low-resolution paths, the problem of blurred target boundaries in remote sensing images is solved, and effective capture of multi-scale contextual information and preservation of high-resolution details are achieved, thereby improving the accuracy and robustness of remote sensing image segmentation. It is particularly suitable for fields such as remote sensing image analysis and medical image analysis.

CN120689626APending Publication Date: 2025-09-23CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510963454.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing remote sensing image semantic segmentation methods have difficulty in effectively balancing the capture of multi-scale contextual information, the preservation of high-resolution spatial details, and the precise perception and enhancement of target boundaries when dealing with man-made targets such as buildings with precise geometric boundaries and diverse scales, resulting in blurred and inaccurate boundaries in the segmentation results.

Method used

This method uses a dual-path neural network architecture, comprising a high-resolution path (HR path) and a low-resolution path (LR path), combined with a boundary-enhanced dual fusion path. This approach achieves accurate segmentation of remote sensing images through global context awareness, multi-scale feature extraction, and boundary-aware modules. The HR path preserves spatial details, the LR path acquires contextual information, and the boundary-enhanced path specifically learns and predicts object boundaries. It also optimizes feature representation through cross-path connections and feature fusion.

Benefits of technology

It significantly improves the accuracy and boundary quality of remote sensing image segmentation, enhances adaptability to target scale changes and robustness to complex scenes, can better capture and process target features of different sizes, and generate clear and sharp target contours. It is suitable for remote sensing image analysis, medical image analysis, and autonomous driving environment perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689626A_ABST
    Figure CN120689626A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception, belongs to the technical field of computer vision and deep learning, and solves the problem that in the existing network design, capture of multi-scale context information and maintenance of high-resolution spatial details cannot be effectively balanced. And the target boundary is accurately perceived and enhanced. Comprising the following steps: S1, acquiring a remote sensing image; s2, establishing a dual-path neural network architecture, wherein the dual-path neural network architecture comprises an HR path, an LR path and a boundary enhancement dual fusion path; and S3, inputting the remote sensing image into the dual-path neural network architecture to complete target segmentation in the remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technologies, and in particular to a remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and computer vision, image semantic segmentation, as a foundational and critical technology, has been widely applied in numerous fields, including autonomous driving, medical image analysis, and remote sensing monitoring, becoming a research focus in both academia and industry. In particular, in the field of remote sensing image analysis, accurate semantic segmentation (e.g., extracting buildings, roads, and water bodies) is crucial for urban planning, resource management, and disaster assessment. However, existing techniques for semantic segmentation in remote sensing imagery still face significant challenges, particularly for man-made objects with precise geometric boundaries and diverse scales, such as buildings. First, the scale of objects in remote sensing imagery varies greatly. Objects of the same category (e.g., buildings) can vary in size from a few pixels to occupying a large area, making it difficult for a single receptive field or fixed-scale network to effectively process them simultaneously. Second, object boundaries are often complex and rich in detail. However, traditional convolutional neural network (CNN)-based segmentation methods often lose a significant amount of spatial detail when downsampling through pooling or strided convolution to expand the receptive field and extract high-level semantic features. This results in blurred and inaccurate boundary conditions in the final segmentation results, making them difficult to meet the requirements of sophisticated applications. In addition, accurate segmentation requires not only attention to local details, but also a full understanding of the global and local contextual information of the target, but simple network structures make it difficult to efficiently capture and utilize both types of information at the same time.

[0003] To address these issues, existing technologies typically employ encoder-decoder architectures based on fully convolutional networks (FCNs), U-Nets (U-Nets), and their variants, introduce dilated convolutions to expand the receptive field (e.g., the DeepLab series), or design feature pyramid networks (FPNs) to fuse multi-scale features. However, these approaches still have shortcomings: simple encoder-decoder structures may not fully recover the fine boundary information lost during downsampling; relying solely on dilated convolutions may result in gridding effects or smoothing of details; and existing multi-scale fusion strategies may not be optimal and often lack specialized mechanisms for processing object boundaries.

[0004] Therefore, how to effectively balance the capture of multi-scale contextual information, the preservation of high-resolution spatial details, and the accurate perception and enhancement of target boundaries in network design has become a key technical bottleneck in improving the semantic segmentation performance of remote sensing images. Summary of the Invention

[0005] The present invention solves the problem that existing network designs cannot effectively balance the capture of multi-scale contextual information, the preservation of high-resolution spatial details, and the accurate perception and enhancement of target boundaries.

[0006] The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in the present invention comprises the following steps:

[0007] Step S1, acquiring remote sensing images;

[0008] Step S2, establishing a dual-path neural network architecture, wherein the dual-path neural network architecture includes an HR path, an LR path, and a boundary enhancement dual fusion path;

[0009] In step S3, the remote sensing image is input into the dual-path neural network architecture to complete the target segmentation in the remote sensing image.

[0010] Furthermore, in the embodiment of the present invention, in step S2, the HR path is specifically:

[0011] The remote sensing images after the detail capture network are divided into three paths. The first path of remote sensing images undergoes global upper and lower aggregation convolution, and is added with the second path of remote sensing images through local upper and lower aggregation convolution to output the first remote sensing feature map. The first remote sensing feature map, the remote sensing image, and the third path of remote sensing images are added through the edge perception module to output the second remote sensing feature map. After the second remote sensing feature map passes through the detail extraction network, it is residually connected with the second remote sensing feature map.

[0012] Furthermore, in the embodiment of the present invention, the edge perception module is specifically:

[0013] Grouped convolution or depth-wise separable convolution is followed by 1x1 convolution, followed by BN and ReLU.

[0014] Furthermore, in the embodiment of the present invention, in step S2, the LR path is specifically:

[0015] After passing through the multi-level semantic feature extraction module in sequence, the remote sensing images are divided into two paths. One path of remote sensing images passes through the semantic feature enhancement module, the global context-aware encoder, the global context enhancement, and the summation before entering the context attention enhancement. The other path of remote sensing images is entered into the cross-path connection.

[0016] Furthermore, in an embodiment of the present invention, the global context-aware encoder is specifically:

[0017] The adaptive average pooling layer is followed by a 1x1 convolutional layer, a BN layer, and a ReLU layer.

[0018] Furthermore, in the embodiment of the present invention, in step S2, the boundary enhancement dual fusion path is specifically:

[0019] The cross-path connection receives the output of the LR path at the multi-level semantic feature extraction module and adds it with the output of the HR path;

[0020] The channel adapter receives the output of the LR path;

[0021] The output of the LR path of the channel adapter and the output data of the cross-path connection sum will be processed and output by the boundary enhancement dual fusion module. The data will be processed and output by the boundary enhancement dual fusion module after feature fusion, boundary perception, feature integration, segmentation and output.

[0022] The feature fusion module fuses the output of the LR path after the channel adapter and the output after the cross-path connection and is divided into two paths. One path of remote sensing feature map undergoes boundary perception, feature integration and segmentation output in sequence, and the other path of remote sensing feature map is fused and feature adjusted through feature integration.

[0023] Furthermore, in the embodiment of the present invention, the cross-path connection is specifically:

[0024] Connect 1x1 convolution, BN, ReLU, EfficientAttention module and upsampling layer in sequence;

[0025] The channel adapter is specifically:

[0026] Connect 1x1 convolution, BN, ReLU and EfficientAttention modules in sequence;

[0027] The boundary perception is specifically:

[0028] Connect the multi-level convolution layer, BN, ReLU, 1x1 convolution layer and Sigmoid in sequence;

[0029] The feature fusion is specifically as follows:

[0030] Concatenate fused_feat and boundary_feat in the channel dimension, and then pass them through a 1x1 convolution layer for feature fusion and channel adjustment;

[0031] The segmentation output is specifically:

[0032] Connect the multi-level convolution layer, BN, ReLU and 1x1 convolution layer in sequence.

[0033] The remote sensing image segmentation system based on dual-path multi-scale attention and boundary perception described in the present invention includes the following modules:

[0034] Module S1, acquiring remote sensing images;

[0035] Module S2, establishing a dual-path neural network architecture, wherein the dual-path neural network architecture includes an HR path, an LR path, and a boundary enhancement dual fusion path;

[0036] Module S3 inputs the remote sensing image into the dual-path neural network architecture to complete the target segmentation in the remote sensing image.

[0037] A computer program product according to the present invention comprises a computer program or instructions, which, when executed by a processor, implements the steps of an image segmentation method based on dual-path multi-scale attention and boundary perception as described in any one of the claims above.

[0038] The present invention provides a computer-readable storage medium, wherein a computer program is stored in the storage medium. When the computer program is run, the image segmentation method based on dual-path multi-scale attention and boundary perception described in any one of the claims above is executed.

[0039] This invention solves the problem that existing network designs cannot effectively balance the capture of multi-scale contextual information, the preservation of high-resolution spatial details, and the accurate perception and enhancement of object boundaries. Specific benefits include:

[0040] 1. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in the present invention significantly improves segmentation accuracy and detail preservation capabilities: through a unique dual-path architecture, the high-resolution path (HR path) uses less downsampling and integrates edge-aware processing units (such as EnhancedEdgeModule or grouped convolution / depth-separable convolution) to effectively retain the fine spatial details and texture information of the image; the low-resolution path (LR path) obtains broad contextual information through downsampling and uses global context processing units (such as adaptive average pooling followed by convolution) to enhance semantic understanding. Compared with the problem of traditional single-path encoder-decoder structures that easily lose details, the present invention can better balance global context understanding and high-resolution detail preservation, thereby comprehensively improving the accuracy of pixel-level segmentation;

[0041] 2. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in this invention enhances adaptability to changes in target scale: the integrated multi-scale feature extraction unit (such as the convolutional layers with different dilation rates in the HR path, MultiDilationModule) and feature pyramid fusion strategy (such as OptimizedFeaturePyramid fusion of multi-level features) enable the network to simultaneously capture and process target features of different sizes. This effectively addresses the challenge of the huge size differences of targets such as buildings commonly found in remote sensing images, enabling the model to achieve good segmentation results for targets of different scales, improving the model's generalization and robustness.

[0042] 3. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in the present invention significantly improves the segmentation accuracy of target boundaries: by introducing an independent boundary perception branch (such as EnhancedBoundaryBranch), which specifically learns and predicts the boundary information of the target, and combines it with the edge perception processing unit in the HR path to enhance the edge features. The final feature integration of the explicitly predicted boundary information and the main segmentation features (such as through splicing and 1x1 convolution fusion) can significantly alleviate the blur and discontinuity problems at the boundaries of traditional segmentation networks, and generate clearer, sharper and more accurate target contours, which is crucial for applications that require precise boundaries, such as building extraction;

[0043] 4. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in the present invention optimizes feature representation and context understanding capabilities: widely used attention mechanisms (such as channel attention EfficientAttention, spatial attention SpatialAttention, contextual attention ContextualAttention, and attention for AdaptiveFeatureFusion) are integrated into multiple links of the network, such as feature extraction, cross-path connection (EnhancedCrossPathConnection), and feature fusion. These attention modules can adaptively enhance the response of key feature channels or spatial regions, suppress irrelevant information interference, and effectively capture long-distance dependencies, thereby improving the network's semantic understanding and feature representation capabilities for complex scenes, further contributing to improved segmentation accuracy;

[0044] 5. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in the present invention improves the segmentation robustness and application potential in complex scenes: The present invention integrates multiple advanced mechanisms such as dual-path collaboration, multi-scale processing, attention optimization, and explicit boundary constraints, so that it exhibits higher robustness when dealing with challenging scenes such as complex backgrounds, dense targets, and changing lighting. This method is not only applicable to tasks such as building extraction and land cover classification in remote sensing image analysis, but its core concept of improving segmentation accuracy and boundary quality also provides valuable technical solutions and improvement ideas for other fields requiring high-precision segmentation (such as medical image analysis and autonomous driving environment perception), and has broad application potential and scalability.

[0045] The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in the present invention has significant advantages over existing technologies in segmentation accuracy, boundary quality, scale adaptability and robustness to complex scenes. It can effectively solve the pain points of traditional segmentation methods in refined segmentation tasks, especially in applications with extremely high requirements on spatial details and boundary accuracy, and has important application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0047] Figure 1 This is a flowchart of remote sensing image segmentation using dual-path multi-scale attention and boundary perception as described in embodiment 1;

[0048] Figure 2 This is a diagram of the remote sensing image segmentation effect of dual-path multi-scale attention and boundary perception described in Implementation Method 1. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe various embodiments of the present invention in conjunction with the accompanying drawings. The embodiments described with reference to the accompanying drawings are exemplary and intended to be used to explain the present invention, but should not be understood as limiting the present invention.

[0050] Implementation 1: A remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in this implementation includes the following steps:

[0051] Step S1, acquiring remote sensing images;

[0052] Step S2, establishing a dual-path neural network architecture, wherein the dual-path neural network architecture includes an HR path, an LR path, and a boundary enhancement dual fusion path;

[0053] In step S3, the remote sensing image is input into the dual-path neural network architecture to complete the target segmentation in the remote sensing image.

[0054] In this embodiment, in step S2, the HR path is specifically:

[0055] The remote sensing images after the detail capture network are divided into three paths. The first path of remote sensing images undergoes global upper and lower aggregation convolution, and is added with the second path of remote sensing images through local upper and lower aggregation convolution to output the first remote sensing feature map. The first remote sensing feature map, the remote sensing image, and the third path of remote sensing images are added through the edge perception module to output the second remote sensing feature map. After the second remote sensing feature map passes through the detail extraction network, it is residually connected with the second remote sensing feature map.

[0056] In this embodiment, the edge perception module is specifically:

[0057] Grouped convolution or depth-wise separable convolution is followed by 1x1 convolution, followed by BN and ReLU.

[0058] In this embodiment, in step S2, the LR path is specifically:

[0059] After passing through the multi-level semantic feature extraction module in sequence, the remote sensing images are divided into two paths. One path of remote sensing images passes through the semantic feature enhancement module, the global context-aware encoder, the global context enhancement, and the summation before entering the context attention enhancement. The other path of remote sensing images is entered into the cross-path connection.

[0060] In this implementation, the global context-aware encoder is specifically:

[0061] The adaptive average pooling layer is followed by a 1x1 convolutional layer, a BN layer, and a ReLU layer.

[0062] In this embodiment, in step S2, the boundary enhancement dual fusion path is specifically:

[0063] The cross-path connection receives the output of the LR path at the multi-level semantic feature extraction module and adds it with the output of the HR path;

[0064] The channel adapter receives the output of the LR path;

[0065] The output of the LR path of the channel adapter and the output data of the cross-path connection sum will be processed and output by the boundary enhancement dual fusion module. The data will be processed and output by the boundary enhancement dual fusion module after feature fusion, boundary perception, feature integration, segmentation and output.

[0066] The feature fusion module fuses the output of the LR path after the channel adapter and the output after the cross-path connection and is divided into two paths. One path of remote sensing feature map passes through the boundary perception, feature integration and segmentation head in sequence, and the other path of remote sensing feature map is fused and channel adjusted after feature integration.

[0067] In this embodiment, the cross-path connection is specifically:

[0068] Connect 1x1 convolution, BN, ReLU, EfficientAttention module and upsampling layer in sequence;

[0069] The channel adapter is specifically:

[0070] Connect 1x1 convolution, BN, ReLU and EfficientAttention modules in sequence;

[0071] The boundary perception is specifically:

[0072] Connect the multi-level convolution layer, BN, ReLU, 1x1 convolution layer and Sigmoid in sequence;

[0073] The feature fusion is specifically as follows:

[0074] Concatenate fused_feat and boundary_feat in the channel dimension, and then pass them through a 1x1 convolution layer for feature fusion and channel adjustment;

[0075] The segmentation output is specifically:

[0076] Connect the multi-level convolution layer, BN, ReLU and 1x1 convolution layer in sequence.

[0077] Based on the technical problems existing in the existing technology, such as Figure 1 As shown, this embodiment proposes a remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception, including the following steps:

[0078] Step S1, dual-path feature extraction: The input image is fed into a high-resolution path (HR path) and a low-resolution path (LR path) in parallel for feature extraction. The HR path aims to preserve spatial details, while the LR path acquires contextual information through downsampling.

[0079] Step S2, cross-path feature interaction: The features extracted at a certain stage of the LR path are processed through the cross-path connection unit (including possible attention enhancement and upsampling), and their information is fused into the corresponding stage of the HR path to achieve interaction of features of different resolutions;

[0080] Step S3, multi-scale feature fusion: the HR path features and LR path features after interaction and channel adjustment are sent to the fusion unit of the Feature Pyramid Network (FPN) to generate a fused feature map containing rich multi-scale information;

[0081] Step S4, boundary-aware prediction: the fused feature map obtained in step S3 is fed into an independent boundary-aware branch to explicitly predict the boundary probability map of the target in the image;

[0082] Step S5, final feature integration: integrating the boundary information predicted in step S4 (e.g., boundary probability map) with the fused feature map in step S3 (e.g., by convolution fusion after splicing in the channel dimension);

[0083] Step S6: Segmentation result generation: The features integrated in step S5 are sent to the segmentation output, and after processing through several convolutional layers, the final pixel-level semantic segmentation result is generated.

[0084] The implementation details for the high-resolution path (HR path) are as follows:

[0085] The HR path (corresponding to the HighResolutionPath class in the code) consists of at least two processing stages (e.g., stage 1 and stage 2). Each stage consists of basic convolutional blocks, such as a 3x3 convolutional layer (nn.Conv2d), a batch normalization layer (nn.BatchNorm2d), and a ReLU activation function (nn.ReLU). An attention module, such as a channel attention module (EfficientAttention), is integrated within or after at least one stage of this path (e.g., stage 1 and stage 2) to adaptively enhance feature channels.

[0086] After the first stage (stage 1) feature extraction, it further includes:

[0087] Edge-aware processing unit (self.edge_aware): It consists of a grouped convolution or depth-wise separable convolution (nn.Conv2d sets groups = base_channels) followed by a 1x1 convolution (nn.Conv2d(...,1)), followed by BN and ReLU, to extract and enhance edge features.

[0088] Multi-scale feature extraction unit: Specifically, it can be composed of at least two parallel convolution sequences (such as self.ms_conv1 and self.ms_conv2). Each sequence contains a 3x3 convolution layer with different convolution kernel dilation rates (for example, dilation=1 and dilation=3), followed by batch normalization and ReLU. The output feature maps of these parallel branches are aggregated by element-wise addition (ms_feat1 + ms_feat2).

[0089] The basic features, edge-aware features and multi-scale aggregation features of stage 1 are combined and enhanced (e.g., by element-wise addition).

[0090] Includes cross-stage connections (self.cross_stage): for example, a 1x1 convolutional layer that transforms the enhanced features of stage1 and fuses them with the output of stage2 through a residual connection (element-wise addition).

[0091] The design of this HR path can effectively preserve the fine spatial details of the image, enhance the target boundary features and capture local contextual information at different scales by maintaining higher resolution, integrating edge-aware modules and multi-scale convolutional aggregation; the application of the attention module further optimizes the feature expression.

[0092] The specific implementation details for the low-resolution path (LR path) are as follows:

[0093] The LR path (corresponding to the LowResolutionPath class in the code) first passes through an initial downsampling module (self.stem), which contains a convolutional layer with a stride of 2 (stride=2), a BN layer and a ReLU to halve the input image resolution.

[0094] The path then includes at least two processing stages (such as stage 1 and stage 2), each of which consists of a basic convolutional block (convolution, batch normalization, ReLU) and an attention module (EfficientAttention). The convolution layer stride of at least one stage (such as stage 2) is 2, which achieves further downsampling to expand the receptive field.

[0095] After the last downsampling stage (stage2), it further includes:

[0096] Global context awareness module (self.global_context): Specifically, the feature map spatial dimension is compressed to 1x1 through an adaptive average pooling layer (nn.AdaptiveAvgPool2d(1)), followed by a 1x1 convolutional layer, a BN layer and a ReLU, and the resulting global context vector is added back to the feature map of stage 2.

[0097] Contextual attention module (self.context_attn): For example, use the ContextualAttention module to further model the contextual relationship and perform attention weighting on the feature map after global context enhancement.

[0098] The LR path effectively obtains the global context information of the image through gradual downsampling, and enhances the understanding of scene semantics and the ability to capture long-distance dependencies through the global context module and context attention mechanism.

[0099] The specific implementation details for path interaction, feature fusion, and boundary processing are as follows:

[0100] In the overall model (corresponding to the AMSCNet class in the code):

[0101] Cross-path connection unit (self.cross_connect): Receives the output of the first stage of the LR path (lr_features[0]), processes it through a sequence of 1x1 convolution, BN, ReLU, EfficientAttention module and upsampling layer (nn.Upsample(scale_factor=2,mode='bilinear')), and then adds the result to the features of the second stage of the HR path (hr_features[1]) element-wise.

[0102] Channel adapter (self.lr_adapter): Receives the output of the last stage of the LR path (lr_features[1]) and adjusts its channel number to match the fusion requirements through a sequence of 1x1 convolution, BN, ReLU and EfficientAttention modules.

[0103] Feature fusion unit (self.fpn): adopts the optimized feature pyramid network (OptimizedFeaturePyramid) structure, receives the adjusted LR features (lr_feat_adapted) and the HR features fused with cross-path information (hr_features[1]) as input, fuses multi-level features through top-down and lateral connections, and outputs the fused feature map (fused_feat).

[0104] Boundary perception branch (self.boundary_branch): Receives fused_feat as input, passes through a sequence of multiple convolutional layers (for example, 3x3), BN layers, ReLU activation functions, and finally outputs a single channel feature map through a 1x1 convolutional layer. The Sigmoid activation function (nn.Sigmoid) is applied to obtain the boundary probability map (boundary_feat).

[0105] Final feature integration (self.fusion): fused_feat and boundary_feat are concatenated on the channel dimension (torch.cat), and then a 1x1 convolution layer (self.fusion) is used for feature fusion and channel adjustment.

[0106] Segmentation output (self.seg_head): Receives the integrated features, passes through a sequence of multiple convolutional layers (for example, 3x3), BN layers, ReLU activation functions, and finally outputs the final segmentation result through a 1x1 convolutional layer. The number of channels is equal to the number of categories (num_classes).

[0107] The information interaction of different resolution paths is optimized through cross-path connections with attention and channel adapters; independent boundary branches provide clear boundary constraints; finally, the boundary information is integrated with the multi-scale features of FPN fusion and jointly input into the segmentation output, thus generating accurate segmentation results with clear boundaries and internal consistency.

[0108] In order to better illustrate the remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception described in this embodiment, the following examples are described in detail:

[0109] like Figure 2 In the comparison diagram of remote sensing image segmentation effects shown in the figure, four groups of test samples are used to verify the segmentation performance of the method in different building distribution scenarios.

[0110] Densely built-up areas (Figure a): For densely packed contiguous buildings, this method effectively captures the edge details of adjacent buildings through a dual-path, multi-scale attention mechanism. This clearly displays the outlines of adjacent white buildings against a black background mask, without any adhesion or over-segmentation. In particular, the boundary perception module accurately identifies eaves and shadow areas, avoiding the boundary blurring common in traditional methods.

[0111] Medium-density building area (Figure b): In a semi-sparse scene with increasing distance between buildings, the multi-scale feature fusion path accurately identifies small and medium-sized buildings in various orientations. The results show that the white mask completely covers each building entity, and no discrete noise points appear in the black background area, verifying the algorithm's strong robustness to background interference.

[0112] Multi-scale mixed regions (Figure c): For a complex scene containing a large factory building, the high-level semantic path in the dual-path architecture effectively integrates global contextual information, ensuring complete segmentation of large buildings. The low-level detail path, through spatial attention weighting, accurately preserves the building's geometric features. The white masks of the building objects in the resulting image show clear right-angled edges and regular contours.

[0113] Independent large building area (Figure d): In the segmentation test of a single large building, the boundary perception module uses a gradient-sensitive convolution kernel to enhance the contrast difference between the roof edge and the background, making the white mask boundary of the large-span building smooth and continuous, and completely preserving the polygonal structural characteristics of the building without edge jaggedness or local missing phenomena.

[0114] In summary, this embodiment innovatively designs a dual-path architecture consisting of a high-resolution path and a low-resolution path. The former focuses on preserving spatial details and enhancing local features using multi-scale feature units and edge-aware mechanisms, while the latter captures global context through downsampling and uses an attention mechanism for feature weighting. The information of the two paths is effectively integrated through carefully designed cross-path connections and feature fusion strategies (such as optimized FPN). What is particularly critical is that the present invention introduces an attention mechanism to dynamically optimize feature representation, and sets an independent boundary-aware branch to explicitly predict and strengthen the target boundary.

[0115] This implementation effectively overcomes the limitations of traditional methods in dealing with scale changes and boundary details through the collaborative work of high- and low-resolution paths, combined with multi-scale feature extraction, attention weighting, and independent boundary prediction and fusion mechanisms, and significantly improves the accuracy and robustness of image semantic segmentation. It is particularly suitable for application scenarios with high requirements on segmentation accuracy and boundary quality, such as refined remote sensing image analysis and building extraction. It has important theoretical research value and broad application prospects.

[0116] The above is a detailed introduction to the remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception proposed in the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception, characterized in that: The following steps are involved: Step S1, acquiring remote sensing images; Step S2, establishing a dual-path neural network architecture, wherein the dual-path neural network architecture includes an HR path, an LR path, and a boundary enhancement dual fusion path; In step S3, the remote sensing image is input into the dual-path neural network architecture to complete the target segmentation in the remote sensing image.

2. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception according to claim 1 is characterized in that: In step S2, the HR path is specifically: The remote sensing images after the detail capture network are divided into three paths. The first path of remote sensing images undergoes global upper and lower aggregation convolution, and is added with the second path of remote sensing images through local upper and lower aggregation convolution to output the first remote sensing feature map. The first remote sensing feature map, the remote sensing image, and the third path of remote sensing images are added through the edge perception module to output the second remote sensing feature map. After the second remote sensing feature map passes through the detail extraction network, it is residually connected with the second remote sensing feature map.

3. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception according to claim 2 is characterized in that: The edge perception module is specifically: Grouped convolution or depth-wise separable convolution is followed by 1x1 convolution, followed by BN and ReLU.

4. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception according to claim 1, characterized in that: In step S2, the LR path is specifically: After passing through the multi-level semantic feature extraction module in sequence, the remote sensing images are divided into two paths. One path of remote sensing images passes through the semantic feature enhancement module, the global context-aware encoder, the global context enhancement, and the summation before entering the context attention enhancement. The other path of remote sensing images is entered into the cross-path connection.

5. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception according to claim 4 is characterized in that: The global context-aware encoder is specifically: The adaptive average pooling layer is followed by a 1x1 convolutional layer, a BN layer, and a ReLU layer.

6. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception according to claim 1, characterized in that: In step S2, the boundary enhancement dual fusion path is specifically: The cross-path connection receives the output of the LR path at the multi-level semantic feature extraction module and adds it with the output of the HR path; The channel adapter receives the output of the LR path; The output of the LR path of the channel adapter and the output data of the cross-path connection sum will be processed and output by the boundary enhancement dual fusion module. The data will be processed and output by the boundary enhancement dual fusion module after feature fusion, boundary perception, feature integration, segmentation and output. The feature fusion module fuses the output of the LR path after the channel adapter and the output after the cross-path connection and is divided into two paths. One path of remote sensing feature map undergoes boundary perception, feature integration and segmentation output in sequence, and the other path of remote sensing feature map is fused and feature adjusted through feature integration.

7. The remote sensing image segmentation method based on dual-path multi-scale attention and boundary perception according to claim 6, characterized in that: The cross-path connection is specifically: Connect 1x1 convolution, BN, ReLU, EfficientAttention module and upsampling layer in sequence; The channel adapter is specifically: Connect 1x1 convolution, BN, ReLU and EfficientAttention modules in sequence; The boundary perception is specifically: Connect the multi-level convolution layer, BN, ReLU, 1x1 convolution layer and Sigmoid in sequence; The feature fusion is specifically as follows: Concatenate fused_feat and boundary_feat in the channel dimension, and then pass them through a 1x1 convolution layer for feature fusion and channel adjustment; The segmentation output is specifically: Connect the multi-level convolution layer, BN, ReLU and 1x1 convolution layer in sequence.

8. A remote sensing image segmentation system based on dual-path multi-scale attention and boundary perception, characterized in that: Includes the following modules: Module S1, acquiring remote sensing images; Module S2, establishing a dual-path neural network architecture, wherein the dual-path neural network architecture includes an HR path, an LR path, and a boundary enhancement dual fusion path; Module S3 inputs the remote sensing image into the dual-path neural network architecture to complete the target segmentation in the remote sensing image.

9. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of an image segmentation method based on dual-path multi-scale attention and boundary perception as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is run, it executes the image segmentation method based on dual-path multi-scale attention and boundary perception described in any one of claims 1 to 7.

Citation Information

Cited By

  • Lesion image processing system and method based on deep learning

    CN121458740A

  • Medical image segmentation method, system and equipment based on structure guidance and medium

    CN121725248A