Detail enhancement and scale selection road extraction network and method for high-resolution remote sensing image
By introducing edge enhancement residual blocks, shallow content guidance spatial attention modules and multi-scale context selection modules into the deep learning road extraction network, the problems of receptive field limitation and insufficient fusion of multi-scale contexts are solved, and the accuracy and integrity of road extraction in high-resolution remote sensing images are achieved.
Patent Information
- Application Number
- CN202510124190.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing deep learning road extraction networks are limited due to the limitation of the receptive field in high-resolution remote sensing images, and the multi-scale context fusion method fails to fully consider adaptive selection and adjustment.
A detailed enhancement and scale selection road extraction network DESSNet is proposed, which enhances the extraction of road detail information and dynamic selection of multi-scale context by introducing edge enhancement residual blocks, shallow content guidance spatial attention modules and multi-scale context selection modules.
It significantly improves the road edge recognition ability, enhances the characterization ability of different types of roads, improves the accuracy and integrity of road extraction, and can extract road information in a relatively complete manner, especially when the road is blocked.
Smart Images

Figure CN120070911A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and particularly relates to a road extraction network and method for detail enhancement and scale selection for high-resolution remote sensing images. Background Art
[0002] With the rapid development of deep learning technology, deep learning has been introduced into the field of road extraction from remote sensing images. Road extraction networks can be constructed based on different deep learning models, such as Convolutional Neural Networks (CNNs), graph neural networks, generative adversarial networks, etc. Among them, the most commonly used is the CNN network.
[0003] Early CNN road extraction was mainly achieved by improving classical semantic segmentation networks in the field of computer vision. The literature " Road Extraction by Deep Residual U-Net.In "IEEE Geoscience and Remote Sensing Letters", Zhang et al. proposed a ResUnet road extraction network by integrating Unet with the residual network ResNet. In the literature "A Novel Road Extraction Method for Remote Sensing Images via Combining High-Level Semantic Feature and Context", Yang et al. proposed an RCNN-Unet road extraction network based on Unet by designing a recurrent CNN (RCNN) unit that can retain more detailed features. In the literature "Improved High-Resolution Remote Sensing Image Rural Road Extraction Method Based on the Fully Convolutional Network", Li et al. introduced the atrous spatial pyramid pooling (ASPP) structure into the skip connections of the Unet network to extract multi-scale features of roads and expand the receptive field of features. In the literature "Road Extraction from High-Resolution Remote Sensing Images Using a Spatial Information-Aware Semantic Segmentation Model", Wu et al. enhanced the ability to perceive spatial information and global context information by introducing coordinate convolution and a global information enhancement module into the residual network ResNet, thus improving the accuracy of road extraction. In the literature "LinkNet with Pretrained Encoder and Dilated Convolution for High Resolution Satellite Imagery Road Extraction", Zhou et al. introduced a dilated convolutional layer between the encoding and decoding of LinkNet, proposed a road extraction network D-LinkNet, and won the first place in the CVPR 2018 road extraction challenge. In the literature "A Multiscale and Multidirection Feature Fusion Network for Road Detection From Satellite Imagery", Wang et al. optimized the LinkNet network using strip convolution and proposed a multi-scale and multi-direction road extraction network MSMDFFNet. Although the above CNN road extraction methods can all effectively extract road information, due to the local connection mechanism of the CNN network, the receptive field is easily limited, which to a certain extent restricts the accuracy of road extraction.
[0004] Regarding the problem of limited receptive fields, the attention mechanism is introduced into CNN road extraction to improve the network's ability to model long-distance correlations and enhance road extraction performance. In the literature "Road Extraction Using a Dual Attention Dilated-LinkNet Based on Satellite Images and Floating Vehicle Trajectory Data", Gao et al. proposed the DAD-LinkNet road extraction network by introducing dual self-attention modules of channels and spaces into the dilated convolutional layer of the D-LinkNet network. In the literature "Toward Lighter But More Accurate Road Extraction With Nonlocal Operations", Wang et al. proposed NL-LinkNet by introducing the non-local attention module into the decoding stage of the LinkNet network, enhancing the network's ability to model global information and thus solving the problem of road occlusion. In the literature "Multiple-Parameters-Guided Squeeze-and-Excitation Integrated D-LinkNet for Road Extraction in Remote Sensing Imagery", Ai et al. first optimized the squeeze and excitation (SE) attention module using variance and coefficient of variation, and then introduced it into the D-LinkNet network to improve the accuracy of road extraction. In the literature "Road Extraction From Satellite Imagery by Road Context and Full-Stage Feature", Yang et al. designed a coordinate dual attention module and introduced it into the full-level feature fusion process to enhance the ability to represent road features.
[0005] Recently, some remote sensing scholars have extracted road information by integrating Transformer and CNN. Transformer adopts the multi-head self-attention mechanism, which can more effectively model the global correlation of images. Transformer and CNN networks can be integrated in two ways: serial and parallel. For the serial method, in the literature "A Novel Road Extraction Method for Remote Sensing Images via Combining High-Level Semantic Feature and Context", Yang et al. first used CNN to extract road features, and then used the Swin-Transformer module to model long-distance correlations to enhance the extracted features. In the literature "Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road Shapes", Wang et al. first used ResNet to extract features, and then proposed an efficient strip Transformer module to model the global context information of the features. In "A Hybrid CNN-Transformer Network for Road Extraction From Satellite Imagery", Liu et al. proposed a hybrid road extraction network RoadCT by alternately using CNN convolutions and Transformer modules. For the parallel method, in "Swin Transformers Make Strong Contextual Encoders for VHR Image Road Extraction", Chen et al. gave a dual-branch encoding module CoSwin by integrating the Swin Transformer module and the residual convolution module in parallel to extract the global and local features of the road simultaneously. In "Dual-path extraction network based on CNN and transformer for accurate building and road extraction", Chen et al. used the CNN network to extract the features of the original image, used the Transformer to extract the features of the image after 2x upsampling of the original image, and then extracted the buildings and roads by fusing the two features.
[0006] Although many deep learning road extraction network models have been proposed, road extraction remains an open problem due to the numerous challenges it faces. The main challenges in road extraction include: ① Roads in remote sensing images are usually slender, complex, and only occupy a very small part of the entire image, which makes them contain a lot of detailed information. However, deep learning models usually need to perform multiple downsamplings, which easily leads to the loss of detailed information. ② Road types and distributions are diverse, and they are often blocked by the surrounding environment, which easily results in discontinuous road extraction. Although this problem can be effectively alleviated by fusing multi-scale context, the existing methods for fusing multi-scale context do not fully consider the adaptive selection and adjustment between multi-scale contexts.
[0007] Based on the above analysis, this patent is based on the classic D-LinkNet network and proposes a Detail-enhancement and scale-selection (DESS) road extraction network DESSNet. First, by introducing an edge detection operator into the residual block through convolution operations, a simple and effective Boundary-enhanced residual block (BECL) is proposed to fully extract the detailed information of the road. Then, a Shallow-content guided attention module (SCGSM) is constructed in the skip connection to use the shallow features to guide the deep features layer by layer, thereby enhancing the detailed information of each deep feature in turn. Finally, a Multi-scale context selection module (MCSM) is designed to adaptively select different scale contexts using complementary dual-branch attention to enhance the representation ability for different types of roads. Summary of the Invention
[0008] The present invention provides a road extraction network and method for detail enhancement and scale selection for high-resolution remote sensing images. Based on the classical D-LinkNet network, a road extraction network DESSNet for detail enhancement and scale selection (DESS) is proposed. First, by means of a convolution operation, an edge detection operator is introduced into the residual block, and a simple and effective boundary-enhanced residual block (BECL) is proposed to fully extract the detail information of the road. Then, a shallow-content guided attention module (SCGSM) is constructed in the skip connection to guide the deep features layer by layer with the shallow features, thereby enhancing the detail information of each deep feature in turn. Finally, a multi-scale context selection module (MCSM) is designed to adaptively select different-scale contexts using complementary dual-branch attention to enhance the representation ability for different types of roads.
[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A road extraction network for detail enhancement and scale selection for high-resolution remote sensing images. The extraction network uses the semantic segmentation neural network D-LinkNet, and based on the encoding and decoding architecture, it includes an encoder based on the boundary-enhanced residual convolutional layer, a shallow-content guided attention module SCGSM, a multi-scale context selection module MCSM, and a decoder; the original image enters the encoder through the semantic segmentation neural network D-LinkNet, and after being processed by the boundary-enhanced residual block, it is processed by the shallow-content guided attention module SCGSM and the multi-scale context selection module MCSM, and then the road information is output by the decoder.
[0010] The above decoder contains a boundary-enhanced residual block module BECL. The boundary-enhanced residual block module BECL uses an unsupervised edge detection algorithm and introduces a Sobel edge detection operator. The boundary-enhanced residual block module BECL identifies the regions with significant pixel changes by calculating the gradient of the image, thereby identifying the boundaries of the road; by adding a boundary-enhanced branch in the encoder, the extraction of road detail information is enhanced, and the accurate extraction of the road edge is realized; The shallow content-guided spatial attention module SCGSM uses shallow features to guide the deep feature information transmitted to the decoder through skip connections, compensating for the lack of spatial information in deep features. To make the dimensions of shallow features the same as those of deep features, detail-preserving downsampling is used to reduce the dimensions of shallow features. At the same time, a strip convolution module is used for deep features to extract strip features, and then the shallow and deep features are fused to obtain deep features guided by shallow content. The multi-scale context selection module MCSM adopts the idea of spatial selection, effectively weights and spatially fuses the features extracted from branches with different dilation rates in the dilated convolution layer, dynamically adjusts the receptive field according to different scales of the road, and selects appropriate context information.
[0011] In the semantic segmentation neural network D-LinkNet, the above-mentioned edge enhancement residual block modular convolutional layer BECL uses ResNe34 as the decoder, and the edge enhancement residual block modular convolutional layer BERCL includes 4 layers.
[0012] In the semantic segmentation neural network D-LinkNet, the above-mentioned shallow content-guided spatial attention module SCGSM contains two branches, a shallow feature branch and a deep feature branch. The input of the shallow feature branch is the output of the previous shallow guidance attention module, and the input of the deep feature branch is the output of the current layer encoder.
[0013] The above-mentioned shallow feature branch includes a detail-preserving downsampling module. Through the detail-preserving downsampling module, the shallow feature map is downsampled to the same size as the deep feature map while retaining the detail features.
[0014] The above-mentioned multi-scale context selection module MCSM contains two parts, a multi-scale feature extraction part and a scale selection part; The multi-scale feature extraction part takes the sum of the features of the last layer encoder and the output of the last hierarchical interaction module as the input, and then inputs it into four parallel branches with different dilation rates. Each branch has a 3×3 convolution, capturing feature information at different scales with dilation rates of 1, 2, 4, and 8 respectively. After the four branches, four different-scale features ( F 1 , F 2 , F 3 , F 4 ) are obtained; then the feature maps of two adjacent scales are concatenated to retain more scale information; the scale selection part is used to consider the feature correlation between adjacent scales.
[0015] Using the extraction method of the detail enhancement and scale selection road extraction network for high-resolution remote sensing images described above, in the edge enhancement residual block module BERCL, the original residual block RB receives the input features E in , the features E in successively pass through a 3×3 convolution, a BN layer, a Relu activation, a 3×3 convolution, and a BN layer, and then are added to the input coming through the identity mapping to obtain E out , and then E out is subjected to Relu activation to obtain the final result. The specific operations are as follows: ; (1) Among them, C represents a 3×3 convolution, Bn represents batch normalization, RL represents the Relu activation operation; In the edge enhancement residual block module BERCL, an edge enhancement branch is added to the original residual block RB. The edge enhancement branch uses the Sobel operator to extract boundary information, and the extracted edge information is added to the output through a residual connection. The edge enhancement branch first reduces the number of channels through a convolution to generate features with 1 channel; secondly, it smooths the feature map through an average pooling layer to reduce noise effects and performs downsampling; then it is input into a Sobel layer to enhance the edge details in the encoder features; finally, the number of channels of the features is restored to twice the original through a convolution module. The specific process is as follows: ; (2) Among them F in and F out respectively represent the input and output of the boundary enhancement branch, Conv 1×1 represents a 1×1 convolution, Avg 2×2 represents a 2×2 average pooling, φ represents the edge detection operation.
[0016] When the above-mentioned Sobel layer enhances the edge details in the encoder features, four different-direction convolutions are used to detect the gradient change between road pixels and background pixels. The specific process of the discrete differential operator containing four 3×3 convolutions is as follows: ; (3) Among them, , , , , f is the input of the Sobel layer, F 0 , F 1 , F 2 , F 3 represent the gradient values of the image in four different directions, F is the output of the Sobel layer. The output of the edge branch is finally added to the output of the original residual structure to obtain the output of the edge-enhanced residual block.
[0017] Under the above details, the downsampling module uses two branches, the convolution branch and the pooling branch, for downsampling operations. The convolution branch introduces the Space-to-depth Conv (SPD-Conv); rearranges the spatial information of the feature map, folds the spatial information within a 2×2 range into the channel information of 1 pixel, and then uses a 1×1 convolution for channel compression and feature fusion. Additionally, a pooling branch is added. First, max-pooling and average-pooling are respectively performed on the features, and with a learnable parameter W, the results of the two poolings are concatenated proportionally. Then, the number of channels is adjusted through a convolution to obtain the output of the pooling branch; finally, the outputs of the two branches are concatenated to obtain the output of the shallow branch. The specific operation can be expressed by the following formula: ; (4) ; (5) ; (6) where, F S1 represents the output of the convolution branch, F S2 represents the output of the pooling branch, SPD represents the Space-to-depth Conv, C represents the 3×3 convolution operation, Cat represents the concatenation operation, W is a learnable parameter, and × represents the dot product operation, Max and Avg respectively represent the max-pooling and average-pooling operations, F S ′ is the output of the shallow branch; The deep branch first uses a bar convolution module for bar feature extraction to enhance road features; the specific operation process of the bar convolution module can be shown as follows: ; (7) ; (8) ; (9) ; (10) In the formula, Conv 1×5 ,Conv 5×1 , Conv 1×7 , Conv 7×1 , Conv 1×9 , Conv 9×1 are convolution operations with convolution kernel sizes of 1×5, 5×1, 1×7, 7×1, 1×9, and 9×1 respectively; F 1 , F 2 , F 3 are the outputs of the three branches respectively, F 4 is the output of the bar-shaped convolution module; after extracting the bar-shaped features through the bar-shaped convolution module, the maximum value and average value of the channel dimension are calculated respectively, and the two pooled features are concatenated and then the channels are changed to 1 through a 1×1 convolution, and the attention map is obtained using the Sigmod function and the attention map is extended to the same dimension as F D ′ through a broadcast operation; the specific operation can be expressed by the following formula: ; (11) In the formula, Max represents the operation of finding the maximum value of the channel dimension, Avg represents the operation of finding the average value of the channel dimension, Concat represents the concatenation operation, Conv represents the 1×1 convolution, σ is the Sigmod function, E represents the broadcast operation, F D ′ is the output of the deep feature layer; finally, the outputs of the two branches are multiplied to obtain the correlation map of the deep and shallow features, and the correlation map can be further used as an attention map to enhance the spatial information in the deep features. The specific operation can be expressed by the following formula: ; (12) The output of the shallow guidance attention module will be input into the next shallow guidance attention module and the decoder at the corresponding level respectively.
[0018] In the above-mentioned multi-scale context selection module MCSM, the multi-scale feature extraction part will obtain four different-scale features (F1 , F 2 , F 3 , F 4 ) The feature maps of two adjacent scales in are concatenated. After concatenation, the features need to pass through two branches respectively. One branch first compresses the features spatially through a 1×1 convolution, compressing the features into 1×H×W, then passes through a Sigmod layer to obtain the weights of the features in the spatial dimension, and then uses the weight map to re-weight the feature map spatially, as shown in the following formula: ; (13) ; (14) In the formula, Cat represents the concatenation operation, C represents the convolution with a convolution kernel of 1×1, Bn represents batch normalization, RL represents the Relu activation function, σ represents the Sigmod function, and × represents the dot product operation; the other branch first passes through a 3×3 convolution, and then performs a Softmax operation on the feature map to obtain two weight vectors W 1 and W 2 and weights the feature map to obtain a weighted feature map V ; finally, the outputs of the two branches are added to obtain F 12 ′ , and the above operations are as follows: ; (15) ; (16) ; (17) where + and × represent element-wise product and addition respectively, φ represents the Softmax activation function, F 12 ′ represents the output of the scale selection part; F 23 ′ is obtained from F 12 ′ and F 3 through the above operations, while F 34 ′ is obtained from F 23′ and F 4 obtained through the above operations; finally, the input feature map is added to the output of the scale selection part through a residual connection; the output of the multi-scale context selection module can be expressed by the following formula: ; (18) where F out represents the output of the scale selection module, F in represents the input of the scale selection module.
[0019] The present invention proposes a road extraction network and method for detail enhancement and scale selection for high-resolution remote sensing images, having the following beneficial effects: 1. First, an edge enhancement branch is designed. This module enhances the edge information in the encoder and decoder features through residual connections, significantly improving the model's ability to identify road edges, which is crucial for road extraction because edge information is often an important part of road features.
[0020] 2. Second, a shallow content-guided spatial attention module is also designed to use shallow features to guide deep features and supplement the spatial information lacking in deep features.
[0021] 3. Finally, a context scale selection module is added between the decoder and the encoder. It can better process road features at different scales. Especially when solving the situation where roads are occluded, it extracts multi-scale context information by dynamically adjusting the receptive field, further improving the overall performance of the model. Compared with five models in recent years on the public datasets CHN6-CUG and Massachusetts Road, the results show that the proposed method has good performance and can extract roads more completely. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The present invention will be further described below with reference to the drawings and embodiments: Figure 1 is the structural diagram of the road extraction network of the present invention; Figure 2 is the structural diagram of the original residual block and the edge enhancement residual block module; Figure 3 is the structural diagram of the shallow content-guided spatial attention module; Figure 4 is the structural diagram of the multi-scale context selection module; Figure 5 is the road prediction map on the CHN6-CUG dataset in the embodiment; Figure 6It is the road prediction map on the Massachusetts Road dataset in the embodiment. Detailed implementation mode
[0023] To make the purpose, technical solution and advantages of the present invention clearer, the following content will combine with the accompanying drawings provided according to the present invention to systematically and completely describe the specific technical solution of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.
[0024] Embodiment 1: As Figure 1 shown, the proposed Detail Enhancement and Scale Selection Road Extraction Network DESSNet adopts an encoding and decoding architecture, mainly including four parts: an encoder based on a boundary enhancement residual block, a Shallow Content Guided Spatial Attention Module SCGSM, a Multi-Scale Context Selection Module MCSM, and a decoder.
[0025] The Boundary Enhancement Residual Block BERCL added to the decoder mainly adopts an unsupervised edge detection algorithm, introducing the Sobel edge detection operator. It identifies regions with significant pixel changes by calculating the gradient of the image, thereby identifying the boundaries of the road. By adding an edge enhancement branch in the encoder, the extraction of road detail information is enhanced, and accurate extraction of road edges is achieved.
[0026] The SCGSM module uses shallow features to guide the deep feature information passed to the decoder through skip connections, making up for the lack of spatial information in deep features. To make the dimensions of shallow features the same as those of deep features, detail-preserving downsampling is used to reduce the dimensions of shallow features. At the same time, a strip convolution module is used for deep features to extract strip features, and then the shallow and deep features are fused to obtain deep features guided by shallow content.
[0027] The MCSM module adopts the idea of spatial selection, effectively weights and spatially fuses the features extracted from branches with different dilation rates in the dilated convolution layer, dynamically adjusts the receptive field according to different scales of the road, and selects appropriate context information.
[0028] 1. Boundary Enhancement Residual Block Module In D-LinkNet, ResNe34 is used as the decoder. It contains 4 layers, including 3, 4, 6, and 3 residual blocks respectively. The residual block utilizes the principle of identity mapping. By directly passing the original input to the output and adding it to the features processed by the convolutional layer, it can alleviate the problems of gradient disappearance and gradient explosion, and can retain low-level features, increase the network depth, and improve the representation ability of the network.Figure 2 Figure (a) shows the original residual block RB, and the input features E in pass through a 3×3 convolution, a BN layer, a Relu activation, a 3×3 convolution, and a BN layer in sequence, and then are added to the input from the identity mapping to obtain E out , and then E out is activated by Relu to obtain the final result. The specific operations are as follows: ; Among them, C represents a 3×3 convolution, Bn represents batch normalization, RL represents the Relu activation operation. Although the original residual block can alleviate the degradation problem, consecutive downsampling will cause some loss of edge details. In high-resolution remote sensing images, road edge features are one of the important features of road information, and clear and complete edge information is an important factor for road extraction. To address the problem of edge detail information loss caused by downsampling, reference
[24] draws on the idea of edge detection in traditional road extraction methods and introduces the edge detection operator - Sobel operator into the road extraction task. By adding a Sobel layer based on residual connection to the Swin-Transformer block to construct the SwinTopology module, the Sobel operator is used to make up for the edge detail loss caused by downsampling, which improves the accuracy of road extraction from remote sensing images. However, the computational complexity of Swin-Transformer is relatively high and requires a large amount of computing resources. To make up for the edge detail loss caused by downsampling without adding too much computational amount and ensure the lightweight of the model, this paper combines the advantages of Resnet and Sobel and proposes an improved residual block - edge-enhanced residual block to enhance the extraction effect of road edges.
[0029] As Figure 2 shown in (b), the proposed edge-enhanced residual block adds an edge enhancement branch to the original residual block. The edge enhancement branch uses the Sobel operator to extract boundary information and adds the extracted edge information to the output through residual connection. The proposed edge enhancement branch first reduces the channel dimension through a convolution to generate features with 1 channel. Secondly, an average pooling layer is used to smooth the feature map to reduce the noise effect and perform downsampling. Then it is input into a Sobel layer to enhance the edge details in the encoder features. Finally, the channel number of the features is restored to twice the original through a convolution module. The specific process is as follows: ; F in is the input of the edge enhancement module, C 1 , C 2 respectively represent 1×1 convolutions that change the number of channels to 1 and 2C, φ is the Sobel layer, F out represents the output of the edge enhancement branch in the encoder.
[0030] The Sobel layer is the core unit of the edge enhancement branch. It uses convolutions in four different directions to detect the gradient changes between road pixels and background pixels.
[0031] 2. Shallow Content Guided Spatial Attention Module In D-LinkNet, by using skip connections, the encoder features are directly passed into the decoder through the skip connections, which can recover the feature loss caused by the upsampling operation and improve the accuracy of road extraction. However, directly passing the encoder features into the decoder through the skip connections, although it can make up for the feature loss caused by upsampling, it ignores the disadvantage that the deep semantic features lack spatial information. During the feature extraction process, the deep feature map E D is smaller and contains accurate road semantic information, which can enhance the model's ability to distinguish roads and backgrounds. The shallow feature map E S contains rich spatial detail information, which can provide the overall structural information of the road. Therefore, the shallow detail features have important guiding significance for the deep semantic features. By guiding the deep semantic features with the shallow detail features, the spatial information lacking in the deep semantic features can be supplemented. Therefore, this paper designs a shallow content guided fusion module to guide the deep semantic features with the shallow detail features and supplement the spatial detail information lacking in the deep semantic features, enabling the network to more accurately identify roads and suppress the interference of background information. The specific structure of the module is as Figure 3 shown. It contains two branches, the shallow feature branch and the deep feature branch. The input of the shallow feature branch is the output of the previous layer's shallow guidance attention module, and the input of the deep feature branch is the output of the current layer's encoder.
[0032] The shallow feature branch includes a detail-preserving downsampling module. Through this module, the shallow feature map is downsampled to the same size as the deep feature map while retaining the detail features. Traditional downsampling operations generally use max pooling or average pooling. However, due to the intricate road structures in remote sensing images, max pooling will lose a lot of detailed semantic information, and average pooling will dilute the detail information in many regions. Therefore, the effects of these two pooling operations are not good. To solve this problem, a detail-preserving downsampling module is proposed. This module uses two branches, a convolutional branch and a pooling branch, for downsampling operations. The convolutional branch introduces the Space-to-depth Conv (SPD-Conv). It rearranges the spatial information of the feature map, folds the spatial information within a 2×2 range into channel information of 1 pixel, and then uses a 1×1 convolution for channel compression and feature fusion. In this way, the image resolution is reduced without causing feature loss. However, considering that some small roads only occupy 1 to 2 pixels in the image, the above method will still cause some road breaks. Therefore, a pooling branch is added. First, max pooling and average pooling are respectively performed on the features. With a learnable parameter W, the results of the two poolings are concatenated proportionally, and then the number of channels is adjusted through a convolution to obtain the output of the pooling branch. Finally, the outputs of the two branches are concatenated to obtain the output of the shallow branch. The specific operation can be expressed by the following formula: ; ; ; Among them, F S1 represents the output of the convolutional branch, F S2 represents the output of the pooling branch, SPD represents the Space-to-depth Conv, C represents the 3×3 convolution operation, Cat represents the concatenation operation, W is a learnable parameter, × represents the dot product operation, Max and Avg represent the max pooling and average pooling operations respectively, F S ′ is the output of the shallow branch.
[0033] The deep branch first uses a bar-shaped convolution module for bar-shaped feature extraction to enhance road features. The specific operation process of the bar-shaped convolution module is as follows: ; ; ; ; In the formula, Conv 1×5 ,Conv 5×1 , Conv 1×7 , Conv 7×1 , Conv 1×9 , Conv 9×1 are convolution operations with convolution kernel sizes of 1×5, 5×1, 1×7, 7×1, 1×9, and 9×1 respectively. F 1 , F 2 , F 3 are the outputs of the three branches respectively. F 4 is the output of the bar convolution module. After extracting bar features through the bar convolution module, the maximum value and average value in the channel dimension are calculated respectively to effectively extract spatial information. In order to enable feature interaction, the two pooled features are concatenated and then the channels are changed to 1 through a 1×1 convolution. The attention map is obtained using the Sigmod function and the attention map is extended to the same dimension as F D ′ through broadcast operation. The specific operation can be expressed by the following formula: ; In the formula, Max represents the operation of finding the maximum value in the channel dimension, Avg represents the operation of finding the average value in the channel dimension, Concat represents the concatenation operation, Conv represents the 1×1 convolution, σ is the Sigmod function, E represents the broadcast operation, F D ′ is the output of the deep feature layer. Finally, the outputs of the two branches are multiplied to obtain the correlation map of the deep and shallow features. The correlation map can be further used as an attention map to enhance the spatial information in the deep features. The specific operation can be expressed by the following formula: ; The output of the shallow guidance attention module will be input into the next layer of the shallow guidance attention module and the decoder at the corresponding level respectively.
[0034] 3. Multi-scale context selection module In the road extraction task, local detailed features contribute to the accurate segmentation of roads, while global context features help reduce misclassification and ensure road connectivity. Rich context features can provide information about the shape, direction, and other valuable features of roads, helping the network more accurately determine whether it is a road, alleviating the problem of road occlusion, and greatly assisting in improving road connectivity. Through the analysis of the two datasets used in this experiment, it can be seen that accurate road extraction requires two conditions: 1) extensive context information is needed. 2) Different-sized roads require context information at different scales. To capture more context information, D-LinkNet adds an atrous convolution layer between the encoder and the decoder, and captures extensive context information by stacking atrous convolutions with different dilation rates to expand the receptive field. Although this can handle attributes such as the narrowness, connectivity, complexity, and long span of roads to a certain extent, its receptive field size is fixed and cannot meet the second condition for accurate road extraction. The road scales are complex and variable, and using a single scale cannot accurately extract road information. Only by dynamically adjusting the size of the context according to the size of the road can the context information be more fully utilized. To solve this problem, a multi-scale context selection module is proposed.
[0035] As Figure 4 shown, the MCSM module consists of two parts, a multi-scale feature extraction part and a scale selection part.
[0036] The multi-scale feature extraction part takes the sum of the features of the last layer of the encoder and the output of the last hierarchical interaction module as the input, and then inputs it into four parallel branches with different dilation rates. Each branch has a 3×3 convolution to capture feature information at different scales with dilation rates of 1, 2, 4, and 8 respectively. After passing through the four branches, four features at different scales ( F 1 , F 2 , F 3 , F 4 ) are obtained. Then, the feature maps of two adjacent scales are concatenated to retain more scale information. Next, the features enter the scale selection part, considering the feature correlation between adjacent scales. This module can effectively weight and spatially fuse features at different scales, and dynamically adjust the receptive field of the feature map. The concatenated features pass through two branches respectively. One branch first passes through a 1×1 convolution to spatially compress the features, compressing the features into 1×H×W, then passes through a Sigmod layer to obtain the weights of the features in the spatial dimension, and then uses the weight map to re-weight the feature map spatially, as shown in the following formula: ; ; In the formula, Cat represents the splicing operation, C represents the convolution with a 1×1 convolution kernel, Bn represents batch normalization, RL represents the Relu activation function, σ represents the Sigmod function, and × represents the dot product operation. Another branch first passes through a 3×3 convolution, and then performs a Softmax operation on the feature map to obtain two weight vectors W 1 and W 2 and weights the feature map to obtain the weighted feature map V . Finally, the outputs of the two branches are added to obtain E 12 ′ , and the above operations are as follows: ; ; ; where + and × represent element-wise product and addition, φ represents the Softmax activation function, F 12 ′ represents the output of the scale selection part. Through the above work, multi-scale features can be effectively aggregated, and the receptive fields of the feature maps at different scales can be automatically adjusted to highlight some important regions at other scales. F 23 ′ is obtained from F 12 ′ and F 3 through the above operations, while F 34 ′ is obtained from F 23 ′ and F 4 through the above operations. Finally, the input feature map is added to the output of the scale selection part through a residual connection. The output of the multi-scale context selection module can be expressed by the following formula: ; where, F out represents the output of the scale selection module, F inRepresents the input of the scale selection module. Through this module, the feature information from different scales is dynamically aggregated by the multi-scale context selection module. The four branches of different scales here can also be extended to multi-scale branches according to the actual task.
[0037] 4. Hybrid Loss Function In the road extraction task, since roads usually only occupy a small part of the remote sensing image, there is a great imbalance between positive and negative samples. Using the Dice loss function can alleviate the problem of sample imbalance and make the model pay more attention to the few positive samples. The Dice loss is calculated by taking the intersection of the predicted pixels and the labeled pixels and then dividing by their total pixels, considering all pixels of a category as a whole, thus being unaffected by a large number of background pixels. Its formula is expressed as: ; where, X represents the ground truth of the road label, Y represents the road label of the prediction result; however, although the Dice loss function is very fast in calculation, it may lead to instability and oscillation in gradient calculation during training. Therefore, the Binary Cross-Entropy (BCE) loss function is added on the basis of the Dice loss function to maintain the stability of model training. The cross-entropy loss function can be expressed as: ; where, N is the number of image pixels; y i represents the true label value of pixel i, taking 1 for the foreground and 0 for the background; 1 x i is the predicted output of pixel i after Sigmoid, with a value range of (0, 1).
[0038] Finally, the formula for the total loss function is: .
[0039] To more objectively evaluate the performance of different road extraction network models, this section conducts comparative experiments and ablation experiments on two datasets for the proposed model to verify the efficiency of the proposed model.
[0040] To test the performance of the DESSNet network, in this section, DESSNet is compared with the DBRANet, RCFSNet, TransRoadNet, and MSMDFF-Net models proposed in the past two years, as well as the classic remote sensing image road extraction model D-LinkNet on the Massachusetts Road and CHN6-CUG datasets. The numerical comparison results are shown in Table 1, and the bold fonts in the table are the optimal results. All models were experimented on the same dataset and experimental environment. As can be seen from the table, the proposed DESSNet in this paper achieved the optimal results on both datasets. In both datasets, the Recall, IoU, and F1 of the method in this paper are the highest. For example, on the CHN6-CUG dataset, the IoU value of the method in this paper is 62.55, which is at least 2.64% higher than other methods. On the Massachusetts Road dataset, the IoU value of the method in this paper is 65.00, which is at least 1.72% higher than other methods.
[0041] Table 1 Qualitative indicators of the experimental results of the dataset
[0042] Figure 5 The prediction maps of five groups of DESSNet and other comparison networks on the CHN6-CUG dataset are shown. It can be clearly seen from the figure that for the CHN6-CUG dataset, the proposed method DESSNet obtained the optimal road extraction results. The performance of the five comparison networks used for comparison is not ideal. For example, in the second row, there are obvious missed detection errors in the road extraction maps of D-LinkNet, MSMDFF-Net, RCFSNet, and DBRANet. Although TranRoadNet did not show obvious missed detection errors, there were many false detection errors. DESSNet did not show obvious missed detection and false detection errors. From the road extraction results in the first row, it can be seen that DESSNet extracts significantly more completely at the edge part than other comparison networks. When the road is blocked by surrounding objects and other networks cannot completely identify the blocked road, DESSNet can more completely identify the blocked road under the combined action of the hierarchical interaction module and the dynamic multi-scale context extraction module. From the fourth and fifth rows, it can be seen that when there are some small roads, the comparison networks often cannot accurately identify the roads and there are large-scale missed detection situations, but DESSNet can effectively solve this problem. It dynamically adjusts the receptive field through the dynamic multi-scale feature extraction module and makes full use of the useful information in the context to help determine the road contour, effectively solving the large-scale missed detection problem.
[0043] Figure 6Shows the prediction maps of five groups of DESSNet and other comparison networks on the Massachusetts Road dataset. It can be clearly seen from the figure that for the Massachusetts Road dataset, the proposed method DESSNet obtains the optimal road extraction results. For example Figure 6 In the first and third rows of Figure 6 , all five groups of methods missed the roads in the dashed boxes, while the proposed DESSNet extracted the corresponding roads more completely. It can be seen from the second row that the five groups of comparison methods have poor performance in extracting road edges and lack a lot of edge detail information, while the proposed DESSNet can extract the road boundaries more completely.
Claims
1. A network for detail enhancement and scale selection of road extraction for high-resolution remote sensing images, characterized by: The extraction network adopts the semantic segmentation neural network D-LinkNet, which is based on the encoding and decoding architecture, including an encoder based on a boundary-enhanced residual convolutional layer, a shallow content-guided spatial attention module SCGSM, a multi-scale context selection module MCSM and a decoder; the original image enters the encoder through the semantic segmentation neural network D-LinkNet and is processed by the boundary-enhanced residual convolutional layer, and then processed by the shallow content-guided spatial attention module SCGSM and the multi-scale context selection module MCSM, and then the road information is output by the decoder.
2. The detail enhancement and scale selection road extraction network for high-resolution remote sensing images according to claim 1, characterized in that: The decoder includes an edge enhancement residual block module BERCL, which adopts an unsupervised edge detection algorithm and introduces a Sobel edge detection operator; The shallow content guides the spatial attention module SCGSM to use shallow features to guide the deep feature information passed to the decoder by skip connection, and uses downsampling to reduce the dimension of shallow features. At the same time, the strip convolution module is used to extract strip features from deep features, and then the deep and shallow features are fused to obtain deep features guided by shallow content. The multi-scale context selection module MCSM weights and spatially fuses the features extracted by branches with different void ratios in the void convolution layer, dynamically adjusts the receptive field according to the different scales of the road, and selects context information.
3. The detail enhancement and scale selection road extraction network for high-resolution remote sensing images according to claim 2, characterized in that: The edge enhanced residual block module convolutional layer BERCL is used in the semantic segmentation neural network D-LinkNet, using ResNe34 as the decoder, and the edge enhanced residual block module BERB includes 4 layers.
4. The detail enhancement and scale selection road extraction network for high-resolution remote sensing images according to claim 3, characterized in that: The shallow content-guided spatial attention module SCGSM in the semantic segmentation neural network D-LinkNet includes two branches, a shallow feature branch and a deep feature branch. The input of the shallow feature branch is the output of the shallow content-guided spatial attention module of the previous layer, and the input of the deep feature branch is the output of the encoder of the current layer.
5. The detail enhancement and scale selection road extraction network for high-resolution remote sensing images according to claim 4, characterized in that: The shallow feature branch includes a detail-preserving downsampling module, through which the shallow feature map is downsampled to the same size as the deep feature map while retaining the detail features.
6. The detail enhancement and scale selection road extraction network for high-resolution remote sensing images according to claim 5, characterized in that: The multi-scale context selection module MCSM comprises two parts, a multi-scale feature extraction part and a scale iteration selection part; The multi-scale feature extraction part takes the features of the last layer encoder and the output of the last level interaction module as input, and then inputs it into four parallel branches with different expansion rates. Each branch has a 3×3 convolution, which captures feature information at different scales with expansion rates of 1, 2, 4, and 8 respectively. After four branches, four features of different scales are obtained ( F 1, F 2, F 3. F 4); The scale selection part is used to consider the feature correlation between adjacent scales.
7. The method for extracting a road network for detail enhancement and scale selection for high-resolution remote sensing images as described in claim 6 is characterized in that: The edge-enhanced residual block modular convolutional layer BERCL uses the original residual block RB to receive the input features. F in ,feature F in It goes through a 3×3 convolution, a BN layer, a Relu activation, a 3×3 convolution, a BN layer, and then adds it to the input through the identity mapping. F out Then, F out Perform Relu activation to get the final result. The specific operations are as follows: ;(1) Where Conv3×3 represents 3×3 convolution, BN represents batch normalization, RL represents Relu activation, Fin and Fout represent the input and output of the residual block respectively; The edge-enhanced residual convolutional layer module BERCL adds an edge enhancement branch to the original residual block RB. The edge enhancement branch uses the Sobel operator to extract boundary information and adds the extracted edge information to the output through the residual connection. The edge enhancement branch first performs channel dimension reduction through a convolution to generate a feature with a channel number of 1; secondly, the feature map is smoothed through an average pooling layer to reduce the impact of noise and downsampled; then it is input into a Sobel layer to enhance the edge details in the encoder feature; finally, the number of channels of the feature is restored to twice the original through a convolution module; the specific process is shown as follows: ;(2) in F in and F out denote the input and output of the boundary enhancement branch, respectively. Conv 1×1 represents 1×1 convolution, Avg 2×2 represents 2×2 average pooling, φ Represents an edge detection operation.
8. The method for detail enhancement and scale selection of roads for high-resolution remote sensing images according to claim 7, characterized in that: When the Sobel layer enhances the edge details in the encoder features, four convolutions in different directions are used to detect the gradient changes between road pixels and background pixels. The specific process of the discrete differential operator including four 3×3 convolutions is described as follows: ;(3) in, , , , , f is the input of the Sobel layer, F 0, F 1, F 2, F 3 represents the gradient value of the image in four different directions. F is the output of the Sobel layer; the output of the edge branch is finally added to the output of the original residual structure to obtain the output of the edge enhanced residual convolution layer.
9. The method for detail enhancement and scale selection of roads for high-resolution remote sensing images according to claim 8, characterized in that: The detail-preserving downsampling module uses a convolution branch and a pooling branch to perform downsampling operations. The convolution branch introduces a spatial depth conversion convolution Space-to-depth Conv, SPD-Conv; the feature map is spatially rearranged, and the spatial information in the 2×2 range is folded into 1-pixel channel information. Then, 1×1 convolution is used to perform channel compression and feature fusion, and a pooling branch is added. First, the features are subjected to maximum pooling and average pooling respectively, and a learnable parameter W is used to splice the two pooling results in proportion. Then, the number of channels is adjusted through a convolution to obtain the output of the pooling branch; finally, the outputs of the two branches are spliced to obtain the output of the shallow branch. The specific operation can be expressed by the following formula: ;(4) ;(5) ;(6) in, F S1 represents the output of the convolution branch, F S2 represents the output of the pooling branch, SPD represents the spatial depth transformation convolution, C represents a 3×3 convolution operation, Cat Represents a splicing operation, W is a learnable parameter, × represents the dot product operation, Max and Avg Represent the maximum pooling and average pooling operations respectively, E S ′ is the output of the shallow branch; The deep branch first uses the strip convolution module to extract strip features and enhance road features. The specific operation process of the strip convolution module can be shown as follows: ;(7) ;(8) ;(9) ;(10) In the formula, Conv 1×5 , Conv 5×1 , Conv 1×7 , Conv 7×1 , Conv 1×9 , Conv 9×1 , are convolution operations with kernel sizes of 1×5, 5×1, 1×7, 7×1, 1×9, and 9×1 respectively; F 1, F 2, F 3 are the outputs of the three branches respectively. F 4 is the output of the strip convolution module; after extracting the strip features through the strip convolution module, the maximum and average values of the channel dimensions are calculated respectively, the two pooled features are concatenated and then the channel is changed to 1 through 1×1 convolution, the attention map is obtained by the Sigmod function and the attention map is expanded to the same as the broadcast operation. E D ′ The same dimension; the specific operation can be expressed by the following formula: ; (11) In the formula, Max Indicates the operation of finding the maximum value of the channel dimension, Avg Indicates the operation of finding the average value of the channel dimension. Concat Represents a splicing operation, Conv represents 1×1 convolution, σ is the Sigmod function, E Indicates a broadcast operation. E D ′ is the output of the deep feature layer; finally, the outputs of the two branches are multiplied to obtain the correlation map of the deep and shallow features. The correlation map can be further used as an attention map to enhance the spatial information in the deep features. The specific operation can be expressed by the following formula: ; (12) The output of the shallow content-guided spatial attention module will be input into the next layer of shallow content-guided spatial attention module and the decoder of the corresponding level respectively.
10. The method for detail enhancement and scale selection of roads for high-resolution remote sensing images according to claim 9, characterized in that: In the multi-scale context selection module MCSM, the multi-scale feature extraction part obtains features of four different scales ( F 1, F 2, F 3. F 4) The feature maps of two adjacent scales are spliced. The spliced features are passed through two branches respectively. One branch first compresses the features spatially through a 1×1 convolution to compress the features into 1×H×W, and then passes through a Sigmod layer to obtain the weight of the features in the spatial dimension. Then, the weight map is used to re-weight the feature map spatially, as shown in the following formula: ; (13) ; (14) In the formula, Cat Represents a splicing operation, C Indicates a convolution with a convolution kernel of 1×1. Bn represents batch normalization, RL represents the Relu activation function, σ represents the Sigmod function, and × represents the dot multiplication operation; the other branch first undergoes a 3×3 convolution, and then performs a Softmax operation on the feature map to obtain two weight vectors W 1 and W 2 And weight the feature map to obtain the weighted feature map V ;Finally, add the outputs of the two branches to get F 12 ′ , the above operations are as follows: ;(15) ;(16) ; (17) where + and × represent element-by-element multiplication and addition, φ represents the Softmax activation function, E 12 ′ Represents the output of the scale selection part; F 23 ′ Depend on F 12 ′ and F 3 is obtained through the above operation, and F 34 ′ Depend on F 23 ′ and F 4 is obtained through the above operations; finally, the input feature map is added to the output of the scale selection part through the residual connection; the output of the multi-scale context selection module can be expressed as follows: ; (18) in, F out represents the output of the scale selection module, F in Represents the input of the scale selection module.
Citation Information
Cited By
Road extraction network construction method, road extraction method, equipment and medium
CN120997685A
A road extraction network construction method, a road extraction method, a device, and a medium
CN120997685B