Road extraction network and method based on multi-scale feature fusion and shallow feature guidance

By introducing a road extraction network with multi-scale multi-directional feature fusion and shallow feature guidance in deep learning methods, the problem that multi-scale and multi-directional features are not fully considered in road extraction is solved, and the accuracy and integrity of road extraction is improved.

CN120411784APending Publication Date: 2025-08-01CHINA THREE GORGES UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510548945.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing deep learning methods do not fully consider the multi-scale and multi-directional characteristics of the road in road extraction, resulting in insufficient extraction accuracy and susceptible to background noise interference.

Method used

D-LinkNet is adopted as the infrastructure, combining the multi-scale multi-directional feature fusion module MSMDFFM and shallow feature guidance attention module SFGAM, and enhance the reconstruction ability of road features through multi-directional convolution and attention mechanisms, and design the multi-directional feature enhancement module MDFEM to improve road extraction accuracy.

Benefits of technology

It improves the accuracy and integrity of road extraction, reduces background noise interference, enhances the ability to reconstruct detailed information, and achieves higher F1-Score and intercombination ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411784A_ABST
    Figure CN120411784A_ABST
Patent Text Reader

Abstract

The invention provides a road extraction network and method based on multi-scale feature fusion and shallow feature guidance, the road extraction network adopts a D-Link Net as an infrastructure, the road extraction network comprises a multi-scale multi-direction feature fusion module MSMDFFM, a multi-direction feature enhancement module MDFEM and a shallow feature guidance attention module SFGAM, and other structural parameters are consistent with the D-Link Net. Firstly, a multi-scale and multi-direction feature fusion module MSMDFFM is designed based on Gabor convolution, local pooling and global direction pooling, so that the characterization capability of the network for road features of different scales and different directions is improved. And then guiding deep features layer by layer by using spatial information of a shallow layer, and designing a shallow feature guidance attention module SFGAM by means of bar convolution, thereby sequentially enhancing the reconstruction capability of each level in the network for detail information. And finally, designing a multi-directional feature enhancement module MDFEM, and complementing the perception capability of road features in different directions by using bar convolution and Gabor convolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing technology, and particularly relates to a road extraction network and method based on multi-scale feature fusion and shallow feature guidance. Background Art

[0002] Given the great success of deep learning in the field of computer vision, it has received increasing attention in the field of remote sensing. Many researchers have applied the semantic segmentation technology in deep learning to road extraction from remote sensing images. The fully convolutional neural network (FullyConvolutional Network, FCN) designed in "Fully convolutional networks for semantic segmentation" and others has been applied to pixel-level semantic segmentation tasks, greatly promoting the rapid development of semantic segmentation technology. Aiming at the problem of limited receptive field of CNN, remote sensing scholars have introduced the attention mechanism into CNN road extraction to model global information of images and improve the performance of the road extraction network. "DDU-Net: Dual-decoder-u-net for road extraction using high-resolution remote sensing images" proposed the DDU-Net road extraction network by introducing a convolutional block attention module (CBAM) between the encoder and the decoder, realizing dual feature calibration in the spatial and channel dimensions. Recently, some remote sensing scholars have integrated Transformer into the direction of CNN road extraction. "RADANet: Road augmented deformable attention network for road extraction from complex high-resolution remote-sensing images" inserted the Transformer module between the encoder and the decoder, combined the Transformer module and deformable convolution to achieve collaborative learning of local geometric details and global context information, and proposed the road extraction model RADANet.

[0003] Although existing deep learning methods have made significant progress in road extraction accuracy, it remains an unsolved problem due to the numerous challenges faced in road extraction. The challenges include: 1) During the feature transfer process of the encoder, simple concatenation or addition methods are mostly used for feature fusion, or the attention mechanism that performs excellently in other tasks is directly transferred, without fully considering the unique properties of roads and the guidance or interaction between different hierarchical features; 2) Roads exhibit irregularities in images, such as changes in width, shape, and direction. Existing methods only address scale changes through multi-scale feature extraction, without fully considering the multi-directional features of roads. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a road extraction network and method based on multi-scale feature fusion and shallow feature guidance, to solve the deficiencies existing in the extraction of roads by existing remote sensing technologies.

[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0006] A road extraction network based on multi-scale feature fusion and shallow feature guidance, where the road extraction network uses D-LinkNet as the basic architecture, including an encoder, a multi-scale multi-directional feature fusion module MSMDFFM, and a decoding part; among them, ResNet34 is used as the backbone network of the encoder, and the input image passes through 3, 4, 6, and 3 residual blocks respectively to extract features of the image. At the same time, each encoding module downsamples the image by a factor of 2; in the MSMDFFM, linear roads are extracted by performing global average pooling in the horizontal and vertical directions, and multi-scale multi-directional roads are extracted through pooling at different scales and multi-directional convolutions; the decoding part includes a multi-directional feature enhancement module MDFEM and a shallow feature guidance attention module SFGAM. The MDFEM uses strip convolution and Gabor convolution to enhance the model's ability to reconstruct road features in different directions; the SFGAM uses strip convolution, attention mechanism, and shallow feature guidance strategy to enhance the network's ability to reconstruct detailed information.

[0007] The above-mentioned shallow feature guidance attention module SFGAM consists of two parts: channel attention and spatial attention.

[0008] In the channel attention part of the above-mentioned shallow feature guidance attention module SFGAM, first, the output feature E of the encoder at the previous level i-1 and the output feature D of the multi-directional feature enhancement module iPerform cascading, and perform global average pooling and global max pooling respectively to extract global information; subsequently, these two sets of features respectively pass through a 1×1 convolution to reduce the number of channels, the relu function, and a second 1×1 convolution to restore the number of channels again, and the two branches are added together and a 1×1 convolution is used to adjust the number of channels, and then pass through the Sigmoid activation to obtain the weight coefficient of the channel feature relationship; finally, multiply this weight by D i to obtain the enhanced feature F c .

[0009] In the spatial attention part of the above-mentioned shallow feature-guided attention module SFGAM, first add E i-1 and D i , perform cascading on the channels through average pooling and max pooling respectively, and extract the linear features of the road through multi-scale bar convolution to better fit the road morphology; finally, pass through a 1×1 convolution to adjust the feature distribution, and calculate the attention weight through Sigmoid, and finally multiply by D i to enhance the spatial feature expression. Finally, add the two attention enhancements and D i together to get the final output result C i-1 .

[0010] The above-mentioned multi-scale multi-directional feature fusion module MSMDFFM consists of two complementary parts. The first part models the global structure through global average pooling in different directions, and the other part performs global averaging at three different scales and uses 3×3 multi-directional convolutions respectively to obtain local detail features at different scales; finally, the module also combines an attention mechanism to enable the model to better select and weight important features.

[0011] The above-mentioned multi-directional feature enhancement module MDFEM extracts the road features in the horizontal and vertical directions through bar convolution and captures various direction information of the road using multi-directional convolution:

[0012] First, perform transposed convolution on the input feature map to increase the resolution, thereby enhancing the ability to retain spatial details during the decoding process; use 3×3 convolution to obtain preliminary road features, providing a basis for subsequent multi-directional feature learning; then input the result into three parallel branches. The three branches can complement each other in road extraction. Use 1×3 convolution, 3×1 convolution, and 3×3 multi-directional convolution to extract more refined road features such as horizontal and vertical bar roads and multi-directional roads in the feature map. This way can reconstruct road features more comprehensively, ensuring the integrity and detail retention of the extraction results; connect the results D1, D2, and D3 of these three parallel branches on the channel and use the ECA module to enhance the expression of relevant road features on the channel, while suppressing irrelevant features. Finally, use 1×1 convolution to reduce the number of channels and use it as the output result of MDFEM.

[0013] Using the above road extraction method based on the road extraction network of multi-scale feature fusion and shallow feature guidance, the mathematical expression of the channel attention part of the shallow feature guidance attention module SFGAM is as follows:

[0014]

[0015] In the formula: Cat is the concatenation operation; GAP is global average pooling; GMP is global max pooling; Conv is the convolution operation; γ is the Relu activation function; σ is the Sigmoid activation function; i is the matrix multiplication operation;

[0016] The mathematical expression of the spatial attention part is as follows:

[0017]

[0018] In the formula: Cat is the concatenation operation; MaxPool is max pooling; AvgPool is average pooling; Conv is the convolution operation; σ is the Sigmoid activation function; i is the matrix multiplication operation.

[0019] The specific processing process of the above multi-scale multi-directional feature fusion module MSMDFFM is as follows:

[0020] In order to obtain complex directional features in the road, multi-directional convolution is designed to capture multi-directional road information within a single receptive field; the multi-directional convolution consists of a group of Gabor filters, and its mathematical expression is as follows:

[0021]

[0022] In the formula, x and y represent the coordinates of pixels respectively; λ refers to the wavelength, and 1 / λ represents the spatial frequency of the cosine function; γ is the spatial aspect ratio; Phase offset; σ is the standard deviation; According to different θ, Gabor filters in different directions can be obtained, and then the Gabor filters in different directions and the convolution kernel are multiplied pixel by pixel to obtain the Gabor convolution kernel containing the direction. By using attention to learn the weight coefficient of each directional Gabor kernel, a convolution operator that is more suitable for the road direction is selected:

[0023]

[0024] Where: w i is the weight coefficient of attention module learning; K is the ordinary convolution kernel; G(θ i ) is the Gabor kernel; i is the pixel-wise product;

[0025] In the global linear branch, global average pooling is first used to perform direction perception in the horizontal and vertical directions to extract the linear structural features of the road; then, the extracted features are upsampled to restore to the original image size and feature cascade is performed; finally, 3×3 convolution is used to further enhance local spatial details, enabling the model to extract more refined road features and obtain the linear feature map E. x :

[0026]

[0027] Where: Cat is the cascade operation; Conv is the convolution operation; up is the bilinear interpolation;

[0028] In the local multi-scale multi-directional branch, the input feature map E is first i 1×1 convolution is performed to reduce the dimension to reduce the amount of calculation; then, three different scales of pooling and 3×3 multi-directional convolution are used to extract local information at different scales. This method can enhance the model's perception of complex road structures; on this basis, bilinear interpolation is performed on the feature maps of each branch to restore them to the original resolution, and the number of channels is adjusted through 1×1 convolution; finally, the processed feature map is added element by element with the input feature map to obtain the feature map E that integrates multi-scale and multi-directional information. s :

[0029]

[0030] Where: Cat is the cascade operation; AvgPool is the average pooling, i is the convolution kernel size, j is the stride; Conv is the convolution operation; up is the bilinear interpolation;

[0031] To further integrate information, a linear feature branch, a multi-scale multi-directional feature branch, and the input feature map are cascaded at the channel dimension to fuse the feature information from all branches to form the final feature map. In addition, to optimize the expression of channel information, an efficient channel attention (ECA) module is introduced. The formula is as follows:

[0032]

[0033] In the formula: Cat is the cascading operation; GAP is the global average pooling; Conv is the convolution operation; i is the matrix multiplication; γ is the Relu activation function; σ is the Sigmoid activation function.

[0034] The mathematical expression of the processing process of the above-mentioned multi-directional feature enhancement module MDFEM is:

[0035]

[0036] In the formula: Cat is the cascading operation; DCOA is the multi-directional convolution; GAP is the global average pooling; Conv is the convolution operation; γ is the Relu activation function.

[0037] A road extraction network and method based on multi-scale feature fusion and shallow feature guidance provided by the present invention first designs a multi-scale multi-directional feature fusion module, namely MSMDFFM (Multi-scale and multi-directional feature fusion module), based on Gabor convolution, local pooling, and global direction pooling, so as to improve the network's representation ability for road features of different scales and different directions. Then, the shallow spatial information is used to guide the deep features layer by layer, and a shallow feature guidance attention module, namely SFGAM (Shallow features guidance attention module), is designed by means of strip convolution, so as to enhance the reconstruction ability of each layer in the network for detail information in turn. Finally, a multi-directional feature enhancement module, namely MDFEM (Multi-directional feature enhancement module), is designed to utilize the complementary perception ability of strip convolution and Gabor convolution for road features in different directions. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention will be further described below in conjunction with the drawings and embodiments:

[0039] Figure 1 It is a schematic diagram of the road extraction network structure based on multi-scale feature fusion and shallow feature guidance of the present invention;

[0040] Figure 2Schematic diagram of the shallow feature-guided attention module SFGAM of the present invention;

[0041] Figure 3 Schematic diagram of the multi-scale multi-directional feature fusion module MSMDFFM of the present invention;

[0042] Figure 4 Schematic diagram of the multi-directional feature enhancement module MDFEM of the present invention;

[0043] Figure 5 Schematic diagram of the road extraction results of different methods in the CHN6-CUG dataset in the embodiments of the present invention;

[0044] Figure 6 Schematic diagram of the road extraction results of different methods in the Massachusetts dataset in the embodiments of the present invention. Detailed implementation manners

[0045] The technical solution of the present invention will be described in detail below with reference to the drawings and embodiments.

[0046] Embodiment 1:

[0047] A road extraction network and method based on multi-scale feature fusion and shallow feature guidance are as follows:

[0048] 1. Network structure

[0049] As Figure 1 shown, the proposed MSNet network uses D-LinkNet as the basic architecture. The main improvements include the multi-scale multi-directional feature fusion module MSMDFFM, the multi-directional feature enhancement module MDFEM, and the shallow feature-guided attention module SFGAM. The remaining structural parameters are the same as those of D-LinkNet. Section 1.1 gives the specific details of the SFGAM module; Section 1.2 details the MSMDFFM module; Section 1.3 introduces the MDFEM module.

[0050] 1.1 Shallow feature-guided attention module SFGAM

[0051] In the road extraction task, remote sensing images often contain a large amount of complex information, where road pixels only account for a small part of the overall pixels, and a large number of background pixels (such as buildings, trees, etc.) may become interference factors. Compared with road information, these redundant information and noises have a higher proportion, seriously interfering with the accuracy of the model. To solve this problem, the attention mechanism can effectively guide the model to focus on road features, thereby reducing the attention to background information. In addition, roads in remote sensing images usually present a slender strip-like structure, and conventional square convolution kernels are likely to introduce redundant information when extracting road features, while strip convolution kernels can more naturally adapt to this geometric characteristic. Finally, the way of directly adding or cascading during the encoder transmission process may cause background noise to be fused with road features together, thus suppressing the extraction of detailed information. The shallow feature guidance strategy effectively improves the ability to extract detailed information through layer-by-layer optimization and fusion. Based on the above analysis, the present invention combines strip convolution, attention mechanism and shallow feature guidance strategy, and proposes a Shallow Feature Guided Attention Module (SFGAM).

[0052] As Figure 2 shown, SFGAM consists of two parts: channel attention and spatial attention. The input of channel attention is concatenated by two groups of features in the channel dimension. This splicing method can retain more semantic information in the feature dimension. The concatenated features will be processed through a dual-branch structure of max pooling and average pooling respectively. This design can retain more detailed information more effectively compared with the traditional single-branch structure. The goal of spatial attention is to highlight key positions, and its addition operation can provide a more effective spatial feature representation. To better adapt to the geometric characteristics of roads, strip convolution kernels of different sizes (1×3, 3×1, 1×5, 5×1, 1×7, and 7×1) are used to construct a multi-scale feature extraction module, enabling the network to more naturally process the slender strip-like structure of roads.

[0053] During the channel enhancement process, first, E i-1 and D i are concatenated, and global average pooling and global max pooling are performed respectively to extract global information. Subsequently, these two groups of features respectively pass through a 1×1 convolution to reduce the number of channels, a relu function, and a second 1×1 convolution to restore the number of channels again. Then the two branches are added and a 1×1 convolution is used to adjust the number of channels, and then through Sigmoid activation to obtain the weight coefficient of the channel feature relationship. Finally, multiply this weight by D i to obtain the enhanced feature F c . This way can strengthen the network's attention to road-related channel features and effectively filter out interference information in the background. Its mathematical expression is as follows:

[0054]

[0055] In the formula: Cat is the concatenation operation; GAP is the global average pooling; GMP is the global maximum pooling; Conv is the convolution operation; γ is the Relu activation function; σ is the Sigmoid activation function; i is the matrix multiplication operation.

[0056] In the spatial enhancement process, first, the output feature E of the encoder in the previous layer i-1 and the output feature D of the multi-directional feature enhancement module i are added together, and then respectively passed through average pooling and maximum pooling and concatenated on the channel, and the linear features of the road are extracted through multi-scale strip convolution to more effectively fit the road morphology. Finally, after a 1×1 convolution to adjust the feature distribution, and the attention weights are calculated through Sigmoid, and finally multiplied by D i to strengthen the spatial feature expression. Finally, the two attention enhancements and D i are added together to obtain the final output result C i-1 . Its mathematical expression is as follows:

[0057]

[0058] In the formula: Cat is the concatenation operation; MaxPool is the maximum pooling; AvgPool is the average pooling; Conv is the convolution operation; σ is the Sigmoid activation function; i is the matrix multiplication operation.

[0059] 1.2 Multi-scale Multi-directional Feature Fusion Module MSMDFFM

[0060] In remote sensing images, roads usually exhibit multi-scale multi-directional features. However, existing methods only consider the multi-scale characteristics through cascaded dilated convolutions or spatial pyramid pooling, without considering the multi-directional characteristics. For this reason, the multi-scale multi-directional feature fusion module MSMDFFM is proposed. As Figure 3 shown, this module consists of two complementary parts. The first part models the global structure (such as the overall extension direction of the road) through global average pooling in different directions, and the other part uses global averaging at three different scales and respectively uses 3×3 multi-directional convolutions to obtain local detail features at different scales (such as boundaries, multi-directional intersections). Finally, the module also combines the attention mechanism, enabling the model to better select and weight important features.

[0061] Multi-directionality is an important feature of road extraction. In order to obtain the complex directional features in the road, the present invention designs multi-directional convolution to capture multi-directional road information within a single receptive field. The multi-directional convolution consists of a set of Gabor filters, which are widely used in 2D image processing. Its mathematical expression is as follows:

[0062]

[0063] where x and y represent the coordinates of the pixel respectively; λ refers to the wavelength, and 1 / λ represents the spatial frequency of the cosine function; γ is the spatial aspect ratio; phase shift; σ is the standard deviation; according to different θ, Gabor filters in different directions can be obtained, and then the Gabor filters in different directions and the convolution kernel are multiplied pixel by pixel to obtain a Gabor convolution kernel containing directions. By using attention to learn the weight coefficients of each directional Gabor kernel, a convolution operator more suitable for the road direction is selected:

[0064]

[0065] where: w i is the weight coefficient learned by the attention module; K is the ordinary convolution kernel; G(θ i ) is the Gabor kernel; i is the pixel-by-pixel multiplication.

[0066] As Figure 3 shown, in the global linear branch, first, global average pooling is adopted to perform direction perception along the horizontal and vertical directions respectively, and the linear structure features of the road are extracted. Subsequently, the extracted features are upsampled to restore to the original image size, and feature concatenation is performed. Finally, 3×3 convolution is used to further enhance the local spatial details, enabling the model to extract more refined road features and obtaining the linear feature map E x .

[0067]

[0068] where: Cat is the concatenation operation; Conv is the convolution operation; up is bilinear interpolation.

[0069] In the local multi-scale multi-direction branch, first, 1×1 convolution is performed on the input feature map E i to reduce the dimensionality and computational complexity. Subsequently, three different scales of pooling and 3×3 multi-directional convolution are used to extract local information at different scales, which can enhance the model's perception ability of complex road structures. On this basis, the feature maps of each branch are bilinearly interpolated to restore to the original resolution, and the number of channels is adjusted by 1×1 convolution. Finally, the processed feature maps are added element by element to the input feature map to obtain the feature map E s that fuses multi-scale and multi-direction information.

[0070]

[0071] where: Cat is the concatenation operation; AvgPool is average pooling, i is the convolution kernel size, j is the stride; Conv is the convolution operation; up is bilinear interpolation.

[0072] To further integrate information, a linear feature branch, a multi-scale multi-direction feature branch, and an input feature map are cascaded at the channel dimension to fuse the feature information from all branches and form the final feature map. In addition, to optimize the expression of channel information, an Efficient Channel Attention (ECA) module is introduced. The ECA module has demonstrated excellent feature calibration capabilities in multiple deep networks, can adaptively calculate the weights of each channel, enhance the expression of important features, and suppress irrelevant features simultaneously, thereby improving the model's feature extraction ability and overall task performance. The formula is as follows:

[0073]

[0074] In the formula: Cat is the concatenation operation; GAP is the global average pooling; Conv is the convolution operation; i is the matrix multiplication; γ is the Relu activation function; σ is the Sigmoid activation function.

[0075] 1.3 Multi-Directional Feature Enhancement Module (MDFEM)

[0076] In the road extraction task, the core function of the decoder is to gradually map low-resolution features back to high-resolution features to restore road information. However, most existing decoders do not fully consider the directional characteristics of roads. Therefore, the present invention proposes a Multi-Directional Feature Enhancement Module (MDFEM), which extracts horizontal and vertical road features through bar-shaped convolutions and captures various directional information of roads using multi-directional convolutions.

[0077] As Figure 4 shown, the specific process of MDFEM is as follows: First, transposed convolution is performed on the input feature map to increase the resolution, thereby enhancing the ability to retain spatial details during the decoding process. A 3×3 convolution is used to obtain preliminary road features, providing a basis for subsequent multi-directional feature learning. Then the result is input into three parallel branches. The three branches can complement each other in road extraction, and use 1×3 convolution, 3×1 convolution, and 3×3 multi-directional convolution respectively to extract more refined road features such as horizontal and vertical bar-shaped roads and multi-directional roads in the feature map. This way can more comprehensively reconstruct road features and ensure the integrity and detail retention of the extraction results. The results D1, D2, and D3 of these three parallel branches are concatenated on the channel dimension and the ECA module is used to enhance the expression of relevant road features in the channel and suppress irrelevant features simultaneously. Finally, a 1×1 convolution is used to reduce the number of channels and serve as the output result of MDFEM. The specific formula is as follows:

[0078]

[0079] Where: Cat is the cascading operation; DCOA is the multi-directional convolution; GAP is the global average pooling; Conv is the convolution operation; γ is the Relu activation function.

[0080] 2. Experimental Results and Analysis

[0081] 2.1 Experimental Data and Design

[0082] CHN6-CUG is a representative large-scale urban satellite remote sensing image road dataset in China, which contains six representative cities in China. The dataset includes 4,511 remote sensing images with a size of 512 pixels × 512 pixels and a spatial resolution of 0.5 m. 3,608 images are used for model training, and 903 images are used for testing and evaluation.

[0083] The Massachusetts road dataset covers various scenarios of remote sensing road datasets in urban, suburban, and rural areas. The dataset includes 1,171 remote sensing images with a size of 1,500 pixels × 1,500 pixels and a spatial resolution of 1 m. The original dataset contains 1,108 training images, 14 validation images, and 49 test images.

[0084] To reduce the training and testing time, the part without road markings in the dataset is removed. Before training the model, data augmentation operations are performed on the training set data, including random horizontal flipping, vertical flipping, rotation, and rotation scaling.

[0085] The Adam optimizer is used for training. The initial learning rate is set to 0.0002, and the batch size (Batchsize) is 4. If the loss function does not decrease for three consecutive epochs, the learning rate will decay to half of the original. The maximum epoch of the network is set to 200. If the loss for six consecutive epochs is greater than the current minimum loss or the learning rate is lower than 0.0000005, the training will stop.

[0086] To verify the effectiveness of the network proposed in the present invention, six road extraction methods such as CARNet, RCFSNet, MSMDFFNet, DBRANet, D-LinkNet, and TransRoadNet are selected for comparative experiments. Precision (P), recall (R), F1-Score (F1), and intersection over union (IoU) are used as evaluation indicators.

[0087] 2.2 Experimental Results and Analysis

[0088] Figure 5 and Figure 6Five groups of road extraction result maps on the CHN6-CUG and Massachusetts datasets are given respectively: (a) and (b) represent the original image and the ground truth label respectively; (c) to (i) represent the road extraction maps of CARNet, RCFSNet, MSMDFF Net, DBRANet, D-LinkNet, TransRoadNet, and the proposed method respectively. The black area is the road area, the white area is the non-road area, the red area is the missed detection area, and the blue area is the false detection area.

[0089] It can be seen from Figure 5 that for the CHN6-CUG dataset, the method of the present invention obtains the road extraction map closest to the ground truth label: FG-MDNet can detect roads of different sizes in the image relatively completely, and the boundaries are relatively accurate; it shows that MSNet has good extraction effects on different roads. However, the performance of the six comparison methods is not ideal: for example, in the box in the first row, there are obvious false detections in the six comparison methods, but the false detections of MSNet are relatively few, which proves its advantage in extraction performance. In the boxes in the second and fourth rows, in the face of complex road scenes, there are obvious breaks in the extraction of multi-directional roads by the other six methods, while the method of the present invention can effectively extract most of the roads with multi-directional characteristics. In the third row, RCFSNet and MSMDFFNet have obvious false detection errors, and the other four methods and the present invention obtain similar results. In the fifth row, there are certain degrees of false detections and missed detections in the road boundary areas of the other methods, while the extraction of the method of the present invention is relatively complete. This is mainly due to the fact that MS MDFFM introduces more context information and multi-directional characteristics in the feature learning process, enabling the network to better understand the relationship between the road and the surrounding background. It can be seen from Figure 6 that for the Massachusetts dataset, the method of the present invention obtains the road extraction map closest to the ground truth label: for example, in the first row, although all seven methods can extract the general outline of the road, compared with the six comparison methods, the extraction result of the method of the present invention is more complete and the boundary is more accurate. In the second row, MSMDFFNet and DBRANet have obvious false detection errors, RCFSNet and TransRoadNet miss small roads, and CARNet obtains similar results to the method of the present invention. In the third row, although the background changes greatly, the method of the present invention can still extract the vast majority of road areas completely. This is due to the fact that SFGAM effectively reduces noise interference and enhances the network's ability to reconstruct detailed information. In the fourth and fifth rows, in the face of road areas covered by vegetation or woods, the prediction results of the method of the present invention have better continuity. This is mainly due to MDFEM, which improves the model's ability to reconstruct the characteristics of roads in different directions and enhances the integrity of road extraction.

[0090] To more objectively evaluate the extraction performance of different road extraction methods, Table 1 presents the quantitative evaluation indicators of the road extraction results of different methods on the CHN6-CUG and Massachusetts datasets. As can be seen from Table 1, on both datasets, the method of the present invention, MSNet, obtains the optimal quantitative indicators for F1-Score and IoU on both datasets. For example, on the CHN6-CUG dataset, the F1-Score and IoU of the proposed method reach 74.26% and 62.04% respectively, which are at least 1.91% and 1.67% higher than those of other comparison methods. On the Massachusetts dataset, the F1-Score and IoU of the proposed method reach 77.73% and 64.97% respectively, which are at least 1.86% and 1.54% higher than those of other comparison methods.

[0091] Table 1 Quantitative analysis indicators of road extraction results of different methods

[0092]

[0093]

[0094] 2.3 Ablation experiments

[0095] To verify the effectiveness of the multi-directional feature fusion module MSMDFFM, the multi-directional feature enhancement module MDFEM, and the shallow feature-guided attention module SFGAM in the method of the present invention, ablation experiments on these three modules are carried out on the CHN6-CUG and Massachusetts datasets in this section. The experimental results are shown in Table 2: where "√" indicates that the module is retained, and "×" indicates that the module is removed.

[0096] From the experimental results of each row in Table 2, it can be seen that removing any module in the present invention will cause a significant decrease in the road extraction indicators, which further verifies the effectiveness of SFGAM, MSMDFFM, and MDFEM in improving the road extraction accuracy. Among the three groups of ablation experiments, when MDFEM is removed, the F1-Scores of the two datasets decrease by 0.68% and 0.52% respectively, and the IoUs decrease by 0.77% and 0.92% respectively, with the most significant decline. This indicates that MDFEM is the key module for improvement, which shows that the multi-directional feature enhancement of MDFEM plays a key role in the complete reconstruction of roads. When SFGAM is removed, the F1-Scores of the two datasets decrease by 0.44% and 0.35% respectively, and the IoUs decrease by 0.39% and 0.63% respectively. When MSMDFFM is removed, the F1-Scores of the two datasets decrease by 0.56% and 0.58% respectively, and the IoUs decrease by 0.71% and 0.98% respectively. It is proved that the SFGAM and MSMDFFM modules can also improve the model performance to a certain extent.

[0097] Table 2 Ablation experiment results of CHN6-CUG and Massachusetts datasets

[0098]

Claims

1. A road extraction network based on multi-scale feature fusion and shallow feature guidance, characterized in that The road extraction network uses D-LinkNet as the basic architecture, including an encoder, a multi-scale multi-direction feature fusion module (MSMDF FM), and a decoding part. Among them, ResNet34 is adopted as the backbone network of the encoder. The input image passes through 3, 4, 6, and 3 residual blocks respectively to extract features of the image. At the same time, each encoding module downsamples the image by a factor of 2. In the MSMDFFM, linear roads are extracted by global average pooling in the horizontal and vertical directions, and multi-scale multi-direction roads are extracted by pooling at different scales and multi-directional convolutions respectively. The decoding part includes a multi-direction feature enhancement module (MDFEM) and a shallow feature guided attention module (SFGAM). The MDFEM uses strip convolution and Gabor convolution to improve the model's ability to reconstruct road features in different directions. The SFGAM uses strip convolution, attention mechanism, and shallow feature guided strategy to enhance the network's ability to reconstruct detailed information.

2. The road extraction network based on multi-scale feature fusion and shallow feature guidance according to claim 1, characterized in that The shallow feature guided attention module (SFGAM) mentioned above consists of two parts: channel attention and spatial attention.

3. The road extraction network based on multi-scale feature fusion and shallow feature guidance according to claim 2, characterized in that, In the channel attention part of the shallow feature-guided attention module SFGAM, first, the output feature E of the encoder at the previous level i-1 and the output feature D of the multi-directional feature enhancement module i are concatenated, and global average pooling and global max pooling are respectively performed to extract global information; Subsequently, these two groups of features are respectively passed through a 1×1 convolution to reduce the number of channels, followed by a ReLU function and a second 1×1 convolution to restore the number of channels again. Then, the two branches are added together, and a 1×1 convolution is used to adjust the number of channels. After that, a Sigmoid activation is performed to obtain the weight coefficients of the channel feature relationship. Finally, this weight is multiplied by D i to obtain the enhanced feature F c .

4. The road extraction network based on multi-scale feature fusion and shallow feature guidance according to claim 3, wherein In the spatial attention part of the shallow feature-guided attention module SFGAM described above, first, E i-1 and D i are added together, and are respectively subjected to average pooling and max pooling and concatenated on the channels, and multi-scale bar convolution is used to extract the linear features of the road to more effectively fit the road morphology; finally, after a 1×1 convolution to adjust the feature distribution, and the attention weights are calculated through Sigmoid, and finally multiplied by D i to strengthen the spatial feature expression; Finally, add the two attentions enhanced and D i to obtain the final output result as C i-1 .

5. The road extraction network based on multi-scale feature fusion and shallow feature guidance according to claim 4, wherein The multi-scale multi-direction feature fusion module (MSMDFFM) consists of two complementary parts. The first part models the global structure through global average pooling in different directions, and the other part performs global averaging at three different scales and uses 3×3 multi-directional convolutions respectively to obtain local detail features at different scales. Finally, the module also combines the attention mechanism, enabling the model to better select and weight important features.

6. The road extraction network based on multi-scale feature fusion and shallow feature guidance according to claim 5, characterized in that, The multi-direction feature enhancement module (MDFEM) extracts road features in the horizontal and vertical directions through strip convolution and captures various direction information of the road using multi-directional convolution: First, transposed convolution is performed on the input feature map to increase the resolution, thereby enhancing the ability to retain spatial details during the decoding process. 3×3 convolution is used to obtain preliminary road features, providing a basis for subsequent multi-direction feature learning. Then the result is input into three parallel branches. The three branches can complement each other in road extraction. 1×3 convolution, 3×1 convolution, and 3×3 multi-directional convolution are respectively used to extract more refined road features such as horizontal and vertical strip roads and multi-directional roads in the feature map. This method can reconstruct road features more comprehensively, ensuring the integrity and detail retention of the extraction result. The results D1, D2, and D3 of these three parallel branches are concatenated on the channel and the ECA module is used to enhance the expression of relevant road features in the channel, while suppressing irrelevant features. Finally, 1×1 convolution is used to reduce the number of channels and used as the output result of the MDFEM.

7. The road extraction method using the road extraction network based on multi-scale feature fusion and shallow feature guidance according to claim 6, characterized in that, The mathematical expression of the channel attention part of the shallow feature guided attention module (SFGAM) is: Where: Cat is the concatenation operation; GAP is the global average pooling; GMP is the global maximum pooling; Conv is the convolution operation; γ is the Relu activation function; σ is the Sigmoid activation function; i is the matrix multiplication operation; The mathematical expression of the spatial attention part is: Where: Cat is the concatenation operation; MaxPool is the max pooling; AvgPool is the average pooling; Conv is the convolution operation; σ is the Sigmoid activation function; i is the matrix multiplication operation.

8. The road extraction method based on multi-scale feature fusion and shallow feature guidance according to claim 7, wherein The specific processing procedure of the multi-scale multi-direction feature fusion module MSMDFFM is as follows: To obtain the complex directional features in the road, multi-directional convolution is designed to capture multi-directional road information within a single receptive field; the multi-directional convolution consists of a group of Gabor filters, and its mathematical expression is as follows: where x and y represent the coordinates of the pixel respectively; λ refers to the wavelength, and 1 / λ represents the spatial frequency of the cosine function; γ is the spatial aspect ratio; phase shift; σ is the standard deviation; according to different θ, Gabor filters in different directions can be obtained, and then the Gabor filters in different directions and the convolution kernel are multiplied pixel by pixel to obtain a Gabor convolution kernel containing directions. By using attention to learn the weight coefficients of each directional Gabor kernel, a convolution operator more suitable for the road direction is selected: Where: w i is the weight coefficient learned by the attention module; K is a common convolution kernel; G(θ i ) is a Gabor kernel; i is the per-pixel product; In the global linear branch, first, global average pooling is adopted to perform direction perception along the horizontal and vertical directions respectively to extract the linear structure features of the road. Subsequently, the extracted features are upsampled to restore to the original image size and feature concatenation is performed. Finally, 3×3 convolution is used to further enhance the local spatial details, enabling the model to extract more refined road features and obtaining the linear feature map E x : Where: Cat is the concatenation operation; Conv is the convolution operation; up is the bilinear interpolation; In the local multi-scale multi-directional branch, first, a 1×1 convolution is performed on the input feature map E i to reduce the dimensionality and computational complexity. Subsequently, pooling with three different scales and 3×3 multi-directional convolution are used to extract local information at different scales, which can enhance the model's perception ability of complex road structures. On this basis, bilinear interpolation is performed on the feature maps of each branch to restore them to the original resolution, and the number of channels is adjusted through 1×1 convolution. Finally, the processed feature maps are added element-wise to the input feature map to obtain the feature map E s that fuses multi-scale and multi-directional information: Where: Cat is the concatenation operation; AvgPool is the average pooling, i is the convolution kernel size, j is the stride; Conv is the convolution operation; up is the bilinear interpolation; To further integrate information, the linear feature branch, the multi-scale multi-direction feature branch, and the input feature map are concatenated in the channel dimension to fuse the feature information from all branches to form the final feature map; in addition, to optimize the expression of channel information, the efficient channel attention ECA module is introduced; the formula is as follows: Where: Cat is the concatenation operation; GAP is the global average pooling; Conv is the convolution operation; i is the matrix multiplication; γ is the Relu activation function; σ is the Sigmoid activation function.

9. The road extraction method based on multi-scale feature fusion and shallow feature guidance according to claim 8, characterized in that The mathematical expression of the processing procedure of the multi-direction feature enhancement module MDFEM is: Where: Cat is the concatenation operation; DCOA is the multi-directional convolution; GAP is the global average pooling; Conv is the convolution operation; γ is the Relu activation function.

Citation Information

Cited By

  • High-resolution remote sensing image road automatic extraction method and system based on multi-branch feature guidance and double-view collaborative decoding

    CN122116177A