Road extraction system and method for remote sensing image
The MSGLF-Net model addresses spatial information loss and redundancy in remote sensing images by integrating global and local features, achieving enhanced road extraction accuracy through a dual-branch encoder and cross-scale information flow with combined loss functions.
Patent Information
- Application Number
- CN202510154545.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-16
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, in the road extraction in remote sensing images, there are problems such as spatial information loss, insufficient global information utilization, and low extraction accuracy caused by low quality redundant features.
The MSGLF-Net model is adopted, and the road extraction model is optimized through the axial attention-convolution neural network encoder module, the cross-scale information flow module and the decoder module, combined with the combined loss function of Dice loss and Focal loss.
The accuracy of remote sensing image road extraction is improved, the utilization of global and local features is enhanced, spatial information redundancy is reduced, and the accuracy of road extraction is improved.
Smart Images

Figure CN120318668A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and in particular to a road extraction system and method for remote sensing images. Background Technique
[0002] The road extraction method based on remote sensing images plays a crucial role in fields such as autonomous driving, emergency response coordination, and smart city development. The roads in remote sensing images have complex tubular topologies and are surrounded by different urban features. This makes the roads present unique characteristics in remote sensing images, including large intra-class differences, a lot of background interference information, and being easily blocked by shadows, which brings challenges to the accurate extraction of roads.
[0003] With the rapid development of computer vision technology, semantic segmentation technology based on deep neural networks has become a key means for road extraction tasks. Existing road extraction methods usually focus on enhancing the model's ability to capture long-range dependencies to address technical challenges. The main technical means include dilated convolution, pooling, attention mechanisms, etc. However, a major challenge faced by these methods in integrating long-range information is the loss of spatial information. This loss of spatial information has a negative impact on pixel-level prediction tasks such as road extraction. To solve this problem, a common strategy is to make extensive use of low-level features in the decoding stage. However, this strategy will introduce spatial redundant information, thus affecting the accuracy of road extraction. Therefore, improving the modeling of long-range dependencies while enhancing spatial information, effectively utilizing global and local features, and flexibly capturing tubular road features in complex scenarios are crucial for improving road extraction performance. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a road extraction system and method for remote sensing images, which solves the problems of low road extraction accuracy caused by the loss of spatial information, insufficient utilization of global information, and low-quality redundant features, and can achieve an improvement in the accuracy of road extraction from remote sensing images.
[0005] The present invention solves its technical problems by adopting the following technical solutions:
[0006] A road extraction system for remote sensing images includes a remote sensing image acquisition module and a remote sensing image processing module, wherein the remote sensing image acquisition module is connected to the remote sensing image processing module. The remote sensing image acquisition module is used to acquire remote sensing images, and the remote sensing image processing module is used to input the remote sensing images into a remote sensing image road extraction model to obtain a road extraction result.
[0007] Moreover, the remote sensing image processing module predicts the road extraction result of the remote sensing image through the MSGLF-Net model, calculates the loss function combining the Dice loss and the Focal loss between the predicted road extraction result and the ground truth label, and performs backpropagation to complete end-to-end training, obtaining the MSGLF-Net model with the optimal road extraction performance. Finally, the optimal MSGLF-Net model is used for road extraction of the remote sensing image.
[0008] A method for extracting roads from a remote sensing image, comprising the following steps:
[0009] Step 1: Obtain a remote sensing image and its corresponding label through the remote sensing image acquisition module, and input the remote sensing image and the corresponding label into the remote sensing image processing module;
[0010] Step 2: The remote sensing image processing module constructs the MSGLF-Net model;
[0011] Step 3: The remote sensing image processing module trains the MSGLF-Net model with the remote sensing image to obtain the optimal MSGLF-Net model;
[0012] Step 4: The remote sensing image processing module processes the remote sensing image with the optimal MSGLF-Net model to obtain the road extraction result of the remote sensing image.
[0013] Moreover, the MSGLF-Net model in the step 2 includes: an axial attention-convolutional neural network encoder module, a cross-scale information flow module, a decoder module, and a classification module. Among them, the axial attention-convolutional neural network encoder module, the cross-scale information flow module, the decoder module, and the classification module are connected in sequence. The axial attention-convolutional neural network encoder module is used for feature extraction and downsampling of the remote sensing image, fusing the feature maps of corresponding scales of two branches to generate fused features of multiple scales; the cross-scale information flow module is used for resampling and fusing the feature maps of different scales from the axial attention-convolutional neural network encoder module to realize information interaction between different-scale features; the decoder module is used for step-by-step upsampling of the feature map output by the cross-scale information flow module, copying and fusing the outputs of different scales of the cross-scale information flow module through the skip connection layer, connecting with the same-scale features of the upsampling path, and jointly passing backward and performing upsampling processing until the original image size is restored; the classification module is used for determining the final pixel classification result.
[0014] Moreover, the axial attention-convolutional neural network encoder module includes an axial attention branch and a convolutional neural network branch in parallel; the axial attention branch is used for capturing global context and long-range dependencies, and the convolutional neural network branch is used for retaining local spatial details and structural information.
[0015] Moreover, the cross-scale information flow module includes a feature resampling stage and a feature fusion stage; in the feature resampling stage, for high-level features with a lower resolution, deconvolution is used to increase the spatial resolution, and then it is fused with the original high-resolution features to reduce the spatial information redundancy in the global features; for low-level features with a higher resolution, standard convolution is used to reduce the spatial resolution, and then it is fused with the original low-resolution features to supplement effective local details to the global features; in the feature fusion stage, the features with the same scale after resampling are fused by first connecting them by depth and then performing convolution to obtain enhanced multi-scale features.
[0016] Moreover, the specific implementation method of step 3 is as follows: use the MSGLF-Net model to calculate the loss function combined with the Dice loss and the Focal loss between the predicted road extraction result and the true label, complete the end-to-end training using backpropagation to obtain the optimal MSGLF-Net model, and use the obtained optimal MSGLF-Net model as the remote sensing image road extraction model to perform road extraction.
[0017] Moreover, the loss function ΔJL is:
[0018] ΔJL = L fl + L dice
[0019] where L fl is the Focal loss, and L dice is the Dice loss.
[0020] The advantages and positive effects of the present invention are:
[0021] 1. The present invention obtains a remote sensing image and its corresponding label through the remote sensing image acquisition module, and at the same time inputs the remote sensing image and its corresponding label into the remote sensing image processing module; the remote sensing image processing module constructs the MSGLF-Net model and trains the MSGLF-Net model through the remote sensing image to obtain the optimal MSGLF-Net model to process the remote sensing image and obtain the road extraction result of the remote sensing image. The present invention constructs the MSGLF-Net model through the axial attention-convolutional neural network encoder module, the cross-scale information flow module, the decoder module and the classification module to extract the roads in the remote sensing image, solves the problem of low road extraction accuracy caused by the loss of spatial information, insufficient utilization of global information and low-quality redundant features, and can improve the road extraction accuracy of the remote sensing image.
[0022] 2. In the process of constructing the MSGLF-Net model of the present invention, by using the axial attention-convolutional neural network encoder module, global and local information is fully captured; the axial attention branch captures global information and long-range dependencies, and the convolutional neural network branch extracts spatial detail features. The features of the two branches are fused to form complementary advantages and jointly generate high-quality feature expressions.
[0023] 3. In the process of constructing the MSGLF-Net model of the present invention, by designing a cross-scale information flow module between the encoder and the decoder, global information and local details are better balanced. Through multi-scale resampling, features at different scales are fully interacted, and then gradually fused during the upsampling process to better use local detail features to guide high-level feature localization while reducing spatial information redundancy in low-level features. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Schematic diagram of the MSGLF-Net model structure constructed for the present invention;
[0025] Figure 2 Schematic diagram of the cross-scale information flow module structure of the present invention;
[0026] Figure 3 Schematic diagram of the experimental results of the Massachusetts road dataset in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] The present invention will be further described in detail below with reference to the accompanying drawings.
[0028] A road extraction system for remote sensing images includes a remote sensing image acquisition module and a remote sensing image processing module. The remote sensing image acquisition module is connected to the remote sensing image processing module. The remote sensing image acquisition module is used to acquire remote sensing images, and the remote sensing image processing module is used to input the remote sensing images into the remote sensing image road extraction model to obtain road extraction results.
[0029] The remote sensing image processing module predicts the road extraction results of the remote sensing images through the MSGLF-Net model, calculates the loss function combined with the Dice loss and the Focal loss between the predicted road extraction results and the true labels, and performs backpropagation to complete end-to-end training to obtain the MSGLF-Net model with the optimal road extraction performance. Finally, the optimal MSGLF-Net model is used for road extraction of remote sensing images.
[0030] An extraction method for a road extraction system of remote sensing images includes the following steps:
[0031] Step 1: Acquire remote sensing images and corresponding labels through the remote sensing image acquisition module, and input the remote sensing images and corresponding labels into the remote sensing image processing module.
[0032] Step 2: The remote sensing image processing module constructs the MSGLF-Net model.
[0033] As Figure 1 shown, the constructed MSGLF-Net model includes: an axial attention-convolutional neural network encoder module, a cross-scale information flow module, a decoder module, and a classification module. Among them, the axial attention-convolutional neural network encoder module, the cross-scale information flow module, the decoder module, and the classification module are connected in sequence. The axial attention-convolutional neural network encoder module is used to extract features and downsample the remote sensing image, fuse the feature maps of corresponding scales of the two branches, and generate fused features of multiple scales; the cross-scale information flow module is used to resample and fuse the feature maps of different scales from the axial attention-convolutional neural network encoder module to achieve information interaction between different-scale features; the decoder module is used to upsample the feature maps output by the cross-scale information flow module step by step, copy and fuse the outputs of different scales of the cross-scale information flow module through skip connection layers, connect with the same-scale features of the upsampling path, and jointly pass backward and perform upsampling processing until the original image size is restored; the classification module is used to determine the final pixel classification result.
[0034] Although convolutional neural networks (CNNs) excel in capturing fine-grained spatial details, their limited receptive fields restrict their ability to effectively capture global context and long-range dependencies. The self-attention mechanism has advantages in integrating global information and capturing long-range dependencies but also faces problems of high computational complexity and loss of spatial information. To fully explore local details and global information, a dual-branch encoder consisting of a convolutional neural network branch and an axial attention branch is designed. The former retains local spatial details and structural information, while the latter captures global context and long-range dependencies.
[0035] The convolutional neural network branch is based on ResNet50, and the convolution is replaced with dynamic snake convolution (DSC) at the 4th layer to enhance the model's ability to flexibly capture tubular road features. The axial attention branch consists of a residual structure with axial attention modules, which is divided into four stages, and each stage includes 1, 2, 4, and 1 axial attention module respectively.
[0036] The reason for adopting the axial attention mechanism lies in its lower computational complexity. The axial attention mechanism decomposes self-attention into independent attention modules along the height and width dimensions, greatly reducing the computational complexity while maintaining the performance of the fully connected self-attention mechanism. In addition, a position bias parameter is introduced as position encoding to supplement the relative spatial relationship between pixels, thereby reducing the spatial structure distortion caused by sequential input. For any given feature map, the width axial attention mechanism with position encoding can be expressed as formula (1):
[0037]
[0038] where S represents the Sigmoid function, q ij , k ij , r ij refer to the query vector, key vector, and value vector at the spatial position (i, j), respectively. are learnable parameters. The calculation method of the attention mechanism along the height axis is similar to that of the width axis.
[0039] Features from the two encoder branches are fused in parallel at different scales through a feature fusion module based on channel and spatial attention. In the feature fusion module, the preliminarily fused features are enhanced through channel and spatial attention mechanisms respectively, so as to perform feature selection in the spatial and channel domains. The enhanced features in the two domains are then added together to generate the fused features from the dual branches.
[0040] As Figure 2 shown, the cross-scale information flow module includes a feature resampling stage and a feature fusion stage; in the feature resampling stage, for the high-level features with lower resolution, deconvolution is used to increase the spatial resolution, and then it is fused with the original high-resolution features to reduce the spatial information redundancy in the global features; for the low-level features with higher resolution, standard convolution is used to reduce the spatial resolution, and then it is fused with the original low-resolution features to supplement effective local details to the global features; in the feature fusion stage, the features of the same scale after resampling are fused in the way of first connecting by depth and then convolving to obtain enhanced multi-scale features.
[0041] If the given remote sensing image different stages of the encoder can provide four different scales of features for the cross-scale information flow module The cross-scale information flow module first resamples each to obtain to match the features of other scales, and the calculation formula is as shown in (2):
[0042]
[0043] where Conv represents the convolution operation, the number of output channels is 2 j-1 C, the convolution kernel size is 3×3, and the stride is 2 j-i ; Conv trans represents the deconvolution, the number of output channels is 2 j-1 C, the convolution kernel size is 3×3, and the stride is 2 i-j. Then, features with the same resolution are concatenated, and then the number of channels is reduced through a convolutional layer, and finally output features with the same size as the input features are obtained.
[0044] The decoder module is based on the decoder structure in UNet*, and the feature fusion method therein is changed from channel-wise concatenation to a feature fusion module based on channel and spatial attention.
[0045] Step 3: Use the MSGLF-Net model to calculate the loss function that combines the Dice loss and the Focal loss between the predicted road extraction result and the ground truth label, and complete end-to-end training using backpropagation to obtain the optimal MSGLF-Net model, and use the obtained optimal MSGLF-Net model as the remote sensing image road extraction model to extract roads.
[0046] The specific implementation method for optimizing the model is as follows: The Dice loss is obtained by calculating the intersection and union of the predicted value of the model and the ground truth label to get the Dice coefficient, and it is converted into the Dice loss value; the Focal loss first determines the class weight parameter through the proportion of each class in the samples; after calculating the cross-entropy loss between the predicted probability value of the model and the ground truth label, it is multiplied by the class weight parameter to get the Focal loss value; the Dice loss and the Focal loss are respectively multiplied by their corresponding weights, and the two are added to get the final loss value, which is used as the objective function for model optimization; the model parameters are updated through the backpropagation algorithm, so that the loss value gradually decreases, and the training ends when the specified number of iterations is reached, thereby obtaining the optimal MSGLF-Net model.
[0047] Among them, in order to make the direction of the loss function decrease more accurately during the model training process, a combined loss function ΔJL that combines the Focal loss and the Dice loss is used to optimize the model parameters:
[0048] ΔJL = L fl + L dice (3)
[0049] Among them, L fl is the Focal loss, and L dice is the Dice loss.
[0050] Step 4: The remote sensing image processing module uses the optimal MSGLF-Net model to process the remote sensing image to obtain the road extraction result of the remote sensing image.
[0051] According to the method of the above-mentioned road extraction system for remote sensing images, experiments were carried out on the Massachusetts road dataset, and the results of the present invention and other algorithms were compared to illustrate the effect of the present invention.
[0052] Experiments were conducted on the Massachusetts Road Dataset, and the performance of MSGLF-Net was compared with that of DeepLabv3+, D-LinkNet, NL-LinkNet, and MACU-Net. As shown in Table 1, MSGLF-Net achieved the highest F1 score (79.23%), exceeding DeepLabv3+ (76.97%), D-LinkNet (77.96%), and MACU-Net (77.38%). Although NL-LinkNet achieved an F1 score of 78.06%, it was still inferior to MSGLF-Net.
[0053] Table 1 Experimental Results on the Massachusetts Road Dataset
[0054]
[0055] As Figure 3 shown, the visualization results of each method on the Massachusetts Road Dataset are presented. The experimental results indicate that MSGLF-Net performs excellently in extracting complex tubular topologies. As shown in the second and third rows, it well enhances the connectivity of the road network extraction results. Additionally, as shown in the first and fourth rows, MSGLF-Net can also achieve more accurate extraction results under complex background interference. Overall, MSGLF-Net demonstrates a powerful road extraction ability in an environment with significant interference, especially in challenging extraction scenarios.
[0056] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes, but is not limited to, the embodiments described in the specific implementation manners. Any other implementation manners derived by those skilled in the art based on the technical solutions of the present invention also fall within the scope of protection of the present invention.
Claims
1. A road extraction system for remote sensing images, characterized in that: It includes a remote sensing image acquisition module and a remote sensing image processing module. The remote sensing image acquisition module is connected to the remote sensing image processing module. The remote sensing image acquisition module is used to acquire remote sensing images, and the remote sensing image processing module is used to input the remote sensing images into a remote sensing image road extraction model to obtain a road extraction result.
2. The road extraction system for remote sensing images according to claim 1, characterized in that: The remote sensing image processing module predicts the road extraction result of the remote sensing image through the MSGLF-Net model, calculates the loss function combining the Dice loss and the Focal loss between the predicted road extraction result and the ground truth label, and performs backpropagation to complete end-to-end training to obtain the MSGLF-Net model with the optimal road extraction performance. Finally, the optimal MSGLF-Net model is used for remote sensing image road extraction.
3. A method for extracting roads from a remote sensing image by the road extraction system according to any one of claims 1 to 2, characterized in that: It includes the following steps: Step 1: Acquire a remote sensing image and its corresponding label through the remote sensing image acquisition module, and input the remote sensing image and its corresponding label into the remote sensing image processing module; Step 2: The remote sensing image processing module constructs the MSGLF-Net model; Step 3: The remote sensing image processing module trains the MSGLF-Net model with the remote sensing image to obtain the optimal MSGLF-Net model; Step 4: The remote sensing image processing module processes the remote sensing image with the optimal MSGLF-Net model to obtain the road extraction result of the remote sensing image.
4. The extraction method of a road extraction system for remote sensing images according to claim 3, characterized in that: In the MSGLF-Net model in Step 2, it includes: an axial attention-convolutional neural network encoder module, a cross-scale information flow module, a decoder module, and a classification module. Among them, the axial attention-convolutional neural network encoder module, the cross-scale information flow module, the decoder module, and the classification module are connected in sequence. The axial attention-convolutional neural network encoder module is used to extract features and downsample the remote sensing image, fuse the feature maps of the corresponding scales of the two branches to generate fused features of multiple scales; the cross-scale information flow module is used to resample and fuse the feature maps of different scales from the axial attention-convolutional neural network encoder module to achieve information interaction between different-scale features; the decoder module is used to upsample the feature maps output by the cross-scale information flow module step by step, copy and fuse the outputs of different scales of the cross-scale information flow module through skip connection layers, connect with the same-scale features in the upsampling path, and jointly pass backward and perform upsampling processing until the original image size is restored; the classification module is used to determine the final pixel classification result.
5. The extraction method of a road extraction system for remote sensing images according to claim 4, characterized in that: The axial attention-convolutional neural network encoder module includes an axial attention branch and a convolutional neural network branch in parallel; the axial attention branch is used to capture global context and long-range dependencies, and the convolutional neural network branch is used to retain local spatial details and structural information.
6. The extraction method of a road extraction system for remote sensing images according to claim 4, characterized in that: The cross-scale information flow module includes a feature resampling stage and a feature fusion stage; In the feature resampling stage, for high-level features with lower resolution, deconvolution is used to increase the spatial resolution, and then it is fused with the original high-resolution features to reduce the redundancy of spatial information in the global features; for low-level features with higher resolution, standard convolution is used to reduce the spatial resolution, and then it is fused with the original low-resolution features to supplement effective local details to the global features; in the feature fusion stage, the features of the same scale after resampling are fused by first connecting them by depth and then performing convolution to obtain enhanced multi-scale features.
7. The extraction method of a road extraction system for remote sensing images according to claim 3, characterized in that: The specific implementation method of step 3 is as follows: Use the MSGLF-Net model to calculate the loss function combined with the Dice loss and the Focal loss between the predicted road extraction result and the ground truth label, and use backpropagation to complete end-to-end training to obtain the optimal MSGLF-Net model, and use the obtained optimal MSGLF-Net model as the remote sensing image road extraction model to extract roads.
8. The extraction method of a road extraction system for remote sensing images according to claim 7, characterized in that: The loss function ΔJL is: ΔJL = L fl +L dice Among them, L fl is the Focal loss, and L dice is the Dice loss.