Remote sensing image road extraction method based on perceptual fusion
Through the remote sensing image road extraction method based on perceptual fusion, the channel graph convolution and pixel edge sensing module are used, combined with Swin Transformer and LinkNet networks, the problem of insufficient global topology and local details of road extraction in remote sensing images is solved, and higher extraction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510649248.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-29
AI Technical Summary
The existing remote sensing image road extraction methods are difficult to capture the global topology and local details simultaneously in complex and changeable land environments, resulting in insufficient extraction accuracy and robustness, especially poor segmentation and identification of narrow roads.
Using a method based on perceptual fusion, a remote sensing image road extraction model for channel graph convolution and pixel edge perception is built. Combined with Swin Transformer and LinkNet networks, a topological consistency loss function that integrates boundary punishment is designed. The global and local feature extraction capabilities of the model are enhanced through sliding window self-attention and cross-window offset window strategies, and feature fusion is optimized through the channel graph convolution module and pixel edge perception module.
It significantly improves the accuracy and robustness of road extraction in remote sensing images, improves accuracy, recall, F1 score and cross-over ratio, and improves 1.41%, 2.66%, 3.44% and 3.46% respectively compared with the existing methods.
Smart Images

Figure CN120388289A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image road extraction, and particularly to a remote sensing image road extraction method based on perceptual fusion. Background Art
[0002] Remote sensing images contain rich environmental information, which is of great value in research fields such as target detection, scene classification, and semantic segmentation. At the same time, its application fields are extremely wide, covering multiple aspects such as urban planning, vehicle navigation, and land resource survey. With the continuous development of remote sensing technology, the acquisition of high-resolution remote sensing images has become more efficient, and the image quality has been significantly improved, providing a more comprehensive data basis for remote sensing image feature interpretation tasks such as road network recognition.
[0003] As one of the typical targets in remote sensing images, roads play an irreplaceable role in connecting different functional areas, promoting social development, and facilitating human activities, and are the key infrastructure supporting urban operation and people's daily life. The accurate extraction of road information not only significantly affects the accurate recognition of other features such as buildings and vehicles in downstream tasks, but is also a key technology in multiple fields such as autonomous driving, map drawing, and disaster emergency response.
[0004] However, roads themselves have complex topological structures and variable geometric features, and their surrounding environments are also very different, which makes it extremely challenging to accurately extract road information from remote sensing images. Although high-quality road label results can be obtained through field data collection and expert manual annotation, this method is not only time-consuming and laborious, but also costly, and it is difficult to support the construction of large-scale road networks. Although the traditional morphological and semi-automatic remote sensing image road extraction methods have alleviated this problem to a certain extent, these methods are often inefficient and the accuracy is difficult to guarantee, and they can no longer meet the modern requirements for efficiency and accuracy.
[0005] With the accumulation of massive high-resolution remote sensing image data and the rapid development of deep learning technology, remote sensing image road extraction methods based on deep neural networks have made significant progress on various remote sensing benchmark datasets and shown good performance. Nevertheless, there are still many problems and challenges in the current field of remote sensing image road extraction.
[0006] The complex and ever-changing terrain environment in remote sensing images, such as building shadows, tree occlusions, and changes in imaging angles, often interferes with road extraction, weakening road visibility and, in turn, affecting its connectivity and integrity. Existing segmentation models, such as UNet and DLinknet, are limited by fixed receptive fields and a single feature fusion mechanism, making it difficult to fully adapt to the changes in multi-level road features in complex backgrounds. In addition, since narrow roads account for a small proportion of pixels in remote sensing images and are often intertwined with surrounding terrain, existing methods find it difficult to accurately segment and identify them, affecting the accuracy and stability of the extraction results. Therefore, under the conditions of complex interference and multi-level roads, how to fully capture the global topological structure and local detail information of roads in remote sensing images remains a key issue that urgently needs to be broken through in the current field of remote sensing image analysis. Summary of the Invention
[0007] In order to overcome the shortcomings of low accuracy and robustness in road extraction in remote sensing images in the existing technology, the present invention provides a remote sensing image road extraction method based on perceptual fusion, which uses a large number of annotated remote sensing image road samples for training and testing to solve the problems of existing segmentation algorithms in maintaining global topology and insufficient local detail accuracy.
[0008] To achieve the aforementioned object of the invention, the technical solutions adopted by the present invention include:
[0009] A remote sensing image road extraction method based on perceptual fusion includes the following steps:
[0010] Step 1: Select finely labeled data from different geographical regions from public remote sensing benchmark datasets, perform ground truth label correction and calibration, and perform data augmentation and expansion on the dataset to ensure input data quality.
[0011] Step 2: Divide the data set after data augmentation into training set and test set with a division ratio of 3:1.
[0012] Step 3: Build a remote sensing image road extraction model based on channel graph convolution and pixel edge perception. The model is based on an encoder-decoder architecture and consists of three parts: an encoding network, a bridging network, and a decoding network. SwinTransformer is used as the feature extraction module. The input image is first divided into patches of fixed size through a patch segmentation-based method. Each patch is mapped to an embedded vector of fixed dimension and serves as input to the subsequent Swin Transformer model for processing. Channel graph convolution modules and pixel edge perception modules are designed in the bridging part and skip connection, respectively. In the main part of the decoder, the decoding structure of the LinkNet network is adopted.
[0013] Step 4: Use the created training set in batches with the Nvidia GPU to train the constructed network model. During the training process, save the network model with the minimum training loss, and optimize the model through the error backpropagation algorithm. During the training process, a topological consistency loss function incorporating boundary penalty is designed, which combines the topological sensitivity loss and the boundary penalty term to optimize the topological structure consistency of road pixels and the uncertainty of the boundary region respectively.
[0014] Step 5: Use the model saved during the training process to test the high-resolution remote sensing images in the test set, and then the result images of road extraction from the remote sensing images can be obtained.
[0015] In the Swin Transformer in Step 3, the core operations include the sliding window self-attention mechanism and the cross-window shifted window strategy, which jointly act on extracting local and global features. The local information in the image is obtained through the sliding window self-attention calculation, while the cross-window shifted window strategy can further enhance the global perception ability of the model to ensure effective interaction of information between different windows. Each Transformer block consists of a self-attention layer and a feed-forward neural network, and extracts the high-order semantic information of the image layer by layer. After being processed by multiple Transformer blocks, the obtained feature map is used for the subsequent remote sensing road extraction task.
[0016] In the channel graph convolution module in Step 3, based on the idea of graph neural network, the channels are regarded as nodes, and feature transfer is carried out by constructing a channel relationship graph, which is divided into the following steps:
[0017] Step 3-2-1) Perform multi-scale feature extraction: The input feature map is evenly divided into 4-8 groups along the channel dimension, and each group uses 3×3 depthwise separable convolutions with different dilation rates (1, 3, 5) for feature extraction. The convolution results of each group are then concatenated to form an enhanced feature map.
[0018] Step 3-2-2) Construct a dynamic graph structure: Generate channel descriptors through a combination operation of global average pooling and max pooling. Based on the channel descriptors, use a two-layer fully connected network to generate a dynamic adjacency matrix, and perform dot product fusion with a preset static adjacency matrix. The static adjacency matrix is initialized according to the channel grouping principle to ensure full connection of channels within the group.
[0019] Step 3-2-3) Perform graph convolution operations: After flattening the feature map, perform convolution operations in the spectral domain, calculate the graph Laplacian matrix and perform spectral decomposition, and then filter the features in the Fourier domain.
[0020] Step 3-2-4) Feature fusion is achieved through residual connection: A learnable gating mechanism is designed to dynamically adjust the fusion ratio of the original features and the enhanced features. The gating is implemented by a 1×1 convolution and a sigmoid activation function, ensuring that the model can adaptively determine how much original feature information to retain.
[0021] In step 3, the pixel edge perception module is embedded at each skip connection and includes the following steps:
[0022] Step 3-3-1) Edge feature generation stage: Align the low-level features and high-level features from the encoder in terms of channels, perform local interaction on the concatenated features through a 3×3 convolution, and then reduce the dimension through a 1×1 convolution and apply sigmoid activation to generate an edge-enhanced feature map;
[0023] Step 3-3-2) Pixel-level attention calculation stage: Perform batch normalization on the main path features and the edge features respectively to ensure the consistency of the feature scales, calculate the similarity matrix between the two by multiplying pixel by pixel, and after sigmoid activation, the value of this matrix ranges from 0 to 1, representing the degree of dependence of each position on the edge features;
[0024] Step 3-3-3) Adaptive feature fusion stage: Dynamically weight and fuse the semantic features and the edge features according to the similarity matrix. In key regions such as road boundaries, the edge features dominate to ensure positioning accuracy; in uniform regions such as inside the road, the semantic features dominate to ensure consistency.
[0025] The formula of the topological consistency loss function for fusing boundary penalty in step 4 is as follows:
[0026] L total =L topo +λL bdry
[0027] Adjust the influence of the topology-sensitive loss term and the boundary penalty term on the model training through the trade-off coefficient λ.
[0028] The formula of the topology-sensitive loss term is as follows:
[0029]
[0030] Where Ω is the image spatial domain, is the road region indicator function, M dilate (x) is the binary mask after morphological dilation of the true road mask, d iter is the adaptive dilation radius, which is dynamically adjusted according to the road width. M bg (x)=1 - M road (x) is the indicator function of the background region. The road pixel weight term is defined as:
[0031]
[0032] Among them, d1(x) represents the Euclidean distance between pixel x and the center of the nearest road pixel, and w0 is the boundary influence factor. is the distance threshold. When the distance exceeds this value, the weight decays to nearly 0; when pixel x is closer to the center line, the weight is higher, and when it is far from the road, the weight decays in a sigmoid manner, thus retaining the main topological structure; the weight term of the background pixel is defined as:
[0033]
[0034] The said d nsp (x) represents the distance from pixel x to the nearest non-structural key area; for pixels adjacent to non-road structures, d nsp (x) is smaller and the weight is reduced, thereby suppressing overfitting of the model in these confusing areas. The said is the normalization constant to prevent the denominator from being too large.
[0035] The formula of the said boundary penalty term is defined as follows:
[0036]
[0037] The said α and β are class weights. A differentiable soft mask mechanism is adopted for correction to adapt to different types of error distributions:
[0038]
[0039] The said y1(x) and y0(x) represent the binary masks of the road and the background in the ground truth (GT); p1(x) and p0(x) represent the probabilities that the model predicts pixel x belongs to the road and the background; η is the confidence adjustment factor, which controls the reward intensity for high-confidence predictions.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] The road extraction method for remote sensing images based on perceptual fusion of the present invention can simultaneously optimize the extraction of the global topological structure and local detail features, significantly improving the accuracy and robustness of road extraction in remote sensing images. At the same time, compared with the existing general semantic segmentation method such as DLinkNet, the precision, recall rate, F1 score, and intersection over union of the four evaluation indicators are respectively improved by 1.41%, 2.66%, 3.44%, and 3.46%. Description of the Drawings
[0042] Figure 1In (a) is the structural diagram of the remote sensing image road extraction network RCGPNet based on channel graph convolution and pixel edge perception proposed by the present invention; (b) is a further illustration of the specific structure of the Decoder block in (a).
[0043] Figure 2 Is the structural diagram of the channel graph convolution module CGCM proposed by the present invention;
[0044] Figure 3 Is the structural diagram of the pixel edge perception module PEAM proposed by the present invention;
[0045] Figure 4 Is the display diagram of the remote sensing image road extraction result corresponding to the remote sensing image road extraction method based on perception fusion of the present invention. Among them, the figures a, b, and c are the original remote sensing images, and the figures a1, b1, and c1 are the road extraction result images respectively;
[0046] Figure 5 Is the display diagram of the remote sensing image road extraction result corresponding to the existing method. Among them, the figures d, e, and f are the original remote sensing images, and the figures d1, e1, and f1 are the road extraction result images respectively. Specific Embodiments
[0047] In the following, the embodiments are exemplary descriptions of the main experimental evidences, rather than limiting the core content and application scope of the present invention disclosed by the quantity of evidences. It should be noted that in all these drawings and corresponding descriptions, only the concepts, principles and representative experimental evidences of the disclosed embodiments of the present invention are shown exemplarily. When the evidence chain is complete, it is not necessary to show all the specific detailed information and extended details of each embodiment listed in the present invention.
[0048] Unless otherwise defined, the technical terms used in the following embodiments have the same meanings as commonly understood by those skilled in the art to which the present invention belongs.
[0049] The main processing steps of the present invention include: data acquisition and errata, data augmentation and enhancement, construction of training set and test set, building of network model, model training, and testing of remote sensing images using the trained network model.
[0050] Step 1: Data acquisition and errata.
[0051] Select data with different geographical regions and fine annotations from the remote sensing public benchmark dataset, perform ground truth label errata and calibration, and perform data enhancement and augmentation on the dataset to ensure the quality of the input data.
[0052] Step 2: Divide the dataset after data enhancement into a training set and a test set, and the division ratio is 3:1.
[0053] Step 3: Build a remote sensing image road extraction model based on perception fusion.
[0054] As Figure 1 shown, this model is based on an encoder-decoder architecture and consists of three parts: an encoding network, a bridging network, and a decoding network. Swin Transformer is used as the feature extraction module (i.e., the encoder). The input image is first divided into small patches of a fixed size through a method based on Patch segmentation. Each patch is mapped to an embedding vector of a fixed dimension and used as input to the subsequent Swin Transformer model for processing. A Channel-wise Graph Convolution Module (CGCM) and a Pixel Edge-Aware Module (PEAM) are designed in the bridging part and skip connections respectively. In the main part of the decoder, the decoding structure of the LinkNet network is adopted.
[0055] 3-1) In Swin Transformer, the core operations include the sliding window self-attention mechanism and the cross-window shifted window strategy, which work together to extract local and global features. The local information in the image is obtained through sliding window self-attention calculation, and the cross-window shifted window strategy (ShiftedWindow) can further enhance the global perception ability of the model, ensuring effective interaction of information between different windows. Each Transformer block consists of a self-attention layer and a multilayer perceptron (MLP), and extracts the high-order semantic information of the image layer by layer. After being processed by multiple Transformer blocks, the obtained feature map can be used for subsequent remote sensing road extraction tasks.
[0056] 3-2) The Channel-wise Graph Convolution Module (CGCM) is designed in the bridging part to further enhance the feature expression ability and improve the perception ability of thin and weak road boundaries, so as to avoid missing local details mainly composed of narrow roads while effectively maintaining the global topology of the road, as shown in the appendix Figure 2 shown for capturing long-range dependencies in the channel dimension. Based on the idea of graph neural networks, this module regards channels as nodes and conducts feature transmission by constructing a channel relationship graph. The specific implementation is divided into the following steps:
[0057] Step 3-2-1: Perform multi-scale feature extraction. The input feature map is evenly divided into 8 groups along the channel dimension, and each group uses 3×3 depthwise separable convolutions with different dilation rates (1, 3, 5) for feature extraction. This step can capture road features under different receptive fields simultaneously. Smaller dilation rates are suitable for extracting local details, while larger dilation rates help to obtain global context information. The convolution results of each group are then concatenated to form an enhanced feature map.
[0058] Step 3-2-2: Construct a dynamic graph structure. A channel descriptor is generated through a combined operation of global average pooling and max pooling, which can characterize the importance of each channel. Based on the channel descriptor, a dynamic adjacency matrix is generated using a two-layer fully connected network and then dot-multiplied and fused with a preset static adjacency matrix. The static adjacency matrix is initialized according to the channel grouping principle to ensure full connection within the group. This design not only retains prior knowledge but also maintains the flexibility of the model.
[0059] Step 3-2-3: Perform graph convolution operations. After flattening the feature map, convolution operations are performed in the spectral domain. First, the graph Laplacian matrix is calculated and spectrally decomposed, and then the features are filtered in the Fourier domain. This process can be regarded as establishing global correlations in the channel dimension, enabling semantically related channels to enhance each other.
[0060] Step 3-2-4: Achieve feature fusion through residual connections. A learnable gating mechanism is designed to dynamically adjust the fusion ratio of the original features and the enhanced features. The gating is implemented by a 1×1 convolution and a sigmoid activation function, ensuring that the model can adaptively decide how much original feature information to retain. This design is particularly beneficial for handling complex scenarios such as shadow occlusion, where more reliance on the original features can be placed when high-level features are unreliable.
[0061] 3-3) The Pixel Edge Awareness Module (PEAM) is embedded at each layer's skip connection as Figure 3 shown, and precisely perceives thin road edges through the following steps:
[0062] Step 3-3-1: In the edge feature generation stage, first, the low-level features and high-level features from the encoder are subjected to channel alignment processing. The low-level features contain rich edge details but more noise, while the high-level features have accurate semantic information but insufficient spatial details. Local interactions are performed on the concatenated features through 3×3 convolutions, followed by 1×1 convolution for dimensionality reduction and sigmoid activation to generate an edge-enhanced feature map. This step effectively highlights the responses in the road boundary area.
[0063] Step 3-3-2: Pixel-level attention calculation stage. Batch normalization is performed on the main path features and edge features respectively to ensure the consistency of feature scales. Then, the similarity matrix between the two is calculated by pixel-wise multiplication. After being activated by sigmoid, the matrix takes values between 0 and 1, indicating the dependence degree of each position on the edge features. Areas with high similarity are usually narrow roads or areas with blurred boundaries, which require stronger edge guidance.
[0064] Step 3-3-3: Adaptive feature fusion stage. The semantic features and edge features are dynamically weighted and fused according to the similarity matrix. In key areas such as road boundaries, the edge features obtain larger weights to ensure positioning accuracy; in uniform areas such as inside the roads, the semantic features dominate to ensure consistency. This adaptive fusion mechanism enables the model to automatically adjust the feature selection strategy according to local characteristics.
[0065] Step 4: Model training
[0066] Using an Nvidia GPU, the constructed network model is trained in batches using the created training set. During the training process, the network model with the smallest training loss is saved, and the model is continuously optimized through the error backpropagation algorithm.
[0067] During the training process, a topological consistency loss function incorporating boundary penalty is designed, which combines the topological sensitivity loss and the boundary penalty term, and optimizes the topological structure consistency of road pixels and the uncertainty of boundary regions respectively. The influence of the two parts of the loss on model training is adjusted by the trade-off coefficient λ:
[0068] L total =L topo +λL bdry
[0069] Among them, the formula for the topological sensitivity loss term is as follows:
[0070]
[0071] Among them, Ω is the image spatial domain, is the road area indicator function; M dilate (x) is a binary mask obtained by morphological dilation of the true road mask, and d iter is the adaptive dilation radius, which is dynamically adjusted according to the road width. M bg (x)=1 - M road (x) is the indicator function of the background area. The road pixel weight term is defined as:
[0072]
[0073] Among them, d1(x) represents the Euclidean distance between pixel x and the center of the nearest road pixel, and w0 is the boundary influence factor. is the distance threshold, and when exceeding this value, the weight decays to close to 0. When pixel x is closer to the centerline, the weight is higher, and when it is far from the road, the weight decays in a sigmoid manner, thereby retaining the main topological structure. The weight term of the background pixel is defined as:
[0074]
[0075] where d nsp (x) represents the distance from pixel x to the nearest non-structural key area (such as a building). is the normalization constant to prevent the denominator from being too large. For pixels near non-road structures, d nsp (x) is smaller and the weight decreases, thereby suppressing overfitting of the model in these confusing areas. Through the above definition, this loss term can effectively enhance the model's learning ability of the road topological structure, reduce breakpoints and false detections, and improve the overall connectivity.
[0076] The formula of the boundary penalty term is defined as follows:
[0077]
[0078] where α and β are class weights, is corrected using a differentiable soft mask mechanism to adapt to different types of error distributions:
[0079]
[0080] where y1(x) and y0(x) represent the binary masks of the road and the background in the ground truth (GT); p1(x) and p0(x) represent the probabilities that the model predicts pixel x belongs to the road and the background; η is the confidence adjustment factor, which controls the reward intensity for high-confidence predictions.
[0081] This penalty term ensures the stability of the model in the boundary region, reduces topological structure breaks, and improves the global connectivity and local accuracy of the prediction results.
[0082] Regarding the setting of training parameters, the Adam optimizer is used, and its momentum parameter is set to the industry standard 0.9. The initial learning rate is set to 2e-4, and the learning rate change strategy adopts a multi-interval decay strategy, decaying to 0.2 times the original learning rate at epochs {90, 110, 130, and 160}. This setting helps the model converge quickly while avoiding falling into local optima. The batch size is set to 8, and the data parallel strategy can be adopted to accelerate training in a multi-GPU environment. To stabilize the training process, a linear learning rate warm-up strategy is used in the first 5 epochs, and a total of 200 epochs are trained.
[0083] Step 5: Test the test set
[0084] Use the model saved during the training process to test the test set, and the corresponding remote sensing image road extraction result picture can be obtained.
[0085] As Figure 4 shown, it is the remote sensing image road extraction result corresponding to the remote sensing image road extraction method based on perceptual fusion of the present invention. As Figure 5 shown, it is the remote sensing image road extraction result corresponding to the existing method, which can intuitively show the advantages of the present invention in improving the global connectivity and local accuracy of remote sensing image road extraction.
[0086] Table 1
[0087]
[0088] As shown in Table 1, it is the quantitative index result, from which it can be directly seen from the data that the present invention has superior performance in specific evaluation indexes. In the four evaluation indexes of precision, recall rate, F1 score and intersection over union, compared with the existing semantic segmentation method UNet, it is improved by 7.2%, 2.16%, 6.23% and 7.42% respectively; compared with the existing general semantic segmentation method DLinkNet, it is improved by 1.41%, 2.66%, 3.44% and 3.46% respectively.
[0089] It should be understood that the above embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several modifications or substitutions can be made, and all of them should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for extracting roads from remote sensing images based on perceptual fusion, characterized in that, It includes the following steps: Step 1: Select data with fine annotations from different geographical regions in the remote sensing public benchmark dataset, perform truth label correction and calibration, and perform data augmentation and expansion on the dataset to ensure the quality of the input data; Step 2: Divide the dataset after data augmentation into a training set and a test set, and the division ratio is 3:1; Step 3: Build a remote sensing image road extraction model based on channel graph convolution and pixel edge perception. The model is based on an encoder-decoder architecture and consists of three parts: an encoding network, a bridging network, and a decoding network; Use SwinTransformer as the feature extraction module. The input image is first divided into fixed-size patches through a Patch-based segmentation method. Each patch will be mapped to an embedding vector of a fixed dimension and used as input to enter the subsequent Swin Transformer model for processing; Channel graph convolution modules and pixel edge perception modules are designed in the bridging part and skip connections respectively; In the main part of the decoder, the decoding structure of the LinkNet network is adopted; Step 4: Use the Nvidia GPU to use the created training set in batches to train the constructed network model. During the training process, save the network model with the smallest training loss, and optimize the model through the error backpropagation algorithm; A topological consistency loss function that combines boundary penalty is designed during the training process, which combines topological sensitive loss and boundary penalty terms, and optimizes the topological structure consistency of road pixels and the uncertainty of the boundary region respectively; Step 5: Use the model saved during the training process to test the high-resolution remote sensing images in the test set, and the result images of remote sensing image road extraction can be obtained.
2. The method for extracting roads from remote sensing images based on perception fusion according to claim 1, wherein In the Swin Transformer in Step 3, the core operations include the sliding window self-attention mechanism and the cross-window shifted window strategy, which work together to extract local and global features; The local information in the image is obtained through sliding window self-attention calculation, and the cross-window shifted window strategy can further enhance the global perception ability of the model to ensure effective interaction of information between different windows; Each Transformer block consists of a self-attention layer and a feed-forward neural network, and extracts the high-order semantic information of the image layer by layer; After being processed by multiple Transformer blocks, the obtained feature maps are used for subsequent remote sensing road extraction tasks.
3. The method for extracting roads from remote sensing images based on perception fusion according to claim 1, characterized in that The channel graph convolution module in Step 3 is based on the idea of graph neural networks, regards channels as nodes, and performs feature transfer by constructing a channel relationship graph, which is divided into the following steps: Step 3-2-1) Perform multi-scale feature extraction: Divide the input feature map evenly into 4-8 groups according to the channel dimension, and each group uses 3×3 depthwise separable convolutions with different dilation rates (1, 3, 5) for feature extraction; The convolution results of each group are then concatenated to form an enhanced feature map; Step 3-2-2) Construct a dynamic graph structure: Generate channel descriptors through a combined operation of global average pooling and max pooling. Based on the channel descriptors, use a two-layer fully connected network to generate a dynamic adjacency matrix, and perform dot product fusion with a preset static adjacency matrix; The static adjacency matrix is initialized according to the channel grouping principle to ensure full connection of channels within the group; Step 3-2-3) Perform graph convolution operations: After flattening the feature map, perform convolution operations in the spectral domain, calculate the graph Laplacian matrix and perform spectral decomposition, and then filter the features in the Fourier domain; Step 3-2-4) Achieve feature fusion through residual connections: Design a learnable gating mechanism to dynamically adjust the fusion ratio of the original features and the enhanced features. The gating is implemented by a 1×1 convolution and a sigmoid activation function to ensure that the model can adaptively determine how much original feature information to retain.
4. The method for extracting roads from remote sensing images based on perception fusion according to claim 1, wherein In step 3, the pixel edge perception module is embedded at each layer jump connection and includes the following steps: Step 3-3-1) Edge feature generation stage: Align the channels of the low-level features and high-level features from the encoder, perform local interaction on the concatenated features through a 3×3 convolution, and then reduce the dimension through a 1×1 convolution and sigmoid activation to generate an edge-enhanced feature map; Step 3-3-2) Pixel-level attention calculation stage: Perform batch normalization on the main path features and edge features respectively to ensure the consistency of the feature scales. Calculate the similarity matrix of the two by multiplying pixel by pixel. After sigmoid activation, the value of this matrix ranges from 0 to 1, indicating the degree of dependence of each position on the edge features; Step 3-3-3) Adaptive feature fusion stage: Dynamically weight and fuse semantic features and edge features according to the similarity matrix; In key areas such as road boundaries, edge features dominate to ensure positioning accuracy; In uniform areas such as inside the road, semantic features dominate to ensure consistency.
5. The method for extracting roads from remote sensing images based on perception fusion according to claim 1, characterized in that The formula for the topological consistency loss function that fuses boundary penalties in step 4 is as follows: L total = L topo + λL bdry Adjust the influence of the topology-sensitive loss term and the boundary penalty term on model training through the trade-off coefficient λ.
6. The method for extracting roads from remote sensing images based on perception fusion according to claim 5, wherein The formula for the topology-sensitive loss term is as follows: where Ω is the image spatial domain, is the road area indication function, M dilate (x) is the binary mask after morphological dilation of the true road mask, d iter is the adaptive dilation radius, dynamically adjusted according to the road width; M bg (x) = 1 - M road (x) is the indication function of the background area; the road pixel weight term is defined as: where d1(x) represents the Euclidean distance between pixel x and the center of the nearest road pixel, w0 is the boundary influence factor, is the distance threshold, and when the distance exceeds this value, the weight decays to near 0; when pixel x is closer to the centerline, the weight is higher, and when it is far from the road, the weight decays in a sigmoid manner, thus retaining the main topological structure; the weight term of the background pixel is defined as: The described d nsp (x) represents the distance from pixel x to the nearest unstructured critical area; for pixels adjacent to non-road structures, d nsp (x) has a reduced weight to suppress overfitting of the model in these confusing areas; the is a normalization constant to prevent the denominator from being too large.
7. The method for extracting roads from remote sensing images based on perception fusion according to claim 5, wherein The formula for the boundary penalty term is defined as follows: where α and β are class weights, which is corrected by a differentiable soft mask mechanism to adapt to different types of error distributions: The y1(x) and y0(x) represent the binary masks of the road and background in the ground truth (GT); p1(x) and p0(x) represent the probabilities that the model predicts that pixel x belongs to the road and background; η is the confidence adjustment factor that controls the reward intensity for high-confidence predictions.
Citation Information
Cited By
Unmanned aerial vehicle aerial image road detection method based on deep learning
CN121170654A
Shielding perception road intelligent extraction method
CN121236618A
An occlusion-aware road intelligent extraction method
CN121236618B
Remote sensing ground feature recognition method and system based on global and dynamic feature perception optimization
CN121392372A
A Remote Sensing Land Cover Identification Method and System Based on Global and Dynamic Feature Perception Optimization
CN121392372B