Remote sensing image road segmentation method combining road detail feature extraction and context information perception

By combining road detail feature extraction and context information perception, the problem of road extraction is easily disturbed and poor continuity in road segmentation in remote sensing image is solved, and more efficient road extraction and more complete road structure recovery are achieved.

CN120107575AActive Publication Date: 2025-06-06TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510029640.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing deep learning methods are susceptible to background, occlusion and road-like information in remote sensing image road segmentation, making it difficult to extract complete and continuous road features.

Method used

A remote sensing image road segmentation method combining road detail feature extraction and context information perception is proposed. By designing a road detail feature extraction module and context information perception module, the continuity and accuracy of road extraction are improved.

Benefits of technology

It effectively alleviates the loss of edge details information caused by continuous downsampling, improves the continuity and integrity of road extraction, and enhances the model's ability to recover road details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107575A_ABST
    Figure CN120107575A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image road segmentation method combining road detail feature extraction and context information perception, and belongs to the field of road segmentation. The problems of shielding and difficulty in distinguishing road levels existing in the current road segmentation network are solved; comprising the following steps of data set construction, model construction, model training and result prediction. A road detail feature extraction module and a context information sensing module are designed. The road detail feature extraction module effectively compensates the problem of loss of spatial information and edge details caused by continuous downsampling through interactive transmission of multi-level feature information, and enhances the understanding of the model for complex scenes and multi-level road structures. The context information sensing module realizes dynamic fusion of long-distance and local multi-scale context information, is helpful for suppressing background interference, shielding effect and confusion of similar road structures, and ensures the integrity and continuity of extracted roads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention provides a remote sensing image road segmentation method combining road detail feature extraction and context information perception, belonging to the technical field of remote sensing image road segmentation. Background Art

[0002] Semantic segmentation is a basic task in the field of computer vision. Its goal is to assign a semantic label category to each pixel in an image. It plays an important role in fields such as autonomous driving, medical images, and remote sensing images. With the development of society, extracting road information from remote sensing image data has gradually become a research hotspot. Road information can be used in many aspects of social life, such as urban planning, geographic information update, etc.

[0003] With the multi-source data and intelligent algorithms, the remote sensing image road extraction method based on deep learning has become the mainstream. Most deep learning methods have improved the ability to restore occluded roads in remote sensing images to varying degrees, but they still lose image details to a greater or lesser extent and ignore the inherent properties of the road. The reason is that most studies have failed to design an extraction module for road features when improving jump connections, and their flexibility is insufficient. In addition, these methods do not use an efficient attention mechanism when extracting context, resulting in certain limitations for the background image of the picture. Taking the above problems and reasons into account, the present invention proposes a remote sensing image road segmentation method that combines road detail feature extraction with context information perception, which improves the road extraction capability while increasing the continuity of the extracted roads. Summary of the invention

[0004] In order to solve the problem that the current road segmentation network is easily disturbed by background, occlusion and road-like information and difficult to extract complete and continuous roads, a remote sensing image road segmentation method combining road detail feature extraction and context information perception is proposed.

[0005] The technical solution adopted by the present invention is: a remote sensing image road segmentation method combining road detail feature extraction and context information perception, comprising the following steps:

[0006] Step 1: Construct a data set: collect high-resolution remote sensing images through remote sensing satellites, divide them into training sets and test sets in proportion, and perform preprocessing and data enhancement operations on all images in the training set. Both the training set and the test set contain original remote sensing images and professionally annotated remote sensing images;

[0007] Step 2: Constructing a remote sensing image road segmentation network model: including:

[0008] Step 2.1: Construction of feature encoding module;

[0009] Step 2.2: Construction of a road detail feature extraction module: The road detail feature extraction module includes a horizontal detail feature extraction module and a vertical detail feature extraction module for fusing multi-level strip features of the road. Both the horizontal detail feature extraction module and the vertical detail feature extraction module include multi-scale strip space attention and effective channel attention for space and channel dimension selection;

[0010] Step 2.3: Construction of context information perception module: The context information perception module extracts multi-scale detail features by designing a selective scanning block preceded by dilated convolution and changing the dilation rate of the dilated convolution in the selective scanning block, which is used for local multi-scale context extraction and global context extraction, and further features screening and fusion are performed in combination with effective channel attention;

[0011] Step 2.4: Construction of feature decoding module;

[0012] The feature encoding module and the feature decoding module realize the step-by-step fusion of features through the road detail feature extraction module; the fused features are extracted by the context information perception module for global and local multi-scale context features; the extracted results are used for decoding by the feature decoding module;

[0013] Step 3: Training remote sensing image road segmentation network model;

[0014] Step 4: Result prediction: Input the preprocessed remote sensing image into the trained remote sensing image road segmentation network model to obtain accurate road segmentation results.

[0015] Furthermore, the feature encoding module and the feature decoding module are each composed of five stages, the feature encoding module includes a downsampling module, encoder 1, encoder 2, encoder 3 and encoder 4 in sequence; the feature decoding module includes decoder 1, decoder 2, decoder 3, decoder 4 and a segmentation head in sequence.

[0016] Furthermore, the downsampling module is a 7×7 convolution operation with a step size of 2 and a 3×3 maximum pooling with a step size of 2; the four encoders are respectively stacked by 3, 4, 6, and 3 residual blocks; each residual block is composed of two 3×3 convolution operations and residual connections.

[0017] Furthermore, each of the four decoders includes a 1×1 convolution that reduces the channel to 1 / 4, a 3×3 transposed convolution with a stride of 2, and a 1×1 convolution that doubles the channel. All three convolutions include batch normalization and activation function ReLU; finally, the output of the decoder is converted into a road extraction result through a segmentation head, which consists of a 3×3 transposed convolution with a stride of 2, a 3×3 convolution, and a 2×2 convolution.

[0018] Furthermore, three road detail feature extraction modules are used to implement step-by-step fusion of the multi-level features output by the encoder. Each road detail feature extraction module receives input from two parts, one is the output of the feature map at the current scale, and the other is the result after feature extraction at the previous scale. These two parts are processed separately and weightedly added as the input of the road detail feature extraction module.

[0019] Furthermore, the lateral detail feature extraction module includes a strip space attention consisting of two lateral strip convolution branches and one convolution branch. The image input to the road detail feature extraction module will be processed by lateral strip convolution kernels with convolution kernels of 1×3 and 1×5, and then layer normalization will be performed; at the same time, the image input to the road detail feature extraction module is normalized by a 3×3 convolution kernel layer, and then average pooling and maximum pooling are performed respectively. The results are added after passing through the MLP layer, and then the added results are subjected to Sigmoid processing to obtain weighted values; the weighted values ​​obtained are then multiplied with the features obtained by the two lateral strip convolution branches, and then added and passed through the output mapping layer and a series of feedforward networks to obtain the final horizontal output, wherein the output mapping layer adopts effective channel attention.

[0020] Furthermore, the longitudinal detail feature extraction module includes two vertical strip convolution branches and one convolution branch. The image input to the road detail feature extraction module will be processed by vertical strip convolution kernels with convolution kernels of 3×1 and 5×1 respectively, and then layer normalization will be performed; at the same time, the image input to the road detail feature extraction module is normalized by the 3×3 convolution kernel layer, and then average pooling and maximum pooling are performed respectively. The results are added after passing through the MLP layer, and then the added results are subjected to Sigmoid processing to obtain weighted values; the weighted values ​​obtained are then multiplied with the features obtained by the two horizontal strip convolution branches, and then added and passed through the output mapping layer and a series of feedforward networks to obtain the final vertical output, wherein the output mapping layer adopts effective channel attention.

[0021] Furthermore, the input of the context information perception module is the result of weighted addition of the output of the encoder 4 and the output of the deepest road detail feature extraction module. The structure of the context information perception module is as follows: it is composed of a position embedding block, 4 groups of dilated convolutions with dilation rates of 1, 3, 5, and 7, 4 2D selective scanning blocks, and effective channel attention, wherein each group of dilated convolutions and 2D selective scanning blocks is followed by element-by-element multiplication of shared weights, and then the four groups of results are concatenated and outputted through effective channel attention.

[0022] Furthermore, the structure of the effective channel attention is as follows: it consists of 1 3×3 convolution, 1 GeLU activation function, 1 layer normalization and efficient channel attention ECA.

[0023] Furthermore, the training process of the model in step 3 is as follows:

[0024] The preprocessed training set data is input into the remote sensing image road segmentation network model. The feature encoding module is initialized using the ResNet-34 weights pre-trained on the ImageNet-1K dataset, and the remaining parameters are randomly initialized. The remote sensing image road segmentation network model is trained until the model converges, and the model parameters are saved after the training is completed.

[0025] The beneficial effects of the present invention compared with the prior art are as follows:

[0026] (1) The present invention uses the residual block of the ResNet-34 network in the encoder part and initializes it with the weights pre-trained on the ImageNet-1K dataset. This strategy not only effectively utilizes the existing knowledge and significantly improves the performance of the encoder, but also lays a solid foundation for the subsequent remote sensing image road segmentation task. This method enables the model to have a certain local feature extraction capability at the beginning of training, thereby accelerating the convergence speed of the model and improving the training efficiency.

[0027] (2) The present invention designs a road detail feature extraction module, which fuses multi-level features through horizontal extraction and vertical extraction, alleviates the problem of edge detail information loss caused by continuous downsampling, reduces the semantic gap between the encoder and the decoder, and makes the extracted road more continuous.

[0028] (3) The present invention designs a context information perception module, which effectively distinguishes road and background information by dynamically extracting and integrating global and local multi-scale context information, making the extracted road structure more complete and continuous.

[0029] (4) The present invention designs a strip detail feature extraction method, which includes two parts: horizontal and vertical feature extraction and channel selection; the horizontal and vertical feature extraction can extract the detail features of the road at the current stage; the channel selection selects the feature map obtained after the horizontal and vertical feature extraction in the channel dimension.

[0030] (5) The present invention designs a long-distance spatial selection, which includes a position embedding block, four groups of dilated convolutions with dilation rates of 1, 3, 5, and 7, four 2D selective scanning blocks, and channel attention. After each group of dilated convolutions and 2D selective scanning blocks, element-wise multiplication of shared weights is set, and then the four groups of results are concatenated to obtain output, thereby achieving spatial selection.

[0031] (6) In the feature decoding module, the present invention upsamples the abstract feature map of the context information perception module to the original image resolution through transposed convolution. Specifically, this process not only effectively restores the details and spatial resolution of the feature map, but also combines the output of the detail feature extraction module, significantly enhancing the retention of road spatial details and context information, and improving the model's recovery of road details. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The present invention will be further described below in conjunction with the accompanying drawings:

[0033] Figure 1 It is a flowchart of the remote sensing image road segmentation method proposed by the present invention;

[0034] Figure 2 It is a structural schematic diagram of the remote sensing image road segmentation network model proposed by the present invention;

[0035] Figure 3 It is a structural schematic diagram of the longitudinal detail feature extraction module proposed by the present invention;

[0036] Figure 4 It is a structural schematic diagram of the horizontal detail feature extraction module proposed by the present invention;

[0037] Figure 5 It is a schematic diagram of the structure of context information perception proposed by the present invention;

[0038] Figure 6 A schematic diagram of the structure of the effective channel attention proposed by the present invention;

[0039] Figure 7 is an example image of a remote sensing image of the DeepGlobe dataset used in an embodiment of the present invention;

[0040] Figure 8 is an example of a remote sensing image with labels in the DeepGlobe dataset used in an embodiment of the present invention;

[0041] Fig. 9 The method of the present invention is used to extract Figure 6 Schematic diagram of the road structure. DETAILED DESCRIPTION

[0042] like Figures 1 to 9 As shown, the present invention provides a remote sensing image road segmentation method combining road detail feature extraction and context information perception, comprising the following steps:

[0043] Step 1: Constructing a dataset: High-resolution remote sensing images of 1024×1024 pixels are collected by remote sensing satellites and divided into training and test sets at a ratio of 8:2. The dataset contains original remote sensing images and professionally annotated data. Subsequently, a series of preprocessing and data enhancement operations are performed on the images and labels of the training set, including horizontal, vertical and joint flipping, rotation, scaling, random noise injection and color adjustment of the images. These enhancement methods increase the diversity of image perspectives and improve the robustness of the model to road structures in different directions and conditions, thereby improving the accuracy and reliability of the remote sensing image road segmentation network. The processed images are used to train the network model to achieve accurate segmentation of road structures in remote sensing images, laying the foundation for subsequent analysis.

[0044] Step 2: Model construction: The remote sensing image road segmentation network model includes a feature encoding module (Encoder), a road details feature extraction module (Road Details Feature Extraction Module, RDFEM), a context-aware fusion module (Context-Aware Fusion Module, CAFM) and a feature decoding module (Decoder). Among them, the feature encoding module and the feature decoding module are each composed of five stages. The feature encoding module includes a downsampling module, encoder 1, encoder 2, encoder 3 and encoder 4 in sequence; the feature decoding module includes decoder 1, decoder 2, decoder 3, decoder 4 and segmentation head in sequence; the feature encoding module and the feature decoding module realize the step-by-step fusion of features through the road details feature extraction module; the fused features are extracted by the context-aware fusion module for global and local multi-scale context features; the extraction results are used for decoding by the feature decoding module.

[0045] The specific steps for building the remote sensing image road segmentation network model are as follows:

[0046] Step 2.1: Construction of feature encoding module: The ResNet-34 network pre-trained on the ImageNet-1K dataset is used to construct the downsampling module, encoder 1, encoder 2, encoder 3 and encoder 4 for feature extraction, and the feature maps of four stages are obtained, which will be used as the input of the bidirectional multi-level road feature dynamic fusion module for bidirectional dynamic fusion.

[0047] The structure of the feature encoding module is as follows:

[0048] The feature encoding module includes a downsampling module, encoder 1, encoder 2, encoder 3 and encoder 4 in sequence; the downsampling module is a 7×7 convolution operation with a stride of 2 and a 3×3 maximum pooling with a stride of 2; the four encoders are respectively stacked by 3, 4, 6, and 3 residual blocks (ResBlock); each residual block is composed of two 3×3 convolution operations and residual connections. The four encoders generate four feature maps respectively, each of which reflects the feature representation of the image at different levels of abstraction. Specifically, the feature map generated by encoder 1 is a shallow feature map, which mainly captures the basic and local detail information in the image; while the feature map generated by encoder 4 is a deep feature map, which contains more abstract and global information, which is particularly critical for performing fine semantic segmentation tasks.

[0049] Step 2.2: Construction of road detail feature extraction module: As the encoder continues to downsample, some image details will gradually be lost. Therefore, it is very important to extract and save the features displayed by each layer of the road in a timely manner. Skip connections play an important role as a carrier between the encoder and decoder. Shallow features contain rich road shapes and road texture features, which play an important role in edge feature recovery. At the same time, the topological structure of the road determines that simple feature extraction cannot be used as an effective supplement to improve model performance. Only by designing a unique and efficient feature extraction module based on the data characteristics can the road shape be restored to a great extent.

[0050] The structure of the road detail feature extraction module is as follows:

[0051] The multi-level features output by the encoder are fused step by step through three road detail feature extraction modules. This module receives input from two parts, one is the output of the feature map at the current scale, and the other is the result of feature extraction at the previous scale. These two parts are processed separately and weighted added as the input of this module. The input will go through two parallel parts, which extract features from the horizontal and vertical directions respectively.

[0052] Figure 3 The schematic diagram of extracting vertical features is shown. The input will then go through three parts. The image will be processed by vertical strip convolution kernels with 3×1 and 5×1 kernels, and then layer normalization will be performed to obtain V1 and V2 respectively. After the input is normalized by the 3×3 convolution kernel layer, average pooling (AP) and maximum pooling (MP) are performed respectively. The results are added after passing through the MLP layer, and then the added results are subjected to Sigmoid processing to obtain weighted values. The weighted values ​​are then multiplied with the features obtained from the two branches, and then added and passed through the output mapping layer (effective channel attention) and a series of feedforward networks to obtain the final vertical output.

[0053] Figure 4 The schematic diagram of extracting horizontal features is shown. The input will then go through three parts. The image will be processed by horizontal strip convolution kernels with kernels of 1×3 and 1×5, and then layer normalization will be performed to obtain H1 and H2 respectively. After the input is normalized by the 3×3 convolution kernel layer, average pooling (AP) and maximum pooling (MP) are performed respectively. The results are added after passing through the MLP layer, and then the added results are subjected to Sigmoid processing to obtain weighted values. The weighted values ​​are then multiplied with the features obtained from the two branches, and then added and passed through the output mapping layer Projection (i.e., effective channel attention, such as Figure 6 As shown) and a series of feed-forward networks obtain the final horizontal output.

[0054] This module fuses information from two scales as input to retain the road information extracted from the previous scale. At the same time, the output of this module not only integrates the road strip features in two directions, but also uses strip convolution kernels of different sizes in each direction to provide detailed road information from a multi-scale perspective. This can effectively serve as a supplement to provide the decoder with comprehensive and accurate road structure features.

[0055] Step 2.3: Construction of contextual information perception module: In the remote sensing image road segmentation task, a large number of road pixels are obscured by surrounding buildings, trees and shadows. The integrity of the extracted road cannot be guaranteed by extracting road features alone. Certain objects in the background have some connection with the road, which can guide the extraction of the road to some extent. Capturing the relationship between the road and the surrounding environment between the encoder and the decoder can improve the accuracy of the prediction results by improving the overall representation of road extraction. The present invention combines the advantages of dilated convolution and 2D selective scanning to design a dilated convolution-guided selective scanning block for dynamically extracting and integrating local multi-scale and global contextual information, effectively distinguishing between road and background information, and making the extracted road structure more complete and continuous; then combined with effective channel attention to filter road interference information, suppress background and other interference information, and finally construct a dual-context dynamic extraction module.

[0056] The structure of the context information perception module is as follows:

[0057] In order to compensate for the information loss caused by step-by-step downsampling, the present invention adopts the strategy of weighted addition of the output of encoder 4 and the output of the deepest road detail feature extraction module as the input of the module. The input passes through the mapping layer to change the number of channels to twice the original channel, and then is divided into two parts in the channel dimension, namely x and z. x is used for subsequent feature extraction and long sequence modeling, and z is multiplied with the results of long sequence modeling as shared weights after passing through the SiLU function. The dilated convolution is matched with position encoding, layer normalization, feedforward network layer and embedded with 2D selective scanning as a spatial attention mechanism to construct a more robust block (selective scanning block preceded by dilated convolution). It is worth noting that due to the need to remember long sequence dependencies, the number of hidden states in the state space equation is usually large, which may lead to channel redundancy, thereby hindering the learning of key channel representations. Therefore, the present invention performs channel attention in the mapping layer in the selective scanning block preceded by the dilated convolution to solve this problem. At the same time, the extraction of multi-scale detail features can be achieved by changing the dilation rate of the dilated convolution in the selective scanning block preceded by the dilated convolution. The present invention uses gradually expanding dilated convolution (dilation rates of 1, 3, 5, and 7, respectively) to extract multi-scale road detail features, and then uses 2D selective scanning to capture detailed features and long-distance contextual dependencies of pixels around the road. The four parallel dilated convolution-led selective scanning blocks are performed on one-fourth of the original input channel, and finally the outputs of the four channels are spliced ​​together, and the output of the module is obtained through the final output mapping layer (effective channel attention), generating a fused feature map with rich detail features and long-range context information. At the same time, the final output of the context information perception module will be used as the input of the decoder to promote more refined reconstruction of prediction results and ensure the accuracy and consistency of the segmentation results.

[0058] Step 2.4: Construction of feature decoding module: The feature decoding module aims to convert the abstract feature map output by the dual context dynamic extraction module back to the same spatial resolution as the original image, so as to achieve accurate semantic segmentation of remote sensing images. In this step, starting from the global-local context extraction features aggregated by the dual context dynamic extraction module, up-sampled by transposed convolution operation, and combined with the output from the corresponding stage of the bidirectional multi-level road feature dynamic fusion module to supplement the detailed boundary information and context information of the road, thereby improving the accuracy and refinement of semantic segmentation.

[0059] The structure of the feature decoding module is as follows:

[0060] Each of the four decoders includes a 1×1 convolution that reduces the channel to 1 / 4, a 3×3 transposed convolution with a stride of 2, and a 1×1 convolution that doubles the channel. All three convolutions include batch normalization and activation function ReLU; finally, the decoder output is converted into a road extraction result through a segmentation head, which consists of a 3×3 transposed convolution with a stride of 2, a 3×3 convolution, and a 2×2 convolution.

[0061] Step 3: Model training: Input the preprocessed training set data into the remote sensing image road segmentation network; use the ResNet-34 model weights pre-trained on the ImageNet-1K dataset to initialize the feature encoding module, and randomly initialize the remaining network parameters; train the remote sensing image road segmentation network model until the model converges; after the training is completed, save the trained remote sensing image road segmentation network model parameters.

[0062] Step 4: Result prediction: Input the preprocessed remote sensing road image into the trained remote sensing image road segmentation network model to obtain the accurate segmentation result of the remote sensing image data.

[0063] The training program is built on the PyTorch framework. The optimizer uses the Adam algorithm, and the loss function combines binary cross entropy and Dice coefficient to alleviate the problem of category imbalance and enhance the model's ability to segment small-scale targets. The initial learning rate of the model is set to 2e-4, the batch size is set to 4, and the maximum training batch is set to 500. The learning rate adjustment mechanism uses the Plateau Learning Rate Scheduler, and sets the attenuation factor to 0.5 and the tolerance period to 5, so as to dynamically adjust the learning rate according to the learning process and avoid local optimal solutions. In addition, an early stopping condition is set in the experiment, that is, when the loss value does not decrease significantly during 8 consecutive epochs, it is determined that the model has reached a convergence state and the training process is terminated in advance. This is to curb the tendency of overfitting in a timely manner and reasonably control the computing cost.

[0064] Among them, the calculation formula of the binary cross entropy loss function is:

[0065]

[0066] In the above formula: y is the true pixel label value, y' is the predicted label pixel value, and N is the number of label categories;

[0067] The calculation formula of the Dice coefficient loss function is:

[0068]

[0069] In the above formula: X is the generated prediction graph, Y is the real label, |X∩Y| is the intersection between the label and the prediction, |X| is the number of elements in the label, and |Y| is the number of elements in the prediction;

[0070] The final semantic segmentation loss function is the sum of the weighted coefficients of the cross entropy loss function and the Dice coefficient loss function, and the calculation formula is:

[0071] L s =L dice +L cross .

[0072] The present invention will be further described below based on specific embodiments.

[0073] The experimental dataset uses the DeepGlobe road extraction dataset, which comes from the remote sensing road extraction challenge held by CVPR2018. The dataset contains 6226 sets of remote sensing images with a size of 1024×1024 pixels and road labels, with an image resolution of 0.5m / pixel. The images are collected by Digital Globe satellites. These images not only cover a wide range of geographical areas such as cities and villages in many countries, but also include various types of roads and street networks. In addition, in order to maintain the authenticity and practicality of the data, the author deliberately did not mark the paths in the farmland. During the recognition, the author did not want the model to make incorrect extractions, making the dataset more challenging and practical, and also providing a broader space for exploration.

[0074] When evaluating the road segmentation performance of the model, the present invention uses four evaluation indicators commonly used in semantic segmentation to evaluate the model. The evaluation results are shown in Table 1. The four evaluation indicators are accuracy, recall, F1 score and intersection over union. These indicators can comprehensively and objectively reflect the performance of the model in the road segmentation task. Figure 6 (original remote sensing road image), Figure 7 (manually annotated remote sensing road segmentation reference image) and Figure 8 The comparison and analysis of (the predicted road segmentation result image obtained after applying the method of the present invention) shows that the method of the present invention has demonstrated excellent performance in road extraction from remote sensing images. Specifically, Figure 8 The predicted segmentation results shown are consistent with Figure 7 This fact strongly confirms the accuracy and reliability of the method in the road segmentation task.

[0075] Accuracy Recall F1 score Intersection and Union 84.58% 84.59% 84.58% 73.29%

[0076] Table 1 Specific indicators on the DeepGlobe road extraction dataset.

[0077] The present invention discloses an innovative remote sensing image road segmentation method, which combines road detail feature extraction and context information perception. The present invention aims to achieve highly accurate segmentation of road areas in high-resolution remote sensing images through advanced deep learning technology combined with specially designed network architecture and modules. First, a detailed manually annotated label image is created based on a high-resolution remote sensing image with RGB three channels, which is used as a standard reference for model training. Then, the complete image data set is divided into a training set and a test set in proportion, and a series of preprocessing operations such as denoising and feature enhancement are performed on the images in the training set to optimize the image quality and lay a good foundation for model training. In the model building stage, the present invention creatively designs a remote sensing image road segmentation method that combines road detail feature extraction and context information perception. Among them, the road detail feature extraction module aggregates the features of the road itself from both horizontal and vertical directions, and passes them to the decoder as detail supplements, which effectively compensates for the loss of spatial information and edge details caused by continuous downsampling, and enhances the model's understanding of complex scenes and multi-level road structures. On the other hand, the context information perception module realizes the fusion of long-distance and local multi-scale context information, which helps to suppress background interference, occlusion effects and confusion of similar road structures, and ensures that the extracted road structure has integrity and continuity. During the model training process, the present invention inputs the preprocessed training image and its corresponding label image into the road segmentation model, and uses the optimization algorithm to continuously adjust the network parameters until the model can accurately identify and segment the road area. When the model reaches a convergence state, the optimal parameter settings are saved for subsequent applications. Finally, the present invention can automatically generate a predicted label image, that is, a segmented road image, by inputting a new remote sensing image into the trained model. Compared with the prior art, the uniqueness of the present invention lies in its road detail feature extraction and the application of context information perception mechanism, which not only improves the connectivity and accuracy of road segmentation, but also enhances the model's ability to distinguish between roads of different levels and restore edge details.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A remote sensing image road segmentation method combining road detail feature extraction and context information perception, characterized by: The following steps are involved: Step 1: Construct a data set: collect high-resolution remote sensing images through remote sensing satellites, divide them into training sets and test sets in proportion, and perform preprocessing and data enhancement operations on all images in the training set. Both the training set and the test set contain original remote sensing images and professionally annotated remote sensing images; Step 2: Constructing a remote sensing image road segmentation network model: including: Step 2.1: Construction of feature encoding module; Step 2.2: Construction of a road detail feature extraction module: The road detail feature extraction module includes a horizontal detail feature extraction module and a vertical detail feature extraction module for fusing multi-level strip features of the road. Both the horizontal detail feature extraction module and the vertical detail feature extraction module include multi-scale strip space attention and effective channel attention for space and channel dimension selection; Step 2.3: Construction of context information perception module: The context information perception module extracts multi-scale detail features by designing a selective scanning block preceded by dilated convolution and changing the dilation rate of the dilated convolution in the selective scanning block, which is used for local multi-scale context extraction and global context extraction, and further features screening and fusion are performed in combination with effective channel attention; Step 2.4: Construction of feature decoding module; The feature encoding module and the feature decoding module realize the step-by-step fusion of features through the road detail feature extraction module; the fused features are extracted by the context information perception module for global and local multi-scale context features; the extracted results are used for decoding by the feature decoding module; Step 3: Training remote sensing image road segmentation network model; Step 4: Result prediction: Input the preprocessed remote sensing image into the trained remote sensing image road segmentation network model to obtain accurate road segmentation results.

2. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 1, characterized in that: The feature encoding module and the feature decoding module are each composed of five stages. The feature encoding module includes a downsampling module, encoder 1, encoder 2, encoder 3 and encoder 4 in sequence; the feature decoding module includes decoder 1, decoder 2, decoder 3, decoder 4 and a segmentation head in sequence.

3. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 2, characterized in that: The downsampling module is a 7×7 convolution operation with a step size of 2 and a 3×3 maximum pooling with a step size of 2; the four encoders are respectively stacked by 3, 4, 6, and 3 residual blocks; each residual block is composed of two 3×3 convolution operations and residual connections.

4. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 2, characterized in that: Each of the four decoders includes a 1×1 convolution that reduces the channel to 1 / 4, a 3×3 transposed convolution with a stride of 2, and a 1×1 convolution that doubles the channel. All three convolutions include batch normalization and activation function ReLU; finally, the output of the decoder is converted into a road extraction result through a segmentation head, which consists of a 3×3 transposed convolution with a stride of 2, a 3×3 convolution, and a 2×2 convolution.

5. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 2, characterized in that: The multi-level features output by the encoder are fused step by step through three road detail feature extraction modules. Each road detail feature extraction module receives input from two parts, one is the output of the feature map at the current scale, and the other is the result after feature extraction at the previous scale. These two parts are processed separately and weightedly added as the input of the road detail feature extraction module.

6. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 5, characterized in that: The horizontal detail feature extraction module includes a strip space attention consisting of two horizontal strip convolution branches and one convolution branch. The image input to the road detail feature extraction module will be processed by the horizontal strip convolution kernels with convolution kernels of 1×3 and 1×5 respectively, and then layer normalization is performed; at the same time, the image input to the road detail feature extraction module is normalized by the 3×3 convolution kernel layer, and then average pooling and maximum pooling are performed respectively. The results are added after passing through the MLP layer, and then the added results are subjected to Sigmoid processing to obtain weighted values; the weighted values ​​obtained are then multiplied with the features obtained by the two horizontal strip convolution branches, and then added and passed through the output mapping layer and a series of feedforward networks to obtain the final horizontal output, wherein the output mapping layer adopts effective channel attention.

7. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 5, characterized in that: The longitudinal detail feature extraction module includes two vertical strip convolution branches and one convolution branch. The image input to the road detail feature extraction module will be processed by the vertical strip convolution kernels with convolution kernels of 3×1 and 5×1 respectively, and then layer normalization is performed; at the same time, the image input to the road detail feature extraction module is normalized by the 3×3 convolution kernel layer, and then average pooling and maximum pooling are performed respectively. The results are added after passing through the MLP layer, and then the added results are subjected to Sigmoid processing to obtain weighted values; the weighted values ​​obtained are then multiplied with the features obtained by the two horizontal strip convolution branches, and then added and passed through the output mapping layer and a series of feedforward networks to obtain the final vertical output, wherein the output mapping layer adopts effective channel attention.

8. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 2, characterized in that: The input of the context information perception module is the result of weighted addition of the output of encoder 4 and the output of the deepest road detail feature extraction module. The structure of the context information perception module is as follows: it consists of a position embedding block, 4 groups of dilated convolutions with dilation rates of 1, 3, 5, and 7, 4 2D selective scanning blocks, and effective channel attention, wherein each group of dilated convolutions and 2D selective scanning blocks is followed by element-by-element multiplication of shared weights, and then the four groups of results are concatenated and outputted through effective channel attention.

9. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 8, characterized in that: The structure of the effective channel attention is as follows: it consists of 1 3×3 convolution, 1 GeLU activation function, 1 layer normalization and efficient channel attention ECA.

10. The method for remote sensing image road segmentation combining road detail feature extraction and context information perception according to claim 1, characterized in that: The training process of the model in step 3 is as follows: The preprocessed training set data is input into the remote sensing image road segmentation network model. The feature encoding module is initialized using the ResNet-34 weights pre-trained on the ImageNet-1K dataset, and the remaining parameters are randomly initialized. The remote sensing image road segmentation network model is trained until the model converges, and the model parameters are saved after the training is completed.

Citation Information

Patent Citations

  • Multi-scale aggregation cloud and cloud shadow identification method, system and device and storage medium

    CN115410081A

  • Remote sensing image road extraction method based on multi-direction space connectivity, storage medium and electronic equipment

    CN118015584A

  • Remote sensing image road extraction method based on dual-channel deep neural network

    CN118397462A

  • Remote sensing image dense road segmentation method using strip-shaped features

    CN118429356A

  • Small target detection method based on Mama feature fusion

    CN118968019A