A Remote Sensing Image Dense Road Segmentation Method Based on Weaving Feature Extraction
By adopting a weaving feature extraction method in remote sensing image road extraction, a U-Net architecture including snake-shaped weaving attention, context information weaving module, etc. was constructed, which solved the problems of extracting road integrity and connectivity under road density, and achieved a more efficient and accurate road extraction effect.
Patent Information
- Application Number
- CN202411602645.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-11
AI Technical Summary
The prior art is difficult to ensure the integrity and connectivity of the extraction road when roads are dense, and the background information in the remote sensing image has many types and a large proportion, which interferes with road extraction.
Using a method based on weaving feature extraction, road features are extracted through horizontal and vertically alternating strip convolutions, and a snake weaving attention, context information weaving module, global information extraction module and multi-scale weaving decoder are designed to construct a remote sensing image road extraction network model of U-Net architecture.
It improves the integrity and connectivity of the extracted roads in dense road environments, enhances the model's understanding and identification of road characteristics, and improves processing efficiency and accuracy.
Smart Images

Figure CN119360349B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of remote sensing image road extraction and provides a remote sensing image dense road segmentation method based on weaving feature extraction. Background Art
[0002] High-resolution remote sensing images can capture rich surface details and provide more accurate visual information, making important contributions to many fields in computer vision, such as target detection and semantic segmentation. Among them, road extraction is one of the representative tasks of semantic segmentation. The extracted road information can provide data support for many fields such as urban planning, autonomous driving, and traffic control, greatly promoting the development and innovation of related technologies. However, in the face of dense roads, these studies are difficult to show good results. In order to extract dense roads, the model must accurately identify and process small-scale detail information at close intervals, which is very difficult. The background information in remote sensing images is of many types and accounts for a large proportion, which will also cause great interference to the extraction of dense roads.
[0003] There are two main approaches to road extraction research: based on traditional machine learning methods and based on deep learning methods. The road extraction method based on deep learning can make full use of a large amount of high-resolution remote sensing image data, and achieves more accurate, efficient and robust road extraction compared to traditional machine learning methods. The road extraction method based on deep learning has made some progress, but these studies are mainly designed to address problems such as blocked roads and small and difficult roads to extract, and can only improve the integrity and connectivity of the extracted roads to a certain extent in the corresponding scenarios. Summary of the invention
[0004] In order to solve the problem that the current model is difficult to ensure the integrity and connectivity of the extracted roads when the roads are dense, the present invention is inspired by the "weaving" behavior in reality and proposes a remote sensing image dense road segmentation method based on weaving feature extraction, that is, continuously using horizontal and vertical alternating strip convolutions to extract road features. The method includes:
[0005] A high-resolution remote sensing image road extraction dataset is obtained, the dataset is divided into a training set and a test set, and a data enhancement operation is performed on the remote sensing images in the training set;
[0006] Inputting the remote sensing images in the training set into the initial remote sensing image road extraction network model to obtain a road segmentation map of the remote sensing images in the training set;
[0007] Based on the difference between the road segmentation map of the remote sensing image in the training set and the labeled example image corresponding to the remote sensing image, performing parameter iteration on the initial remote sensing image road extraction network model to obtain the remote sensing image road extraction network model;
[0008] Inputting the test set images into the remote sensing image road extraction network model, and outputting accurate segmentation results of the remote sensing image data;
[0009] The remote sensing image road extraction network model is a U-Net architecture, including an encoder, a snake weaving attention, a context information weaving module, a global information extraction module and a multi-scale weaving decoder;
[0010] The images in the training set are first input into the encoder part in the remote sensing image road extraction network model, and the encoder quickly extracts the feature map and outputs the feature map to the next layer of encoder and snake weaving attention;
[0011] The snake weaving attention adds the feature map of the encoder and the previous layer input, performs average pooling to reduce random noise, and then uses strip convolution and dynamic snake convolution of different sizes to extract features from the feature map in both horizontal and vertical directions, and finally fuses the features in different directions and outputs the feature map to the context information weaving module;
[0012] The context information weaving module adds the feature map of the encoder and the previous layer input, respectively uses standard convolution, dilated convolution and depth-separable convolution to extract and fuse the feature maps, learns multi-scale context information, and then uses average pooling operations in both horizontal and vertical directions to perform weaving context learning on the obtained feature maps containing multi-scale context information, and then adds the feature maps of the encoder and the weaving module and sends them to the global information extraction module;
[0013] The global information extraction module performs a global average pooling operation on the feature map to highlight the global information of the feature map, and inputs the feature map after highlighting the global information into RepViT Blocks for global information modeling;
[0014] The multi-scale braided decoder module performs a deconvolution operation on the feature map, reproduces the road information using the horizontal and vertical strip depth-separable convolution, and then converts the feature map into a probability value of each pixel belonging to a certain category through upsampling and Sigmoid activation function, and outputs an accurate road segmentation map.
[0015] The encoder module has five stages, including:
[0016] The first stage consists of two layers. The first layer consists of convolution with a kernel size of 3×3 and a stride of 2, batch normalization operation, and SiLU activation function. The second layer consists of two FusedMBConv blocks to quickly extract shallow feature maps.
[0017] The second stage consists of encoder 1, which contains four FusedMBConv blocks. Encoder 1 inputs the obtained feature map to encoder 2 and the first snake-weaving attention block to learn road details;
[0018] The third stage consists of encoder 2, which contains four FusedMBConv blocks. Encoder 2 inputs the obtained feature map to encoder 3 and the second snake-weaving attention block to learn road details;
[0019] The fourth stage consists of encoder 3, which contains two sets of stacked convolutional blocks, consisting of 6 MBConv blocks and 9 MBConv blocks respectively. Encoder 3 inputs the obtained feature map to encoder 4 and the first context information weaving module;
[0020] The fifth stage consists of encoder 4, which contains two layers. The first layer consists of 15 MBConv blocks, and the second layer consists of convolution with a kernel size of 3×3 and a stride of 2, batch normalization operation, and SiLU activation function. Encoder 4 inputs the obtained feature map into the decoder and global information extraction module;
[0021] The five stages of the encoder output five batches of feature maps, wherein the feature map generated in the first stage is the shallowest feature map, and the feature map generated in the last stage is the deepest feature map.
[0022] The snake-weaving attention module is set in the shallow part of the jump connection in the U-Net architecture;
[0023] The serpentine weaving attention module first uses 1×1 convolution and depth-separable convolution to process the feature map of the corresponding layer encoder, directly adds the processing result to the feature map of this layer as the input of the serpentine weaving attention, and then uses a convolution of size 13×13 and a step size of 1 and an average pooling of size 13×13 to process the input; the processed feature map is transmitted to two parallel branches for horizontal and vertical weaving learning respectively, wherein the horizontal learning branch consists of 1×15, 1×13, The vertical learning branch consists of 1×11 strip convolutions and 13×13 dynamic snake convolutions learned along the X-axis. The vertical learning branch consists of 15×1, 13×1, and 11×1 strip convolutions and 13×13 dynamic snake convolutions learned along the Y-axis. The feature maps obtained by horizontal and vertical learning are added element by element to the three batches of feature maps of the original input. The addition results are sequentially subjected to 1×1 convolution, batch normalization, and ReLU activation function operations to obtain the output of the snake weaving attention module.
[0024] The context information weaving module is arranged in the deep part of the jump connection in the U-Net architecture;
[0025] The context information weaving module first uses 1×1 convolution and depth-separable convolution to process the feature map of the previous layer, and directly adds the processed feature map to the feature map of the current layer as the input of the context information weaving module;
[0026] The context information weaving module is divided into two stages. The first stage performs multi-scale context learning on the input feature map, and the second stage performs weaving context learning in the horizontal and vertical directions.
[0027] The first stage includes three branches, the first is a standard convolution of 5×5 size, the second is a dilated convolution of 7×7 size with a dilation factor of 2 and a dilated convolution of 9×9 size with a dilation factor of 3, and the third is a depthwise separable convolution of 11×11 and 13×13. The convolution results of the three branches are fused to obtain a feature map containing multi-scale context information, which is then processed with 1×1 convolution, batch normalization and ReLU activation function in sequence. The processed result is multiplied and added with the original input and input into the second stage.
[0028] The second stage performs 1×H and W×1 average pooling operations on the input feature map, and then multiplies it with the input feature map respectively to make the vertical and horizontal contextual information more prominent; then the two multiplication results are added together, and the added result is processed with 1×1 convolution, batch normalization and ReLU activation function, and then directly added to the input of the second stage to obtain the output of the contextual information weaving module.
[0029] The global information extraction module is placed in the bottleneck part of the U-Net architecture;
[0030] The global feature extraction module includes two global feature learning branches, and the inputs are the output of the second context information weaving module and the output of encoder 4;
[0031] The global feature extraction module first performs global average pooling on two batches of feature maps to extract the global information of the feature maps, and then multiplies the results of the global average pooling with the input feature maps to enhance the global information of the feature maps; then the inputs of the branches are added to the multiplication results, and input into RepViT Blocks for global modeling, the output feature maps of RepViT Blocks are concatenated with the original input, and then 1×1 convolution, batch normalization, and ReLU activation function are used to reduce the dimension of the feature maps to obtain the final output of the global feature extraction module.
[0032] The multi-scale braided decoder module includes a multi-scale braided decoder and an upsampling module corresponding to multiple stages of the encoder module;
[0033] The multi-scale weaving decoder first upsamples the feature map processed by the snake weaving attention or context information weaving module through inverse convolution, so that the length and width of the feature map become twice the input, and the number of channels is halved; the upsampling result is input into two parallel branches for weaving learning, one branch learning is composed of 1×5, 1×3 strip depth separable convolution, reproducing the horizontal road information of the feature map, and the other branch is composed of 5×1, 3×1 strip depth separable convolution, reproducing the vertical road information of the feature map; the outputs of the two branches are spliced with the upsampling results, and then the dimension is reduced through 1×1 convolution, batch normalization, and ReLU activation function to obtain the output of the multi-scale weaving decoder;
[0034] The upsampling module uses inverse convolution to upsample the output feature map of the multi-scale braided decoder, uses two ReLU activation functions and two convolutions to adjust the number of feature map channels, and finally converts the feature map into a probability value of each pixel belonging to a certain category through a Sigmoid activation function to obtain a road segmentation map.
[0035] The data enhancement operation is to perform data enhancement by performing image flipping, random rotation, resizing, adding random noise, and light and dark transformation operations on the images and labels in the training set.
[0036] The method of the present invention has the following beneficial effects compared with the prior art:
[0037] 1. In the encoder part, the method of the present invention uses the EfficientNetV2 network parameters pre-trained on the ImageNet-1K dataset to initialize the encoder module, and randomly initializes the remaining network parameters, achieving faster training speed and higher parameter efficiency, and being more efficient and accurate when processing complex dense roads.
[0038] 2. The method of the present invention designs a novel attention mechanism called snake weaving attention at the shallow jump connection of the model, and uses strip convolution to extract features from the feature map in both horizontal and vertical directions, and finally fuses the features in different directions. This feature extraction mode similar to weaving behavior can better understand the spatial and structural relationship between dense roads, and strip convolution can reduce the computational complexity of the model, improve efficiency and running speed. Snake weaving attention also uses dynamic snake convolution to allow the model to adaptively focus and match the road morphology.
[0039] 3. The method of the present invention designs a context information weaving module at the deeper jump connection of the model. The context information weaving module uses the complementary advantages of three methods: standard convolution, multi-scale dilated convolution, and large-scale channel separable convolution, to better extract the context information of the road from different levels. Then use 1×W and H×1 pooling to extract the context information in both horizontal and vertical directions, continue to "weave" the feature map, strengthen the important road features in different directions, and greatly improve the extraction effect of dense roads.
[0040] 4. The method of the present invention designs a global information extraction module that can learn the global information of the road in the bottleneck part of the model, and uses the global average pooling operation to extract the overall feature representation of the feature map, effectively capturing the global information of the entire feature map, which can help the model understand the overall scene instead of just focusing on local details. In addition, the novel RepViT Blocks are used in the global information extraction module to enhance the model's perception of global features without significantly increasing the computational burden, which can more effectively integrate global information and extract features, and is very helpful in improving the connectivity of extracted dense roads.
[0041] 5. The method of the present invention designs a multi-scale weaving decoder at the end of the model, and uses strip convolution in the horizontal and vertical directions to perform multi-scale feature extraction on the feature map of each stage of the decoder, so that the decoder can effectively capture the characteristics of the road structure and improve the decoder's ability to reconstruct the road, thereby helping the model to more accurately identify and track the direction of the road, which is helpful for extracting dense roads. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flowchart of a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0043] Figure 2 It is a schematic diagram of the network structure of a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0044] Figure 3 It is a structural diagram of FusedMBConv in a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0045] Figure 4 It is a structural diagram of MBConv in a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0046] Figure 5It is a structural schematic diagram of a serpentine weaving attention for efficiently extracting road strip features in a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0047] Figure 6 It is a structural schematic diagram of a context information weaving module in a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0048] Figure 7 It is a structural schematic diagram of a global information extraction module capable of highlighting road strip features in a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0049] Figure 8 It is a schematic structural diagram of a multi-scale weaving decoder in a remote sensing image dense road segmentation method based on weaving feature extraction provided by an embodiment of the present invention.
[0050] Fig. 9 This is an example of a remote sensing image of the DeepGlobe dataset used in the first embodiment of the present invention.
[0051] Fig.10 This is an example of a remote sensing image with labels in the DeepGlobe dataset used in the first embodiment of the present invention.
[0052] Fig.11 It is a schematic diagram of road results extracted from the DeepGlobe data set of Example 1 using the road segmentation method of the present invention.
[0053] Fig.12 This is an example of a remote sensing image of the Massachusetts dataset used in the second embodiment of the present invention.
[0054] Fig.13 This is an example of a remote sensing image with labels in the Massachusetts dataset used in the second embodiment of the present invention.
[0055] Fig.14 This is a schematic diagram of road results extracted from the Massachusetts data set in Example 2 using the road segmentation method of the present invention. DETAILED DESCRIPTION
[0056] The best mode of implementation of the present invention is described below by way of example. It should be understood that the specific embodiments herein are used to explain the present invention in detail and should not be construed as limiting the present invention. It should be noted that various changes and modifications may be made under the premise of following the principles and core scope of the present invention, and these changes should all be deemed to be within the scope of protection of the present invention. In conjunction with the accompanying drawings, the specific implementation steps of the present invention are described in detail.
[0057] like Figure 1 As shown, the present invention provides a remote sensing image dense road segmentation method based on weaving feature extraction, comprising the following steps:
[0058] Step 1: Get two high-resolution remote sensing image road extraction datasets, DeepGlobe and Massachusetts, and the remote sensing images to be processed. Use the high-resolution remote sensing image datasets with a size of 1024×1024 pixels and the high-resolution remote sensing image datasets with a size of 1500×1500 pixels as training sets. The datasets contain original remote sensing images and manually annotated high-quality road label data. Perform preprocessing and multiple data enhancement operations on the remote sensing images in the training set.
[0059] The data augmentation method uses image flipping, random rotation, resizing, adding random noise, and light and dark transformation operations to enhance the images and labels in the training set.
[0060] Data enhancement can help generate more samples of rare categories, thereby balancing the sample distribution between categories, thereby improving the performance of the model on unbalanced data and reducing model overfitting.
[0061] Step 2: inputting the remote sensing images in the training set into the initial remote sensing image road extraction network model to obtain a road segmentation map of the remote sensing images in the training set;
[0062] The remote sensing image road extraction network model includes an encoder part, a snake weaving attention, a context information weaving module, a global information extraction module and a multi-scale weaving decoder. The entire remote sensing image road extraction network model is as follows Figure 2 As shown, it is the U-Net architecture.
[0063] In a specific embodiment, the encoder module is initialized using the network parameters of EfficientNet V2 pre-trained on the ImageNet-1K dataset, and the remaining network parameters are randomly initialized. EfficientNet V2 achieves faster training speed and higher parameter efficiency than EfficientNet V1 by using training-aware neural architecture search and scaling strategies, making the present invention more efficient and accurate in processing complex dense roads.
[0064] In a specific embodiment, the encoder module consists of 5 stages, each of which outputs a feature map of a corresponding size. The detailed structure is as follows:
[0065] The first stage consists of two layers. The first layer consists of a convolution with a kernel size of 3×3 and a stride of 2 (one convolution for every two pixels moved), a batch normalization operation, and a SiLU activation function. The second layer consists of two FusedMBConv blocks. The FusedMBConv blocks are as follows: Figure 3 As shown, it is used to output the shallowest feature map;
[0066] The second stage consists of encoder 1, which contains four FusedMBConv blocks. Encoder 1 inputs the obtained feature map to encoder 2 and the first snake-weaving attention block to learn road details;
[0067] The third stage consists of encoder 2, which contains four FusedMBConv blocks. Encoder 2 inputs the obtained feature map to encoder 3 and the second snake-weaving attention block to learn road details;
[0068] The fourth stage consists of encoder 3, which contains two sets of stacked convolutional blocks, consisting of 6 MBConv blocks and 9 MBConv blocks respectively. Encoder 3 inputs the obtained feature map to encoder 4 and the first context information weaving module. The MBConv blocks are as follows: Figure 4 As shown, two sets of stacked MBConv blocks are designed to gradually refine the representation of the input features, so that the encoder can better learn and retain the key information in the image;
[0069] The fifth stage consists of encoder 4, which contains two layers. The first layer consists of 15 MBConv blocks, and the second layer consists of convolution with a kernel size of 3×3 and a stride of 2, batch normalization operation, and SiLU activation function. Encoder 4 inputs the obtained feature map into the decoder and global information extraction module;
[0070] In the early stages of the network, the feature map size of the input image is large. Figure 3 The FusedMBConv block shown combines the 3x3 convolution and 1x1 point-by-point convolution processes into a single convolution operation, which is more suitable for processing large-size feature maps. Therefore, more FusedMBConv blocks are used in the first three stages. In the subsequent stages, as the feature map size decreases and the number of channels increases, Figure 4 The MBConv blocks shown can better balance computational efficiency and performance, so they are more commonly used in the last two stages of the encoder.
[0071] The five stages of the encoder output a total of five batches of feature maps, where the feature maps produced in the first stage are considered to be the shallowest feature maps, and the feature maps produced in the last stage are considered to be the deepest feature maps.
[0072] The snake-weaving attention module is set in the shallow part of the jump connection in the U-Net architecture. The detailed structure is as follows Figure 5 As shown:
[0073] The input of the snake weaving attention is the feature map of the current layer and the feature map of the previous layer transmitted by the jump connection. Since the dimensions of the two do not match, it is necessary to first use 1×1 convolution and depthwise separable convolution to process them so that the feature map of the previous layer has the same dimension as the feature map transmitted by the current layer. Then, the processing result is directly added to the feature map of the current layer as the input of the snake weaving attention. Then, a convolution of size 13×13 and a stride of 1 and an average pooling of size 13×13 are used to process the input. By averaging the area, the random noise in the image is reduced and the road features are highlighted.
[0074] Next, the processed feature maps are transferred to two parallel branches for horizontal and vertical weaving learning respectively: the horizontal learning branch consists of 1×15, 1×13, 1×11 strip convolutions and a dynamic snake convolution of size 13×13 learned along the X-axis, and the vertical learning branch consists of 15×1, 13×1, 11×1 strip convolutions and a dynamic snake convolution of size 13×13 learned along the Y-axis. The feature maps obtained by horizontal and vertical learning are added element by element with the original input feature maps to fuse information to form a richer and more discriminative feature representation;
[0075] Among them, using strip convolutions of different sizes can better capture long and continuous edge features. Longer convolution kernels have larger receptive fields and are suitable for capturing global information; while shorter convolution kernels can capture local detail information more finely. By combining these convolutions of different scales, the model can achieve a good balance between the global and local.
[0076] Strip convolution first extracts the directional information of the road, reducing the noise or unnecessary details in the image, thereby optimizing the input of dynamic snake convolution. Dynamic snake convolution further enhances the network's ability to adapt to curves and irregular structures, especially those curved roads and complex edges. This convolution can dynamically adapt to the direction of the road, allowing the model to accurately follow curves and curved paths, and improve the perception of non-straight features.
[0077] Finally, the addition result is subjected to 1×1 convolution, batch normalization, ReLU activation function and other operations in sequence, and the result is the output of snake-weaving attention; among them, 1×1 convolution can adjust and recombine the channel information of the feature map to reduce the computational complexity; batch normalization reduces the internal covariate shift phenomenon by standardizing the input features; ReLU activation function introduces nonlinear capabilities to the network, allowing the model to capture complex nonlinear features, and can also help suppress negative values to ensure that the model outputs sparse activations, thereby improving computational efficiency and the generalization ability of the model. The whole process can effectively prevent overfitting problems and improve the robustness of the model.
[0078] The snake weaving attention is set in the shallow part of the jump connection, and the output of the encoder of the corresponding layer is weighted by the snake weaving attention to learn the road detail information rich in the shallow feature map. Then the feature map processed by the snake weaving attention is input into the snake weaving attention of the next layer or the context information weaving module that can highlight the multi-scale context information. The feature map is also passed to the multi-scale weaving decoder of the corresponding layer;
[0079] This module can significantly improve the representation of road features by specifically designing the attention mechanism to mimic the strip-like structure of the road. By imitating the "weaving" behavior, the module uses horizontal and vertical strip convolutions to simulate the natural shape of the road, and adapts to the slender and tortuous characteristics of different roads through dynamic snake convolutions, thereby achieving high-precision feature capture. This approach increases the model's attention to road features and greatly enhances the model's ability to accurately identify and extract road structures in complex backgrounds.
[0080] Contextual information plays a vital role in road extraction tasks and can significantly improve the performance and adaptability of the model. Effective use of contextual information can help the model better understand how various elements in the image are related to each other. For example, roads often appear together with elements such as vehicles and traffic signs. If the model can use this relevant information to assist in inferring the location of roads, it can more accurately distinguish between roads and non-road areas, especially in scenes with dense roads and closely intertwined foreground and background.
[0081] In current road extraction research, the means of extracting contextual information are relatively simple, such as extracting multi-scale information, using dilated convolution to expand the receptive field, etc. Although these methods can achieve certain results, they still face some limitations when dealing with complex traffic scenes.
[0082] In a specific embodiment, the context information weaving module of the present invention is set in the deep part of the jump connection in the U-Net architecture, and the detailed structure is as follows: Figure 6 As shown:
[0083] The input of the context information weaving module is the feature map of the current layer and the feature map of the previous layer transmitted by the jump connection. Since the dimensions of the two do not match, it is necessary to first use 1×1 convolution and depth-separable convolution to process the feature map of the previous layer so that the dimension of the feature map of the previous layer is the same as that of the feature map transmitted by the current layer; then the processed feature map and the feature map of the current layer are directly added as the input of the context information weaving module;
[0084] The context information weaving module of the present invention is divided into two stages. The first stage performs multi-scale context learning on the input feature map, and the second stage performs weaving context learning in the horizontal and vertical directions.
[0085] Among them, there are three branches of multi-scale context information. The first one is a standard convolution of size 5×5, the second one is a dilated convolution of size 7×7 with a dilation factor of 2 and a dilated convolution of size 9×9 with a dilation factor of 3, and the third one is a depth-wise separable convolution of 11×11 and 13×13.
[0086] The 5*5 standard convolution in the first branch can make full use of all pixel information in the receptive field. The dilated convolution in the second branch expands the receptive field by introducing holes between the elements of the convolution kernel, capturing a wider range of contextual information without increasing the computational complexity. The third branch is the 11×11 and 13×13 depthwise separable convolutions, which can expand the receptive field while having a small number of parameters and are easy to calculate.
[0087] The learning results of the three branches are spliced together to complement each other’s strengths, and the contextual information of the road is better extracted from different levels. Then, 1×1 convolution, batch normalization, and ReLU activation functions are used in sequence, and the processed results are multiplied with the original input and then added;
[0088] The result of the addition is subjected to the second stage of weaving context learning, and the input of the weaving context learning is subjected to H×1 and 1×W average pooling operations, that is, the feature map is averaged horizontally and vertically respectively to extract the context information in the horizontal and vertical directions, and the feature map is further subjected to "weaving" learning, and then multiplied with the input feature map respectively to make the vertical and horizontal context information more prominent, strengthen the important road features in different directions, and greatly improve the extraction effect of dense roads; then the two multiplication results are added, and the addition result is processed by 1×1 convolution, etc., and then directly added to the input of the second stage to obtain the output of the context information weaving module;
[0089] The context information weaving module receives the output of the previous layer of snake weaving attention or the output of the previous layer of context information weaving module, uses multiple convolution operations to learn the context information of the deeper feature map, and uses pooling operations in both horizontal and vertical directions to perform weaving context learning on the obtained feature map containing multi-scale context information. The feature map processed by the context information weaving module is then passed to the next layer of context information weaving module or the global information extraction module, and the feature map is also passed to the multi-scale weaving decoder of the corresponding layer.
[0090] Global information refers to information obtained from the entire image or a large area, which helps to capture the structure, pattern and contextual relationship in the image. Roads are usually highly structured and continuous, that is, many roads run through the entire image. Therefore, effective use of global information is crucial to understanding and reconstructing the dense and continuous road network structure in the image. Therefore, the present invention designs a global information extraction module and places it in the bottleneck part of the U-Net architecture to extract and utilize global information. Placing the global information extraction module in the bottleneck part of the U-Net can effectively utilize the maximum receptive field and the deepest high-level semantic features of this position, thereby optimizing the capture and processing of global information, allowing the module to more comprehensively understand the overall structure of the image. This approach can significantly improve the accuracy of the model, especially in dealing with tasks such as road extraction that require strong structural recognition.
[0091] In a specific embodiment, the global feature extraction module of the present invention is set in the bottleneck part of the U-Net architecture, and the detailed structure is as follows: Figure 7 As shown:
[0092] The global feature extraction module includes two global feature learning branches, whose inputs are the output of the second context information weaving module and the output of encoder 4, respectively. The dimensions of the two batches of feature maps are the same;
[0093] First, global average pooling is performed on the two batches of feature maps to extract the global information of the feature maps. The results of global average pooling are then multiplied with the input feature maps to highlight the global information of the feature maps. The inputs of the branches are then added to the multiplication results. The addition results are input into RepViT Blocks for global modeling. The output of RepViT Blocks is concatenated with the original input. The feature maps are then reduced in dimension using 1×1 convolution to obtain the final output of the global feature extraction module. The output of the global feature extraction module is transmitted to the decoder for road reproduction.
[0094] The global information extraction module performs global average pooling operations on the output of the deepest context information weaving module and the output of the deepest encoder to highlight the global information of the feature map and enhance the model's understanding of the integrity and continuity of the road. The global information extraction module inputs the feature map after highlighting the global information into RepViT Blocks for global information modeling. Without significantly increasing the computational burden, the model's perception of global features is enhanced, and global information integration and feature extraction can be performed more effectively, which is of great help in improving the connectivity of extracted dense roads. The feature map processed by the module is passed to the multi-scale weaving decoder.
[0095] In the U-Net architecture, the decoder gradually refines the feature map by upsampling and feature fusion, thereby accurately restoring the details and boundaries of the road, and finally obtaining the result of road extraction. In the current road extraction research, most models are still using simple inverse convolution for upsampling. In order to improve the effect of road reconstruction, the present invention designs a multi-scale weaving decoder module. In deep learning road extraction, multi-scale information is very important because it can help the model capture different levels of features from subtle road edges to vast road networks. This multi-level integration of information greatly improves the accuracy and robustness of road extraction, especially in complex and changeable road environments.
[0096] In a specific embodiment, Figure 8 As shown, the detailed structure of the multi-scale weaving decoder 1 is as follows:
[0097] The multi-scale weaving decoder 1 first upsamples the feature map processed by the snake weaving attention or context information weaving module through inverse convolution, so that the length and width of the feature map become twice that of the input and the number of channels is halved;
[0098] The upsampling result is input into two parallel branches for weaving learning. One branch learns to reproduce the horizontal road information of the feature map, and the other branch learns to reproduce the vertical road information of the feature map. The first branch is composed of 1×5 and 1×3 strip depth-separable convolutions, and the second branch is composed of 5×1 and 3×1 strip depth-separable convolutions.
[0099] The outputs of the two branches are concatenated with the upsampling results, and then the dimension is reduced by 1×1 convolution. The dimension reduction results are batch normalized and ReLU activated to obtain the output of the multi-scale braided decoder. The output of the multi-scale braided decoder is passed to the next multi-scale braided decoder.
[0100] At the end of the decoder module, the feature map is upsampled using the inverse convolution of the upsampling module, and the number of feature map channels is adjusted using two ReLU activation functions and two convolutions. Finally, after processing with the Sigmoid activation function, the road segmentation map is obtained.
[0101] The multi-scale weaving decoder module includes multi-scale weaving decoders corresponding to multiple stages of the encoder module and the final upsampling stage. The processed feature maps are input into the multi-scale weaving decoders of the corresponding stages respectively, and the road is reconstructed by using weaving feature learning and strip convolution in different directions, and finally an accurate road segmentation map is output.
[0102] Step 3: Based on the difference between the road segmentation map of the training set remote sensing image and the labeled example image corresponding to the training set remote sensing image, iterate the parameters of the initial remote sensing image road extraction network model to obtain the remote sensing image road extraction network model;
[0103] In a specific embodiment, the images in the training set are first input to the encoder part in the remote sensing image road extraction network model, and the encoder quickly extracts the feature map and outputs the feature map to the next layer of encoder and snake weaving attention;
[0104] The snake weaving attention adds the feature maps of the encoder and the previous layer input, performs average pooling to reduce random noise, and then uses strip convolution and dynamic snake convolution of different sizes to extract features from the feature maps in both horizontal and vertical directions. Finally, the features in different directions are fused and the feature maps are output to the context information weaving module.
[0105] The context information weaving module adds the feature maps of the encoder and the previous layer input, extracts and fuses the feature maps using standard convolution, dilated convolution and depthwise separable convolution, learns multi-scale context information, and then uses average pooling operations in both horizontal and vertical directions to perform weaving context learning on the feature maps containing multi-scale context information. The feature maps of the encoder and the weaving module are then added and sent to the global information extraction module.
[0106] The global information extraction module performs global average pooling on the feature map to highlight the global information of the feature map, and inputs the feature map after highlighting the global information into RepViT Blocks for global information modeling;
[0107] The multi-scale weaving decoder module performs a deconvolution operation on the feature map, and uses the horizontal and vertical strip depth-separable convolution to reproduce the road information. Then, through upsampling and Sigmoid activation function, the feature map is converted into the probability value of each pixel belonging to a certain category, and an accurate road segmentation map is output.
[0108] By performing parameter iteration on the initial remote sensing image road extraction network model, an optimal remote sensing image road extraction network model is obtained.
[0109] Step 4: inputting the remote sensing image to be processed into the remote sensing image road extraction network model, and outputting the accurate segmentation result of the remote sensing image data;
[0110] In order to further verify the feasibility and effectiveness of the present invention, the present invention is based on Figure 8-11 The embodiment 1 and Figure 12-14 The experiment was conducted in Example 2 shown in the figure, and the experimental results are shown in Tables 1 and 2, where Table 1 is the specific indicators on the DeepGlobe road extraction dataset, and Table 2 is the specific indicators on the Massachusetts road extraction dataset.
[0111]
[0112] Table 1
[0113]
[0114] Table 2
[0115] The architecture of the remote sensing road extraction network was built using the deep learning framework Pytorch1.13.1. The network was tested on the DeepGlobe dataset and the Massachusetts road dataset, and the four main evaluation indicators in the field of semantic segmentation, accuracy, recall, intersection-over-union, and harmonic mean, were used to measure the performance of the model in the road segmentation task. Among them, the accuracy is used to evaluate the accuracy of the model prediction, the recall rate evaluates the coverage of the model, the harmonic mean is the average of the accuracy and recall rate, and the intersection-over-union ratio provides an indicator that directly measures the overlap of the segmented area, which is one of the most commonly used performance evaluation criteria in image segmentation tasks. In the embodiment of the present invention, the accuracy, recall rate, intersection-over-union ratio and harmonic mean in the two road extraction datasets are all high, which shows that the present invention has good performance in both embodiments.
[0116] exist Fig. 9 In the first embodiment shown, a road image in the DeepGlobe dataset before using this method is shown. Fig.10 The corresponding manually annotated road segmentation map is shown. Fig.11 is the predicted road segmentation map obtained after applying the method proposed in this study. Fig.12 , Fig.13 and Fig.14The road images, manually annotated images and predicted road segmentation images of the Massachusetts dataset before and after processing are respectively shown. By comparing the manually annotated road segmentation image and the predicted road segmentation image obtained using the present method, it can be seen that the method provided by the present invention can segment roads very well, and even images with densely distributed roads are not much different from the manually annotated images.
[0117] Compared with the prior art, the present invention includes a serpentine weaving attention that greatly improves the model's attention to roads by imitating "weaving" behavior, a context information weaving module that can make full use of multi-scale context information, a global information extraction module that uses global average pooling operations and novel RepViT Blocks to enhance road integrity and connectivity, and a multi-scale weaving decoder that uses multi-scale strip convolutions to enhance the decoder's reconstruction capability. The present invention is suitable for the road segmentation task of remote sensing images, can efficiently and accurately segment densely distributed roads, and is particularly suitable for road extraction in complex environments.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for dense road segmentation in remote sensing images based on weaving feature extraction, characterized in that: include: Obtain a remote sensing image training set and remote sensing images to be processed, and perform data enhancement operations on the remote sensing images in the remote sensing image training set; The remote sensing image training set includes remote sensing images and label example images corresponding to the remote sensing images; Inputting the remote sensing images in the remote sensing image training set into the initial remote sensing image road extraction network model to obtain a road segmentation map of the remote sensing images in the remote sensing image training set; Based on the difference between the road segmentation map of the remote sensing image in the remote sensing image training set and the labeled example image corresponding to the remote sensing image in the remote sensing image training set, performing parameter iteration on the initial remote sensing image road extraction network model to obtain the remote sensing image road extraction network model; Inputting the remote sensing image to be processed into the remote sensing image road extraction network model, and outputting a road segmentation map of the remote sensing image to be processed; The remote sensing image road extraction network model is a U-Net architecture, including an encoder module, a snake weaving attention module, a context information weaving module, a global information extraction module and a multi-scale weaving decoder module, wherein: The serpentine weaving attention module is used to extract features from the feature map using strip convolutions and dynamic serpentine convolutions of different sizes in both horizontal and vertical directions, and fuse features in different directions to obtain shallow feature maps that adapt to the slender and tortuous characteristics of different roads; the serpentine weaving attention module is set in the shallow part of the jump connection in the U-Net architecture; wherein, The serpentine weaving attention module first uses 1×1 convolution and depth-separable convolution to process the feature map of the previous layer of the serpentine weaving attention module, directly adds the processed feature map and the encoder feature map of the layer where the serpentine weaving attention module is located as the input of the serpentine weaving attention module, and then uses a convolution of size 13×13 and a stride of 1 and an average pooling of size 13×13 to process the input; The snake-weaving attention module transmits the processed feature map to two parallel branches, performs horizontal and vertical weaving learning respectively, adds the feature maps obtained by horizontal learning and vertical learning to the three batches of original input feature maps element by element, and performs 1×1 convolution, batch normalization, and ReLU activation function operations on the addition results in sequence to obtain the output feature map of the snake-weaving attention module; wherein the horizontal learning branch is composed of 1×15, 1×13, 1×11 strip convolutions and a dynamic snake-shaped convolution learned along the X-axis with a size of 13×13, and the vertical learning branch is composed of 15×1, 13×1, 11×1 strip convolutions and a dynamic snake-shaped convolution learned along the Y-axis with a size of 13×13; The context information weaving module is used to learn multi-scale context information using standard convolution, dilated convolution and depth-separable convolution, and then use average pooling operations in both horizontal and vertical directions to perform weaving context learning to obtain a middle-level feature map with multi-scale context information; The global information extraction module is used to highlight the global information of the feature map, perform global information modeling, and obtain a deep feature map with global information.
2. The method according to claim 1, characterized in that The remote sensing image road extraction network model also includes an encoder module, and the encoder module is used to obtain a feature image according to the remote sensing image in the remote sensing image training set; The encoder module consists of five stages, including: The first stage consists of two layers. The first layer consists of a convolution with a kernel size of 3×3 and a stride of 2, a batch normalization operation, and a SiLU activation function. The second layer consists of two FusedMBConv blocks for fast feature map extraction. The second stage consists of encoder 1, which contains four FusedMBConv blocks. Encoder 1 inputs the obtained feature map to encoder 2 and the first snake weaving attention block to learn road details; The third stage consists of encoder 2, which contains four FusedMBConv blocks. Encoder 2 inputs the obtained feature map to encoder 3 and the second snake-weaving attention block to learn road details; The fourth stage consists of encoder 3, which contains two sets of stacked convolution blocks, consisting of 6 MBConv blocks and 9 MBConv blocks respectively. Encoder 3 inputs the obtained feature map to encoder 4 and the first context information weaving module; The fifth stage consists of an encoder 4, which contains two layers. The first layer consists of 15 MBConv blocks, and the second layer consists of convolution with a kernel size of 3×3 and a stride of 2, a batch normalization operation, and a SiLU activation function. The encoder 4 inputs the obtained feature map into the multi-scale weaving decoder module and the global information extraction module; The five stages of the encoder module are arranged from shallow to deep in the remote sensing image road extraction network model architecture.
3. The method according to claim 1, characterized in that The remote sensing image road extraction network model also includes a multi-scale weaving decoder module, which is used to obtain an accurate segmentation result of the remote sensing image according to the feature image, shallow feature map, middle feature map and deep feature map output by the encoder module; The multi-scale braided decoder module includes a multi-scale braided decoder and an upsampling module corresponding to multiple stages of the encoder module; The multi-scale weaving decoder first upsamples the feature map processed by the snake weaving attention module or the context information weaving module through inverse convolution, so that the length and width of the feature map become twice that of the input, and the number of channels is halved; The multi-scale braided decoder inputs the upsampling results into two parallel branches for braided learning; one branch learning consists of 1×5, 1×3 strip depth separable convolutions to reproduce the horizontal road information of the feature map, and the other branch consists of 5×1, 3×1 strip depth separable convolutions to reproduce the vertical road information of the feature map; The multi-scale braided decoder concatenates the outputs of the two branches with the upsampling result, and then performs dimensionality reduction through 1×1 convolution, batch normalization, and ReLU activation function to obtain an output feature map of the multi-scale braided decoder; The upsampling module uses inverse convolution to upsample the output feature map of the multi-scale braided decoder, uses two ReLU activation functions and two convolutions to adjust the number of feature map channels, and finally converts the feature map into a probability value of each pixel belonging to a certain category through a Sigmoid activation function to obtain a road segmentation map of the remote sensing image.
4. The method according to claim 1, characterized in that: The context information weaving module is arranged in the deep part of the jump connection in the U-Net architecture; The context information weaving module first uses 1×1 convolution and depth-separable convolution to process the feature map of the previous layer of the context information weaving module, and directly adds the processed feature map to the encoder feature map of the layer where the context information weaving module is located as the input of the context information weaving module; The context information weaving module is divided into two stages. The first stage performs multi-scale context learning on the input feature map, and the second stage performs weaving context learning in the horizontal and vertical directions. The first stage includes three branches, the first is a standard convolution of 5×5 size, the second is a dilated convolution of 7×7 size with a dilation factor of 2 and a dilated convolution of 9×9 size with a dilation factor of 3, and the third is a depthwise separable convolution of 11×11 and 13×13. The context information weaving module fuses the convolution results of the three branches to obtain a feature map containing multi-scale context information, and then processes it with 1×1 convolution, batch normalization and ReLU activation function in sequence, multiplies the processing result with the original input and then adds them, and inputs them into the second stage; The context information weaving module performs H×1 and 1×W average pooling operations on the input feature map in the second stage, and then multiplies the input feature map respectively to make the vertical and horizontal context information more prominent; The two multiplication results are then added together, and the added result is processed with 1×1 convolution, batch normalization and ReLU activation function, and then directly added to the input of the second stage to obtain the output feature map of the context information weaving module.
5. The method according to claim 1, characterized in that The global information extraction module is placed in the bottleneck part of the U-Net architecture; The global information extraction module includes two global feature learning branches, and the inputs are the output feature map of the context information weaving module in the previous layer of the global information extraction module and the output feature map of the encoder of the corresponding layer of the global information extraction module; The global information extraction module first performs global average pooling on two batches of feature maps to extract the global information of the feature maps, and then multiplies the results of the global average pooling with the input feature maps to enhance the global information of the feature maps; then the inputs of the branches are added to the multiplication results, and input into RepViT Blocks for global modeling, the output feature maps of RepViT Blocks are concatenated with the original input, and then the feature maps are reduced in dimension using 1×1 convolution, batch normalization, and ReLU activation function to obtain the output feature maps of the global information extraction module.
6. The method according to claim 1, characterized in that The data enhancement operation is to perform data enhancement on the remote sensing images in the remote sensing image training set and the label example images corresponding to the remote sensing images by image flipping, random rotation, size scaling, adding random noise, and light and dark transformation operations.
Citation Information
Patent Citations
Remote sensing image road segmentation method combining channel attention mechanism and multilayer axial Transform feature fusion structure
CN118351538A
Remote sensing image dense road segmentation method using strip-shaped features
CN118429356A