Road information extraction method based on self-attention model and convolutional neural network
By combining the self-attention model and convolutional neural network, road information in remote sensing images is extracted, and road segmentation problems in the existing technology are solved, and better road segmentation connectivity and accuracy are achieved.
Patent Information
- Application Number
- CN202310244138.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-03-14
AI Technical Summary
The prior art has problems of poor segmentation and fragmentation when extracting road information from remote sensing images, making it difficult to achieve fast and accurate road segmentation.
The road information extraction method based on the self-attention model and the convolutional neural network is adopted to extract global information through the self-attention model, and the convolutional neural network extracts local information, and combines the pixel connectivity structure to improve the connectivity of road segmentation.
It improves the connectivity and accuracy of road segmentation, alleviates the phenomenon of road segmentation results in remote sensing road data, and improves the overall effect of the model on road extraction.
Smart Images

Figure CN116452987B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a road information extraction method based on a self-attention model and a convolutional neural network. Background Art
[0002] Roads are an important part of geographic information systems. Obtaining timely and complete road information plays an important role in digital city construction, public transportation, and unmanned vehicle driving. In recent years, with the rapid development of remote sensing technology, remote sensing images have been greatly improved in terms of spatial and spectral resolution. Extracting roads from high-resolution images has gradually become a research hotspot. However, the manual road extraction method not only has a long cycle, but is also easily affected by human subjective factors. Massive satellite remote sensing data is generated every day, and it is obviously not feasible to rely entirely on manpower to extract roads. In order to quickly extract road information from remote sensing images, people have done a lot of research on using remote sensing images for road extraction, and have formed many methods with different extraction accuracy. These traditional methods can be divided into two categories according to different extraction tasks. The first category relies on expert knowledge, road geometric features and shape features, and extracts road skeletons through template matching, knowledge-driven and other algorithms. This type of method has the disadvantages of high computational complexity and low degree of automation. The second category uses object-oriented thinking to detect all road areas in remote sensing images through methods such as image segmentation and support vector machines to obtain road information. This type of method is affected by problems such as building shadows and uneven road grayscale changes, which leads to a large number of road breaks. In addition, the road shapes in remote sensing images are complex and the scales are different, so the road information extraction effect is not good.
[0003] In order to segment roads quickly and accurately, the current method of extracting roads based on deep learning has gradually become an efficient and automated solution. However, in the road segmentation task for satellite remote sensing, since the roads have a large span and are relatively narrow, the overall segmentation effect of the final model is usually poor and fragmented. Summary of the invention
[0004] The purpose of the present invention is to solve the problem of poor segmentation and fragmentation in the prior art of extracting road information from remote sensing images, and provide a road information extraction method based on a self-attention model and a convolutional neural network. The method adopts a self-attention model structure to extract global information to improve the fragmentation problem of road segmentation, and adopts a convolutional neural network structure to extract local information to improve the segmentation effect of the road edge, and can improve the connectivity of the road segmentation.
[0005] The purpose of the present invention is mainly achieved through the following technical solutions:
[0006] The road information extraction method based on the self-attention model and convolutional neural network includes the following steps:
[0007] Step S1, selecting remote sensing road data as raw data, and preprocessing the raw data;
[0008] Step S2, using the encoder of the road segmentation model to perform image spatial resolution compression and feature extraction on the original data image; wherein the road segmentation model includes an encoder and a decoder, the encoder includes an image segmentation module, four lightweight self-attention structures and a feature information fusion module, the image segmentation module is used to input the preprocessed original data image information, perform image spatial resolution compression and feature extraction to obtain the divided image and output it to the lightweight self-attention structure; the lightweight self-attention structure introduces a scaling factor to reduce time complexity, introduces a space reduction calculation to compress the output dimension, the lightweight self-attention structure is used to extract image features from the image segmentation module output image, and the feature information fusion module is used to fuse the output features of the four lightweight self-attention structures to obtain the final output result of the encoder;
[0009] Step S3, a decoder based on the road segmentation model generates a pixel connectivity structure prediction result and a road segmentation result; wherein the decoder is implemented based on a convolutional neural network and uses road labels and pixel connectivity structure labels for learning;
[0010] Step S4, reversely derive the output result with the same resolution as the input image based on the pixel connectivity structure prediction result, and combine the road segmentation result to obtain the final output. The lightweight self-attention structure of the present invention introduces a scaling factor and a space reduction calculation on the basis of the self-attention structure to reduce the amount of calculation, thereby realizing the lightweight function of the self-attention structure.
[0011] When the present invention is applied, the encoder in the road segmentation model uses a module composed of four lightweight self-attention structures to extract the feature information of the input image, and then performs feature fusion to obtain the final output result of the encoder structure, and then generates the prediction result of the pixel connectivity structure and the road segmentation result, which can improve the comprehensive expression ability of the network's high-level semantic information and low-level detail information. The present invention uses a lightweight self-attention structure to extract image features, which can reduce the amount of calculation introduced, introduce a scaling factor to reduce the overall time complexity of the self-attention structure, and also reduce a lot of computational costs for the overall road segmentation model, and can also provide rich long-distance dependency capture and global context information for the overall model. The present invention uses road segmentation results and pixel connectivity results, which are responsible for segmenting and connectivity detection of roads, respectively, and quickly achieve the final result prediction.
[0012] Furthermore, the image segmentation module adopts a convolution structure with a step size of 4 and a convolution kernel size of 7×7, the number of channels of the image size of the original data is 3, the number of channels of the image size output by the image segmentation module is 32, and the image size of the original data B×3×H×W is converted into an image output of size B×32×H / 4×W / 4 by the image segmentation module, where B is the image batch, H is the height of the image, and W is the width of the image. The road segmentation model of the present invention uses the image segmentation module to perform a feature extraction on the image and retain the connection information between different areas of the image.
[0013] Furthermore, the lightweight self-attention structure introduces a scaling factor to reduce the time complexity. Specifically, the lightweight self-attention structure introduces a scaling factor to reduce the overall time complexity of the lightweight self-attention structure from O(N 2 ) is reduced to Where r is the scaling factor, O(N 2 ) is the time complexity of the self-attention structure, To reduce the time complexity of the self-attention structure, N = H × W, where H is the height of the image and W is the width of the image;
[0014] The lightweight self-attention structure introduces space reduction calculation to compress the output dimension. Specifically: The lightweight self-attention structure introduces space reduction calculation to QW j q and KW j w The dimension is compressed, where the space reduction calculation formula is:
[0015] SR(x)=Norm(Reshape(x,r)·W)
[0016] Among them, Q, K and V correspond to the input of self-attention respectively, j is the index, x is the input, which refers to Q, K, V respectively, SR(x) is the scaling of x, Attention(Q, K, V) is the self-attention calculation, SR(Q) is the scaling of Q, T is the transposition, and SR(K) is T To scale K and then transpose, Reshape(x, r) converts the dimension of x from HW×C to Norm means standardization, Cr 2 To represent this value, is the dimension of the word matrix, the first dimension is Cr^2, the second dimension is C, C is the number of categories, H is the height of the image, and W is the width of the image.
[0017] Furthermore, the feature information fusion module fuses the output features of the four lightweight self-attention structures to obtain the final output result of the road segmentation model, including:
[0018] First, the output of each lightweight self-attention structure is upsampled to the input image resolution. Then the four output results are spliced according to the channels, and finally a convolutional neural network with batch normalization, a convolution kernel size of 1, and a stride of 1 is used to operate on the spliced results to obtain the final output result of the road segmentation model.
[0019] Furthermore, the step S3 of generating the pixel connectivity structure prediction result includes:
[0020] Step S31, initializing the pixel spacing distance d; initializing the pixel connectivity structure label element values to be 0, and the shape is: 8×H×W, where 8 represents the 8 orientations of the current pixel point, H is the height of the image, and W is the width of the image;
[0021] Step S32, starting from the first pixel in the upper left corner, according to the principle of from left to right and from top to bottom, search for the current pixel and the upper left pixel with a distance d. If both are road target pixels, the label of the current pixel is recorded as 1;
[0022] Step S33, repeating step S32, traversing the eight directions of upper left, upper, upper right, left, right, lower left, lower and lower right, and obtaining a pixel connectivity structure label of a shape of 8×H×W, which respectively represents the connectivity relationship of the eight directions;
[0023] The pixel point labels corresponding to the generated pixel connectivity structure labels only contain 0 and 1, where 1 indicates the existence of connectivity and 0 indicates the absence of connectivity.
[0024] Furthermore, the step S4 obtains output results in eight directions with the same resolution as the input image.
[0025] In summary, compared with the prior art, the present invention has the following beneficial effects: the present invention combines the respective advantages of convolutional neural networks and self-attention models, which are used for local information and global information extraction respectively, improves the stability and robustness of the overall model structure, and proposes the use of pixel connectivity structure to predict the connectivity of roads, alleviates the phenomenon of disconnected road segmentation results in remote sensing road datasets, and improves the overall effect of the model on road extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:
[0027] Figure 1 A flowchart of a specific embodiment of the present invention;
[0028] Figure 2A framework diagram of a network architecture in a specific embodiment of the present invention;
[0029] Figure 3 A partial image of remote sensing road data used in a specific embodiment of the present invention;
[0030] Figure 4 An example diagram of an existing self-attention structure and a self-attention structure using space reduction in a specific embodiment of the present invention;
[0031] Figure 5 A schematic diagram of a process of generating a connectivity structure label in a specific embodiment of the present invention;
[0032] Figure 6 A schematic diagram of a segmentation result obtained by reversely deducing the output of the pixel connectivity structure in a specific embodiment of the present invention;
[0033] Figure 7 A schematic diagram comparing the remote sensing image segmentation results of a specific embodiment of the present invention and the prior art;
[0034] Figure 8 A comparison diagram of segmentation results of three different parameter versions in a specific embodiment of the present invention;
[0035] Fig. 9 A comparison diagram of segmentation results in a specific embodiment of the present invention when a method of generating pixel connectivity structure labels is not used and a method of generating pixel connectivity structure labels is used;
[0036] Fig.10 This is a comparison diagram of a feature map obtained by intermediate network calculation when a specific embodiment of the present invention is applied with the prior art. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.
[0038] Example:
[0039] like Figure 1 and Figure 2As shown, the road information extraction method based on the self-attention model and the convolutional neural network includes the following steps: step S1, selecting remote sensing road data as raw data, and preprocessing the raw data; step S2, using the encoder of the road segmentation model to perform image spatial resolution compression and feature extraction on the raw data image; wherein the road segmentation model includes an encoder and a decoder, the encoder includes an image division module, 4 lightweight self-attention structures and a feature information fusion module, the image division module is used to input the preprocessed raw data image information, perform image spatial resolution compression and feature extraction to obtain the divided image and output it to the lightweight self-attention structure; the lightweight self-attention structure introduces a compression The time complexity is reduced by a factor, and a space reduction calculation is introduced to compress the output dimension. The lightweight self-attention structure is used to extract image features from the image segmentation module output image. The feature information fusion module is used to fuse the output features of the four lightweight self-attention structures to obtain the final output result of the encoder; step S3, a decoder based on the road segmentation model generates a pixel connectivity structure prediction result and a road segmentation result; wherein the decoder is implemented based on a convolutional neural network and uses road labels and pixel connectivity structure labels for learning; step S4, reverse deduction is performed based on the pixel connectivity structure prediction result to obtain an output result with the same resolution as the input image, and the final output is obtained in combination with the road segmentation result.
[0040] This embodiment uses existing cropping and data enhancement methods to preprocess the raw data. The purpose of image cropping is to remove the area outside the study, reduce the complexity of calculation, and effectively improve the calculation speed; the purpose of data enhancement is to enhance the visual effect, improve the image quality and highlight the required information, which is conducive to analysis and interpretation or further processing. It mainly enriches the diversity of the data set from the perspective of prior knowledge and improves the robustness of the model during subsequent data processing. The road segmentation model of this embodiment is defined as the Seg-Road model.
[0041] The image segmentation module of this embodiment is different from the Swin Transformer and other models that use the image segmentation module to cut the input image and add relevant position embedding information. The image segmentation module of this embodiment uses a convolution structure with a step size of 4 and a convolution kernel size of 7×7. The number of channels of the image size of the original data is 3, and the number of channels of the image size output by the image segmentation module is 32. The image size of the original data B×3×H×W is converted to an image output of size B×32×H / 4×W / 4 by the image segmentation module, where B is the image batch, H is the height of the image, and W is the width of the image. Compared with the prior art that directly cuts the image, the Seg-Road model of this embodiment extracts features of the image once through the image segmentation module, and retains the connection information between different areas of the image during this feature extraction.
[0042] The lightweight self-attention structure of this embodiment introduces a scaling factor to reduce the time complexity. Specifically, the lightweight self-attention structure introduces a scaling factor to reduce the overall time complexity of the lightweight self-attention structure from O(N 2 ) is reduced to Where r is the scaling factor, O(N 2 ) is the time complexity of the self-attention structure, To reduce the time complexity of the self-attention structure, N = H × W, where H is the height of the image and W is the width of the image. This embodiment reduces the overall time complexity of the self-attention mechanism by introducing a scaling factor, which also reduces a lot of computational costs for the overall Seg-Road model, and can also provide the overall model with rich long-distance dependency capture and global context information. Especially when the input scale of the model is large, the computational speed advantage of the improved self-attention mechanism is more obvious.
[0043] Usually Transformer-based uses Vanilla Transformer for image feature extraction, using Q, K, and V as inputs, and their dimensions are all N×C, where N=H×W. d k The value is C, which is used to prevent Q and K T The inner product is too large.
[0044]
[0045]
[0046] Softmax is used to adjust QW j q (KW j w ) T The shape is HW×HW for dimension normalization, which is used to perceive the connection between two pixels, and introduces global perceptual attention to the overall model. Compared with the CNN model, it has more global attention. In order to avoid the large amount of computational cost introduced by directly using Vanilla Transformer. The lightweight self-attention structure of this embodiment introduces space reduction calculation to compress the output dimension. Specifically: The lightweight self-attention structure introduces space reduction calculation to QW j q and KW j w The dimension can be compressed to reduce the amount of calculation introduced. The space reduction calculation formula is:
[0047] SR(x)=Norm(Reshape(x,r)·W)
[0048]
[0049] Among them, Q, K and V correspond to the input of self-attention respectively, j is the index, x is the input, which refers to Q, K, V respectively, SR(x) is the scaling of x, Attention(Q, K, V) is the self-attention calculation, SR(Q) is the scaling of Q, T is the transposition, and SR(K) is T To scale K and then transpose, Reshape(x, r) converts the dimension of x from HW×C to Norm means standardization, Cr 2 To represent this value, is the dimension of the word matrix, the first dimension is Cr^2, the second dimension is C, C is the number of categories, H is the height of the image, and W is the width of the image.
[0050] In the road segmentation model of this embodiment, the encoder uses 4 blocks composed of lightweight self-attention structures to extract feature information of the input image, which are respectively recorded as Block 1, Block 2, Block 3 and Block 4. Since semantic segmentation not only needs to predict the pixel category, but also needs to determine the pixel position of the pixel target, in order to accurately segment the roads in the remote sensing image, the road segmentation model uses the output results of the 4 blocks in the encoder for feature fusion, thereby improving the comprehensive expression ability of the network's high-level semantic information and low-level detail information.
[0051] like Figure 4 As shown in FIG. 1 , in order to further simplify the computational cost of the model, the feature information fusion module of this embodiment fuses the output features of the four lightweight self-attention structures to obtain the final output result of the road segmentation model, including: firstly upsampling the output result of each lightweight self-attention structure to the input image resolution. Then the four output results are spliced according to the channels, and finally a convolutional neural network with batch normalization, a convolution kernel size of 1, and a stride of 1 is used to operate on the spliced results to obtain the final output result of the road segmentation model.
[0052] Usually in semantic segmentation tasks, it is difficult to accurately classify edge pixels due to the lack of rich object information and the difficulty in defining edge areas. In road segmentation tasks, because the shape of the road is usually narrow and the span is large, there are a lot of edge segmentation problems, which leads to discontinuities in the prediction results. Although the road segmentation task has many difficulties, it has good connectivity. Therefore, in order to improve the accuracy of road segmentation, this embodiment proposes to use a pixel connectivity structure to improve the fragmentation problem of road segmentation. Figure 5As shown, step S3 of this embodiment generates the pixel connectivity structure prediction result, including: step S31, initializing the pixel interval distance d; initializing the pixel connectivity structure label element values to 0, and the shape is: 8×H×W, where 8 represents the 8 orientations of the current pixel, H is the height of the image, and W is the width of the image; step S32, starting from the first pixel in the upper left corner, according to the principle of from left to right and from top to bottom, find the current pixel and the upper left pixel with a distance of d. If both are road target pixels, the label of the current pixel is recorded as 1; step S33, repeating step S32, traversing the 8 directions of upper left, upper, upper right, left, right, lower left, lower and lower right, and obtaining a pixel connectivity structure label of 8×H×W shape, which respectively represents the connectivity relationship of the 8 orientations, and is used to improve the fragmentation problem of road segmentation in remote sensing images. Among them, the pixel point labels corresponding to the generated pixel connectivity structure label only contain 0 and 1, 1 indicates that there is a connectivity phenomenon, and 0 indicates that there is no connectivity phenomenon. In step S32 of this embodiment, the pixel connectivity structure of this embodiment is trained using pixel connectivity structure labels from left to right and from top to bottom, specifically from left to right and from top to bottom of the road label image of the remote sensing image.
[0053] In this embodiment, the pixel connectivity structure is used to provide certain connectivity information for the model, and the pixel spacing distance r=2 is used in the actual road segmentation model to assist the road segmentation in the remote sensing image. During the training phase of the pixel connectivity structure, because the corresponding label elements only contain 0 and 1, 1 represents the existence of a connectivity relationship, the BCE loss is used as the loss function, as shown in the following formula, where y i represents the connectivity label, p i Represents the prediction result, and C represents the number of H×W pixels.
[0054]
[0055] In the training of the segmentation branch in the pixel connectivity structure, no class balance loss such as Focal Loss is introduced, but the Binary Cross Entropy loss function is also used to achieve the final pixel-by-pixel classification task. In the training of the pixel connectivity structure in the road segmentation model, BCE loss is used. In order to ensure the optimization priority of the segmentation branch, the overall loss function is shown below, where α is taken as 0.2 in actual training; where Loss = L seg +αL con .
[0056] like Figure 6As shown, step S4 of this embodiment uses the prediction result of the pixel connectivity structure label to perform reverse deduction to generate the final prediction result, which is the reverse process of generating labels from the pixel connectivity structure, and predicts the output results of 8 channels (orientations) with the same resolution as the input image. The road segmentation model of this embodiment is mainly composed of two segmentation branches, the conventional semantic segmentation branch and the pixel connectivity structure, which are responsible for the segmentation and connectivity detection of the road respectively. Among them, the pixel connectivity structure mainly participates in multi-task training and provides connectivity information for the model. However, in order to make full use of the information of the pixel connectivity structure prediction results, it is proposed to use the pixel connectivity structure prediction results to perform reverse deduction to generate the final prediction result. The road segmentation model of this embodiment predicts the pixel connectivity structure to obtain the output results of 8 channels (orientations) with the same resolution as the input image. Taking the output result of the upper left corner orientation as an example, as shown Figure 6 As shown in the left figure, the road segmentation result obtained by reverse mapping is as follows Figure 6 As shown in the figure on the right, the output results of the eight directions are reverse mapped and the union is taken as the connectivity output result of the model.
[0057] In order to verify the effectiveness of the road segmentation model proposed in this embodiment for segmenting roads in remote sensing images, this embodiment uses the DeepGlobe remote sensing image dataset for testing. The DeepGlobe road dataset is a set of high-resolution remote sensing image road datasets proposed by the 2018 DeepGlobe Road Extraction Challenge, of which there are 6226 images with labeled data. The image scenes are cities, villages, wilderness, seaside, tropical rainforest and other scenes in Thailand, India and Indonesia. Some images in the dataset are as follows: Figure 3 As shown. This embodiment also uses the existing DeepRoadMapper, LinkNet34, D-LinkNet, PSPNet, RoadCNN, CoANet and CoANet-UB models for comparison. Among them, the Seg-Road proposed in this embodiment includes three versions with different parameter amounts. The parameters (parameters) and detection speeds (Frames Per second, FPS) of the three different parameter amount versions are shown in Table 1, among which Seg-Road-s has the least parameters and the fastest calculation speed; Seg-Road-m has fewer parameters and faster calculation speed, and Seg-Road-l has more parameters and slower calculation speed.
[0058] Table 1 Comparison of parameter quantities and detection speeds of three different parameter quantity versions of this embodiment
[0059] Seg-Road-s Seg-Road-m Seg-Road-l Parameters(Mb) 4.18 14.12 28.67 FPS 98 81 42
[0060] Specifically, the relevant evaluation indicators are explained as follows. This embodiment selects Mean Intersection over Union (MIoU), Precision, Recall and F1 as the main evaluation indicators. Among them, Precision refers to the ratio of the number of pixels correctly predicted as roads to the number of pixels whose prediction results are roads, and the calculation formula is shown in Formula (1). Recall refers to the ratio of the number of pixels correctly predicted as roads to the number of pixels that are actually road labels, and the calculation formula is shown in Formula (2). F1 is a comprehensive evaluation of precision and recall, and the calculation formula is shown in Formula (3). Among them, TP refers to correctly predicted positive examples, FP refers to incorrectly predicted positive examples, and FN refers to incorrectly predicted negative examples.
[0061]
[0062]
[0063]
[0064] MIOU is a classic and authoritative measurement indicator for semantic segmentation. Its calculation method is shown in formula (4), where K represents the number of categories and pij can be understood as the number of targets of category i predicted as targets of category j.
[0065]
[0066] The detailed evaluation index results are shown in Table 2:
[0067] Table 2 Comparison of the effectiveness of the Seg-Road model of this embodiment and the existing model for segmenting roads in remote sensing images
[0068]
[0069]
[0070] It can be clearly observed from the above table that compared with DeepRoadMapper, LinkNet34, D-LinkNet, PSPNet, RoadCNN, CoANet and CoANet-UB models, Seg-Road-s, Seg-Road-m and Seg-Road-l have achieved significant improvements. Excellent road segmentation effect is also obtained in CoANet-UB, and Seg-Road-m and Seg-Road-l proposed in this embodiment have achieved better segmentation effect than CoANet-UB under the premise of less parameters.
[0071] Secondly, this example randomly selects four remote sensing images in the test set and uses PSPNet, LinkNet32, CoANet-UB and Seg-Road to predict the results. The results are as follows: Figure 7 As displayed. Figure 7 In the figure, (a) is the input remote sensing image, (b) is the true value of the data set, (c) is the calculation result of PSPNet, (d) is the calculation result of LinkNet32, (e) is the calculation result of CoANet-UB, and (f) is the calculation result of Seg-Road-l. It can be clearly observed that compared with CoANet-UB and Seg-Road, the detection results of PSPNet and LinkNet32 without connectivity branches have a lot of fragmentation, and the segmentation effect at the edge of the road is not good. The overall segmentation results of CoANet-UB are better connected, but compared with Seg-Road, there are more FPs, so the overall evaluation index is lower than Seg-Road. It is worth mentioning that Seg-Road has better connectivity for road segmentation in remote sensing images, fewer FPs, and the highest accuracy.
[0072] In addition, Seg-Road contains three version models with different parameter quantities: Seg-Road-s, Seg-Road-m and Seg-Road-l. The overall structure of the model is roughly the same, and only the number of channels in the intermediate calculation process and the number of space reduction transformations used in each module are different. Compared with Seg-Road-m and Seg-Road-l, Seg-Road-s has a faster calculation speed and can meet the usage scenarios with high real-time requirements. The Seg-Road-s proposed in this embodiment also has a faster segmentation speed and better segmentation effect, and is more suitable for segmentation scenarios with lower real-time requirements. In order to more clearly show the difference in segmentation effects of the three version models, a remote sensing image was randomly selected from the verification set for segmentation. The effect is shown in the figure. Figure 8 As displayed. Figure 8 From left to right, the first picture is the segmentation result of Seg-Road-s, the second picture is the segmentation result of Seg-Road-m, and the third picture is the segmentation result of Seg-Road-l. It can be observed that the three Seg-Roads with different parameter values have good connectivity performance. Seg-Road-l detects the detailed features of the road better and has relatively fewer FPs.
[0073] In Seg-Road, the pixel connectivity structure is used to enhance the network's perception of the connectivity between road pixels in remote sensing images. On the contrary, there is no connectivity prediction in segmentation models such as LinkNet34. Therefore, it is speculated that the connectivity structure provides a more accurate segmentation effect for the model as a whole. In order to verify this conjecture, taking Seg-Road-s as an example, the connectivity structure is canceled, and only the segmentation branch is used to retrain and test the DeepGlobe data under the same hyperparameters. The evaluation index obtained is: 72.46% MIOU. Therefore, this conjecture is more fully verified. The connectivity structure proposed in this embodiment will bring excellent segmentation performance for the segmentation of roads in remote sensing images, and even for targets of similar style to be segmented. Fig. 9 The figure shows the segmentation results of Seg-Road and Seg-Road without connectivity structure. The left figure is the segmentation result of Seg-Road without connectivity structure, and the right figure is the segmentation result of Seg-Road with connectivity structure. It can be observed that the segmentation results of Seg-Road without connectivity structure for roads in remote sensing images are more fragmented, while Seg-Road has stronger overall connectivity and better segmentation effect on road details.
[0074] Convolutional neural networks are generally understood to have better local information extraction capabilities, while networks based on self-attention models are generally understood to have better global information extraction capabilities. The Seg-Road proposed in this embodiment uses a spatial reduction transform in the encoder to extract features of the input image, and uses a convolutional neural network in the encoder for multi-scale feature fusion and prediction of the final result. In contrast, many current semantic segmentation models use only convolutional neural networks and self-attention models, resulting in a large deviation between global information and local information, and the segmentation effect of the overall image model. In order to demonstrate the advantages of the combined self-attention model and convolutional neural network, this embodiment visualizes the feature map of the intermediate calculated image, such as Fig.10 As shown, Fig.10 From left to right, the first picture is the input remote sensing image, the second picture is the feature map obtained by the intermediate calculation of PSPNet, the third picture is the feature map obtained by the intermediate calculation of CoANet, and the fourth picture is the feature map obtained by the intermediate calculation of Seg-Road. It can be clearly observed that the global connectivity information and local detail information of the intermediate feature map are well utilized, and the road features extracted by PSPNet and CoANet-UB are more accurate and clear. Therefore, in comparison, the Seg-Road proposed in this embodiment has achieved better segmentation results on the DeepGlobe dataset, and it is believed that it can achieve better prediction results in similar tasks.
[0075] In summary, in order to solve the problem that the road segmentation results in remote sensing images are usually discontinuous, this embodiment proposes a new semantic segmentation model Seg-Road based on the self-attention model and convolutional neural network. And Seg-Road proposes to use the pixel connectivity structure to enhance the network's perception of the connection information of road pixels, which to a certain extent optimizes the problem of fragmentation of road segmentation in remote sensing images in the current semantic segmentation model. By using the DeepGlobe dataset for verification, the results show that Seg-Road has achieved the current highest level of results, MIOU 82.06%, F1 91.43%, precision 90.05%, recall 92.85%, surpassing segmentation models such as LinkNet, D-LinkNet, PSPNet and CoANet. In addition, this embodiment proposes three versions of Seg-Road with different parameter quantities, namely Seg-Road-s, Seg-Road-m and Seg-Road-l, and all have achieved good segmentation effects, which can be selected according to the requirements of real-time segmentation in different application scenarios. Finally, compared with the existing semantic segmentation models, the Seg-Road proposed in this embodiment has higher practical application value for road segmentation in remote sensing images.
[0076] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A road information extraction method based on a self-attention model and a convolutional neural network, characterized in that: The following steps are involved: Step S1, selecting remote sensing road data as raw data, and preprocessing the raw data; Step S2, using the encoder of the road segmentation model to perform image spatial resolution compression and feature extraction on the original data image; wherein the road segmentation model includes an encoder and a decoder, the encoder includes an image segmentation module, four lightweight self-attention structures and a feature information fusion module, the image segmentation module is used to input the preprocessed original data image information, perform image spatial resolution compression and feature extraction to obtain the divided image and output it to the lightweight self-attention structure; the lightweight self-attention structure introduces a scaling factor to reduce time complexity, introduces a space reduction calculation to compress the output dimension, the lightweight self-attention structure is used to extract image features from the image segmentation module output image, and the feature information fusion module is used to fuse the output features of the four lightweight self-attention structures to obtain the final output result of the encoder; Step S3, a decoder based on the road segmentation model generates a pixel connectivity structure prediction result and a road segmentation result; wherein the decoder is implemented based on a convolutional neural network and uses road labels and pixel connectivity structure labels for learning; Step S4, reversely deriving the pixel connectivity structure prediction result to obtain an output result with the same resolution as the input image, and combining it with the road segmentation result to obtain the final output; The lightweight self-attention structure introduces a scaling factor to reduce the time complexity. Specifically, the lightweight self-attention structure introduces a scaling factor to reduce the overall time complexity of the lightweight self-attention structure from O(N 2 ) is reduced to Where r is the scaling factor, O(N 2 ) is the time complexity of the self-attention structure, To reduce the time complexity of the self-attention structure, N = H × W, where H is the height of the image and W is the width of the image; The lightweight self-attention structure introduces space reduction calculation to compress the output dimension. Specifically: The lightweight self-attention structure introduces space reduction calculation to QW j q and The dimension is compressed, where the space reduction calculation formula is: SR(x)=Norm(Reshape(x,r)·W) Among them, Q, K and V correspond to the input of self-attention respectively, j is the index, x is the input, which refers to Q, K, V respectively, SR(x) is the scaling of x, Attention(Q, K, V) is the self-attention calculation, SR(Q) is the scaling of Q, T is the transposition, and SR(K) is T To scale K and then transpose, Reshape(x, r) converts the dimension of x from HW×C to Norm means standardization, Cr 2 is a numerical value, is the dimension of the word matrix, the first dimension is Cr^2, the second dimension is C, C is the number of categories, H is the height of the image, and W is the width of the image; The step S3 of generating the pixel connectivity structure prediction result comprises: Step S31, initializing the pixel spacing distance d; initializing the pixel connectivity structure label element values to be 0, and the shape is: 8×H×W, where 8 represents the 8 orientations of the current pixel point, H is the height of the image, and W is the width of the image; Step S32, starting from the first pixel in the upper left corner, according to the principle of from left to right and from top to bottom, search for the current pixel and the upper left pixel with a distance d. If both are road target pixels, the label of the current pixel is recorded as 1; Step S33, repeating step S32, traversing the eight directions of upper left, upper, upper right, left, right, lower left, lower and lower right, and obtaining a pixel connectivity structure label of a shape of 8×H×W, which respectively represents the connectivity relationship of the eight directions; The pixel point labels corresponding to the generated pixel connectivity structure labels only contain 0 and 1, where 1 indicates the existence of connectivity and 0 indicates the absence of connectivity.
2. The road information extraction method based on the self-attention model and convolutional neural network according to claim 1, characterized in that: The image division module adopts a convolution structure with a step size of 4 and a convolution kernel size of 7×7. The number of channels of the image size of the original data is 3, and the number of channels of the image size output by the image division module is 32. The image size B×3×H×W of the original data is converted into an image output of size B×32×H / 4×W / 4 by the image division module, where B is the image batch, H is the height of the image, and W is the width of the image.
3. The road information extraction method based on the self-attention model and convolutional neural network according to claim 1, characterized in that: The feature information fusion module fuses the output features of the four lightweight self-attention structures to obtain the final output results of the road segmentation model, including: First, the output of each lightweight self-attention structure is upsampled to the input image resolution. Then the four output results are spliced according to the channels, and finally a convolutional neural network with batch normalization, a convolution kernel size of 1, and a stride of 1 is used to operate on the spliced results to obtain the final output result of the road segmentation model.
4. The road information extraction method based on the self-attention model and convolutional neural network according to claim 1, characterized in that: The step S4 obtains output results in eight directions with the same resolution as the input image.
Citation Information
Patent Citations
Remote sensing image automatic road extraction method based on deep convolutional neural network
CN112749578A
Remote sensing image road segmentation method combining intensive attention and parallel upsampling
CN114092824A