Enhanced Method for Urban Street View Semantic Segmentation Based on Multidimensional Attention Mechanism
Through the multi-dimensional attention fusion module MAFM, the contradiction between accuracy and speed in semantic segmentation of urban street scenes is solved, and attention weight is extracted using strip pooling operation, which improves the segmentation accuracy and speed of strip targets in urban street scenes.
Patent Information
- Application Number
- CN202210692153.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-17
AI Technical Summary
The existing technology has a contradiction between accuracy and computing speed in semantic segmentation of urban street scenes. The traditional attention mechanism is complex in computing and occupies a large amount of GPU memory, making it difficult to efficiently handle complex scenes in urban street scenes.
The multi-dimensional attention fusion module MAFM is constructed to extract the attention weights on the height and width of the feature map through strip pooling operations, combine channel domain and spatial domain attention, reduce the calculation complexity and improve segmentation accuracy, and adapt to strip target objects in urban street scenes.
With the reduction of the number of parameters, the accuracy and speed of semantic segmentation of urban street scenes are improved, and the segmentation effect of strip objects such as roads, high-rise buildings, and street lights is adapted to the complex scenes in urban street scenes, especially the segmentation effect of strip objects such as roads, high-rise buildings, and street lights.
Smart Images

Figure CN115035298B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the fields of artificial intelligence and image processing, and specifically relates to an enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism in an urban context. Background Art
[0002] Semantic image segmentation is a basic task in computer vision. Traditional segmentation mainly extracts low-level features of pictures and then performs segmentation, such as threshold segmentation method, edge detection method, region segmentation method, etc. This stage is generally unsupervised learning, and the segmented results lack semantic annotation. Image semantic segmentation based on deep learning can perform semantic division according to labels, and has the advantages of batch processing and multi-classification, and has been widely applied in various fields. Such as biomedical, drone aerial photography, image editing, etc. Urban scene image semantic segmentation takes urban street scene images as the research object to understand the complex street scenes and traffic conditions in the city, and thus analyze and obtain road condition information. This technology is of great significance for realizing potential application fields such as autonomous driving, robot sensing, and image processing in the city.
[0003] Introducing the soft attention mechanism is one of the effective means to enhance image context correlation and establish pixel long-range dependence. In the current research related to the attention mechanism, the structure can be roughly divided into three categories: channel attention, spatial attention, and hybrid attention. Channel attention uses global pooling to extract channel features, with few parameters. For example, the SE module in SENet obtains a global receptive field through global average pooling, emphasizes the weights of different channels, and proves the necessity of channel attention for result improvement. ECANet continues this theory and proposes a non-dimensionality-reducing local cross-channel interaction strategy, significantly reducing the complexity of the model. However, such operations ignore the attention of the pixels themselves and lose segmentation details. Spatial attention is usually combined with multi-scale input and pyramid structures. The feature map expands the receptive field through convolutional kernels of different sizes to capture context correlation and strengthen the correlation between pixels in the same frame and between pixels in different frames. For example, CBAM captures spatial attention by combining average pooling and max pooling; the non-local block in the non-local neural network combines all dimensions except the channel, and establishes the relationship between the current pixel and all other pixels through dot product operations. Although such methods ensure accuracy, the dot product operation will introduce a large amount of computation and occupy a large amount of GPU memory at the same time. Hybrid attention combines channel and spatial attention at the same time. For example, DANet merges the dimensions except the number of channels through a reshape operation, then calculates the similarity between all pixels and all pixels through matrix dot product operations, and then fuses with channel attention, with a very high spatial complexity. Therefore, a balance needs to be made between computing resources and computing accuracy. Summary of the Invention
[0004] The purpose of this application is to provide an enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism. Aiming at the contradiction between the segmentation accuracy and the operation speed of the traditional attention mechanism, a multi-dimensional attention fusion module MAFM is constructed to reduce the computational burden brought by ordinary two-dimensional convolution operations, and fuse the attention in the channel domain and the spatial domain with only a very small increase in the number of parameters.
[0005] To achieve the above purpose, the technical solution of this application is as follows:
[0006] An enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism, including:
[0007] Obtain an urban street scene image and input it into the backbone network ResNet101 to extract the low-level feature map output by the first residual block of the backbone network ResNet101 and the high-level feature map output by the fourth residual block;
[0008] Input the extracted high-level feature map into the Atrous Spatial Pyramid Pooling module and the multi-dimensional attention fusion module respectively, and add the outputs of the Atrous Spatial Pyramid Pooling module and the multi-dimensional attention fusion module element-wise to obtain the first feature map;
[0009] After connecting the low-level feature map with the first feature, input it into the multi-dimensional attention fusion module again to obtain the second feature;
[0010] Input the feature obtained by connecting the low-level feature map with the first feature into the first convolutional layer of the decoding module. Add the output feature of the first convolutional layer and the second feature element-wise, and then pass through the second convolutional layer of the decoding module to output the enhanced semantic segmentation image;
[0011] Among them, the multi-dimensional attention fusion module performs the following operations:
[0012] Extract the attention weights in the height direction of the high-level feature map, and multiply them element-wise with the input high-level feature map to obtain the first-stage feature map;
[0013] Extract the attention weights in the width direction of the high-level feature map, and multiply the attention weights in the width direction and the first-stage feature map element-wise to obtain the second-stage feature map;
[0014] Perform global pooling operation on the high-level feature map in the channel dimension to obtain the channel-domain feature map;
[0015] Pass the second-stage feature map through a convolutional operation to obtain the spatial-domain feature map;
[0016] Fuse the spatial-domain feature map and the channel-domain feature map to obtain the feature map output by the multi-dimensional attention fusion module.
[0017] Further, the convolutional layer in the backbone network ResNet101 includes three 3×3 convolutions.
[0018] Further, the extraction of the attention weights in the height direction of the high-level feature map includes:
[0019] Perform strip pooling operation on the width of the input high-level feature map, fuse the long-range information in the width direction, integrate the height features of each channel, perform dimensionality reduction operation on the height features of each channel, and obtain a two-dimensional tensor of the channel in the height direction;
[0020] Perform average pooling on the two-dimensional tensor of the channel in the height direction, and then use the sigmoid function for multi-label problems to calculate a probability distributed on [0,1], and obtain a two-dimensional tensor with attention weights in the height direction;
[0021] Upscale the two-dimensional tensor with attention weights in the height direction to obtain the attention weights in the height direction.
[0022] Further, the extraction of the attention weights in the width direction of the high-level feature map includes:
[0023] Perform strip pooling operation on the height of the input high-level feature map, fuse the long-range information in the height direction, integrate the width features of each channel, perform dimensionality reduction operation on the width features of each channel, and obtain a two-dimensional tensor of the channel in the width direction;
[0024] Perform average pooling on the two-dimensional tensor of the channel in the width direction, and then use the sigmoid function for multi-label problems to calculate a probability distributed on [0,1], and obtain a two-dimensional tensor with attention weights in the width direction;
[0025] Upscale the two-dimensional tensor with attention weights in the width direction to obtain the attention weights in the width direction.
[0026] Further, the method for enhancing urban street scene semantic segmentation based on a multi-dimensional attention mechanism further includes
[0027] Calculate the output loss of the third residual block in the backbone network ResNet101;
[0028] Calculate the final output loss of the decoding module;
[0029] Set corresponding weights for the output loss of the third residual block and the final output loss of the decoding module respectively, and calculate the weighted joint loss to complete network training.
[0030] An enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism proposed in this application. Aiming at the shape characteristics of strip-shaped objects such as roads, high-rise buildings, street lamps, and fences in urban street scenes, a strip-based dimensional attention mechanism SPDA is proposed. It uses strip pooling to extract single-dimensional feature weights, captures long-range context semantic associations, and through dimensionality reduction operations, reduces the spatial complexity of weight calculation from square to linear, requiring less memory for calculation. The lightweight design of the module allows this module to be inserted into various network structures. Based on the strip pooling-based attention mechanism, it can better adapt to a large number of strip-shaped target objects in urban street scenes without affecting the discrimination of other objects. The multi-dimensional attention fusion module that combines the channel domain and the spatial domain fuses the attention of the channel domain and the spatial domain with only a small increase in the number of parameters. The lightweight design of the module allows this module to be inserted into various network structures, achieving higher-quality image segmentation prediction results. Description of the Drawings
[0031] Figure 1 It is a flowchart of the enhanced method for urban street scene semantic segmentation based on the multi-dimensional attention mechanism in this application;
[0032] Figure 2 It is a schematic diagram of the overall network structure of an embodiment of this application;
[0033] Figure 3 It is a schematic diagram of the structure of the multi-dimensional attention fusion module of an embodiment of this application;
[0034] Figure 4 It is a schematic diagram of the SPDA structure of an embodiment of this application. Detailed Embodiment
[0035] In order to make the purpose, technical solutions and advantages of this application clearer, the following further details this application in combination with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.
[0036] In one embodiment, as Figure 1 shown, an enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism is proposed, including:
[0037] Step S1, obtain an urban street scene image, input it into the backbone network ResNet101, and extract the low-level feature map output by the first residual block of the backbone network ResNet101 and the high-level feature map output by the fourth residual block.
[0038] As Figure 2As shown in the figure, in this embodiment, ResNet101 with better effects is used as the backbone network. ResNet101 includes five parts, namely conv1, conv2_x, conv3_x, conv4_x, and conv5_x, which can also be expressed as layer0 - layer4. Among them, conv1 is a 7×7 convolution, usually referred to as the convolutional layer, while conv2_x, conv3_x, conv4_x, and conv5_x are residual blocks respectively, corresponding to 3, 4, 23, and 3 blocks respectively, which are called the first residual block to the fourth residual block.
[0039] In a specific embodiment, in this embodiment, 1 7×7 convolution in the convolutional layer is replaced by 3 3×3 convolutions.
[0040] For high - resolution input images, using 3 3×3 convolutions can significantly reduce the parameters while ensuring the same receptive field, enabling the feature maps with regular properties themselves to more easily learn a generalizable feature space.
[0041] In this application, the low - level feature maps output by the first residual block of the backbone network ResNet101 and the high - level feature maps output by the fourth residual block are respectively extracted as the feature maps for subsequent processing.
[0042] In a specific embodiment, since the depth of the third residual block (23 blocks) is much larger than the other groups, in order to better supervise the segmentation quality and accelerate the network convergence, an auxiliary loss is added after the third residual block.
[0043] Step S2: Input the extracted high - level feature maps into the Atrous Spatial Pyramid Pooling (ASPP) module and the Multi - Dimensional Attention Fusion Module (MAFM) respectively, and add the outputs of the ASPP module and the MAFM module element - by - element to obtain the first feature map.
[0044] In this step, the high - level feature maps are respectively input into the Atrous Spatial Pyramid Pooling (ASPP) module and the Multi - Dimensional Attention Fusion Module (MAFM). Before the feature maps are input into the MAFM module, the number of channels is first adjusted, while the number of channels in the original network remains unchanged when input into the ASPP. After adding the output feature maps of the ASPP and the MAFM and compressing the number of channels, the local and global information is integrated to obtain the first feature map.
[0045] The multi - dimensional attention fusion module in this embodiment is as Figure 3 shown and performs the following operations:
[0046] Step 21: Extract the attention weights in the height direction of the high - level feature maps, multiply them element - by - element with the input high - level feature maps to obtain the first - stage feature maps;
[0047] Step 22: Extract the attention weights in the width dimension of the high-level feature map, multiply the attention weights in the width dimension element-wise with the first-stage feature map to obtain the second-stage feature map;
[0048] Step 23: Perform global pooling operation on the high-level feature map in the channel dimension to obtain the channel-domain feature map;
[0049] Step 24: Pass the second-stage feature map through a convolution operation to obtain the spatial-domain feature map;
[0050] Step 25: Fuse the spatial-domain feature map and the channel-domain feature map to obtain the feature map output by the multi-dimensional attention fusion module.
[0051] Specifically, extract the attention weights in the height dimension of the high-level feature map, as Figure 4 shown, including:
[0052] Step 211: Perform strip pooling operation on the width of the input high-level feature map, fuse the long-range information in the width, integrate the height features on each channel, and perform dimensionality reduction operation on the height features on each channel to obtain a two-dimensional tensor of the channel in the height dimension.
[0053] That is, for the input high-level feature map X ∈ R C×W×H , perform the width strip pooling operation to obtain:
[0054]
[0055] where, W0 = 1.
[0056] Then perform a squeeze dimensionality reduction operation on X C×H , delete the width dimension of the three-dimensional feature map, and finally obtain a two-dimensional tensor S C×H ∈ R C×H , representing the information set of a certain channel in the height.
[0057] Step 212: Perform average pooling on the two-dimensional tensor of the channel in the height dimension, and then use the sigmoid function for multi-label problems to calculate a probability distributed in [0, 1] to obtain a two-dimensional tensor with attention weights in the height.
[0058] It is expressed by the formula as follows:
[0059]
[0060]
[0061] The obtained two-dimensional tensor with attention weights in the height is denoted as
[0062] Step 213: Upscale the two-dimensional tensor with height-wise attention weights to obtain the height-wise attention weights.
[0063] It should be noted that upscaling a two-dimensional tensor means replicating the two-dimensional tensor, and the number of replications is the size of the original high-level feature map in the third dimension, which is the width in this embodiment, so that the finally obtained feature map has the same scale as the original feature map.
[0064] In this embodiment, the operations corresponding to Step 212 and Step 213 are also denoted as the SPDA operation, as Figure 3 shown.
[0065] Similarly, to extract the attention weights on the width of the high-level feature map, it includes:
[0066] Step 221: Perform a strip pooling operation on the height of the input high-level feature map to fuse the long-range information in the height direction, integrate the width features on each channel, and perform a dimensionality reduction operation on the width features on each channel to obtain a two-dimensional tensor of the channels in the width direction.
[0067] Step 222: Perform average pooling on the two-dimensional tensor of the channels in the width direction, and then use the sigmoid function for multi-label problems to calculate a probability distributed on [0, 1] to obtain a two-dimensional tensor with width-wise attention weights.
[0068] Step 213: Upscale the two-dimensional tensor with width-wise attention weights to obtain the width-wise attention weights.
[0069] In one embodiment, the attention weights on the height of the high-level feature map are multiplied element-wise with the input high-level feature map to obtain the first-stage feature map, which is expressed as:
[0070]
[0071] Among them, "mul" represents element-wise multiplication of tensors.
[0072] In one embodiment, the attention weights on the width of the high-level feature map are multiplied element-wise with the width-wise attention weights and the first-stage feature map to obtain the second-stage feature map, which is expressed as:
[0073]
[0074] In one embodiment, a global pooling operation is performed on the channels of the high-level feature map to obtain the channel-domain feature map, which is expressed as:
[0075]
[0076] By obtaining the average value of W×H elements in a single channel, the features of each channel are mapped to a single number, and then the weights of each channel are calculated using the sigmoid function to obtain the channel-domain feature map:
[0077]
[0078] In one embodiment, the second-stage feature map is subjected to a convolution operation to obtain the spatial-domain feature map. Specifically, the second-stage feature map is processed by a 3x3 convolution, and the number of output channels is the same as the input, resulting in the spatial-domain feature map.
[0079] In one embodiment, the spatial-domain feature map and the channel-domain feature map are fused to obtain the feature map output by the multi-dimensional attention fusion module, which is expressed as:
[0080]
[0081] where X att is the feature map finally output by the MAFM. The overall number of parameters of the MAFM is small, the calculation is relatively simple, and it can be flexibly added to any part of any backbone network.
[0082] Step S3: After connecting the low-level feature map with the first feature, it is input into the multi-dimensional attention fusion module again to obtain the second feature.
[0083] The operation of the multi-dimensional attention fusion module in this step is the same as that of the multi-dimensional attention fusion module in the previous step, and will not be elaborated here.
[0084] Step S4: The feature obtained by connecting the low-level feature map with the first feature is input into the first convolutional layer of the decoding module. The output feature of the first convolutional layer is element-wise added to the second feature, and then passes through the second convolutional layer of the decoding module to output the image with enhanced semantic segmentation.
[0085] As Figure 2 shown, the decoding module in this embodiment includes two 3×3 convolutions. After connecting the low-level feature map with the first feature, one branch is input into the multi-dimensional attention fusion module to obtain the second feature. The other branch is input into the first convolutional layer, and the output feature of the first convolutional layer is element-wise added to the second feature. The added feature map is then input into the second convolutional layer of the decoding module to output the image with enhanced semantic segmentation.
[0086] The technical solution of this application inserts the MAFM module into the encoding-decoding network based on the ResNet-101 backbone network to construct the spatial-channel attention semantic segmentation network MANet, realizing the enhancement of semantic segmentation of urban street scenes.
[0087] In a specific embodiment, the method for enhancing urban street scene semantic segmentation based on a multi-dimensional attention mechanism of this embodiment further includes
[0088] Calculating the output loss of the third residual block in the backbone network ResNet101;
[0089] Calculating the final output loss of the decoding module;
[0090] Corresponding weights are respectively set for the output loss of the third residual block and the final output loss of the decoding module, and the weighted joint loss is calculated to complete network training.
[0091] The loss function of the network model of this embodiment includes the output loss of the third residual block and the final output loss. The weights of the two loss functions are 0.4 and 0.6 respectively. The cross-entropy function is respectively used as the loss function, and the optimizer is the SGD optimizer to complete network training.
[0092] The method for enhancing urban street scene semantic segmentation based on a multi-dimensional attention mechanism of this application uses a strip-based dimensional attention mechanism to obtain the attention weights on the height and width of the feature map respectively. The attention mechanism based on strip pooling can better adapt to the target objects in urban street scenes. After fusing the attention in the spatial domain and the channel domain in MAFM, this module can be added to different positions of different backbone networks, which is flexible and convenient. MAFM uses fewer parameters and has a simple model. Its application can produce better prediction results for objects with a greater dependence on remote context.
[0093] The above-described embodiments only represent several implementation manners of this application. The description is relatively specific and detailed, but it cannot be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application shall be subject to the appended claims.
Claims
1. An enhanced method for semantic segmentation of urban street scenes based on a multi-dimensional attention mechanism, characterized in that, The enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism includes: Obtain an urban street scene image and input it into the backbone network ResNet101 to extract the low-level feature map output by the first residual block of the backbone network ResNet101 and the high-level feature map output by the fourth residual block; Input the extracted high-level feature map into the Atrous Spatial Pyramid Pooling (ASPP) module and the multi-dimensional attention fusion module respectively, and add the outputs of the ASPP module and the multi-dimensional attention fusion module element-wise to obtain the first feature; After connecting the low-level feature map with the first feature, input it into the multi-dimensional attention fusion module again to obtain the second feature; Input the feature obtained by connecting the low-level feature map with the first feature into the first convolutional layer of the decoding module. Add the output feature of the first convolutional layer and the second feature element-wise, and then pass through the second convolutional layer of the decoding module to output the image after enhanced semantic segmentation; Among them, the multi-dimensional attention fusion module performs the following operations: Extract the attention weights in the height direction of the high-level feature map, multiply them element-wise with the input high-level feature map to obtain the first-stage feature map; Extract the attention weights in the width direction of the high-level feature map, multiply the attention weights in the width direction and the first-stage feature map element-wise to obtain the second-stage feature map; Perform global pooling operation on the high-level feature map in the channel dimension to obtain the channel-domain feature map; Pass the second-stage feature map through a convolutional operation to obtain the spatial-domain feature map; Fuse the spatial-domain feature map and the channel-domain feature map to obtain the feature map output by the multi-dimensional attention fusion module.
2. The method for enhancing urban street view semantic segmentation based on a multi-dimensional attention mechanism according to claim 1, characterized in that, The convolutional layers in the backbone network ResNet101 include three 3×3 convolutions.
3. The method for enhancing urban street view semantic segmentation based on a multi-dimensional attention mechanism according to claim 1, characterized in that The extraction of the attention weights in the height direction of the high-level feature map includes: Perform strip pooling operation on the width of the input high-level feature map to fuse the long-range information in the width direction, integrate the height features on each channel, perform dimensionality reduction operation on the height features on each channel to obtain the two-dimensional tensor of the channel in the height direction; Perform average pooling on the two-dimensional tensor of the channel in the height direction, and then use the sigmoid function for multi-label problems to calculate a probability distributed on [0, 1] to obtain the two-dimensional tensor with attention weights in the height direction; Perform dimensionality increase on the two-dimensional tensor with attention weights in the height direction to obtain the attention weights in the height direction.
4. The method for enhancing urban street view semantic segmentation based on a multi-dimensional attention mechanism according to claim 1, wherein The extraction of the attention weights in the width direction of the high-level feature map includes: Perform strip pooling operation on the height of the input high-level feature map to fuse the long-range information in the height direction, integrate the width features on each channel, perform dimensionality reduction operation on the width features on each channel to obtain the two-dimensional tensor of the channel in the width direction; Perform average pooling on the two-dimensional tensor of the channel in the width direction, and then use the sigmoid function for multi-label problems to calculate a probability distributed on [0, 1] to obtain the two-dimensional tensor with attention weights in the width direction; Perform dimensionality increase on the two-dimensional tensor with attention weights in the width direction to obtain the attention weights in the width direction.
5. The method for enhancing urban street scene semantic segmentation based on a multi-dimensional attention mechanism according to claim 1, characterized in that The enhanced method for urban street scene semantic segmentation based on a multi-dimensional attention mechanism further includes Calculate the output loss of the third residual block in the backbone network ResNet101; Calculate the final output loss of the decoding module; Set corresponding weights for the output loss of the third residual block and the final output loss of the decoding module respectively, and calculate the weighted joint loss to complete network training.
Citation Information
Patent Citations
Multi-resolution semantic segmentation method and device based on attention pyramid
CN114359297A
System and method for boundary aware semantic segmentation
US20210089807A1