Real-time Semantic Segmentation Method for Road Scene Images Based on Improved U-Net
By resetting the convolution module to depth-separable convolution in the U-Net network and introducing channel domain attention and multi-scale spatial attention mechanism, the existing real-time semantic segmentation network has been solved, and efficient road scene image segmentation is achieved and suitable for mobile deployment.
Patent Information
- Application Number
- CN202310080417.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-01-11
AI Technical Summary
The existing real-time semantic segmentation network has poor processing capabilities for image detail features, is slow, and is difficult to deploy on mobile.
A real-time semantic segmentation method based on improved U-Net is proposed. The convolution module is reset to a depth-separable convolution in the sampling stage, and a channel domain attention mechanism is introduced to optimize the up-down sampling fusion stage to enhance spatial information capture capabilities through the multi-scale spatial attention cascade module.
Average crossover ratio of 70.6% and 67.9% is achieved on Cityscapes and CamVid datasets, improving the accuracy of object boundaries and improving recognition accuracy while ensuring real-time efficiency, making it suitable for deployment on mobile devices.
Smart Images

Figure CN116363358B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road condition scene detection, and in particular to a real-time semantic segmentation method for road scene images based on an improved U-Net. Background Art
[0002] Road condition scene detection is a key concern in the field of driverless. Real-time semantic segmentation technology can segment road driving scenarios, helping the driverless system to understand the road conditions and make decisions in a timely manner. In the field of image processing, semantic segmentation technology is commonly used to achieve image analysis. The existing image semantic segmentation methods cannot well balance accuracy and efficiency, which greatly affects their application in the actual environment. While achieving faster image semantic segmentation, ensuring its accuracy will enable the segmentation network to be truly applied in the actual scenario, especially for some special application scenarios with limited computing resources. Therefore, whether it is possible to quickly implement image semantic segmentation, improve the speed of semantic segmentation, and at the same time ensure the accuracy of the semantic segmentation model is also the key issue to be solved in current image semantic segmentation.
[0003] Traditional semantic segmentation tasks mainly classify using the original pixel features of images, and these pixel features include brightness, texture, color, etc. With the development of convolutional neural networks (CNNs), a variety of fully convolutional-based semantic segmentation algorithms have been successively proposed, solving the deficiencies of traditional semantic segmentation such as poor accuracy and information loss. Long et al. [4]The fully convolutional neural network (FCN) was proposed to achieve low-high level feature fusion in an encoder-decoder manner, realizing higher-precision image semantic segmentation. Subsequently, Ronneberger et al. proposed a symmetric semantic segmentation model, U-Net, which adopted an encoding and decoding structure. Through feature fusion in the contracting path and the expanding path, it achieved good results in medical image segmentation applications. When dealing with some small targets, such as pedestrians, trees or other non-homogeneous targets with similar shapes in urban road scenes, the U-Net network performs better than other deep neural networks. Badrinarayanan et al. proposed the SegNet model, which solved the problem of losing the receptive field and detailed information in the process of semantic segmentation by the FCN. This model upsamples the low-resolution feature map through the pooling index method in the decoder stage, and has great advantages in inference and calculation time. Aiming at the problem that the fully convolutional symmetric semantic segmentation network ignores pixel spatial information, new dilated networks emerged. Chen et al. proposed DeepLabv1, which used VGG-16 as the backbone network and introduced dilated convolution and fully connected conditional random field (CRF) to improve the model's ability to capture details. DeepLabv2 replaced the backbone network with ResNet-101 and added the atrous spatial pyramid pooling (ASPP) module to integrate multi-scale feature information. DeepLabv3 applied dilated convolution in the cascaded module and improved the ASPP. DeepLabv3+ used the spatial pyramid pooling module (SPP) for the deep network structure. Since these networks are more refined, it also means more computational complexity and the number of parameters. Semantic segmentation is regarded as dense classification of images in the field of autonomous driving and requires higher real-time performance. Therefore, the DeepLab series of networks are not suitable for mobile devices with limited storage and computing resources. Currently, there have been many lightweight networks achieving high precision and real-time performance. Zhang et al. proposed shuffleNet, which adopted two operations, group convolution and channel shuffle, to balance speed and precision. Paszke et al. proposed ENet, which adopted an early downsampling strategy to reduce computation, optimized details using an asymmetric structure, and added non-linear activation and dilated convolution to achieve a high running speed. Fan et al. believed that adding extra branches in BiSeNet increased the running time of the network, and proposed the STDC (short-term dense concatenate) module, which can extract multi-scale features with fewer parameters. At the same time, it improved the multi-path structure of BiSeNet, extracting low-level detailed features while reducing the network's computational complexity.
[0004] However, the existing real-time semantic segmentation networks have poor processing ability for image detail features, slow speed, and are difficult to be deployed on mobile devices. Summary of the Invention
[0005] In view of the problems that the existing real-time semantic segmentation network has poor processing ability for image detail features, slow speed, and is difficult to be deployed on mobile devices, the present invention proposes a real-time semantic segmentation method for road scene images based on an improved U-Net. In the sampling stage, the convolution module is reset to reduce the number of model parameters and improve the real-time inference speed. After convolution, channel domain attention is added to more efficiently extract context features and improve the classification of pixel categories. The upsampling and downsampling fusion stage is optimized by adding a multi-scale spatial attention mechanism to supplement the ability to capture spatial information, enhance the network's feature representation, improve the accuracy of object boundaries, and ensure real-time efficiency and increase recognition accuracy with a small number of parameters.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A real-time semantic segmentation method for road scene images based on an improved U-Net, comprising:
[0008] Step 1: In the sampling stage, replace the standard convolution in the original U-Net structure with depthwise separable convolution to obtain an optimized convolution block, and introduce an attention module to enhance the feature acquisition ability and supplement context information, obtaining a channel domain attention depth convolution module structure;
[0009] Step 2: In the upsampling and downsampling fusion stage, add a multi-scale spatial attention cascade module, obtain a low-dimensional space by focusing on key regions and fuse it with high-dimensional features after transformation, obtaining an improved U-Net;
[0010] Step 3: Perform real-time semantic segmentation on road scene images based on the improved U-Net.
[0011] Further, the depthwise separable convolution first performs an n×n convolution on each channel of the input feature map respectively, outputs n channel-separated features, and then performs a 1×1 convolution to obtain the output.
[0012] Further, the attention module includes two operations: compression and excitation. Compression is a global average pooling process. After the compression operation, the feature map becomes a 1×1×C vector. The excitation operation consists of two fully connected layers. The output vector of compression is used as the input vector of excitation. After transformation, it is multiplied by the original feature map to obtain the final output result with the same size as the original feature, where C is the number of channels.
[0013] Further, the input feature of the multi-scale spatial attention cascade module is the result of the upsampling stage process, and different information feature representations are obtained through two pooling operations.
[0014] Compared with the prior art, the present invention has the following beneficial effects:
[0015] In the present invention, by resetting the convolutional module in the sampling stage, the number of model parameters is reduced, and the real-time inference speed is improved. After convolution, channel domain attention is added to more efficiently extract the context features of road scene images and improve the classification of pixel categories. The upsampling and downsampling fusion stage is optimized, and a multi-scale spatial attention mechanism is added to supplement the ability to capture spatial information, enhance the network's feature representation, improve the accuracy of object boundaries, and ensure real-time efficiency and increase recognition accuracy with a small number of parameters. The average intersection over union (mIoU) of the proposed method on the Cityscapes and CamVid datasets reaches 70.6% and 67.9% respectively, showing obvious advantages compared with other lightweight networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 FIG. is a basic flowchart of a real-time semantic segmentation method for road scene images based on an improved U-Net according to an embodiment of the present invention.
[0017] Figure 2 FIG. is a schematic diagram of the original U-Net structure according to an embodiment of the present invention;
[0018] Figure 3 FIG. is a schematic diagram of the CADCM structure according to an embodiment of the present invention;
[0019] Figure 4 FIG. is a schematic diagram of the DSConv structure according to an embodiment of the present invention;
[0020] Figure 5 FIG. is a schematic diagram of the SE attention module structure according to an embodiment of the present invention;
[0021] Figure 6 FIG. is a schematic diagram of the MSACM structure according to an embodiment of the present invention;
[0022] Figure 7 FIG. is a schematic diagram of the CSAU-Net structure according to an embodiment of the present invention;
[0023] Figure 8 FIG. shows the effect comparison of adding different modules according to an embodiment of the present invention;
[0024] Figure 9 FIG. is a visualization of the comparison of different methods on the Cityscapes dataset according to an embodiment of the present invention;
[0025] Figure 10 FIG. is a visualization of the comparison of different methods on the CamVid dataset according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further explained below with reference to the drawings and specific embodiments:
[0027] As Figure 1As shown in the figure, a real-time semantic segmentation method for road scene images based on an improved U-Net includes:
[0028] Step 1: In the sampling stage, replace the standard convolution in the original U-Net structure with depthwise separable convolution to obtain an optimized convolution block, and introduce an attention module to enhance the feature acquisition ability and supplement context information, resulting in a channel-domain attention depth convolution module structure;
[0029] Step 2: In the upsampling and downsampling fusion stage, add a multi-scale spatial attention cascade module, obtain a low-dimensional space by focusing on key regions and fuse it with high-dimensional features after transformation to obtain an improved U-Net;
[0030] Step 3: Perform real-time semantic segmentation on road scene images based on the improved U-Net
[0031] Specifically, the following is a detailed introduction:
[0032] 1 Channel-domain attention depth convolution module structure
[0033] Context information is generally understood as perceiving and applying some or all of the information that can affect the objects in a scene and an image. For image information, it is represented as capturing the interaction information between different objects, and using the interaction information between the object and the scene as a condition to identify and process new targets. The reasonable use of context information can better complete the task. Based on the U-Net structure, the present invention proposes a channel-domain attention depth convolution module structure (CADCM). In the original U-Net structure (as Figure 2 shown), during the feature extraction stage in the upsampling and downsampling process, only two standard convolutions are used to extract image features, and they are output to the next stage in the form of pooling or transposed convolution.
[0034] Figure 3 For the optimized convolution block, replace the standard convolution with depthwise separable convolution (DSConv) and introduce an attention (SE) module to enhance the feature acquisition ability and supplement context information. As can be seen from Figure 2 the figure, each layer consists of three feature blocks. The blue arrows represent convolution operations, and the red and green ones are pooling and transposed convolution. This process only involves simple operations, which easily leads to information loss and lack of context information. Figure 4This is the detailed structure of DSConv. Depthwise separable convolution first performs an n×n convolution on each channel of the input feature map separately, outputting n channel-separated features, and then performs a 1×1 convolution to obtain the output. The operation of depthwise separable convolution achieves the same performance as standard convolution while ensuring fewer parameters.
[0035] After sampling the last feature map, add an SE module (as Figure 5 shown), enabling the network to focus on channel feature information and allocate weights to bring a significant improvement in effect with minimal additional computational cost.
[0036] The SE module consists of two main operations, one is squeeze and the other is excitation. Squeeze is a global average pooling process. After the squeeze operation, the feature map becomes a 1×1×C vector. The excitation operation consists of two fully connected layers. The output vector of the squeeze is used as the input vector of the excitation. After transformation, it is multiplied by the original feature map to obtain the final output with the same size as the original feature, where C is the number of channels. The squeeze operation is expressed as:
[0037]
[0038] In the formula, F sq is the squeeze function, u c ∈R H×W is the input two-dimensional feature matrix, and H and W represent the height and width of the feature map information respectively. The excitation operation is expressed as:
[0039] s = F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z)) (2)
[0040] In the formula, σ represents the Sigmoid activation function, and δ represents the ReLU activation function, where represent certain elements respectively. After two operations, the output of the excitation function is multiplied by the original feature dimension to obtain the final output, which is expressed as:
[0041]
[0042] Among them u c ∈R H×W , F scale represents the product on the channels.
[0043] 2 Spatial fusion
[0044] Consists of Figure 2It can be seen that the original U-Net upsampling and downsampling process uses a replication connection method to combine low- and high-dimensional features. Although this process can utilize different feature information of low and high dimensions, the process is rough and prone to the problem of spatial information loss. Therefore, a multi-scale spatial attention cascading module (MSACM) is proposed. This module obtains the low-dimensional space by focusing on the key areas and fuses it with the high-dimensional features after transformation. This method can effectively supplement the missing spatial information of the original network and filter out non-key features, effectively improving the network performance. The structure of MSACM is as Figure 6 shown. Figure 6 In [Figure 1], the input feature is the result of the upsampling stage process. After two pooling operations, different information feature representations are obtained. The operations are as follows:
[0045]
[0046]
[0047] F c represents the input feature map. After two transformations and then passing through the Sigmoid activation function, a spatial feature map is obtained. Then, it is multiplied by the original image to obtain the final output feature. The final network structure diagram is as Figure 7 shown.
[0048] 3 Experiments and Analysis
[0049] 3.1 Dataset and Experimental Environment
[0050] The two datasets used are Cityscapes and CamVid. The Cityscapes dataset is a large dataset of urban street scenes, focusing on the semantic understanding of urban streetscapes. It contains 50 cities, 5,000 finely annotated images, and 19 classification categories, with 2,975 images for training, 500 for the validation set, and 1,525 for testing. The CamVid dataset is an image segmentation dataset of roads and driving scenes, containing 700 finely annotated images, 11 categories, 367 for training, 101 for validation, and 233 for testing. All training rounds are set to 300, the initial learning rate is 5×10-4, the batch size for Cityscapes is 6, the batch size for CamVid is 8, the Adam optimizer and the Ploy learning strategy with a power of 0.9 are adopted. The experimental software environment is PyTorch, and the programming language is Python 3.7. The hardware environment is an Intel Core i7 CPU, a GeForce RTX 3090 GPU, 64G of memory, and the WIN10 operating system is used. The evaluation metrics adopted are the mean intersection over union (mIoU) and the frames per second (FPS). The specific expressions are as follows:
[0051]
[0052] In the formula, k represents the number of categories, and p ij represents the number of objects in the i-th category classified into the j-th category.
[0053]
[0054] In the formula, N represents the number of images, and T j represents the time to process the j-th image.
[0055] 3.2 Ablation Experiment
[0056] To verify the effectiveness of each module of CSAU-Net, individual experiments were conducted on each module. The dataset used was CamVid, and the specific results are shown in Table 1. The overall experiment was divided into 4 parts. First was the original U-Net, then only the CADCM module was added and only the MSACM module was added, and finally the results with all modules added.
[0057] Table 1 Comparison of the effects of adding each module
[0058]
[0059] As can be seen from Table 1, the accuracy of the original U-Net is about 64.6%. After adding only the CADCM module, it is improved by 0.7 percentage points. Because this module replaces ordinary convolution with depthwise separable convolution and introduces channel attention, it can capture image context information. Secondly, adding only the MSACM module improves by 1.3% compared to U-Net, and the improvement effect of focusing on the spatial information and key region information of the image is more obvious. Finally, integrating the two modules reaches an accuracy of 67.9%, which is an improvement of 3.1 percentage points compared to the original U-Net, and also achieves a calculation speed of 60 frames per second. Figure 8 Shows the effect comparison of adding different modules.
[0060] From Figure 8 it can be seen that CSAU-Net has a greater improvement in dealing with details compared to U-Net. The street lights and road divisions in the figure are clear, thanks to the optimization of the proposed method for context information and spatial information. Generally speaking, the proposed CSAU-Net has an obvious improvement in effect.
[0061] 3.3 Cityscapes Comparison Results
[0062] On the Cityscapes dataset, several lightweight networks in recent years such as SegNet, CGNet, and LEDNet were selected for comparative analysis, and the results are shown in Table 2.
[0063] Table 2 Comparison Results of Different Methods on the Cityscapes Dataset
[0064]
[0065] As can be seen from Table 2, the proposed CSAU-Net achieves the highest index score with an accuracy of 70.6%, which is an improvement of 5.4 percentage points compared to the original U-Net, and is 7.9%, 4.3%, and 0.8% higher than SegNet, CGNet, and LEDNet respectively. The index analysis shows that the proposed optimization method has a good effect. CSAU-Net has 60 frames per second on the FPS index. Although it does not reach the best, it also has a small difference from LEDNet with 71 frames per second and has a speed advantage compared to other networks. The visualization results are as Figure 9 shown.
[0066] Figure 9It is divided into 7 columns. The first column is the original image, the second column is the corresponding labeled image, and the 3rd, 4th, 5th, and 6th columns are the comparison methods respectively. The last column is the proposed CSAU-Net method. It can be seen from the figure that CSAU-Net has a clearer division and better recognition effect for objects such as roads, pedestrians, and utility poles compared with other methods. There are significant object confusions in the visualization of U-Net. Through the optimization and supplementation of context and spatial information, the final effect is effectively improved. Through the overall experimental analysis, the proposed method has advantages in both accuracy and real-time performance.
[0067] 3.4 Comparison Results on CamVid
[0068] Related experiments were also carried out on the CamVid dataset, and the comparison results of different methods are shown in Table 3.
[0069] Table 3 Comparison Results of Different Methods on the CamVid Dataset
[0070]
[0071] It can be seen from Table 3 that CSAU-Net achieved the best score of 67.9%, and the FPS was also 60 frames / s. It increased by 1.4 percentage points compared with LEDNet and 2.3% compared with CGNet. Although it did not reach the best in terms of real-time performance, it still had advantages compared with other methods. The visualization effect is as Figure 10 shown.
[0072] From Figure 10 it can be seen that CSAU-Net can basically correctly recognize different objects, and the boundaries between different objects are relatively clear, especially for small objects such as street lights, trees, and signs. There are disconnection problems in the overall recognition of U-Net, such as discontinuous roads and street lights. Other networks such as CGNet have large-scale recognition errors, and plants and buildings are recognized confusingly. Overall, CSAU-Net can achieve better results through the optimization of context and spatial information.
[0073] In summary, for the real-time road scene semantic segmentation task, the present invention proposes a new network structure CSAU-Net based on U-Net, which resets the convolutional module, adds channel domain attention, and extracts the key information of the image context more efficiently. The upsampling and downsampling fusion stage is optimized to supplement the ability to capture spatial information, and a multi-scale spatial attention mechanism is added to enhance the network's feature representation. While maintaining fewer parameters, real-time efficiency is maintained and recognition accuracy is increased. Relevant experiments are carried out on two datasets, Cityscapes and CamVid, achieving accuracies of 70.6% and 67.9% respectively, and the inference speed reaches 60 frames / s, achieving a high segmentation accuracy while ensuring the inference speed. Moreover, CSAU-Net belongs to a lightweight network and can be deployed on mobile devices.
[0074] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A real-time semantic segmentation method for road scene images based on improved U-Net, characterized in that, it includes: Step 1: In the sampling stage, replace the standard convolution in the original U-Net structure with depthwise separable convolution to obtain an optimized convolution block, and introduce an attention module after sampling the last feature map to enhance the feature acquisition ability and supplement context information, obtaining a channel-domain attention depth convolution module structure; the attention module contains two operations: compression and excitation. Compression is a global average pooling process. After the compression operation, the feature map becomes a 1×1×C vector. The excitation operation consists of two fully connected layers. The output vector of compression is used as the input vector of excitation. After transformation, it is multiplied by the original feature map to obtain the final output result with the same size as the original feature, where C is the number of channels; Step 2: In the upsampling and downsampling fusion stage, add a multi-scale spatial attention cascade module. By focusing on key regions, a low-dimensional space is obtained and transformed and fused with high-dimensional features to obtain an improved U-Net; the input feature of the multi-scale spatial attention cascade module is the result of the upsampling stage process. Different information feature representations are obtained through two pooling operations. The operations are as follows: Among them F c represents the input feature map. After two transformations and then passing through the Sigmoid activation function, the spatial feature map is obtained, and then multiplied by the original image to obtain the final output feature; Step 3: Perform real-time semantic segmentation on road scene images based on the improved U-Net.
2. The real-time semantic segmentation method for road scene images based on improved U-Net according to claim 1, characterized in that, for each channel of the input feature map, the depthwise separable convolution first performs an n×n convolution separately to output n channel-separated features, and then performs a 1×1 convolution to obtain the output.
Citation Information
Patent Citations
Real-time semantic segmentation method based on context attention mechanism and information fusion
CN112541503A
Remote sensing image semantic segmentation method based on pyramid segmentation attention module
CN113807210A