A lightweight real-time semantic segmentation method for urban roads based on deep learning

By adopting a lightweight real-time semantic segmentation method of deep learning in urban road scenarios, using the fusion of encoder-decoder structure and cross-attention, the problems of insufficient generalization ability and high computational complexity in the existing technology are solved, and efficient and accurate semantic segmentation effect is achieved.

CN119048758BActive Publication Date: 2025-05-13NANJING CHENYUAN ROBOT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411248293.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-05-13
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

The existing technology lacks generalization ability of semantic segmentation in urban complex road scenarios, and the computational complexity of deep learning models is high, resulting in slow real-time inference speed and large computing resources consumption.

Method used

The lightweight real-time semantic segmentation method of urban roads based on deep learning is adopted, and an end-to-end encoder-decoder structure is adopted, combining cross-attention fusion, upsampling and loss function supervision training to optimize the multi-scale feature extraction and fusion of the model.

Benefits of technology

It improves the accuracy of classification prediction for complex road scenarios, reduces the computational complexity, improves segmentation efficiency, and makes applications possible in real-time systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048758B_ABST
    Figure CN119048758B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight real-time semantic segmentation method for urban roads based on deep learning. A real-time urban road scene semantic segmentation model (EMFANet_AsSC model) is used for semantic segmentation. The EMFANet_AsSC model is an end-to-end encoder-decoder structure, which can accurately (72.4% mIoU on the Cityscapes test set) segment various objects in complex urban road scenes in real time (140FPS), while maintaining low computational overhead and storage overhead (0.9M parameters). The encoder side includes three stages: Stage 1, Stage 2 and Stage 3. Two dual attention aggregation units on the decoder side are responsible for guiding and optimizing the multi-scale features extracted by the encoder side, and at the same time, they are fully fused and the final semantic segmentation map is output through an upsampler.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a light-weight real-time semantic segmentation method for urban roads based on deep learning, belonging to image processing technology. Background Art

[0002] Nowadays, image semantic segmentation technology based on convolutional neural networks (CNNs) is widely used in computer vision tasks, such as autonomous driving and robot vision. Convolutional neural networks are one of the key technologies to achieve intelligence, safety and efficiency, and are crucial for real-time environmental perception and decision-making.

[0003] At present, the generalization ability of traditional semantic segmentation methods for urban scenes is not high, which is not enough to accurately distinguish a large number of complex patterns in real data. Although various large models based on deep learning can achieve high segmentation accuracy, they have high memory and computing costs, which reduces mobility and portability. For example, current deep models such as FCN, ResNet and VGGNet have achieved extremely high classification accuracy in urban street semantic segmentation tasks. However, the underlying networks of these models, such as VGGNet and ResNet, are complex to build and have many parameters, resulting in slow real-time reasoning and high consumption of computing resources. On the other hand, SegNet, ENet, ESPNet and ERFNet have explored real-time lightweight networks. Although the model parameters have been greatly compressed, the segmentation ability has been significantly reduced, and the overall effect is not good. Some other methods propose to use multi-scale information to improve the network structure, such as the two-branch structure proposed by BiseNet and the single-branch structure of DFANet, as well as the use of an asymmetric encoder-decoder structure to optimize the reasoning speed and segmentation ability, but there is still room for optimization in terms of calculation. Therefore, how to improve the segmentation ability while keeping the number of parameters small has become a key issue in real-time semantic segmentation methods. Summary of the invention

[0004] Purpose of the invention: In order to improve the accuracy of lightweight real-time semantic segmentation methods in complex urban road scenes, while reducing the computational complexity and making it possible to apply them in real-time systems, the present invention provides a lightweight real-time semantic segmentation method for urban roads based on deep learning.

[0005] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:

[0006] A lightweight real-time semantic segmentation method for urban roads based on deep learning, comprising the following steps:

[0007] Step 1: Take the finely annotated complex road scene image as the input image, and preprocess the input image to obtain the original image;

[0008] Step 2, divide the original image into training set and test set according to proportion;

[0009] Step 3: Input the training set into the real-time urban road scene semantic segmentation model, and generate the final segmentation prediction map through cross-attention fusion and upsampling;

[0010] Step 4: Use the loss function to supervise the real-time urban road scene semantic segmentation model to obtain a trained real-time urban road scene semantic segmentation model;

[0011] Step 5. Input the test set into the trained real-time urban road scene semantic segmentation model to obtain the segmentation prediction map.

[0012] Specifically, the finely annotated complex road scene images are derived from the Cityscapes dataset and the Camvid dataset; the input images are preprocessed, and the preprocessing operations include horizontal flipping, randomly changing the input image size, and cropping; for the input images derived from the Cityscapes dataset, the category weights are kept consistent with the Cityscapes dataset; for the input images derived from the Camvid dataset, the category weights are adjusted according to the following formula:

[0013]

[0014] Where: W i represents the adjusted category weight of category i, h represents the set hyperparameter, which is generally set to 1.12, and P i Represents the weight of category i in the Camvid dataset.

[0015] Specifically, the real-time urban road scene semantic segmentation model is an end-to-end encoder-decoder structure;

[0016] The encoder end includes three stages: Stage 1, Stage 2 and Stage 3. At the beginning of each of the three stages, a downsampler is used to downsample the input features. At the end of each of the three stages, a channel attention layer is used to optimize the extracted multi-scale features. At the same time, a space-channel optimization unit is used to extract the spatial features and background features of the original image and use deep separable convolution for fine-tuning and fusion. The output of the channel attention layer is added to the output of the space-channel optimization unit as the output of the corresponding stage. The output of the space-channel optimization unit is used to compensate for the information loss in the downsampling process, thereby improving high precision, thereby compensating for and optimizing the extracted multi-scale features. In Stag In the e1 stage, the input feature is the original image, and three 3×3 standard convolutions are set between the downsampler and the channel attention layer. The multi-scale features output by the channel attention layer in this stage are called low-scale features; in the Stage2 stage, the input feature is the output of the Stage1 stage, and four symmetric attention residual units are set between the downsampler and the channel attention layer. The multi-scale features output by the channel attention layer in this stage are called medium-scale features; in the Stage3 stage, the input feature is the output of the Stage2 stage, and eight symmetric attention residual units are set between the downsampler and the channel attention layer. The multi-scale features output by the channel attention layer in this stage are called high-scale features;

[0017] The decoder first guides and optimizes the multi-scale features extracted by the encoder in sequence through two dual attention aggregation units, and then uses an upsampler to upsample and output the final segmentation prediction map; the first dual attention aggregation unit performs feature fusion on the output of the Stage 3 stage and the mid-scale features, the second dual attention aggregation unit performs feature fusion on the output of the first dual attention aggregation unit and the low-scale features, and uses the upsampler to upsample the output of the second dual attention aggregation unit to obtain the final segmentation prediction map;

[0018] The loss is calculated after upsampling the output of Stage 3 using an upsampler. The calculated loss is called the auxiliary loss. The loss is calculated for the final segmentation prediction map. The calculated loss is called the main loss. The auxiliary loss and the main loss are weighted summed to get the total loss.

[0019] Specifically, the downsampler includes a 3×3 standard convolution branch and a 2×2 maximum pooling branch. The input of the downsampler passes through the 3×3 standard convolution branch and the 2×2 maximum pooling branch respectively, and then the outputs of the two branches are channel-joined to obtain the output of the downsampler.

[0020] Specifically, the input of the channel attention layer is sequentially subjected to a global average pooling operation, a compression operation, a transposition operation, a 1×1 standard convolution, a transposition operation, a decompression operation, and a Sigmoid activation function, and then multiplied by the input of the channel attention layer, and the result is used as the output of the channel attention layer; the channel attention layer is described as:

[0021] Y C (X C )=δ(f Unsq (f Trans (C 1×1 (f Trans (f sq (f AvgPool ))))))

[0022] Where: X C represents the input of the channel attention layer, Y C represents the output of the channel attention layer, f AvgPool represents the global average pooling operation, f sq represents the compression operation, f Trans represents the transpose operation, C 1×1 represents a standard convolution with a kernel size of 1×1, f Unsq represents the decompression operation, and δ represents the Sigmoid activation function.

[0023] Specifically, the space-channel optimization unit includes a maximum pooling branch and a global average pooling branch. The input of the space-channel optimization unit passes through the maximum pooling branch and the average pooling branch respectively, and then the outputs of the two branches are channel-joined, and then sequentially pass through a 3×3 depth-separable convolution and a 1×1 standard convolution, and the result is used as the output of the space-channel optimization unit; the space-channel optimization unit is described as:

[0024] Y S (X ori )=C 1×1 (DC 3×3 (Concat(f MaxPool ,f AvgPool )))

[0025] Where: X ori represents the original image, Y S represents the output of the space-channel optimization unit, f MaxPool represents the maximum pooling operation, f AvgPool represents the global average pooling operation, Concat represents the channel concatenation operation, DC 3×3 represents a depthwise separable convolution with a kernel size of 3×3, C 1×1 Represents a standard convolution with a kernel size of 1×1;

[0026] For the maximum pooling branch and the average pooling branch, r is used to represent the ratio of the output of the space-channel optimization unit to the original image, the pooling kernel of the maximum pooling branch and the average pooling branch are both k, the pooling step size of the maximum pooling branch and the average pooling branch are both s, and (r, k, s) is used as the setting parameter of the space-channel optimization unit; for the space-channel optimization units in the three stages of Stage1, Stage2 and Stage3, the setting parameters (r, k, s) are (2, 2, 2), (4, 4, 4) and (8, 8, 8), respectively.

[0027] Specifically, the input of the symmetric attention residual unit first passes through a channel attention layer and a 3×3 depth-separable convolution, and then is equally separated into two parts through the Split function, one part is input into the standard decomposition convolution, and the other part is input into the depth-level decomposition convolution, and then the outputs of the standard decomposition convolution and the depth-level decomposition convolution are channel-joined, and then pass through a 3×3 depth-separable convolution, and the output of the 3×3 depth-separable convolution is added to the input of the symmetric attention residual unit and then a channel confusion operation is performed, and the result is used as the output of the symmetric attention residual unit;

[0028] The standard decomposition convolution includes a 3×1 standard convolution, a 1×3 standard convolution, a channel attention layer, a 3×1 standard convolution and a 1×3 standard convolution. The depth-level decomposition convolution includes a 3×1 depth-separable convolution, a 1×3 depth-separable convolution, a channel attention layer, a 3×1 depth-separable convolution and a 1×3 depth-separable convolution. The expansion rates of the standard convolution and the depth-separable convolution are both d; in the four symmetric attention residual units in the Stage2 stage, the expansion rate d is 1, 2, 5, and 9 respectively; in the eight symmetric attention residual units in the Stage3 stage, the expansion rate d is 1, 2, 5, 9, 2, 5, 9, and 17 respectively.

[0029] Specifically, the dual attention aggregation unit includes a left block and a right block; in the left block, the lower-level scale feature is subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution and then a global average pooling operation is performed to obtain the lower-level output feature of the left block, and the higher-level scale feature is subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution and then input into the Sigmoid activation function to obtain the higher-level output feature of the left block, and the lower-level output feature of the left block and the higher-level output feature of the left block are multiplied and then input into the upsampler for upsampling to obtain the output feature of the left block; in the right block In the figure, the lower-level scale features are subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution to obtain the lower-level output features of the right block. The higher-level scale features are subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution, and then input into the upsampler for upsampling, and then input into the Sigmoid activation function to obtain the higher-level output features of the right block. The lower-level output features of the right block and the higher-level output features of the right block are multiplied to obtain the output features of the right block; the output features of the left block and the output features of the right block are added, and the result is used as the output of the dual attention aggregation unit;

[0030] In the first dual attention aggregation unit, the lower-level scale features are the mid-level scale features, and the higher-level scale features are the high-level scale features; in the second dual attention aggregation unit, the lower-level scale features are the low-level scale features, and the higher-level scale features are the output of the first dual attention aggregation unit.

[0031] Beneficial effects: The lightweight real-time semantic segmentation method for urban roads based on deep learning provided by the present invention can not only improve the classification prediction accuracy of complex road scenes in the real world, but also reduce the computational complexity and greatly improve the segmentation efficiency, making it possible to apply it in real-time systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a schematic diagram of the structure of the real-time urban road scene semantic segmentation model;

[0033] Figure 2 It is a structural diagram of the downsampler;

[0034] Figure 3 It is a structural diagram of CAM Layer;

[0035] Figure 4 It is a structural schematic diagram of the SBU unit;

[0036] Figure 5 It is a structural schematic diagram of the ARU unit;

[0037] Figure 6 Schematic diagram of the structure of the DAAU unit. DETAILED DESCRIPTION

[0038] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0039] A lightweight real-time semantic segmentation method for urban roads based on deep learning is adopted. Figure 1 The real-time urban road scene semantic segmentation model (EMFANet_AsSC model) shown in the figure performs semantic segmentation. The EMFANet_AsSC model is an end-to-end encoder-decoder structure that can segment various objects in complex urban road scenes in real time (140FPS) and accurately (72.4% mIoU on the Cityscapes test set) while maintaining low computational overhead and storage overhead (0.9M parameters). The encoder side includes three stages: Stage 1, Stage 2, and Stage 3. The decoder side consists of two dual attention aggregation units that are responsible for guiding and optimizing the multi-scale features extracted by the encoder side, and at the same time fully fusing them and outputting the final semantic segmentation map through the upsampler.

[0040] The present invention is further described below in conjunction with specific implementation steps.

[0041] Step 1: Get the original image X ori

[0042] In order to improve the generalization ability of the EMFANet_AsSC model, this case uses finely annotated complex road scene images as input images. The input images come from the Cityscapes dataset and the Camvid dataset. Both datasets contain rich urban road scene images. Among them: the categories that can be segmented in the Cityscapes dataset include roads, vehicles, pedestrians, trees, skies, buildings, etc., up to 30 categories, and the Camvid dataset also has up to 11 categories that can be segmented.

[0043] The input image is preprocessed to obtain the original image, which is used as the input of the encoder. The preprocessing operations include horizontal flipping, random change of input image size, cropping, etc., so that the scale range of the original image is 0.75, 1.0, 1.25, 1.5, 1.75 or 2.0. In order to speed up the reasoning speed of the EMFANet_AsSC model, the original image from the Cityscapes dataset is cropped from the original 1024×2048 to 512×1024, and the original image from the Camvid dataset is cropped from 720×960 to 360×480. Preprocessing the input image can further improve the robustness of the spatial position and multi-scale objects in the image.

[0044] In order to alleviate the class imbalance problem of the input images from the Camvid dataset, the class weights need to be adjusted as follows:

[0045]

[0046] Where: W i represents the adjusted category weight of category i, h represents the set hyperparameter, which is set to 1.12, and P i Represents the weight of category i in the Camvid dataset.

[0047] For input images from the Cityscapes dataset, the category weights remain consistent with the Cityscapes dataset.

[0048] Step 2: Divide the original image into training set and test set in proportion

[0049] Step 3: Stage 1

[0050] In the Stage 1 stage, a downsampler is used to downsample the original image, and then after three 3×3 standard convolutions, a channel attention layer is input to obtain low-level scale features, and a space-channel optimization unit (SBU unit) is used to extract the original image X ori The spatial features and background features of the network are fine-tuned and fused using depthwise separable convolution to obtain low-level optimization features. The low-level scale features are added to the low-level optimization features as the output of Stage 1.

[0051] like Figure 2 As shown in FIG. 1 , the downsampler includes a 3×3 standard convolution branch and a 2×2 maximum pooling branch. The input of the downsampler passes through the 3×3 standard convolution branch and the 2×2 maximum pooling branch respectively, and then the outputs of the two branches are channel-joined to obtain the output of the downsampler.

[0052] like Figure 3 As shown in the figure, it is a lightweight and efficient CAM Layer structure. The input of the CAM Layer is sequentially subjected to global average pooling operation, compression operation, transposition operation, a 1×1 standard convolution, transposition operation, decompression operation and a Sigmoid activation function, and then multiplied by the input of the CAM Layer. The result is used as the output of the CAM Layer. The CAM Layer is described as:

[0053] Y C (X C )=δ(f Unsq (f Trans (C 1×1 (f Trans (f sq (fAvgPool ))))))

[0054] Where: X C Represents the input of CAM Layer, the size is C×H×W, C represents the number of channels, H represents X C High, W means X C Width; Y C Represents the output of CAM Layer, with a size of C×1×1; f AvgPool represents the global average pooling operation; f sq Indicates compression operation; f Trans Represents the transpose operation; C 1×1 represents a standard convolution with a kernel size of 1×1; f Unsq represents the decompression operation; δ represents the Sigmoid activation function.

[0055] In Stage 1, the output Y of CAM Layer C Recorded as low-level scale features.

[0056] like Figure 4 As shown in the figure, the SBU unit includes a maximum pooling branch and a global average pooling branch. The input of the SBU unit passes through the maximum pooling branch and the average pooling branch respectively, and then the outputs of the two branches are channel-joined. Then, a 3×3 depth-separable convolution and a 1×1 standard convolution are sequentially performed, and the result is used as the output of the SBU unit. The SBU unit is described as:

[0057] Y S (X ori )=C 1×1 (DC 3×3 (Concat(f MaxPool ,f AvgPool )))

[0058] Where: X ori Represents the original image, size H ori ×W ori ×3,H ori Represents X ori High, W ori Represents X ori In this example, the original image size from the Cityscapes dataset is 512×1024×3, and the original image size from the Camvid dataset is 360×480×3; S represents the output of the SBU unit; f MaxPool represents the maximum pooling operation; f AvgPool Represents the global average pooling operation; Concat represents the channel concatenation operation, DC 3×3represents a depth-wise separable convolution with a kernel size of 3×3; C 1×1 Represents a standard convolution with a kernel size of 1×1.

[0059] For the maximum pooling branch and the average pooling branch, r is used to represent the difference between the output of the SBU unit and the original image X ori The ratio of the maximum pooling branch and the average pooling branch is k, the pooling step size of the maximum pooling branch and the average pooling branch is s, and (r, k, s) is used as the setting parameter of the SBU unit.

[0060] The SBU unit can compensate for the loss of characteristic information in downsampling to a certain extent, thereby improving the accuracy. In the Stage 1 stage, the output Y of the SBU unit is S Denoted as a low-level optimization feature, the parameters (r, k, s) are set to (2, 2, 2).

[0061] Through the downsampling of the downsampler, we can get the original image X ori The deep semantic information of the original image X can be extracted through three 3×3 standard convolutions. ori The shallow features of the image are obtained by optimizing the outputs of three 3×3 standard convolutions through the CAM Layer to obtain low-level scale features; the low-level scale features and low-level optimized features are added, and the output of the SBU unit is used to make up for the information lost during the downsampling process, thereby making up for and optimizing the extracted low-level scale features, thereby improving the segmentation accuracy. The low-level scale features will be fused with the high-level scale features at the decoding end to make up for the detailed information lost by the high-level scale features, better guide the expression of the final fused features, and thus improve the segmentation accuracy.

[0062] Step 4: Stage 2

[0063] In the Stage 2 stage, a downsampler is used to downsample the output of the Stage 1 stage, and then after passing through four symmetric attention residual units (ARU units), a CAM Layer is input to obtain the mid-scale features, and an SBU unit is used to extract the original image X ori The spatial features and background features are fine-tuned and fused using depthwise separable convolution to obtain intermediate optimized features. The intermediate scale features are added to the intermediate optimized features as the output of Stage 2.

[0064] In Stage 2, the structures of the downsampler, CAM Layer, and SBU unit are the same as those in Stage 1. The difference is that in this stage: ① the input of the downsampler is the output of Stage 1; ② the output Y of the CAM Layer is converted to CRecorded as low-level scale features; ③ The setting parameters (r, k, s) of the SBU unit are (4, 4, 4), and the output Y S Denoted as the intermediate optimization feature,.

[0065] like Figure 5 As shown, the input of the ARU unit first passes through a CAM Layer and a 3×3 depth-separable convolution, and then is equally separated into two parts through the Split function. One part is input into the standard decomposition convolution, and the other part is input into the depth-level decomposition convolution. Then, the outputs of the standard decomposition convolution and the depth-level decomposition convolution are channel-joined, and then a 3×3 depth-separable convolution is passed. The output of the 3×3 depth-separable convolution is added to the input of the ARU unit and then a channel confusion operation is performed. The result is used as the output of the ARU unit. The standard decomposition convolution includes a 3×1 standard convolution, a 1×3 standard convolution, a CAM Layer, a 3×1 standard convolution, and a 1×3 standard convolution. The depth-level decomposition convolution includes a 3×1 depth-separable convolution, a 1×3 depth-separable convolution, a CAM Layer, a 3×1 depth-separable convolution, and a 1×3 depth-separable convolution. The expansion rate of the standard convolution and the depth-separable convolution is d.

[0066] The ARU unit adopts a bilateral asymmetric residual structure design. The left branch uses standard decomposition convolution, and the right branch uses depth-level decomposition convolution. It reduces the redundant information between the two branches and thus reduces the number of parameters. At the same time, the ARU unit also makes extensive use of CAM Layer to optimize feature expression through attention.

[0067] In the four ARU units in Stage 2, the expansion rates d are 1, 2, 5, and 9 respectively.

[0068] Step 5: Stage 3

[0069] In the Stage 3 stage, a downsampler is used to downsample the output of the Stage 2 stage, and then after passing through eight ARU units, a CAM Layer is input to obtain high-level scale features, and an SBU unit is used to extract the original image X ori The spatial features and background features are fine-tuned and fused using depthwise separable convolution to obtain high-level optimization features. The high-level scale features are added to the high-level optimization features as the output of Stage 3.

[0070] In Stage 3, the structures of the downsampler, CAM Layer, ARU unit, and SBU unit are the same as those in Stage 2. The difference is that in this stage: ① the input of the downsampler is the output of Stage 2; ② the output of the CAM Layer YC Recorded as high-level scale feature; ③ The setting parameters (r, k, s) of the SBU unit are (8, 8, 8), and the output Y S Recorded as high-level optimization features; ④ Among the eight ARU units, the expansion rates d are 1, 2, 5, 9, 2, 5, 9, and 17, respectively, which are completely different from the expansion rates d of the four ARU units in Stage 4, aiming to obtain contextual information in different ranges.

[0071] Step 6: Decoder

[0072] The decoder first uses two dual attention aggregation units (DAAU units) to guide and optimize the multi-scale features extracted by the encoder in sequence, and then uses an upsampler to upsample and output the final segmentation prediction map; the first DAAU unit fuses the output of Stage 3 and the mid-scale features, and the second DAAU unit fuses the output of the first DAAU unit and the low-scale features, and uses the upsampler to upsample the output of the second DAAU unit to obtain the final segmentation prediction map.

[0073] like Figure 6 The figure shows the structure of the DAAU unit, including a left block and a right block; in the left block, the lower-level scale features are subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution and then a global average pooling operation is performed to obtain the lower-level output features of the left block, and the higher-level scale features are subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution and then input into the Sigmoid activation function to obtain the higher-level output features of the left block, and the lower-level output features of the left block and the higher-level output features of the left block are multiplied and then input into the upsampler for upsampling to obtain the output features of the left block; in the right block In the figure, the lower-level scale features undergo a 3×3 depth-separable convolution and a 1×1 standard convolution to obtain the lower-level output features of the right block. The higher-level scale features undergo a 3×3 depth-separable convolution and a 1×1 standard convolution, first input into the upsampler for upsampling, and then input into the Sigmoid activation function to obtain the higher-level output features of the right block. The lower-level output features of the right block and the higher-level output features of the right block are multiplied to obtain the output features of the right block; the output features of the left block and the output features of the right block are added, and the result is used as the output of the DAAU unit.

[0074] In the first DAAU unit, the lower-level scale features are the mid-level scale features, and the higher-level scale features are the high-level scale features; in the second DAAU unit, the lower-level scale features are the low-level scale features, and the higher-level scale features are the output of the first dual attention aggregation unit.

[0075] The DAAU unit is extremely lightweight and efficient, and can achieve a 1.3% improvement in segmentation accuracy with a slight overhead of only 0.03M parameters. The key to the design of the DAAU unit is the use of a dual aggregation strategy. Its left and right blocks use Sigmoid to activate the attention weights of higher-level features at different scales and perform dual fusion with lower-level detail information, thereby achieving the effect of the attention mechanism. At the same time, the two DAAU units on the encoder side also make full use of semantic information and detail information during the fusion process, so the decoder side can be very lightweight and efficient.

[0076] Step 7: Supervised Training

[0077] The loss is calculated after upsampling the output of Stage 3 using an upsampler. The calculated loss is called the auxiliary loss. The loss is calculated for the final segmentation prediction map. The calculated loss is called the main loss. The auxiliary loss and the main loss are weighted summed to get the total loss.

[0078] The upsampler uses traditional bilinear interpolation upsampling to obtain the final semantic segmentation map of the original resolution. In order to accelerate and stabilize model training, auxiliary loss and main loss are used when calculating the total loss. The auxiliary loss can strengthen the EMFANet_AsSC model's attention to deep features. The calculation formula of the total loss is expressed as:

[0079] Loss = Loss1 + α Loss2

[0080] Among them: Loss represents the total loss, Loss1 represents the main loss, Loss2 represents the auxiliary loss, α represents the weight of the auxiliary loss, and after experimental analysis, α is optimally set to 0.5.

[0081] The total loss is used to perform supervised training on the EMFANet_AsSC model to obtain the trained EMFANet_AsSC model.

[0082] Step 8: Model testing

[0083] The test set is input into the trained EMFANet_AsSC model to obtain the segmentation prediction map. The evaluation indicators of this method are reasoning speed (FPS), segmentation accuracy (mIoU) and parameter quantity (M). It can be seen from Table 1 that the method of the present invention has a certain improvement in segmentation accuracy compared with similar algorithms, and the algorithm parameter quantity overhead is small, and it also has a more obvious advantage in running speed.

[0084] Table 1 Comparison of indicators of different algorithms on the Cityscapes test set

[0085]

[0086] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A lightweight real-time semantic segmentation method for urban roads based on deep learning, characterized by: The steps include: Step 1: Take the finely annotated complex road scene image as the input image, and preprocess the input image to obtain the original image; Step 2, divide the original image into training set and test set according to proportion; Step 3: Input the training set into the real-time urban road scene semantic segmentation model, and generate the final segmentation prediction map through cross-attention fusion and upsampling; Step 4: Use the loss function to supervise the real-time urban road scene semantic segmentation model to obtain a trained real-time urban road scene semantic segmentation model; Step 5: Input the test set into the trained real-time urban road scene semantic segmentation model to obtain the segmentation prediction map; The real-time urban road scene semantic segmentation model is an end-to-end encoder-decoder structure; The encoder end includes three stages: Stage 1, Stage 2 and Stage 3. At the beginning of each of the three stages, a downsampler is used to downsample the input features. At the end of each of the three stages, a channel attention layer is used to optimize the extracted multi-scale features. At the same time, a space-channel optimization unit is used to extract the spatial features and background features of the original image and a depth-separable convolution is used for fine-tuning and fusion. The output of the channel attention layer is added to the output of the space-channel optimization unit as the output of the corresponding stage. In the Stage 1 stage, the input feature is the original image, and three 3×3 standard convolutions are set between the downsampler and the channel attention layer. The multi-scale features output by the channel attention layer in this stage are called low-scale features; in the Stage 2 stage, the input feature is the output of the Stage 1 stage, and four symmetric attention residual units are set between the downsampler and the channel attention layer. The multi-scale features output by the channel attention layer in this stage are called medium-scale features; in the Stage 3 stage, the input feature is the output of the Stage 2 stage, and eight symmetric attention residual units are set between the downsampler and the channel attention layer. The multi-scale features output by the channel attention layer in this stage are called high-scale features; The decoder first guides and optimizes the multi-scale features extracted by the encoder in sequence through two dual attention aggregation units, and then uses an upsampler to upsample and output the final segmentation prediction map; the first dual attention aggregation unit performs feature fusion on the output of the Stage 3 stage and the mid-scale features, the second dual attention aggregation unit performs feature fusion on the output of the first dual attention aggregation unit and the low-scale features, and uses the upsampler to upsample the output of the second dual attention aggregation unit to obtain the final segmentation prediction map; Use the upsampler to upsample the output of Stage 3 and then calculate the loss. The calculated loss is called auxiliary loss. The loss is calculated for the final segmentation prediction map, and the calculated loss is called the main loss; the auxiliary loss and the main loss are weighted summed to get the total loss.

2. The method for lightweight real-time semantic segmentation of urban roads based on deep learning according to claim 1 is characterized in that: The finely annotated complex road scene images are derived from the Cityscapes dataset and the Camvid dataset. The input images are preprocessed, and the preprocessing operations include horizontal flipping, randomly changing the input image size, and cropping. For the input images derived from the Cityscapes dataset, the category weights are kept consistent with those of the Cityscapes dataset; for the input images derived from the Camvid dataset, the category weights are adjusted according to the following formula: Where: W i represents the adjusted category weight of category i, h represents the set hyperparameter, P i Represents the weight of category i in the Camvid dataset.

3. The method for lightweight real-time semantic segmentation of urban roads based on deep learning according to claim 1, characterized in that: The downsampler includes a 3×3 standard convolution branch and a 2×2 maximum pooling branch. The input of the downsampler passes through the 3×3 standard convolution branch and the 2×2 maximum pooling branch respectively, and then the outputs of the two branches are channel-joined to obtain the output of the downsampler.

4. The method for lightweight real-time semantic segmentation of urban roads based on deep learning according to claim 1, characterized in that: The input of the channel attention layer is sequentially subjected to a global average pooling operation, a compression operation, a transposition operation, a 1×1 standard convolution, a transposition operation, a decompression operation, and a Sigmoid activation function, and then multiplied by the input of the channel attention layer. The result is used as the output of the channel attention layer. The channel attention layer is described as: Y C (X C )=δ(f Unsq (f Trans (C 1×1 (f Trans (f sq (f AvgPool )))))) Where: X C represents the input of the channel attention layer, Y C represents the output of the channel attention layer, f AvgPool represents the global average pooling operation, f sq represents the compression operation, f Trans represents the transpose operation, C 1×1 represents a standard convolution with a kernel size of 1×1, f Unsq represents the decompression operation, and δ represents the Sigmoid activation function.

5. The method for lightweight real-time semantic segmentation of urban roads based on deep learning according to claim 1, characterized in that: The space-channel optimization unit includes a maximum pooling branch and a global average pooling branch. The input of the space-channel optimization unit passes through the maximum pooling branch and the average pooling branch respectively, and then the outputs of the two branches are channel-joined, and then sequentially pass through a 3×3 depth-separable convolution and a 1×1 standard convolution, and the obtained result is used as the output of the space-channel optimization unit; the space-channel optimization unit is described as: Y S (X ori )=C 1×1 (DC 3×3 (Concat(f MaxPool ,f AvgPool ))) Where: X ori represents the original image, Y S represents the output of the space-channel optimization unit, f MaxPool represents the maximum pooling operation, f AvgPool represents the global average pooling operation, Concat represents the channel concatenation operation, DC 3×3 represents a depthwise separable convolution with a kernel size of 3×3, C 1×1 Represents a standard convolution with a kernel size of 1×1; For the maximum pooling branch and the average pooling branch, r is used to represent the ratio of the output of the space-channel optimization unit to the original image, the pooling kernel of the maximum pooling branch and the average pooling branch are both k, the pooling step size of the maximum pooling branch and the average pooling branch are both s, and (r, k, s) is used as the setting parameter of the space-channel optimization unit; for the space-channel optimization units in the three stages of Stage1, Stage2 and Stage3, the setting parameters (r, k, s) are (2, 2, 2), (4, 4, 4) and (8, 8, 8), respectively.

6. The method for lightweight real-time semantic segmentation of urban roads based on deep learning according to claim 1, characterized in that: The input of the symmetric attention residual unit first passes through a channel attention layer and a 3×3 depth-separable convolution, and is then equally separated into two parts through the Split function. One part is input into the standard decomposition convolution, and the other part is input into the depth-level decomposition convolution. The outputs of the standard decomposition convolution and the depth-level decomposition convolution are then channel-joined, and then pass through a 3×3 depth-separable convolution. The output of the 3×3 depth-separable convolution is added to the input of the symmetric attention residual unit and then a channel confusion operation is performed. The result is used as the output of the symmetric attention residual unit; The standard decomposition convolution includes a 3×1 standard convolution, a 1×3 standard convolution, a channel attention layer, a 3×1 standard convolution and a 1×3 standard convolution. The depth-level decomposition convolution includes a 3×1 depth-separable convolution, a 1×3 depth-separable convolution, a channel attention layer, a 3×1 depth-separable convolution and a 1×3 depth-separable convolution. The expansion rates of the standard convolution and the depth-separable convolution are both d; in the four symmetric attention residual units in the Stage2 stage, the expansion rate d is 1, 2, 5, and 9 respectively; in the eight symmetric attention residual units in the Stage3 stage, the expansion rate d is 1, 2, 5, 9, 2, 5, 9, and 17 respectively.

7. The method for lightweight real-time semantic segmentation of urban roads based on deep learning according to claim 1, characterized in that: The dual attention aggregation unit includes a left block and a right block; in the left block, the lower-level scale feature is subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution and then a global average pooling operation is performed to obtain the lower-level output feature of the left block; the higher-level scale feature is subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution and then input into a Sigmoid activation function to obtain the higher-level output feature of the left block; the lower-level output feature of the left block and the higher-level output feature of the left block are multiplied and then input into an upsampler for upsampling to obtain the output feature of the left block; in the right block, The lower-level scale features are subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution to obtain the lower-level output features of the right block. The higher-level scale features are subjected to a 3×3 depth-separable convolution and a 1×1 standard convolution, and then input into the upsampler for upsampling, and then input into the Sigmoid activation function to obtain the higher-level output features of the right block. The lower-level output features of the right block and the higher-level output features of the right block are multiplied to obtain the output features of the right block. The output features of the left block and the output features of the right block are added, and the result is used as the output of the dual attention aggregation unit. In the first dual attention aggregation unit, the lower-level scale features are mid-level scale features, and the higher-level scale features are high-level scale features; In the second dual attention aggregation unit, the lower-level scale features are the low-level scale features, and the higher-level scale features are the output of the first dual attention aggregation unit.

Citation Information

Patent Citations

  • Lightweight network real-time semantic segmentation method based on attention mechanism

    CN112330681A