An asymmetric road scene semantic segmentation system with multi-scale feature extraction

CN117746364BActive Publication Date: 2026-09-29ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311693642.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2026-09-29
Estimated Expiration
2043-12-06

AI Technical Summary

Technical Problem

[0006]上述语义分割方法在单方面的性能取得了较好的结果,但对语义分割实时性和准确性的平衡依然有待提升

Benefits of technology

(1)本发明通过不同特征提取模块对不同层次(1/2、1/4、1/8)的特征图进行特征提取,可以有效地提取上下文特征,然后通过特征融合模块将不同特征结果进行融合,最后使用解码器加强对空间细节的恢复,模型结构简单,大大减少计算量,提升解算速度,并且通过多层次的信息提取和融合,提升特征提取的准确度,从而整体方案实现对语义分割实时性和准确性的平衡。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117746364B_ABST
    Figure CN117746364B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-scale feature extraction asymmetric road scene semantic segmentation systems, including encoder and decoder, encoder includes standard convolution module, first feature fusion module, first RCAM module, DEB module, second feature fusion module, second RCAM module, TEB module and third feature fusion module, standard convolution module reduces the size of original input image by half and average pooling to 1 / 2 size, and the original input image is merged;Then the obtained feature map is down-sampled to 1 / 4 of the original picture resolution, and DEB module and RCAM module are used for feature extraction respectively;Finally, the feature map is down-sampled to 1 / 8, and feature extraction is carried out by TEB module and second RCAM module respectively, and the decoder upsamples the output results to recover and fuse;The application has the advantages that the balance between real-time and accuracy of semantic segmentation is achieved, and the semantic segmentation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and more specifically to an asymmetric road scene semantic segmentation system with multi-scale feature extraction. Background Technology

[0002] Semantic segmentation refers to predicting the category of each pixel in an input image to provide richer information for scene understanding based on the semantics of each pixel. Semantic segmentation has wide applications in real-world fields such as autonomous driving, geological monitoring, and remote sensing. Deep learning-based semantic segmentation algorithms outperform traditional methods in terms of accuracy and inference speed, making them increasingly popular in semantic segmentation research. While some semantic segmentation algorithms have made significant progress in accuracy in recent years, their complex network structures, often due to deep convolutional layers and large channel counts, hinder their deployment and application on mobile devices with limited memory and computing power. Conversely, some segmentation algorithms offer faster inference speeds and fewer parameters, but their accuracy is not ideal. Therefore, in real-time applications, designing a lightweight and efficient real-time semantic segmentation algorithm that balances accuracy, inference speed, and parameter count is more important than pursuing networks with only single performance goals.

[0003] FCN was the first to apply convolutional neural networks to the field of semantic segmentation. However, due to its limited receptive field size and simple structure, the network could not effectively utilize contextual information. Therefore, subsequent research has mostly focused on increasing the receptive field size and extracting feature information to improve segmentation accuracy. For example, the paper "Yang Lili, Chen Yan, Tian Weize, et al. Improved UNet segmentation method for field roads [J]. Transactions of the Chinese Society of Agricultural Engineering, 2021, 37(09): 185-191. doi:10.11975 / j.issn.1002-6819.2021.09.021." improves segmentation accuracy by changing the convolutional kernel to increase the receptive field; Deeplabv3+ adopts an encoding-decoding structure, which encodes contextual information through the backbone network and ASPP module, and adds a relatively simple decoder for accuracy optimization; DecoupleSegNet adopts a dual-stream framework and achieves high-resolution information sharing between different branches to achieve segmentation of low-resolution images; the paper "ZHANG Zhijie, PANG Yanwei. CGNet: Cross-guidance network for semantic segmentation [J]. Science China Information Sciences, 2020, 63(2): 120104. doi: 10.1007 / s11432-019-2718-7.》Zhang et al. unified edge and saliency information in CGNet, modeled the intrinsic information between them, and then used a cross-guided module to optimize feature information extraction.

[0004] High-precision network models often involve deeper convolutions and higher computational and memory consumption, limiting their widespread application on resource-constrained mobile devices. Therefore, lightweight real-time semantic segmentation algorithms have gained more attention, aiming to find the optimal balance between accuracy, network model complexity, and inference speed. Currently, real-time semantic segmentation networks mainly fall into two architectures: lightweight backbone architectures and multi-branch architectures. The paper "Liu Yun, Lu Chengzhe, Li Shijie, et al. Lightweight Semantic Segmentation Based on Efficient Multi-Scale Feature Extraction [J]. Chinese Journal of Computers, 2022, 45(07):1517-1528. doi: 10.11897 / SP.J.1016.2022.01517." proposes a segmentation model called MiniNet. This network achieves high-precision semantic segmentation with fewer parameters by fusing multi-scale features from small-scale feature maps and balancing network depth and the number of convolutional channels by adding spatial pyramid convolution and pooling modules. DABNet uses a depth-asymmetric bottleneck layer to extract feature information and effectively recovers spatial information by integrating features of different depths. This network does not contain upsampling layers, which can further reduce the number of model parameters. DFANet designs a backbone network based on a lightweight extended space and improves the segmentation performance of the model by using a cross-level feature aggregation module. In recent years, multi-branch architectures have attracted attention due to the problem of information loss during downsampling in semantic segmentation networks. CaBiNet employs a dual-branch structure, and this network effectively acquires the relevance of contextual information by adding a global information aggregation module to the semantic branch; the cascaded network ICNet acquires semantic and spatial information from multiple images at different resolutions, and integrates features through a cascade strategy to gradually restore and refine the segmentation results; DDRNet, based on the dual-branch structure, cascades two different branches and adds an improved pyramid module to the high-resolution branch to extract contextual information.

[0005] To further refine feature maps, some studies have introduced attention mechanisms into computer vision, with channel attention and spatial attention being two important components. SENet assigns weights to different channels based on the loss through a fully connected network, thereby distinguishing the importance of each channel; CBAM, in addition to channel attention, incorporates a spatial attention mechanism in a concatenated manner to weight pixels in the feature map; FANet modifies the self-attention mechanism to obtain a fast spatial attention mechanism, capturing spatial context information with relatively low computational cost.

[0006] The aforementioned semantic segmentation methods have achieved good results in terms of performance in one aspect, but there is still room for improvement in balancing the real-time performance and accuracy of semantic segmentation. Summary of the Invention

[0007] The technical problem to be solved by this invention is how to achieve a balance between the real-time performance and accuracy of semantic segmentation, and improve the semantic segmentation effect.

[0008] This invention solves the above-mentioned technical problems through the following technical means: an asymmetric road scene semantic segmentation system with multi-scale feature extraction, including an encoder and a decoder. The encoder includes a standard convolution module, a first feature fusion module, a first RCAM module, a DEB module, a second feature fusion module, a second RCAM module, a TEB module, and a third feature fusion module. The standard convolution module reduces the size of the original input image by half and averages the original input image to 1 / 2 size. The two are then merged through the first feature fusion module. The resulting feature map is then downsampled to 1 / 4 of the original image resolution, and features are extracted using the DEB module and the RCAM module respectively. The feature map obtained from the above operations is then averaged with the feature map of the original image to 1 / 4 size and merged through the second feature fusion module. Finally, the feature map is downsampled to 1 / 8, and features are extracted using the TEB module and the second RCAM module respectively. The resulting feature map is then merged with the feature map of the original image to 1 / 8 size through the third feature fusion module. The outputs of the first to third feature fusion modules are respectively input to the decoder. The decoder upsamples and restores each output result and merges them to output the final feature map.

[0009] Furthermore, the standard convolutional module includes three 3×3 convolutional layers. One 3×3 convolutional layer with a stride of 2 is used to reduce the size of the original input image by half and expand the number of channels to 32. Then, two 3×3 convolutional layers are used to extract contextual information and average pool the original input image to half its size before merging it with the image.

[0010] Furthermore, the first to third feature fusion modules have the same structure, each including a fusion unit and a pointwise multiplication unit. The fusion unit receives the output of the previous stage and the original input image after average pooling to 1 / 2, 1 / 4 or 1 / 8. The output of the fusion unit is connected to the pointwise multiplication unit, and the output of the pointwise multiplication unit is used as the output of the entire feature fusion module.

[0011] Furthermore, the DEB module includes a 3×3 convolutional unit, a 1×1 convolutional unit, a 3×3 depthwise separable convolutional unit, a 3×1 depthwise separable dilated convolutional unit, and a 1×3 depthwise separable dilated convolutional unit. The input feature map is processed by the 3×3 convolutional unit and then fed into the 3×3 depthwise separable convolutional unit and the 3×1 depthwise separable dilated convolutional unit, respectively. By using the 3×3 depthwise separable convolutional unit, local information is extracted while preserving spatial information. The 3×1 depthwise separable dilated convolutional unit with a dilation rate of 2 and the 1×3 depthwise separable dilated convolutional unit are used sequentially to extract the contextual information of the feature map while reducing the computational load. Then, the two branch information are fused by the 1×1 convolutional unit and residually connected with the input feature map. Channel shuffling operation is used to enhance the information interaction between channels.

[0012] Furthermore, the TEB module includes a 3×3 convolutional unit, a 1×1 convolutional unit, and three depthwise separable dilatational convolutional branches. Each depthwise separable dilatational convolutional branch includes a 3×1 depthwise separable dilatational convolutional unit and a 1×3 depthwise separable dilatational convolutional unit. The input feature map is fed into the three depthwise separable dilatational convolutional branches after passing through the 3×3 convolutional unit. Then, the information from the three branches is fused through the 1×1 convolutional unit and residually connected to the input feature map. A more comprehensive feature representation is obtained through channel shuffling.

[0013] Furthermore, the first RCAM module and the second RCAM module have the same structure. The first RCAM module includes an efficient attention module (ECA) and an average pooling layer. The feature map is input to the efficient attention module (ECA) and the average pooling layer. The output of the average pooling layer is residually concatenated with the output of the efficient attention module (ECA) to serve as the output of the first RCAM module.

[0014] Furthermore, the ECA module uses 1×1 convolutions instead of fully connected layers after global average pooling to avoid dimensionality reduction and capture cross-channel interactions. The kernel size is changed through an adaptive function, as follows:

[0015] in, Indicates the kernel size. Indicates the channel dimension. It is 2. =1, This indicates taking the nearest odd number.

[0016] Furthermore, the decoder includes three 1×1 convolutional units, a SAM attention module, and a 3×3 depthwise separable convolutional unit. The first to third feature fusion modules output 1 / 2, 1 / 4, and 1 / 8 feature maps, respectively. The 1 / 8 feature map is input into a 1×1 convolutional unit and then upsampled by a factor of 2. Simultaneously, spatial information is extracted from the 1 / 4 feature map through a 1×1 convolutional unit and the SAM attention module. The two outputs are merged, and then an upsampling operation is performed by a 3×3 depthwise separable convolutional unit. Finally, the 1 / 2 feature map is added point by point to recover the spatial information, and the final feature map is output through upsampling.

[0017] Furthermore, the SAM attention module performs max pooling and average pooling operations along the channel direction, then performs concatenation and convolution operations on the output, and finally uses an activation function to determine the weights to generate effective feature descriptions. This process is represented by the following formula:

[0018] in, This represents the resulting spatial feature map. Indicates the input feature map, This represents a standard convolution with kernel k. Represents the Concat merge operation. and These represent average pooling and max pooling operations, respectively. This indicates the Sigmoid activation function.

[0019] Furthermore, a dataset is constructed to train the semantic segmentation network. Stochastic gradient descent is used as the optimizer to update parameters, with a batch size of 8 and a learning rate decay policy formula of...

[0020] in, This represents the current learning rate. This represents the initial learning rate. and These represent the current iteration number and the maximum iteration number, respectively. The preset index is set to 0.9; Network accuracy is evaluated using the average intersection-over-union ratio (AUC). The formula for AUC is:

[0021] in, This indicates the number of pixels that are correctly classified. This represents the number of pixels with actual category i and predicted category j. This represents the number of pixels with actual category j and predicted category i. The total number of categories for classifying pixels.

[0022] The advantages of this invention are: (1) This invention extracts features from feature maps at different levels (1 / 2, 1 / 4, 1 / 8) through different feature extraction modules, which can effectively extract context features. Then, the different feature results are fused through the feature fusion module. Finally, the decoder is used to enhance the restoration of spatial details. The model structure is simple, greatly reducing the amount of computation and improving the solution speed. Furthermore, through multi-level information extraction and fusion, the accuracy of feature extraction is improved. Thus, the overall scheme achieves a balance between the real-time performance and accuracy of semantic segmentation.

[0023] (2) In the encoder, the improved DEB module and channel attention module (RCAM module) are used to extract features from different images respectively. The outputs of different modules and the pooling map of the original image are connected through the feature fusion module to obtain richer contextual information. In the decoder structure, a lightweight multi-scale decoder using spatial attention mechanism is designed to effectively recover spatial details under the condition of a small number of parameters, so that the network can achieve a good balance in various indicators.

[0024] (3) The DEB and TEB modules of this invention perform batch normalization and PReLU operations before each convolution operation to help improve the convergence effect of the model. Both modules use depthwise separable convolution, so after connecting different branches, channel dependencies are restored by using 1×1 pointwise convolution, thereby reducing model parameters while maintaining model performance. The two modules use different receptive fields for feature maps at different levels to extract features, so as to improve the feature extraction efficiency with a smaller number of parameters.

[0025] (4) In the RCAM module of this invention, the ECA module uses 1×1 convolutions instead of fully connected layers after global average pooling to avoid dimensionality reduction and capture cross-channel interactions. The kernel size is changed through an adaptive function, the input feature map is averaged, and residual connections are made with the ECA module to enhance the acquisition of global spatial information. Therefore, this module can better improve model performance with a smaller number of parameters.

[0026] (5) The present invention designs an asymmetric decoder to upsample and recover the output of the encoder. The main idea of ​​the multi-scale feature decoder is to combine feature maps of different resolutions and fuse low-level feature maps with high-level feature maps to recover spatial information and obtain richer spatial information. Multi-level feature information is used to improve performance. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the overall structure of an asymmetric road scene semantic segmentation system for multi-scale feature extraction disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the principle of the feature extraction module in an asymmetric road scene semantic segmentation system with multi-scale feature extraction disclosed in an embodiment of the present invention. Figure 2 (a) is a schematic diagram of the DEB module. Figure 2 (b) is a schematic diagram of the TEB module; Figure 3 This is a schematic diagram of the RCAM module principle in an asymmetric road scene semantic segmentation system for multi-scale feature extraction disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of the decoder principle in an asymmetric road scene semantic segmentation system with multi-scale feature extraction disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram comparing ablation experiment results in an asymmetric road scene semantic segmentation system for multi-scale feature extraction disclosed in an embodiment of the present invention. Figure 6 This is a diagram showing a comparison of visualization results of different algorithms in an asymmetric road scene semantic segmentation system for multi-scale feature extraction disclosed in an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] like Figure 1 As shown, this invention provides an asymmetric road scene semantic segmentation system with multi-scale feature extraction, including an encoder and a decoder. The encoder includes a standard convolutional module, a first feature fusion module, a first RCAM module, a DEB module, a second feature fusion module, a second RCAM module, a TEB module, and a third feature fusion module. The feature fusion module is abbreviated as FFM. The semantic segmentation network is abbreviated as LMANet.

[0030] The original input image is halved in size using a standard 3×3 convolution with a stride of 2, and the number of channels is expanded to 32. Two more standard 3×3 convolutions are then used to extract rich contextual information. The original input image is then averaged and pooled to half its size before being merged with it. The resulting feature map is then downsampled to 1 / 4 of the original image resolution, and features are extracted using a dual-branch extract bottleneck (DEB) module and a residual channel attention (RCAM) module. The feature map obtained from the above operations is then merged with the feature map of the original image, which has been averaged and pooled to 1 / 4 of its size, through a feature fusion module (FFM). Finally, the feature map is downsampled to 1 / 8 of its size, and features are extracted using a three-branch extract bottleneck (TEB) module and RCAM. The resulting feature map is then merged with the feature map of the original image, which has been averaged and pooled to 1 / 8 of its size, through an FFM module. In the decoder stage, the Multiscale Feature Decoder (MFD) introduces a Spatial Attention Module (SAM) to enhance the recovery of spatial information, effectively improving the segmentation performance. The specific structure of the semantic segmentation network is shown in Table 1.

[0031] Table 1. Specific Structure of LMANet

[0032] In real-time semantic segmentation algorithms, pursuing a wide receptive field to capture richer contextual information is a common approach. However, maintaining the same receptive field size throughout the network can limit model efficiency for feature maps at different levels. Therefore, using receptive fields of different scales at different stages to extract features can obtain effective information at different levels. Thus, this invention designs two different modules, the DEB module and the TEB module, to extract features from feature maps of different resolutions, with the following structures: Figure 2 As shown.

[0033] like Figure 2As shown in (a), the DEB module includes a 3×3 convolutional unit, a 1×1 convolutional unit, a 3×3 depthwise separable convolutional unit, a 3×1 depthwise separable dilated convolutional unit, and a 1×3 depthwise separable dilated convolutional unit. The input feature map is fed into the 3×3 convolutional unit and then into the 3×3 depthwise separable convolutional unit and the 3×1 depthwise separable dilated convolutional unit, respectively. By using the 3×3 depthwise separable convolutional unit, local information is extracted while preserving spatial information. The 3×1 depthwise separable dilated convolutional unit with a dilation rate of 2 and the 1×3 depthwise separable dilated convolutional unit are used in sequence to extract the contextual information of the feature map while reducing the amount of computation. Then, after fusing the information of the two branches through the 1×1 convolutional unit, a residual connection is made with the input feature map, and the information interaction between the channels is enhanced by channel shuffling operation. The DEB module is primarily used for extracting contextual information from low-level feature maps. Its left branch uses a 3×3 depthwise separable convolution to extract local information while preserving spatial information, while the right branch uses a depthwise separable dilated convolution with an dilation rate of 2 to extract contextual information from the feature map while reducing computational cost. The information from both branches is then fused with the input feature map via a residual connection, and channel shuffling enhances the information exchange between channels. Here, DDConv represents depthwise separable dilated convolution, Conv represents convolution, and DConv represents depthwise separable convolution.

[0034] like Figure 2 As shown in (b), the TEB module includes a 3×3 convolutional unit, a 1×1 convolutional unit, and three depthwise separable dilated convolutional branches. Each depthwise separable dilated convolutional branch includes a 3×1 depthwise separable dilated convolutional unit and a 1×3 depthwise separable dilated convolutional unit. The input feature map is fed into the three depthwise separable dilated convolutional branches after passing through the 3×3 convolutional unit. Then, the information from the three branches is fused through the 1×1 convolutional unit and residually connected to the input feature map. A more comprehensive feature representation is obtained through channel shuffling. The TEB module is mainly used for extracting contextual information from high-level feature maps. This module meets the requirement of adapting a single module to different receptive fields through a multi-branch structure. Each branch also uses depthwise separable dilated convolution to reduce the number of model parameters. After merging all branches, a residual connection is made with the input feature map. The module also incorporates channel shuffling to obtain a more effective and comprehensive feature representation.

[0035] In the feature extraction process of the two modules above, batch normalization and PReLU operations were performed before each convolution operation to improve the model's convergence performance. Both modules use depthwise separable convolutions, so after connecting different branches, 1×1 pointwise convolutions are used to restore channel dependencies, thereby reducing model parameters while maintaining model performance. The DEB and TEB modules use different receptive fields for feature maps at different levels to extract features, aiming to improve feature extraction efficiency with a smaller number of parameters.

[0036] like Figure 3 As shown, in order to obtain richer channel information, this invention introduces a channel attention module RCAM, namely the first RCAM module and the second RCAM module, while extracting information from feature maps at different levels. The first RCAM module and the second RCAM module have the same structure. The first RCAM module includes an efficient attention module ECA and an average pooling layer. The feature map is input to the efficient attention module ECA and the average pooling layer. The output of the average pooling layer is residually concatenated with the output of the efficient attention module ECA to serve as the output of the first RCAM module. This module mainly consists of the Efficient Channel Attention (ECA) module and average pooling layers, as disclosed in the paper "WANG Qilong, WUBanggu, ZHU Pengfei, et al. ECA-Net: Efficient Channel Attention for DeepConvolutional Neural Networks[C] / / 2020 IEEE / CVF Conference on ComputerVision and Pattern Recognition, Seattle, USA, 2020: 11531-11539. doi:10.1109 / CVPR42600.2020.01155.". Since dimensionality reduction in the SE module has side effects and is inefficient in capturing relationships between channels, the ECA module uses 1×1 convolutions instead of fully connected layers after global average pooling to avoid dimensionality reduction and capture cross-channel interactions. The convolution kernel size is changed through an adaptive function, allowing for more interaction between multiple channels. The adaptive function is as follows:

[0037] in, Indicates the kernel size. Indicates the channel dimension. It is 2. =1, This indicates that the nearest odd number is taken. The input feature map is then subjected to average pooling and residually connected to the ECA module to enhance the acquisition of global spatial information. Therefore, this module can improve model performance with a smaller number of parameters.

[0038] like Figure 4 As shown, in the encoder-decoder structure, the encoder is responsible for generating dense feature maps, while the decoder is responsible for upsampling the feature maps to the original resolution. In this model, an asymmetric decoder is designed to upsample and recover the encoder's output. The main idea of ​​the multi-scale feature decoder is to combine feature maps of different resolutions to obtain richer spatial information. The decoder includes three 1×1 convolutional units, a SAM attention module, and a 3×3 depthwise separable convolutional unit. The first to third feature fusion modules output 1 / 2, 1 / 4, and 1 / 8 feature maps, respectively. The 1 / 8 feature map is input into a 1×1 convolutional unit and then upsampled by a factor of 2. Simultaneously, spatial information is extracted from the 1 / 4 feature map using a 1×1 convolutional unit and the SAM attention module. The two outputs are merged, and then an upsampling operation is performed using a 3×3 depthwise separable convolutional unit. Finally, spatial information is recovered by adding the 1 / 2 feature map point by point, and the final feature map is output. The SAM attention module performs max pooling and average pooling operations along the channel direction, then performs concatenation and convolution operations on the output, and finally uses an activation function to determine the weights to generate effective feature descriptions. This process is represented by the following formula:

[0039] in, This represents the resulting spatial feature map. Indicates the input feature map, This represents a standard convolution with kernel k. Represents the Concat merge operation. and These represent average pooling and max pooling operations, respectively. This represents the Sigmoid activation function. This model uses multi-level feature information in the decoder to improve performance, that is, fusing low-level feature maps with high-level feature maps to recover spatial information.

[0040] The following is an experimental analysis of semantic segmentation networks: 1) Dataset and Experiment Setup This invention primarily uses the Cityscapes dataset to verify the network model's balance between real-time performance, accuracy, and parameter count, and the CamVid dataset to verify the model's generalization ability. Both datasets are commonly used for validating real-time semantic segmentation models. The Cityscapes dataset, as one of the main datasets, contains 5000 finely annotated images and 20000 coarsely annotated images from street scenes of different cities. The finely annotated images consist of 2975 finely annotated images, 500 validation images, and 1525 test images. In this invention, the images in this dataset are cropped from a resolution of 1024×2048 to a quarter size for training and speed evaluation. The CamVid dataset, as an auxiliary dataset for this invention, contains 701 images, including 367 images for training, 101 images for validation, and 233 images for testing, with a resolution of 720×960. This experiment uses a 2080Ti GPU with CUDA version 10.0 and an Intel Xeon Platinum 8255C @ 2.50GHz CPU. In this study, stochastic gradient descent is used as the optimizer to update parameters, the batch size is set to 8, and the learning rate decay strategy is formulated as follows:

[0041] in, This represents the current learning rate. This represents the initial learning rate. and These represent the current iteration number and the maximum iteration number, respectively. The preset index is set to 0.9. Furthermore, this invention uses three metrics—parameter quantity, inference speed (FPS), and mIoU—to reflect the real-time performance and accuracy of the network model. Here, mIoU represents the average intersection-union ratio, and the formula for the average intersection-union ratio is...

[0042] in, This indicates the number of pixels that are correctly classified. This represents the number of pixels with actual category i and predicted category j. This represents the number of pixels with actual category j and predicted category i. The total number of categories for classifying pixels.

[0043] 2) Ablation experiment As described above, the model as a whole employs an asymmetric encoding and decoding structure. Feature extraction and attention modules are used at different stages of the decoder, and a lightweight and efficient decoder structure is designed to more effectively acquire spatial and contextual information under conditions of low parameters and high speed, achieving better segmentation results. To verify the effectiveness of the modules and their interrelationships within the model, ablation experiments were conducted on the Cityscapes dataset, and the results are shown in Table 2.

[0044] Table 2 Ablation Experiment Results of Different Modules

[0045] In the experiments, a deep asymmetric bottleneck network was used as the base network for comparison. As shown in the table, when the low-level feature extraction module was replaced with the DEB module, the number of parameters and inference speed remained almost unchanged compared to the base network, while its mIoU value increased to 71.2%. Adding the TEB module, due to its additional branch compared to the original feature extraction module and the different dimensionality settings in each module to obtain multi-scale feature information, had a certain impact on the number of parameters and inference speed. However, with the other two indicators remaining relatively similar, the mIoU value after adding the TEB module was significantly improved. In the FFM module, the input feature map was processed by adding a channel attention module, and compared with the base network, its number of parameters and inference speed remained almost unchanged, while the mIoU value increased by 0.9 percentage points. The multi-scale feature decoder MFD designed in this invention, with the introduction of a spatial attention mechanism, only increased the number of model parameters by 0.01M, and the inference speed was almost unaffected, but the mIoU value was improved, indicating that this decoder can effectively process multi-scale features and recover spatial details well.

[0046] This invention selects some visualization results from the Cityscapes validation set for comparison to perform a visual analysis of the performance of each module in the model, such as... Figure 5 As shown in the figure, when using different feature extraction modules and channel attention modules, the segmentation results are more obvious in edge details and contours, such as the sign in scene 1 and the wall in scene 3; while when using the decoder MFD, due to the addition of the spatial attention module, the segmentation results are better for small target objects, such as distant targets on the road in scenes 2 and 3. This shows that each module in LMANet has a certain degree of effectiveness.

[0047] To verify the effectiveness of the internal structure of the module, ablation experiments were conducted on the dimensionality of the feature extraction module, the selection of the attention module in RCAM, and the effectiveness of the connection method in the MFD structure. The experimental results are shown in Tables 3 to 6.

[0048] Table 3 Experimental results for different dimensions in the TEB module

[0049] As shown in Table 3, various scenarios were considered when setting the dimensions of the two branches in the TEB module. Removing one branch slightly optimized the model's parameter count and inference speed, but significantly reduced the mIoU value. Furthermore, the data in the first two rows indicate that setting different dimensions to extract multi-scale features can effectively improve the segmentation performance. When the dimensions of the two branches in the TEB module were set to other sizes or the module order was changed, the model achieved better segmentation results with TEB modules having branch dimensions of {2,2,4,4,8,8} and {4,4,8,8,16,16} than with other choices in the experiment, without a significant change in the parameter count. This demonstrates the effectiveness of setting the dimensions in the TEB module.

[0050] Table 4 Ablation Experiment Results of RCAM

[0051] Table 4 compares the different attention modules used to replace the ECA module in RCAM, and also compares the pooling methods in the residual connections. The table shows that using CBAM and SE attention mechanisms slightly increases the number of model parameters and decreases inference speed. However, the model with ECA achieves the best segmentation performance. Furthermore, in RCAM, average pooling results in a 0.7 percentage point higher mIoU value than max pooling, making global average pooling more suitable. Ablation experiments were also conducted on the connection methods of feature maps at different scales in the MFD structure, and the results are shown in Table 5.

[0052] Table 5 Experimental results of different connection methods in MFD

[0053] In the table, the high-level fusion method represents the fusion of the 1 / 4 and 1 / 8 feature maps, while the low-level fusion method represents the fusion of the 1 / 2 feature map with the feature map obtained from the high-level fusion. As shown in Table 5, this invention uses three methods—connection, pointwise multiplication, and pointwise addition—in arbitrary combinations. When connection and pointwise addition are used in MFD, the model's mIoU value reaches its highest value with almost no change in the number of parameters and inference speed, demonstrating the effectiveness of the multi-scale feature map fusion method in the MFD structure.

[0054] In addition, this invention also explored the impact of the number of DEB modules and TEB modules on the balance between model real-time performance and accuracy, and the results are shown in Table 6.

[0055] Table 6 Experimental results for different numbers of extraction modules

[0056] In the table, m and n represent the number of DEB modules and TEB modules, respectively. Experimental results show that as the number of modules increases, the model achieves higher accuracy at the cost of slower inference speed and a larger number of parameters, with the number of TEB modules having a more significant impact on the cost. However, the data in the first two rows indicates that increasing only the number of a certain type of module does not significantly improve model performance. To strike a proper balance between parameters and accuracy, this invention selects 3 and 6 as the values ​​for m and n, respectively.

[0057] 3) Network model performance comparison Based on the proposed semantic segmentation network, performance tests were conducted on the Cityscapes dataset, and the results were compared with existing lightweight semantic segmentation algorithms. The comparison results are shown in Table 7.

[0058] Table 7 compares the performance of LMANet with different networks on the Cityscapes dataset.

[0059] The quantitative analysis in the table shows that the model of this invention achieves a good balance among different metrics. On the Cityscapes dataset, it achieves an inference speed of 111 FPS and a mIoU of 74.5%, with only 0.85M parameters. Compared to ENet and ESPNet, which have the lowest parameter counts in the table, LMANet shows a significant improvement in accuracy. Compared to DABNet, which also uses an asymmetric encoding / decoding structure, this model improves its mIoU value by 4.2 percentage points with minimal impact from parameter count and inference speed. Furthermore, compared to the Hyperseg-S model, which has the highest mIoU value in the table, LMANet has fewer than one-tenth the number of parameters as Hyperseg-S and also shows a significant improvement in inference speed. Therefore, LMANet achieves a good balance between parameter count, inference speed, and accuracy.

[0060] To further verify LMANet's generalization ability, this invention also tested LMANet on the CamVid dataset, and the comparison results are shown in Table 8.

[0061] Table 8 compares the performance of LMANet with different networks on the CamVid dataset.

[0062] As shown in Table 8, compared with other network models, LMANet's inference speed can meet the requirements of real-time segmentation tasks. Compared with ESPNet, which has the fastest inference speed in the table, its mIoU value has been significantly improved. Compared with BiseNetv2, which has a higher accuracy, LMANet has improved in terms of inference speed and parameter count. This shows that the network can better balance the relationship between accuracy and real-time performance, and also indicates that the network has strong generalization ability.

[0063] This invention compares the segmentation accuracy of LMANet with that of some typical networks on the Cityscapes dataset for different categories, and the comparison results are shown in Table 9. The results show that the proposed algorithm outperforms the comparison algorithms in segmentation accuracy for all categories, which demonstrates that the multi-scale feature extraction and recovery method of this network can effectively improve the model's ability to distinguish between different categories.

[0064] Table 9. Accuracy (IoU / %) ​​of different networks for each category

[0065] Continued from Table 9: Accuracy (IoU / %) ​​of Different Network Categories

[0066] To intuitively demonstrate the segmentation performance of different network models, this invention performs a qualitative analysis of LMANet on the Cityscapes dataset and provides a visual comparison with the segmentation performance of some other models. The experimental results are as follows: Figure 6 As shown in the comparison figures, in scenarios one and four, LMANet performs better segmentation of small, distant objects compared to other algorithms. In scenarios two and three, the proposed algorithm provides clearer segmentation outlines for traffic lights, pedestrians, fences, and warning signs, while other algorithms show less significant segmentation results. In scenarios five and six, the proposed algorithm demonstrates better segmentation of roads obscured by pedestrians and the edges of walls on both sides of the road. This indicates that while most segmentation algorithms are relatively accurate for larger objects, the proposed algorithm is more effective at segmenting complex scenes. Furthermore, due to the use of feature extraction modules at different scales and a decoder MFD with a spatial attention module, LMANet achieves more accurate spatial detail recovery compared to other networks.

[0067] By employing the above technical solutions, this invention addresses the challenge of balancing segmentation accuracy, model parameter count, and inference speed in road scene semantic segmentation algorithms. It proposes LMANet, an asymmetric lightweight semantic segmentation network model based on an encoder-decoder structure. This algorithm extracts features from feature maps at different levels through different feature extraction modules, effectively extracting contextual features. Then, a feature fusion module fuses the different feature results, and finally, a decoder (MFD) enhances the recovery of spatial details. Experiments on the Cityscapes dataset show that LMANet achieves 74.5% mIoU with 0.85M parameters and an inference speed of 111 FPS. Furthermore, on the CamVid dataset, LMANet still maintains a good balance between accuracy, speed, and parameter count. These experimental results demonstrate that the algorithm effectively meets the requirements of real-time semantic segmentation algorithms and possesses a certain degree of effectiveness.

[0068] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An asymmetric road scene semantic segmentation system with multi-scale feature extraction, characterized in that, The system includes an encoder and a decoder. The encoder comprises a standard convolution module, a first feature fusion module, a first RCAM module, a DEB module, a second feature fusion module, a second RCAM module, a TEB module, and a third feature fusion module. The standard convolution module reduces the size of the original input image by half and averages the original input image to half its size. The two are then merged by the first feature fusion module. The resulting feature map is then downsampled to 1 / 4 of the original image resolution, and features are extracted using the DEB and RCAM modules respectively. The feature map obtained from the above operations is then averaged with the original image and averaged to 1 / 4 of its size, and merged by the second feature fusion module. Finally, the feature map is downsampled to 1 / 8 of its size, and features are extracted using the TEB and second RCAM modules respectively. The resulting feature map is then averaged with the original image and averaged to 1 / 8 of its size by the third feature fusion module. The outputs of the first to third feature fusion modules are input to the decoder, which upsamples and restores each output and then merges them to output the final feature map. The DEB module includes a 3×3 convolutional unit, a 1×1 convolutional unit, a 3×3 depthwise separable convolutional unit, a 3×1 depthwise separable dilated convolutional unit, and a 1×3 depthwise separable dilated convolutional unit. The input feature map is processed by the 3×3 convolutional unit and then by the 3×3 depthwise separable convolutional unit and the 3×1 depthwise separable dilated convolutional unit. The 3×3 depthwise separable convolutional unit extracts local information while preserving spatial information. The 3×1 depthwise separable dilated convolutional unit with a dilation rate of 2 and the 1×3 depthwise separable dilated convolutional unit are then used sequentially. The TEB module extracts contextual information from the feature map while reducing computational cost. This information is then fused with the two branch information via a 1×1 convolutional unit and residually connected to the input feature map. Channel shuffling enhances information exchange between channels. The TEB module includes a 3×3 convolutional unit, a 1×1 convolutional unit, and three depthwise separable dilated convolutional branches. Each depthwise separable dilated convolutional branch includes a 3×1 depthwise separable dilated convolutional unit and a 1×3 depthwise separable dilated convolutional unit. The input feature map is processed by the 3×3 convolutional unit and then input into the three branches respectively. A depthwise separable dilated convolutional branch is used, and then the information from the three branches is fused through a 1×1 convolutional unit before being residually concatenated with the input feature map. A channel shuffling operation is then performed to obtain a more comprehensive feature representation. The first and second RCAM modules have the same structure. The first RCAM module includes an efficient attention module (ECA) and an average pooling layer. The feature map is input to the ECA and the average pooling layer, and the output of the average pooling layer is residually concatenated with the output of the ECA to serve as the output of the first RCAM module. The decoder includes three 1×1 convolutional units and one SAM attention module. The system includes a force module and a 3×3 depthwise separable convolutional unit. The first to third feature fusion modules output 1 / 2, 1 / 4, and 1 / 8 feature maps, respectively. The 1 / 8 feature map is input into a 1×1 convolutional unit and then upsampled by 2x. Simultaneously, a 1×1 convolutional unit and the SAM attention module extract spatial information from the 1 / 4 feature map. The two outputs are merged and then upsampled by a 3×3 depthwise separable convolutional unit. Finally, the system adds the 1 / 2 feature map point by point to recover the spatial information and outputs the final feature map through upsampling.

2. The asymmetric road scene semantic segmentation system for multi-scale feature extraction according to claim 1, characterized in that, The standard convolutional module includes three 3×3 convolutional layers. One 3×3 convolutional layer with a stride of 2 is used to reduce the size of the original input image by half and expand the number of channels to 32. Then, two 3×3 convolutional layers are used to extract contextual information and average pool the original input image to half its size before merging it with the image.

3. The asymmetric road scene semantic segmentation system for multi-scale feature extraction according to claim 1, characterized in that, The first to third feature fusion modules have the same structure, each including a fusion unit and a pointwise multiplication unit. The fusion unit receives the output of the previous stage and the original input image after average pooling to 1 / 2, 1 / 4 or 1 / 8. The output of the fusion unit is connected to the pointwise multiplication unit, and the output of the pointwise multiplication unit is used as the output of the entire feature fusion module.

4. The asymmetric road scene semantic segmentation system for multi-scale feature extraction according to claim 1, characterized in that, The ECA module uses 1×1 convolutions instead of fully connected layers after global average pooling to avoid dimensionality reduction and capture cross-channel interactions. The kernel size is changed using an adaptive function, as follows: in, Indicates the kernel size. Indicates the channel dimension. It is 2. =1, This indicates taking the nearest odd number.

5. The asymmetric road scene semantic segmentation system for multi-scale feature extraction according to claim 1, characterized in that, The SAM attention module performs max pooling and average pooling operations along the channel direction, then performs concatenation and convolution operations on the output, and finally uses an activation function to determine the weights to generate effective feature descriptions. This process is represented by the following formula: in, This represents the resulting spatial feature map. Indicates the input feature map, This represents a standard convolution with kernel k. Represents the Concat merge operation. and These represent average pooling and max pooling operations, respectively. This indicates the Sigmoid activation function.

6. The asymmetric road scene semantic segmentation system for multi-scale feature extraction according to claim 1, characterized in that, We construct a dataset, train the semantic segmentation network, use stochastic gradient descent as the optimizer to update parameters, set the batch size to 8, and use the learning rate decay policy formula as follows: in, This represents the current learning rate. This represents the initial learning rate. and These represent the current iteration number and the maximum iteration number, respectively. The preset index is set to 0.9; Network accuracy is evaluated using the average intersection-over-union ratio (AUC). The formula for AUC is: in, This indicates the number of pixels that are correctly classified. This represents the number of pixels with actual category i and predicted category j. This represents the number of pixels with actual category j and predicted category i. The total number of categories for classifying pixels.