Lightweight network-based automatic driving real-time semantic segmentation enhancement method

The real-time semantic segmentation model constructed by a lightweight network, combined with the depth-separable convolution and inverse residual structure, solves the problem of insufficient real-time and spatial information utilization in autonomous driving, and achieves faster and more accurate image segmentation, especially the segmentation of small-sized objects and edges.

CN120388339APending Publication Date: 2025-07-29FUJIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277703.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art lacks real-time semantic segmentation in automatic driving, especially in complex scenarios, because the use of spatial information of intermediate feature layer is lacking, resulting in insufficient image segmentation and accuracy.

Method used

A real-time semantic segmentation model is constructed using a lightweight network, using MobileNetV2 network, a hollow space convolution pooled pyramid module and a dual-branch feature extraction network, combining deep separable convolution and inverse residual structures, enhancing feature extraction capabilities through the ACBA module and enriching spatial location information.

Benefits of technology

It improves the real-time and accuracy of autonomous driving image segmentation, especially the segmentation ability of small-sized objects and edges, reduces information loss, and enhances the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388339A_ABST
    Figure CN120388339A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving real-time semantic segmentation enhancement method based on a lightweight network, which is characterized in that a lightweight and efficient coding path is built by using depth separable convolution in combination with an inverse residual structure to carry out feature extraction, so that parameters and operand are greatly reduced, and the real-time performance of automatic driving image segmentation is enhanced. The design of the double-branch structure enables abstract deep-layer features to be sequentially spliced and fused with middle-shallow-layer features, enriches the spatial position information of feature layers, and effectively improves the segmentation precision of the model, especially for small-size objects and edges. The ACBA module combines the advantages of the cavity convolution and the asymmetric convolution block ACB, so that the receptive field is expanded, the information extraction of the middle-shallow layer features of the two branches is effectively enhanced, and the information loss is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving assistance, and particularly to a method for enhancing real-time semantic segmentation for autonomous driving based on a lightweight network. Background Art

[0002] With the development of economic technology and the improvement of people's living standards, autonomous driving assistance technology has begun to be gradually applied to people's production and life, bringing endless convenience to people's production and life. Therefore, ensuring the safety and reliability of the autonomous driving process has become the research focus of researchers.

[0003] Semantic segmentation is to segment the images generated in a specific scene, that is, to classify each pixel in the image according to the predefined semantic categories, so as to achieve the purpose of segmenting the image. Real-time semantic segmentation is the core part of the autonomous driving process; therefore, real-time semantic segmentation is particularly important in the autonomous driving process.

[0004] An implementation method of unmanned street view image segmentation based on DeeplabV3+ in the prior art, due to the fusion of the mixed pooling module MPM and the dual attention network DANet, results in insufficient real-time performance of image segmentation. A prior patent on a method for real-time semantic segmentation of autonomous driving in a traffic scene, although it strengthens the deep features, lacks the utilization of the spatial information of the intermediate feature layers. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for enhancing real-time semantic segmentation for autonomous driving based on a lightweight network, so as to improve the real-time segmentation ability in complex-scene driving.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A method for enhancing real-time semantic segmentation for autonomous driving based on a lightweight network, which includes the following steps:

[0008] Step 1, obtain an autonomous driving training data set;

[0009] Step 2, construct an initial model for real-time semantic segmentation based on the MobileNetV2 network, the atrous spatial pyramid pooling (ASPP) module, and the dual-branch feature extraction network;

[0010] MobileNetV2 is used to extract deep features from the input image. One branch of the dual-branch feature extraction network performs 3 times of 2-fold downsampling on the input image to obtain a feature map with 1 / 8 size of the input image; the other branch of the dual-branch feature extraction network performs 2 times of 2-fold downsampling on the input image to obtain a feature map with 1 / 4 size of the input image; the ASPP module performs multi-scale feature fusion on the deep features to obtain fused features. After 2-fold upsampling, the fused features are concatenated and fused with the feature layer extracted by one branch of the dual-branch feature extraction network to enrich the spatial location information, and then after 2-fold upsampling, they are concatenated and fused with the feature layer extracted by the other branch of the dual-branch feature extraction network to enrich the spatial location information and semantic information of the feature layer; finally, the obtained features are upsampled 4 times to restore to the original image size to complete the image segmentation task.

[0011] Step 3: Use the training dataset to train the initial real-time semantic segmentation model to obtain the real-time semantic segmentation model.

[0012] Step 4: Obtain the image data information around the target vehicle in real time, and use the obtained real-time semantic segmentation model to perform semantic segmentation to obtain the semantic segmentation result.

[0013] Step 5: Realize the automatic driving of the target vehicle in real time according to the semantic segmentation result.

[0014] Furthermore, the initial real-time semantic segmentation model adopts an encoder-decoder structure.

[0015] Furthermore, in step 2, a lightweight and efficient encoding path MobileNetV2 network is built using depthwise separable convolution combined with the inverted residual structure for feature extraction; that is, the input features are first dimensionally increased through a 1*1 convolution, then the depthwise separable convolution is used to extract features, and finally the features are dimensionally reduced through a 1*1 convolution to obtain the deep features extracted by the MobileNetV2 network.

[0016] Specifically, the depthwise separable convolution includes depth convolution and pointwise convolution; the depth convolution independently performs convolution operations on each channel of the input layer to obtain feature maps of different channels, and then combines the feature maps of all channels to obtain a feature map with the same number of channels as the input channels; the pointwise convolution performs weighted combination on the feature map output by the depth convolution in the depth direction to generate a new feature map. This decomposition of the depthwise separable convolution can greatly reduce the computational complexity while maintaining a high classification accuracy.

[0017] The inverted residual structure first performs channel upsampling through a 1×1 convolution, then uses depthwise separable convolutions to extract features, and finally performs channel downsampling through a 1×1 convolution. This design allows for an increase in the model's non-linear representation ability while maintaining computational efficiency. Since the computational cost of depthwise separable convolutions is relatively low, even when operating in high-dimensional spaces, it does not significantly increase the computational burden. Using a linear bottleneck layer can prevent non-linearity from destroying too much information, and combining it with the inverted residual structure can enhance the model's linear representation ability, enabling the model to more effectively extract and utilize feature information when dealing with complex tasks.

[0018] Furthermore, an ACBA module is connected to the output end of each downsampling in the dual-branch feature extraction network; the ACBA module includes an asymmetric convolution block ACB and a dilated convolution. The dilated convolution is used to expand the receptive field, and the asymmetric convolution module is used to enhance the feature extraction ability and reduce information loss during the extraction process.

[0019] Specifically, one branch of the dual-branch feature extraction network includes three ACBA modules; the other branch of the dual-branch feature extraction network includes two ACBA modules.

[0020] Furthermore, the asymmetric convolution block ACB uses the additivity of convolutions to add horizontal and vertical asymmetric convolution kernels to the standard square convolution. The asymmetric convolution block ACB includes a 3×3 convolution, a 3×1 convolution, and a 1×3 convolution. The outputs of each convolution are fused through parallel computing without increasing the computational time.

[0021] Specifically, the asymmetric convolution block ACB can enhance the model's robustness to rotational distortion and strengthen the central skeleton part of the square convolution kernel. Furthermore, the dilated convolution can effectively expand the receptive field of the convolutional layer without increasing the number of parameters and computational complexity, thereby capturing more context information. However, traditional dilated convolutions also have some drawbacks. Although the dilation rate expands the receptive field of the convolutional layer, it is also prone to local information loss, resulting in problems such as inaccurate segmentation of small objects and discontinuous segmentation.

[0022] Therefore, combining ACB and dilated convolution to form an ACBA module can make full use of the advantages of both. It not only uses the dilated convolution to expand the receptive field of feature extraction but also uses the asymmetric convolution block ACB to capture more context information, effectively alleviating the information loss problem and enhancing the segmentation ability of small objects and the robustness of the model.

[0023] The present invention adopts the above technical solutions and has the following technical advantages: (1) The encoder uses a linear bottleneck layer combined with an inverted residual structure to enhance the linear expression ability of the model. By replacing the Xception structure with a separable convolution design, the number of model parameters is greatly reduced, and the real-time performance of segmentation is improved. (2) The present invention constructs a dual-branch structure for feature fusion. After multi-scale feature fusion of the deep features extracted by the backbone network through the ASPP module, they are first concatenated and fused with branch 1 to enrich the spatial information of the feature layer. After upsampling, they are concatenated and fused with branch 2 to further enrich the space of the feature layer. Finally, they are upsampled to restore the original image size. (3) The asymmetric convolution block (ACB) is combined with atrous convolution. The atrous convolution is used to expand the receptive field without increasing the number of parameters, and the asymmetric convolution is used to capture more context information, effectively enriching the feature extraction of the two branches and reducing information loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The following further elaborates on the present invention in detail with reference to the drawings and specific embodiments;

[0025] Figure 1 Schematic diagram of the TBAA-Deeplabv3+ network framework adopted by a method for enhancing real-time semantic segmentation of autonomous driving based on a lightweight network according to the present invention;

[0026] Figure 2 Schematic diagram of the depthwise separable convolution structure according to the present invention;

[0027] Figure 3 Schematic diagram of the feature extraction principle of branch 1 of the dual-branch feature extraction network according to the present invention;

[0028] Figure 4 Schematic diagram of the feature extraction principle of branch 2 of the dual-branch feature extraction network according to the present invention;

[0029] Figure 5 Schematic diagram of the structure of the asymmetric convolution block ACB according to the present invention;

[0030] Figure 6 Schematic diagram of the comparison of feature extraction of two convolutions before and after the image is flipped vertically according to the present invention;

[0031] Figure 7 Schematic diagram of the comparison of visualization results of different models on the Cityscapes validation set;

[0032] Figure 8 Schematic diagram of the comparison of visualization results of different models on the surrounding environment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application.

[0034] As Figures 1 to 8 shown in one of them, the present invention discloses a method for enhancing real-time semantic segmentation of autonomous driving based on a lightweight network, which includes the following steps:

[0035] Step 1, obtain an autonomous driving training data set;

[0036] Step 2, construct an initial model for real-time semantic segmentation based on the MobileNetV2 network, the Atrous Spatial Pyramid Pooling (ASPP) module, and a dual-branch feature extraction network;

[0037] MobileNetV2 is used to extract deep features from the input image. One branch of the dual-branch feature extraction network performs 3 times of 2-fold downsampling on the input image to obtain a feature map with a size of 1 / 8 of the input image; the other branch of the dual-branch feature extraction network performs 2 times of 2-fold downsampling on the input image to obtain a feature map with a size of 1 / 4 of the input image; the ASPP module performs multi-scale feature fusion on the deep features to obtain fused features. After the fused features are upsampled by 2 times, they are concatenated and fused with the feature layer extracted by one branch of the dual-branch feature extraction network to enrich the spatial location information. Then, after upsampling by 2 times, they are concatenated and fused with the feature layer extracted by the other branch of the dual-branch feature extraction network to enrich the spatial location information and semantic information of the feature layer; finally, the obtained features are upsampled by 4 times to restore to the original image size to complete the image segmentation task.

[0038] Step 3, train the initial model for real-time semantic segmentation with the training data set to obtain a real-time semantic segmentation model;

[0039] Step 4, obtain the image data information around the target vehicle in real time, and perform semantic segmentation using the obtained real-time semantic segmentation model to obtain a semantic segmentation result;

[0040] Step 5, realize the autonomous driving of the target vehicle in real time according to the semantic segmentation result.

[0041] Furthermore, as Figure 1 shown, in order to improve the real-time segmentation ability in complex-scene driving, the TBAA-Deeplabv3+ model of the method for enhancing real-time semantic segmentation of autonomous driving based on a lightweight network in this case adopts an encoder-decoder structure.

[0042] Furthermore, a lightweight and efficient encoding path MobileNetV2 is built using depthwise separable convolutions combined with an inverted residual structure for feature extraction, greatly reducing the number of parameters and the amount of computation, and enhancing the real-time performance of autonomous driving image segmentation. That is, the input features first go through a 1×1 convolution for channel dimensionality increase, then use depthwise separable convolutions to extract features, and finally go through a 1×1 convolution for channel dimensionality reduction to obtain the deep features extracted by the MobileNetV2 network.

[0043] The deep features extracted by the lightweight backbone network MobileNetV2 enter the ASPP module for multi-scale feature fusion. After being upsampled by a factor of 2, they are concatenated and fused with the feature layer extracted through branch 1 to enrich their spatial location information. Then, after being upsampled by a factor of 2 again, they are concatenated and fused with the feature layer extracted through branch 2 to further enrich the spatial location information and semantic information of the feature layer. Finally, the feature layer is upsampled by a factor of 4 to restore it to the original image size to complete the image segmentation task.

[0044] Specifically, different from traditional convolutions, depthwise separable convolutions decompose the operation into two independent steps: depthwise convolution and pointwise convolution, as Figure 2 shown. Depthwise convolution performs convolution operations independently on each channel of the input layer, and then combines all the feature maps to obtain a feature map with the same number of channels as the input channels. Pointwise convolution then performs weighted combination on the feature map in the depth direction to generate a new feature map. This decomposition can greatly reduce the computational complexity while maintaining a high classification accuracy.

[0045] Furthermore, using a linear bottleneck layer can prevent the nonlinearity from destroying too much information, and combining it with an inverted residual structure can enhance the linear expression ability of the model, enabling the model to more effectively extract and utilize feature information when dealing with complex tasks. The inverted residual structure first goes through a 1×1 convolution for channel dimensionality increase, then uses depthwise separable convolutions to extract features, and finally goes through a 1×1 convolution for channel dimensionality reduction again. This design allows for an increase in the nonlinear expression ability of the model while maintaining computational efficiency. Since the computational amount of depthwise separable convolutions is relatively low, even when operating in a high-dimensional space, it will not significantly increase the computational burden.

[0046] Therefore, in order to meet the real-time requirements of autonomous driving semantic segmentation, a lightweight and efficient MobileNetV2 is selected to be built using a linear bottleneck combined with an inverted residual structure as the encoding path.

[0047] Furthermore, an ACBA module is connected to the output end of each downsampling of the double-branch feature extraction network; the ACBA module includes an asymmetric convolution block ACB and a dilated convolution. The dilated convolution is used to expand the receptive field, and the asymmetric convolution module is used to enhance the feature extraction ability and reduce information loss during the extraction process.

[0048] Specifically, one branch of the dual-branch feature extraction network includes three ACBA modules; the other branch of the dual-branch feature extraction network includes two ACBA modules. As Figure 3 shown in Branch 1, the input image is downsampled by a factor of 2 and then passes through the ACBA module. The dilated convolution is used to expand the receptive field, and at the same time, the asymmetric convolution module is used to enhance the feature extraction ability and reduce information loss during the extraction process. Then, the operations of 2x downsampling and the ACBA module are repeated twice for feature extraction. Finally, Branch 1 obtains a feature map with a size of 1 / 8 of the original Figure 1 size. As Figure 4 shown in Branch 2, similar to Branch 1, the input image is downsampled by a factor of 2 and then passes through the ACBA module. The dilated convolution is used to expand the receptive field, and at the same time, the asymmetric convolution module is used to enhance the feature extraction ability and reduce information loss during the extraction process. Then, the operations of 2x downsampling and the ACBA module are performed once for feature extraction. Finally, Branch 2 obtains a feature map with a size of 1 / 4 of the original Figure 1 size.

[0049] Furthermore, as Figure 5 shown, the ACBA module is designed by combining ACB and dilated convolution. The asymmetric convolution block ACB uses the additivity of convolution to add horizontal and vertical asymmetric convolution kernels in the standard square convolution. The asymmetric convolution block ACB includes 3*3 convolution, 3*1 convolution, and 1*3 convolution. The outputs of each convolution are fused through parallel computing without increasing the computing time.

[0050] The asymmetric convolution block ACB can enhance the robustness of the model to rotational distortion and strengthen the central skeleton part of the square convolution kernel. Taking the 1*3 convolution as an example, as Figure 6 (a) figure shows, the two red rectangular frames represent the feature extraction operations before and after the image is flipped vertically. It can be seen that the 1*3 convolution can still extract the correct features after the image is flipped. When only 3*3 convolution kernels are used during the training phase, as Figure 6 (b) figure shows, after the image is flipped vertically, the extracted features are obviously different. Therefore, introducing a horizontal convolution kernel such as 1*3 can improve the robustness of the model to vertical flipping of the image, and the same applies to the 3*1 convolution kernel in the vertical direction.

[0051] Furthermore, dilated convolution can effectively expand the receptive field of the convolutional layer without increasing the number of parameters and computational complexity, thereby capturing more context information. However, traditional dilated convolution also has some drawbacks. Although the dilation rate expands the receptive field of the convolutional layer, it is also prone to local information loss, resulting in problems such as inaccurate segmentation of small objects and discontinuous segmentation. Therefore, combining ACB with dilated convolution to form the ACBA module can make full use of the advantages of both. It not only uses dilated convolution to expand the receptive field of feature extraction but also uses the asymmetric convolutional block ACB to capture more context information, effectively alleviating the problem of information loss and enhancing the segmentation ability of small objects and the robustness of the model.

[0052] Effect description: The proposed TBAA-Deeplabv3+ model of the present invention is applied to the autonomous driving scene segmentation task. To comprehensively evaluate the performance of the model, mIoU, mPA, FPS, Parameters, and GFlops are selected as evaluation metrics, and experiments are carried out on the Cityscapes dataset. Performance evaluation is carried out on the Cityscapes validation set, and the comparison of each evaluation metric is shown in Table 1.

[0053] Table 1 Comparison of each evaluation metric of the model before and after improvement on the Cityscapes validation set

[0054] Par / M GFlops FPS mPA / % mIoU / % Deeplabv3+ 54.71 667.93 29.15 77.40 68.25 TBAA-Deeplabv3+ 7.76 345.31 43.86 77.51 66.78

[0055] The IoU of different classification predictions of the model before and after improvement on the Cityscapes validation set is counted, as shown in Table 2.

[0056] Table 2 IoU (%) of different classification predictions of the Cityscapes validation set by the model before and after improvement

[0057] Classes models Deeplabv3+ TBAA-Deeplabv3+ road 96 96 sidewalk 76 74 building 87 87 wall 49 51 fence 46 46 pole 43 43 traffic light 50 50 traffic sign 59 60 vegetation 88 88 terrain 60 60 sky 92 92 person 66 66 rider 51 48 car 90 90 truck 74 67 bus 81 77 train 75 69 motorcycle 49 43 bicycle 64 63 mIoU 68.25 66.78

[0058] The performance comparison of different networks on the Cityscapes dataset is carried out, and the results are shown in Table 3.

[0059] Table 3 Comparison of evaluation metrics of different networks on the Cityscapes dataset

[0060] Par / M GFlops FPS mPA / % mIoU / % BiseNet 6.60 146.74 65.19 70.78 62.97 PIDNet 7.58 193.35 59.61 72.57 63.76 PSPNet 8.79 235.43 54.46 73.83 65.39 SCTNet 5.42 117.45 75.35 75.41 65.82 Deeplabv3+ 54.71 667.93 29.15 77.40 68.25 TBAA-Deeplabv3+ (the present invention) 7.76 345.31 43.86 77.51 66.78

[0061] The predictions of the model before and after improvement on the Cityscapes validation set are visualized, and some visualization results are as Figure 7 shown.

[0062] Images of the surrounding environment were also taken, and predictions were made using the Deeplabv3+ and TBAA-Deeplabv3+ models trained on the Cityscapes dataset. The visualization results are as shown in Figure 8 shown.

[0063] The present invention adopts the above technical solutions. By using depthwise separable convolutions in combination with an inverted residual structure to build a lightweight and efficient encoding path for feature extraction, the number of parameters and the amount of computation are greatly reduced, enhancing the real-time performance of autonomous driving image segmentation. The design of the dual-branch structure enables the abstract deep features to be successively concatenated and fused with the middle and shallow features, enriching the spatial location information of the feature layers and effectively improving the segmentation accuracy of the model, especially for small-sized objects and edges. Different from traditional dilated convolutions, the designed ACBA module combines the advantages of dilated convolutions and the asymmetric convolution block ACB, expanding the receptive field and effectively enhancing the information extraction of the middle and shallow features of the two branches, reducing information loss.

[0064] Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The components of the embodiments of the present application described and illustrated herein can generally be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

Claims

1. A method for enhancing real-time semantic segmentation of autonomous driving based on a lightweight network, characterized in that: It includes the following steps: Step 1: Obtain an autonomous driving training dataset; Step 2: Construct an initial real-time semantic segmentation model based on the MobileNetV2 network, the Atrous Spatial Pyramid Pooling (ASPP) module, and a dual-branch feature extraction network; Among them, MobileNetV2 is used to extract deep features from the input image. One branch of the dual-branch feature extraction network performs 3 times of 2-fold downsampling on the input image to extract a feature map with a size of 1 / 8 of the input image; the other branch of the dual-branch feature extraction network performs 2 times of 2-fold downsampling on the input image to extract a feature map with a size of 1 / 4 of the input image. The Atrous Spatial Pyramid Pooling (ASPP) module performs multi-scale feature fusion on the deep features to obtain fused features. After the fused features are upsampled by 2 times, they are concatenated and fused with the feature layer extracted by one branch of the dual-branch feature extraction network to enrich the spatial location information. Then, after upsampling by 2 times again, they are concatenated and fused with the feature layer extracted by the other branch of the dual-branch feature extraction network to enrich the spatial location information and semantic information of the feature layer. Finally, the obtained features are upsampled by 4 times to restore to the original image size to complete the image segmentation task; Step 3: Use the training dataset to train the initial real-time semantic segmentation model to obtain a real-time semantic segmentation model; Step 4: Real-time obtain the image data information around the target vehicle, and use the obtained real-time semantic segmentation model to perform semantic segmentation to obtain a semantic segmentation result; Step 5: Realize the autonomous driving of the target vehicle in real time according to the semantic segmentation result.

2. The real-time semantic segmentation enhancement method for autonomous driving based on a lightweight network according to claim 1, characterized in that: The initial real-time semantic segmentation model adopts an encoder-decoder structure.

3. An enhanced method for real-time semantic segmentation of autonomous driving based on a lightweight network according to claim 1, characterized in that: In Step 2, a lightweight and efficient encoding path MobileNetV2 network is built using depthwise separable convolutions combined with an inverted residual structure for feature extraction; that is, the input features are first dimensionally expanded through a 1*1 convolution, then the depthwise separable convolution is used to extract features, and finally the channels are dimensionally reduced through a 1*1 convolution to obtain the deep features extracted by the MobileNetV2 network.

4. An enhanced method for real-time semantic segmentation of autonomous driving based on a lightweight network according to claim 3, characterized in that: The depthwise separable convolution includes a depth convolution and a pointwise convolution; the depth convolution independently performs convolution operations on each channel of the input layer to obtain feature maps of different channels, and then combines all the channel feature maps to obtain a feature map with the same number of channels as the input channels; the pointwise convolution performs weighted combination on the feature map output by the depth convolution in the depth direction to generate a new feature map.

5. A method for enhancing real-time semantic segmentation of autonomous driving based on a lightweight network according to claim 1, characterized in that: An Asymmetric Convolution Bottleneck with Atrous (ACBA) module is connected to the output end of each downsampling of the dual-branch feature extraction network; the ACBA module includes an Asymmetric Convolution Block (ACB) and an atrous convolution. The atrous convolution is used to expand the receptive field, and the asymmetric convolution module is used to enhance the feature extraction ability.

6. The enhanced method for real-time semantic segmentation of autonomous driving based on a lightweight network according to claim 5, wherein: The Asymmetric Convolution Block (ACB) includes a 3*3 convolution, a 3*1 convolution, and a 1*3 convolution. The outputs of each convolution are fused through parallel computing and then output after passing through the ReLU function.