Automatic driving real-time semantic segmentation method and readable storage medium
Patent Information
- Application Number
- CN202311817869.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-12-27
AI Technical Summary
[0005]为了克服现有技术无法兼顾准确性和实时性的问题的不足,本发明提供了一种高精度低时延的自动驾驶实时语义分割方法及可读存储介质
[0009]The real-time semantic segmentation method and readable storage medium for autonomous driving provided by this invention employs structural reparameterization and deep supervised training techniques to design the model and the model training process. The aim is to improve the segmentation effect of the model without increasing inference latency. At the same time, a well-designed lightweight pyramid pooling module is inserted at an appropriate position, which not only takes into account the accuracy and running speed in autonomous driving scenarios, but also has a small number of model parameters, fast inference speed, high accuracy, and low latency. The method has strong universality, low cost, and high reliability.
Smart Images

Figure CN117788818B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a real-time semantic segmentation method for autonomous driving and a readable storage medium. Background Technology
[0002] Semantic segmentation is a fundamental task in computer vision, aiming to assign each pixel in an input image to a specific category label, thereby achieving scene parsing. In the field of autonomous driving, semantic segmentation technology is widely used for road scene understanding to identify key elements such as roads, vehicles, pedestrians, and traffic lights, thus enabling safe and reliable autonomous driving.
[0003] To meet the high standards required in the field of autonomous driving, researchers have proposed many different real-time semantic segmentation methods in recent years. Based on their research approaches, these methods can be categorized into module lightweighting, architecture lightweighting, and transfer learning schemes. Module lightweighting involves designing lightweight module structures, such as combining lightweight attention, asymmetric convolution, or depthwise separable convolution. These schemes attempt to reduce the complexity of the model or module with minimal accuracy loss, resulting in lower parameter and computational costs. However, the feature extraction capabilities of these lightweight modules are not as good as ordinary convolution. Some methods try to improve this by adjusting channels, but this severely impacts the model's inference speed. Therefore, this approach is not outstanding in terms of either accuracy or speed. Architecture lightweighting involves designing lightweight model architectures. Currently, the two mainstream architectures are lightweight encoding / decoding architectures and dual-resolution architectures. The architecture uses an encoder to extract features and a decoder to restore features. However, the architecture is prone to losing low-level features and boundary information, resulting in unsatisfactory segmentation results. The dual-resolution architecture uses two branches with different resolutions to capture high-level and low-level features respectively, while also achieving good inference speed. However, it requires the design of a corresponding feature fusion module, otherwise feature confusion will occur. The transfer learning scheme pre-learns massive amounts of data during the training phase and then makes some fine-tuning for the required scenario to achieve good results. However, learning from massive amounts of data requires a lot of resources, so its universality is weaker compared to the other two schemes.
[0004] In autonomous driving scenarios, real-time performance and accuracy are both indispensable. However, most existing real-time semantic segmentation methods cannot achieve a good balance between these two aspects. Therefore, designing a real-time semantic segmentation method with high accuracy and low latency is of great research significance for the fields of autonomous driving and intelligent transportation. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies that cannot simultaneously achieve both accuracy and real-time performance, this invention provides a high-precision, low-latency real-time semantic segmentation method for autonomous driving and a readable storage medium.
[0006] The technical solution of the present invention: In a first aspect, embodiments of the present invention provide a real-time semantic segmentation method for autonomous driving, comprising the following steps: Step S1. Obtain the autonomous driving dataset and divide it into a training set and a validation set in a ratio of 7:3; Step S2. Design a reparameterized feature extraction module using a reparameterization method and build a dual-resolution backbone network. The dual-resolution backbone network consists of a transition layer, a detail branch, a semantic branch, and a bidirectional fusion layer. The transition layer rapidly reduces the resolution of the input features of the dual-resolution backbone network to 1 / 8, and its output features are input to the detail branch and the semantic branch to extract different features respectively. The input features of the dual-resolution backbone network will always maintain a high resolution of 1 / 8 through the detail branch, and the input features through the semantic branch will be downsampled four times to a resolution of 1 / 64. The bidirectional fusion layer divides the detail branch and the semantic branch into three stages, and feature fusion is performed in each stage. Step S3. Insert a fast aggregation pyramid pooling module at the end of the semantic branch, and add an auxiliary segmentation head that only exists during the training phase at the end of the semantic branch in the dual-resolution backbone network to construct the model; Step S4. Train and validate the model using a deep supervised training strategy to obtain the model with the highest validation accuracy; Step S5. Use the model with the highest verification accuracy to perform real-time prediction, visualization, and feedback of the results for the acquired images or video streams.
[0007] In a second aspect, embodiments of the present invention provide a data processing apparatus, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the real-time semantic segmentation method for autonomous driving as described above.
[0008] Thirdly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, characterized in that the computer-executable instructions are used to execute the real-time semantic segmentation method for autonomous driving as described above. Beneficial effects
[0009] The real-time semantic segmentation method and readable storage medium for autonomous driving provided by this invention employs structural reparameterization and deep supervised training techniques to design the model and the model training process. The aim is to improve the segmentation effect of the model without increasing inference latency. At the same time, a well-designed lightweight pyramid pooling module is inserted at an appropriate position, which not only takes into account the accuracy and running speed in autonomous driving scenarios, but also has a small number of model parameters, fast inference speed, high accuracy, and low latency. The method has strong universality, low cost, and high reliability. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating the method and system of the present invention; Figure 2 This is a structural diagram of the feature extraction module designed using the reparameterization method in this invention; Figure 3 This is a diagram of the dual-resolution network structure constructed based on the RepRes (reparameterized residual module) module in this invention. Figure 4 This is a structural diagram of the fast aggregation pyramid pooling module of the present invention; Figure 5 This is a comparison prediction graph of the present invention with other advanced algorithms on the Cityscapes dataset. Detailed Implementation
[0011] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0012] This invention discloses a real-time semantic segmentation method and system for autonomous driving, the implementation flowchart of which is shown below. Figure 1 As shown, it includes the following steps: Step S1. Obtain the autonomous driving dataset and randomly divide it into a training set and a validation set in a ratio of 7:3; Step S2. Design a feature extraction module using a reparameterization method and build a segmentation backbone network based on a dual-resolution architecture, which is called a dual-resolution backbone network; The dual-resolution backbone network consists of a transition layer, a detail branch, a semantic branch, and a bidirectional fusion layer. The transition layer rapidly reduces the resolution of the input features of the dual-resolution backbone network to 1 / 8, and its output features are input to the detail branch and the semantic branch to extract different features respectively. The input features of the dual-resolution backbone network will always maintain a high resolution of 1 / 8 through the detail branch, and the input features through the semantic branch will be downsampled four times to a resolution of 1 / 64. The bidirectional fusion layer divides the detail branch and the semantic branch into three stages, and feature fusion is performed in each stage.
[0013] The feature extraction module designed using a reparameterization method is a multi-branch residual module. This module serves as a fundamental component of each layer of the dual-resolution backbone network, such as... Figure 2 As shown, the training and inference phases have different structures. Specifically, the training phase structure is a residual block containing two convolutional components. When training ends, it is transformed into an inference structure using a reparameterization method. The inference structure is a residual block containing two 3×3 convolutions and two non-linear activation layers (ReLU).
[0014] The feature extraction module designed based on the reparameterization method consists of two convolutional components and residual connections. Each convolutional component has three branches. The features input to the convolutional components are passed through 3×3 convolution and batch normalization layers, 1×1 convolution and direct mapping, respectively. The resulting features are added and fused, and then passed through batch normalization layers and non-linear activation layers. The output features of the two convolutional components are obtained by adding the three branches of the second convolutional component and adding an additional batch normalization layer. The resulting output features are fused with the input features of the module to obtain the output features of the feature extraction module designed based on the reparameterization method. During the inference stage, the three branches of the convolutional components are merged into a single branch using the reparameterization method.
[0015] The reparameterization method used in the convolutional component follows these steps: Step S2.1 Merge the branches of the 3×3 convolutional component and the batch normalization layer; Specifically, it includes the following: The three branches of the convolution component in the feature extraction module are merged to reduce runtime memory usage. For input feature x, the formulas for convolution Conv(x) and batch normalization BN(x) include: Conv(x)=w(x)+b Where w(x) represents the convolution kernel and b represents the bias of the convolution; Where γ represents the scaling factor of the batch normalization layer, and β represents the offset factor. represents the variance, and m represents the mean of the batch normalization layer.
[0016] Combining the formulas for convolution Conv(x) and BN(x), we obtain set up The result after merging is calculated. Conv'(x)=w'(x)+b' Where Conv'(x) represents the merged convolution, w'(x) represents the merged convolution kernel, and b' represents the merged bias.
[0017] Step S2.2 Expand both the 1×1 convolutional branch and the direct mapping branch of the convolutional component into 3×3 convolutions, and then merge the three branches into a single 3×3 convolution.
[0018] Specifically, it includes the following: For the 1×1 convolution branch of the convolution component, the 1×1 convolution kernel is padded with 0s to become a 3×3 convolution, and its kernel and bias are represented as w. 1×1 b 1×1 ; For the direct mapping branch of the convolution component, construct a 1×1 convolution with the identity matrix as the kernel, then pad it with zeros to convert it into a 3×3 convolution. The kernel and bias are represented as w. identity b identity Three parallel 3×3 convolutions can be combined into a single 3×3 convolution, as shown in the formula. w'=w 3×3 +w 1×1 +w identity b' = b 3×3 +b 1×1 +b identity Conv'(x)=w'(x)+b' The kernel and bias of the 3×3 convolution branch are represented as w. 3×3 b 3×3 w'(x) represents the merged convolution kernel, b' represents the merged bias, and Conv'(x) is the merged convolution. The merged structure is a residual block-like structure, which is easy to deploy.
[0019] Step S2.3 merges the merged 3×3 convolution with the batch normalization layer.
[0020] A reparameterized feature extraction module, named RepresBlock, is designed using a reparameterization method. A dual-resolution backbone network is then constructed based on this module, as shown in the diagram below. Figure 3As shown, specifically, given an input image, the Stem stage of the network quickly reduces it to 1 / 4 of the initial resolution to ensure the model's lightweight nature. Stage-1 is a transitional stage before the two branches. Stage-2 begins to split into two branches in parallel, where the detail branch D always maintains a high resolution of 1 / 8, and the semantic branch is downsampled by a factor of 2 and the channels are doubled after each stage to extract high-level semantic information; In both Stage-2 and Stage-3, both branches enhance feature interaction through a bidirectional fusion layer; Stage-4 utilizes the FAPPM module to acquire high-level semantic information at multiple scales and the global receptive field, and completes the final fusion of the two branches. Finally, there is only a simple segmentation head, mainly composed of two 3×3 convolutions, which reduces the number of channels of the 1 / 8 resolution features generated by the detail branch to the number of categories, and finally upsamples to the initial resolution size, thereby completing the final pixel-level prediction.
[0021] Step S3. Based on the dual-resolution backbone network built in step S2, insert the designed fast aggregation pyramid pooling module at the end of the semantic branch, and add a segmentation head at the end of the dual-resolution backbone network to complete the model construction. The detailed structure of the designed fast aggregation pyramid pooling module FAPPM is as follows: Figure 4 As shown, the average pooling layer parameters are (kernel=5, stride=2), convolution refers to 3×3 convolution and batch normalization layer, and bilinear interpolation upsampling operator is used for upsampling. The input feature resolution of the fast aggregation pyramid pooling module is 1 / 64 of the original image. The input features are processed through four 5×5 average pooling layers and a 1×1 convolution to extract feature information at five different scales. These features are then adjusted to 1 / 64 resolution through 3×3 convolution and bilinear interpolation upsampling operations, aggregated along the channel dimension, and added to the original input features for fusion. Finally, a 1×1 convolution is applied to adjust the input to the required size, thus obtaining the output features of the fast aggregation pyramid pooling module.
[0022] Step S4. Train and validate the model obtained in Step S3 using a deep supervised training strategy, as follows: Figure 3 As shown; Before training begins, the pre-trained weight parameters of the model backbone are loaded, with ImageNet as the pre-training dataset. In the first 6 rounds of training, the backbone gradient of the model is frozen for training, except for FAPPM and SegHea segmentation. During the training phase, a deep supervised training strategy with multiple losses is employed. The aforementioned deep supervised training strategy specifically includes adding an auxiliary segmentation head to the detail branch based on the model obtained in step S3, extracting boundary features from the output using the Canny boundary extraction operator, and adding boundary loss for constraint. The total loss during the model training phase is expressed as: Loss = L S +λ0L auxS +λ1L SB Among them, L S L represents the loss of the final predicted output. auxS λ0 and L represent the auxiliary segmentation loss and its weights, respectively. SB λ0 and λ1 represent the boundary loss and its weights, respectively, and are typically set as follows: λ0 = 0.4, λ1 = 20. During the training phase, stochastic gradient descent (SGD) is used to optimize the model; During the training phase, a multi-step learning rate decay strategy is used to adjust the learning rate, which can be expressed as follows: Where iter represents the current iteration number, and lr iter The learning rate is represented by iter, and Gamma is the preset decay rate.
[0023] Step S5. Use the model with the highest verification accuracy from step S4 to perform real-time prediction, visualization, and feedback of the acquired images or video streams. This specifically includes the following steps: Use the model with the highest validation accuracy to perform real-time inference on the input image or video stream; Block out categories or areas in the inference results that are not of interest to the user; Based on user visualization needs, the inference results are visualized and output to the terminal in real time, such as... Figure 5 The figure shows the visualization results of different models based on the Cityscapes dataset. The solid boxes represent the differences in predictions for certain categories by different models. It can be seen from the figure that even advanced algorithms from the last two years show inaccurate predictions for some complex categories, particularly for "eroded" telephone poles and buses. In addition, there are varying degrees of confusion issues for categories with similar colors. The reparameterized dual-resolution network of this invention does not have these problems and exhibits more accurate boundary predictions.
[0024] Additionally, one embodiment of this application provides a data processing apparatus, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor and the memory can be connected via a bus or other means.
[0025] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0026] The non-transitory software program and instructions required to implement the real-time semantic segmentation method for autonomous driving described in the above embodiments are stored in memory. When executed by a processor, the real-time semantic segmentation method for autonomous driving described in the above embodiments is executed, for example, the method described above is executed. Figure 1 The method steps S1 to S5 are described above. Furthermore, one embodiment of this application provides a computer-readable storage medium storing computer-executable instructions that are executed by a processor or controller, for example, by a processor in the above-described device embodiment, causing the processor to execute the real-time semantic segmentation method for autonomous driving described above, for example, executing the methods described above. Figure 1 Method steps S1 to S5.
[0027] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A real-time semantic segmentation method for autonomous driving, comprising the following steps: Step S1. Obtain the autonomous driving dataset and divide it into a training set and a validation set in a ratio of 7:3; Step S2. Design a reparameterized feature extraction module using a reparameterization method and build a dual-resolution backbone network. The dual-resolution backbone network consists of a transition layer, a detail branch, a semantic branch, and a bidirectional fusion layer. The transition layer rapidly reduces the resolution of the input features of the dual-resolution backbone network to 1 / 8, and its output features are input to the detail branch and the semantic branch to extract different features respectively. The input features of the dual-resolution backbone network will always maintain a high resolution of 1 / 8 through the detail branch, and the input features through the semantic branch will be downsampled four times to a resolution of 1 / 64. The bidirectional fusion layer divides the detail branch and the semantic branch into three stages, each stage... Feature fusion is performed; the feature extraction module is a multi-branch residual module, which is a basic component of each layer of the dual-resolution backbone network; the feature extraction module consists of two convolutional components and residual connections. Each convolutional component has three branches. The features input to the convolutional components are respectively passed through 3×3 convolution and batch normalization layer, 1×1 convolution and direct mapping. The resulting features are added and fused, and then passed through batch normalization layer and non-linear activation layer. The output features of the two convolutional components are fused with the input features of the module to obtain the output features of the feature extraction module designed based on the reparameterization method. During the inference stage, the three branches of the convolutional components are merged into a single branch through the reparameterization method. Step S3. A fast aggregation pyramid pooling module is inserted at the end of the semantic branch, and an auxiliary segmentation head that only exists during the training phase is added to the end of the semantic branch in the dual-resolution backbone network to construct the model; the input feature resolution of the fast aggregation pyramid pooling module is 1 / 64 of the original image. The input features are processed through four 5×5 average pooling layers and a 1×1 convolution to extract five different scales of feature information. After being adjusted to 1 / 64 resolution by 3×3 convolution and bilinear upsampling, they are aggregated along the channel and added to the original input features for fusion. Then, they are adjusted to the required input size by 1×1 convolution to obtain the output features of the fast aggregation pyramid pooling module. Step S4. Train and validate the model using a deep supervised training strategy to obtain the model with the highest validation accuracy; Step S5. Use the model with the highest verification accuracy to perform real-time prediction, visualization, and feedback of the results for the acquired images or video streams.
2. The real-time semantic segmentation method for autonomous driving according to claim 1, characterized in that, The reparameterization method comprises the following steps: Step S2.1 Merge the branches of the 3×3 convolutional component and the batch normalization layer; Step S2.2 Expand the convolutions of the remaining two branches of the convolution component into 3×3 convolutions, and then merge the three branches into a single 3×3 convolution; Step S2.3 merges the merged 3×3 convolution with the batch normalization layer.
3. The real-time semantic segmentation method for autonomous driving according to claim 2, characterized in that, In step S2.1, merging the branches of the 3×3 convolutional component and the batch normalization layer specifically includes the following:
4. The real-time semantic segmentation method for autonomous driving according to claim 2, characterized in that, In step S2.2, expanding the convolutions of the remaining two branches of the convolution component into 3×3 convolutions, and then merging the three branches into a single 3×3 convolution, specifically includes the following:
5. The real-time semantic segmentation method for autonomous driving according to claim 1, characterized in that, In step S4, the training and validation of the model using a deep supervised training strategy specifically includes the following: Training begins by loading the pre-trained weights of the model backbone; The backbone gradient of the model is frozen for the first six rounds of training. During the training phase, a deep supervised training strategy with multiple losses is employed. During the training phase, stochastic gradient descent is used to optimize the model; During the training phase, a multi-step learning rate decay strategy is used to adjust the learning rate, as expressed by the following formula:
6. The real-time semantic segmentation method for autonomous driving according to claim 5, characterized in that, The deep supervised training strategy adds an auxiliary segmentation head to the semantic branch of the dual-resolution backbone network based on the model obtained in step S3. Simultaneously, it uses the Canny boundary extraction operator to extract boundary features from the output and adds boundary loss for constraint. The total loss during model training is expressed as:
7. A data processing apparatus, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the real-time semantic segmentation method for autonomous driving as described in any one of claims 1 to 6.
8. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the real-time semantic segmentation method for autonomous driving as described in any one of claims 1 to 6.