Improved Lane Detection Method Based on Ultra-Fast Structure-Aware Deep Network
Through the improved ultra-fast structural perception deep network, the Resnet backbone network and attention mechanism are used to solve the fitting problem of lane line detection in curves and missing situations, and high-precision and fast lane line detection are achieved, suitable for autonomous driving.
Patent Information
- Application Number
- CN202211287914.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-20
AI Technical Summary
The existing lane line detection methods are inaccurate in the fitting of the anchors in curves, and the missing lane line parts are inaccurate, which affects the accuracy of autonomous driving.
The Resnet backbone network is used to extract multi-scale feature maps, and combined with SE and CBAM attention mechanisms, the auxiliary segmentation and classification prediction network is improved, and the pyramid pooling is used to replace the full connection layer to build an improved ultra-fast structural perception deep network.
It improves the accuracy and generalization ability of lane line detection, can identify missing and blocked lane lines, ensures accurate fitting of lane line positions in curves, and improves the recognition speed and accuracy of autonomous driving.
Smart Images

Figure CN115690707B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of driverless technology, and specifically to an improved lane line detection method based on an ultra-fast structure-aware deep network. Background Art
[0002] Lane line detection is one of the functional components of the perception module in driverless technology and plays an important role in the driverless process. Many intelligent detection institutions are conducting extensive research on the Advanced Driver Assistance System (ADAS), such as Baidu's driverless cars, Tesla's Google's self-driving cars, etc. In addition, lane line detection also plays a very important role in road assisted driving and parking assistance systems.
[0003] The research on lane line detection has always received much attention. In the current research results, the methods can be roughly divided into traditional detection methods based on features and models and deep learning methods. However, lane detection still faces many challenges. For example, the lane lines are not clear enough due to occlusion, and the color and shape of the lane lines obtained are affected by weather and light conditions. Under such harsh conditions, it is particularly important to build a network that ensures fast calculation with low parameters and improves the detection accuracy of the network as much as possible.
[0004] The existing technical solution - the lane line detection method of the ultra-fast structure-aware deep learning network: comprehensively updates the original pixel segmentation detection problem, and regards the lane line detection process as a selection problem based on rows using global features. This method has the receptive field of the entire image. Compared with pixel segmentation based on a limited receptive field, it can learn clues and information at different positions to solve the problem of no clues. Based on the proposed loss function formula, this method proposes a structural loss and utilizes the prior features of the lane.
[0005] This method uses feature fusion as an auxiliary segmentation network, but the main network based on predicted row classification does not use multi-scale feature fusion. In actual tests, the row anchors classified by the network during driving and turning cannot accurately fit the position of the end lane line, and there is also a situation where the row anchors fit to the wrong lane line. Therefore, the present invention is an improvement based on this method, aiming to reduce the fitting error of the row anchors in actual tests. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] In view of the deficiencies of the prior art, the present invention provides an improved lane line detection method based on an ultra-fast structure-aware deep network. The established improved lane line detection model has the advantages of fast speed, high precision, and high generalization ability. It can identify missing, occluded, and unmarked lane lines in the driving road to achieve the integrity of lane line detection, and has the advantages of high-speed recognition and applicability to autonomous driving. It solves the problem that in actual tests, the row anchors classified by the network learning during the driving turn cannot accurately fit the position of the end lane line, and there is also the problem that the row anchors are fitted to the wrong lane line.
[0008] (2) Technical solution
[0009] To achieve the advantages of fast speed, high precision, and high generalization ability of the established improved lane line detection model, and to be able to identify missing, occluded, and unmarked lane lines in the driving road to achieve the integrity of lane line detection, and high-speed recognition for the purpose of autonomous driving, the present invention provides the following technical solution:
[0010] An improved lane line detection method based on an ultra-fast structure-aware deep network, comprising the following steps:
[0011] S1. Build a backbone network
[0012] S1.1. Use Resnet as the backbone network to obtain feature maps of different dimensions of the input image as a feature extractor. Improve the depth of the overall model network, expand the search space of the model, and fully extract the information of the entire image;
[0013] S1.2. Unify the sizes of the multi-scale feature maps extracted by Resnet-18. The basic feature map selected in this article is the second-layer output X2. After convolution processing of the Resnet third-layer output X3, perform upsampling to double the resolution, and after convolution processing of the Resnet fourth-layer output X4, perform upsampling to quadruple the resolution;
[0014] S2. Build an improved auxiliary segmentation network
[0015] S2.1. First perform global information embedding on the spliced multi-scale feature maps. Use global average pooling to generate channel statistics to compress the spatial information of the fused multi-scale feature maps (Squeeze) into a channel descriptor. Computationally, the statistical result Z is generated by contracting U (the fused multi-scale feature maps) through its spatial dimensions H×W;
[0016] Use the information aggregated by the Squeeze operation to perform the Excitation operation to capture the channel dependencies;
[0017] The fusion result with the channel attention mechanism is fed into the auxiliary segmentation network model for processing to calculate the auxiliary loss value;
[0018] S2.2. Use cross - entropy as the loss function for auxiliary segmentation. The auxiliary network serves as an auxiliary branch of the entire model to improve the network learning accuracy. It does not participate in the test calculation to reduce the computational complexity and improve the running speed;
[0019] S3. Build an improved row - based classification prediction network
[0020] S3.1. Fuse the multi - scale feature maps with unified channel numbers through the CBAM attention mechanism, and feed the result into the row - based classification network;
[0021] S3.2. The fusion result uses pyramid pooling with three different sizes of pooling operations, 1x1, 2x2, and 3x3, to obtain feature maps of multiple sizes. Then, perform "1x1 Conv" on these feature maps of different sizes again to reduce the number of channels, and then perform bilinear interpolation for upsampling to obtain feature maps of the same size, and splice them on the channels;
[0022] S3.3. Process the input fused feature maps through full connection. Next, reshape the fully - connected layer and then perform network training to model the positional relationship of lane points, and construct a loss function according to the structure of the lane lines for learning;
[0023] S4. Train the improved ultra - fast structure - aware deep network
[0024] S4.1. Train the organized Tusimple dataset through the improved deep network model.
[0025] Preferably, in step S1.1, Resnet - 18 is used as the backbone network. The number of channels output by the four - layer network of Resnet - 18, and at the same time, the number of feature maps. There are N convolutional kernels forming N feature maps, which are 64, 128, 256, and 512 respectively. The resolution of the second - layer output X2 is 1 / 2 of the original scale map, the resolution of the third - layer output X3 is 1 / 4 of the original scale map, and the resolution of the fourth - layer output X4 is 1 / 8 of the original scale map.
[0026] Preferably, in step S1.2, the number of channels of the X2, X3, and X4 layers are each 128. Finally, the feature maps with unified scales are feature - stitched to form a feature block, and the total number of channels of the stitched feature block is 384.
[0027] Preferably, in step S2.1, each convolutional kernel learns only a local receptive field for each channel. After the Squeeze operation, each unit of the output U can utilize the context feature information.
[0028] Preferably, in the step S2.1, the c-th element of Z is calculated as follows:
[0029]
[0030] Preferably, in the step S2.1, by using a simple gating mechanism and sigmoid activation function:
[0031]
[0032] Restrict the complexity of the model, form a bottleneck containing two fully connected layers (FC) around the non-linearity to parameterize the gating mechanism, one Relu, and one layer for increasing the dimension to return to the channel dimension of the transformed output U.
[0033] Preferably, in the above formula, where δ is the Relu function, W1 ∈ R C / r×C and W2 ∈ R C×C / r .
[0034] Preferably, in the step S2.1, the final output of the block is obtained by rescaling U through the activation function s:
[0035]
[0036] Preferably, in the step S3.2, pyramid pooling has obvious advantages compared with the original global average pooling GAP. It can generate feature maps of different scales, finally flatten and splice them, and then pass them into the fully connected layer for classification. Using pyramid pooling can utilize the contextual relationship between different sub-regions and improve the accuracy of the model.
[0037] (III) Beneficial Effects
[0038] Compared with the prior art, the present invention provides an improved lane line detection method based on an ultra-fast structure-aware deep network, having the following beneficial effects:
[0039] 1. The improved lane line detection method based on the ultra-fast structure-aware deep network solves the problem of inaccurate row anchor fitting in the case of curves in the actual test of lane line recognition in the prior art solution. While ensuring a high-speed recognition effect, it improves the training of the network for multi-scale features of images, and at the same time adds two attention mechanisms, SE and CBAM, to improve the attention of the model to key features in the feature map during the training process and enhance the network recognition performance.
[0040] 2. The improved lane line detection method based on the ultra-fast structure-aware deep network solves the problem of inaccurate fitting of the missing part of the lane line in the original dataset picture. In the line classification network, the original fully connected layer is improved to a pyramid pooling fully connected layer, which can use pooling of different sizes to increase the receptive field and add multi-scale fusion based on the attention mechanism to enhance the depth of the network and further learn the global features, effectively improving the original inaccurate fitting effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Flowchart of the present invention;
[0042] Figure 2 Overall architecture diagram of the present invention;
[0043] Figure 3 CBAM attention module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Please refer to Figures 1-3 , the improved lane line detection method based on the ultra-fast structure-aware deep network includes the following steps:
[0046] S1. Build a backbone network
[0047] S1.1. Use Resnet as the backbone network to obtain feature maps of different dimensions of the input picture as a feature extractor, enhancing the depth of the overall model network, expanding the search space of the model, and fully extracting the information of the entire picture. Resnet-18 is used as the backbone network. The number of output channels of the four-layer network of Resnet-18, and at the same time the number of feature maps. There are N convolutional kernels forming N feature maps, which are 64, 128, 256, and 512 respectively. The resolution of the second-layer output X2 is 1 / 2 of the original scale map, the resolution of the third-layer output X3 is 1 / 4 of the original scale map, and the resolution of the fourth-layer output X4 is 1 / 8 of the original scale map.
[0048] S1.2. Unify the sizes of the multi-scale feature maps extracted by Resnet-18. The base feature map selected in this paper is the output X2 of the second layer. After performing convolutional processing on the output X3 of the third layer of Resnet and then performing upsampling to double the resolution, and after performing convolutional processing on the output X4 of the fourth layer of Resnet and then performing upsampling to quadruple the resolution; respectively take the number of channels of the X2, X3, and X4 layers as 128. Finally, perform feature concatenation on the feature maps with unified scales to form a feature block, and the total number of channels of the concatenated feature block is 384.
[0049] S2. Build an improved auxiliary segmentation network
[0050] S2.1. First perform global information embedding on the concatenated multi-scale feature maps. Use global average pooling to generate channel statistics to compress the spatial information of the fused multi-scale feature maps (Squeeze) into a channel descriptor. Each convolutional kernel learns only a local receptive field for each channel. After the Squeeze operation, each unit of the output U can utilize the context feature information. Computationally, the statistical result Z is generated by contracting U (the fused multi-scale feature maps) through its spatial dimensions H×W; the c-th element of Z is calculated by the following method:
[0051]
[0052] Utilize the information aggregated by the Squeeze operation to perform the Excitation operation to capture the channel dependencies; by using a simple gating mechanism and the sigmoid activation function:
[0053]
[0054] where δ is the Relu function, W1 ∈ R C / r×C and W2 ∈ R C×C / r . Limit the complexity of the model and form a bottleneck containing two fully connected layers (FC) around the non-linearity to parameterize the gating mechanism. One Relu and one layer for dimension increase are used to return to the channel dimension of the transformed output U. The final output of the block is obtained by rescaling U through the activation function s:
[0055]
[0056] The fusion result with the channel attention mechanism is passed into the auxiliary segmentation network model for processing to calculate the auxiliary loss value;
[0057] S2.2. Use cross-entropy as the loss function for auxiliary segmentation. The auxiliary network serves as an auxiliary branch of the entire model to improve the learning accuracy of the network. It does not participate in the test calculation to reduce the computational complexity and improve the running speed;
[0058] S3. Build an improved row-based classification prediction network
[0059] S3.1. Fuse the multi-scale feature maps with unified number of channels through the CBAM attention mechanism, and input the result into the row-based classification network;
[0060] S3.2. Use pyramid pooling for the fusion result to obtain feature maps of multiple sizes by performing pooling operations of three different sizes, 1x1, 2x2, and 3x3, then perform "1x1 Conv" again on these feature maps of different sizes to reduce the number of channels, and then perform bilinear interpolation for upsampling to obtain feature maps of the same size, and splice them on the channels; Pyramid pooling has obvious advantages compared with the original global average pooling GAP. It can generate feature maps of different scales, finally flatten and splice them, and then input them into the fully connected layer for classification. Using pyramid pooling can utilize the contextual relationship between different sub-regions and improve the model accuracy.
[0061] S3.3. Fully connect the input fusion feature map after processing, and then reshape the fully connected layer and perform network training to model the positional relationship of lane points, and construct a loss function according to the structure of the lane line for learning;
[0062] S4. Train the improved ultra-fast structure-aware deep network
[0063] S4.1. Train the sorted Tusimple dataset through the improved deep network model.
[0064] This method improves the auxiliary segmentation network that uses multi-scale feature fusion in the existing solution into a feature fusion network based on the SE channel attention mechanism.
[0065] Improve the input feature map of the row classification network in the existing solution into a multi-scale feature fusion map based on the CBAM attention mechanism.
[0066] Add pyramid pooling to the row classification network in the existing solution to further expand the receptive field and utilize the global feature information after feature fusion, and utilize the prior information of the global scene
[0067] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An improved lane line detection method based on an ultra-fast structure-aware deep network, characterized in that, It includes the following steps: S1. Build a backbone network S1.
1. Use Resnet as the backbone network to obtain feature maps of different dimensions of the input image, serving as a feature extractor to increase the depth of the overall model network, expand the search space of the model, and fully extract the information of the entire image; S1.
2. Unify the sizes of the multi-scale feature maps extracted by Resnet-18. The base feature map selected in this paper is the output X2 of the second layer. After convolution processing of the output X3 of the third layer of Resnet, perform upsampling to double the resolution. After convolution processing of the output X4 of the fourth layer of Resnet, perform upsampling to quadruple the resolution; S2. Build an improved auxiliary segmentation network S2.
1. First, perform global information embedding on the concatenated multi-scale feature maps. Use global average pooling to generate channel statistics and squeeze the spatial information of the fused multi-scale feature maps into a channel descriptor. Computationally, the statistical result Z is generated by contracting U through its spatial dimensions H×W, where U is the fused multi-scale feature map; Use the information aggregated by the Squeeze operation to perform an Excitation operation to capture channel dependencies; The fused result with the channel attention mechanism is passed into the auxiliary segmentation network model for processing to calculate the auxiliary loss value; S2.
2. Use cross-entropy as the loss function for auxiliary segmentation. The auxiliary network serves as an auxiliary branch of the entire model to improve the learning accuracy of the network, does not participate in the test calculation to reduce the computational complexity, and improve the running speed; S3. Build an improved row-based classification prediction network S3.
1. Fuse the multi-scale feature maps with unified channel numbers through the CBAM attention mechanism, and the result is passed into the row-based classification network; S3.
2. The fused result uses pyramid pooling to obtain feature maps of multiple sizes by using three different sizes of pooling operations of 1x1, 2x2, and 3x3, and then perform 1x1 Conv on these feature maps of different sizes to reduce the number of channels, and then perform bilinear interpolation for upsampling to obtain feature maps of the same size, and concatenate them on the channel; S3.
3. After processing the passed-in fused feature maps, perform full connection. Next, reshape the fully connected layer and then perform network training to model the positional relationship of lane points, and construct a loss function according to the structure of the lane line for learning; S4. Train the improved ultra-fast structure-aware deep network S4.
1. Train the sorted Tusimple dataset through the improved deep network model.
2. The improved lane line detection method based on an ultra-fast structure-aware deep network according to claim 1, wherein In the step S1.1, Resnet-18 is used as the backbone network. The number of channels output by the four-layer network of Resnet-18, and at the same time the number of feature maps. There are N convolutional kernels forming N feature maps, which are 64, 128, 256, and 512 respectively. The resolution of the second layer output X2 is 1 / 2 of the original scale map, the resolution of the third layer output X3 is 1 / 4 of the original scale map, and the resolution of the fourth layer output X4 is 1 / 8 of the original scale map.
3. The improved lane line detection method based on the ultra-fast structure-aware deep network according to claim 1, wherein In the step S1.2, the number of channels of the X2, X3, and X4 layers is 128 respectively. Finally, the feature maps with unified scales are subjected to feature splicing to form a feature block, and the total number of channels of the feature block after splicing is 384.
4. The improved lane line detection method based on an ultra-fast structure-aware deep network according to claim 1, wherein In the step S2.1, each channel learned by each convolutional kernel has only a local receptive field. After the Squeeze operation, each unit of the output U can utilize the feature information of the context.
Citation Information
Patent Citations
Real-time lane line detection method and system for learning context information by adopting attention mechanism
CN112241728A
Semantic segmentation method and system based on low-illumination complex road scene
CN113902915A