A lane detection method based on deep learning
By adopting deep learning methods in lane line detection, the network structure is simplified, the calculation amount and parameter amount are reduced, and the enhanced receptive field module and CBAM module are introduced, the existing lane line detection methods are solved, and efficient and real-time lane line detection is achieved.
Patent Information
- Application Number
- CN202210441263.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-04-25
AI Technical Summary
The existing lane line detection methods have complex detection processes, large parameters and calculations, making it difficult to meet the real-time requirements of autonomous driving.
The lane line detection method based on deep learning is adopted to simplify the network structure through data augmentation and feature fusion modules, reduce the amount of calculation and parameter, and introduce enhanced receptive field modules and CBAM modules into the encoding and decoding networks to improve detection speed and accuracy.
While ensuring accuracy, the speed of lane line detection is significantly improved, meeting the real-time requirements of autonomous driving.
Smart Images

Figure CN114913493B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle automatic driving, and specifically relates to a lane line detection method based on deep learning. Background Art
[0002] Autonomous driving technology has been a hot topic in recent years. With the rapid development of automotive industry technology and artificial intelligence, autonomous driving technology has gradually moved from science fiction to reality. The main research contents of autonomous driving technology are: environmental perception, positioning navigation, path planning, and motion control; among them, environmental perception is the use of multiple sensors to detect and process the road traffic environment, helping autonomous driving vehicles understand the surrounding environment information and provide traffic environment information for control algorithms; lane line detection is an important part of environmental perception. The vehicle obtains road images through the camera and detects the lane line information of the current road to complete a series of assisted driving behaviors of the vehicle, including lane keeping, adaptive cruise and other auxiliary functions.
[0003] The lane detection method based on deep learning relies on big data. The model learns autonomously to obtain the characteristics of the lane, clusters them using a clustering algorithm, and finally fits the lane using a polynomial. It can achieve good accuracy in most scenarios on the road, and the algorithm is robust. However, the detection process of the above method is relatively complicated, and the number of parameters and calculations are large. At the same time, the requirements for computer hardware are high, which makes it difficult to meet the real-time requirements of autonomous driving. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a lane line detection method based on deep learning, so as to solve the problems that the existing lane line detection method has a relatively complex detection process, a large number of parameters and a large amount of calculation, and at the same time has high requirements on computer hardware, and it is difficult to meet the real-time requirements of autonomous driving; the method of the present invention improves the detection speed while ensuring the accuracy, thereby meeting the real-time requirements.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A lane line detection method based on deep learning of the present invention comprises the following steps:
[0007] Step S1: Obtain Tusimple image dataset;
[0008] Step S2: perform data enhancement on the lane line images in the Tusimple image dataset, and adjust the resolution of the enhanced lane line images to 512×256 (width×height), and use the adjusted images as the training dataset for the lane line detection neural network model;
[0009] Step S3: Build a lane detection neural network model, determine the loss function, and use the training data set in step S2 to train the lane detection neural network model until convergence to obtain the optimal model;
[0010] Step S4: Load the best model parameters, input the road image into the best model, and obtain the point sets determined as different lane lines;
[0011] Step S5: Use quadratic polynomials to fit lane lines of different categories, and overlay the fitted lane lines on the original image to achieve visualization of lane line detection.
[0012] Furthermore, the data enhancement in step S2 includes: rotation, horizontal flipping, random cropping and adding Gaussian white noise.
[0013] Furthermore, the neural network model consists of an encoding network, a decoding network, an enhanced receptive field module, a CBAM module and two feature fusion modules; the encoding network includes a preprocessing module and five residual layers connected in sequence, and the decoding network includes three convolutional upsampling modules and an output module connected in sequence.
[0014] Furthermore, the preprocessing module includes: a convolution layer with a convolution kernel size of 7×7, a step size of 2, and a padding of 3, and a maximum pooling layer with a kernel of 3×3, a step size of 1, and a padding of 1. The input image resolution of the preprocessing module is 512×256 (width×height), and the width and height of the output image are halved.
[0015] Furthermore, each residual layer is composed of two residual blocks, and each residual block is composed of two branches. The first branch contains two depth-wise separable convolutions with a convolution kernel size of 3×3; the second branch is a convolution layer with a convolution kernel size of 1×1, and the second branch is used to ensure that the resolution and dimension of the input feature map and the output feature map are the same; the second residual layer and the third residual layer add a channel attention mechanism, and the fourth residual layer and the fifth residual layer introduce a hole convolution, and the expansion rates are 2 and 4 respectively; in the above encoding network process, the feature maps (out1, out2 and out5) obtained by outputting the first residual layer, the second residual layer and the fifth residual layer are output, and the feature map (out5) output by the fifth residual layer is enhanced by After the wild module, the feature map with attention weight is obtained by the CBAM module and then enters the decoding network; the feature map (out5) output by the fifth residual layer passes through the first convolution upsampling module and then passes through the first feature fusion module, and the first feature fusion module is also connected to the feature map output by the second residual layer. The end of the first feature fusion module is connected to the second convolution upsampling module, and the feature map passing through the second convolution upsampling module passes through the second feature fusion module, and the second feature fusion module is also connected to the feature map output by the first residual layer. The end of the first feature fusion module is connected to the third convolution upsampling module, and the feature map passing through the third convolution upsampling module passes through the output module, and finally a feature map with six channels is obtained.
[0016] Furthermore, the enhanced receptive field module consists of four parallel branches, the first branch is a 1×1 convolution, which is equivalent to the residual structure in the residual network, the second branch is a 3×3 convolution with a dilation rate of 3, the third branch is two 3×3 convolutions, and the dilation rates are 3 and 6 respectively, and the fourth branch is the global maximum pooling. The results of the second branch and the third branch are fused through a 1×1 convolution, and then fused with the first channel and the fourth channel. There is a 1×1 convolution at the input and output of the enhanced receptive field module, which is used to reduce and restore the number of channels, reduce the amount of calculation in the four branches, and speed up the network operation speed.
[0017] Furthermore, the CBAM module includes a channel attention and a spatial attention. The input is multiplied by itself to obtain a new feature map after generating the input weight through the channel attention. The weight of the new feature map is then multiplied by itself to obtain the output through the spatial attention. The output result enters the first convolution upsampling module in the decoding network.
[0018] Furthermore, the feature fusion module includes two inputs, the first input comes from the decoding network, and the second input comes from the encoding network; the first input is calculated through spatial attention to obtain a weight with an attention mechanism, and the second input is multiplied by the weight to obtain a new feature map, which is then fused with the initial first input, that is, concatenated in the channel dimension, and the result is continuously output to the decoding network to enter the convolution upsampling module.
[0019] Furthermore, the three convolution upsampling modules each include, in sequence, a 1×1 ordinary convolution, an upsampling, and a depth-separable convolution with a convolution kernel size of 3×3.
[0020] Furthermore, the output module includes a 1×1 ordinary convolution and a depthwise separable convolution with a convolution kernel size of 3×3 and 6 output channels; the depthwise separable convolution operation is followed by Batch Normalization and ReLU nonlinear activation function processing.
[0021] Furthermore, the loss functions in step S3 are cross entropy loss function and OHEM loss function. The network used in the lane line detection neural network model is a multi-classification semantic segmentation network, including background and five lane lines. The cross entropy loss function is used for training, and the maximum number of iterations is set to 100, the initial learning rate is 1e-2, and the learning rate adjustment strategy is an exponential decay adjustment strategy. The training is stopped after completing 100 training times, and the OHEM loss function is used to continue training. The samples are arranged according to the size of the cross entropy loss, and the samples with large losses are screened out, and the loss is used for back propagation calculation. The maximum number of iterations is set to 100, the initial learning rate is 1e-4, and the training is stopped when the loss value reaches a stable value.
[0022] Furthermore, the road image in step S4 contains lane lines and the number of lane lines does not exceed five. After the road image is input into the optimal model, the output obtained is a feature map of six channels, that is, each pixel in the feature map corresponds to six categories (background and five lane lines). After softmax is performed on the feature map, a lane line pixel probability map is obtained, and the prediction points (x, y) corresponding to each lane line are searched to form a point set [(x1, y1), (x2, y2), ... (x n ,y n )], where (x i ,y i ), i = 1, 2, ... n represents the pixel points that are divided into a lane line.
[0023] Furthermore, in step S5, the quadratic polynomial y=a1x 2+a2x+b, the point set [(x1,y1),(x2,y2),……(x n ,y n )] for fitting, where a1, a2, and b are the parameters to be solved, a1 is the coefficient of the quadratic term in the quadratic polynomial, a2 is the coefficient of the linear term in the quadratic polynomial, and b is the constant term in the quadratic polynomial. After solving the coefficients a1, a2, and b using the point set, the quadratic polynomial y=a1x is fitted on the input image. 2 +a2x+b draws lane lines, and selects different colors to draw different lane lines to achieve visualization of lane line detection.
[0024] Beneficial effects of the present invention:
[0025] The method of the present invention proposes a network structure, including an encoding network and a decoding network. A residual structure is used in the encoding network and the 3×3 convolution therein is replaced by a depth-separable convolution. The structure is simple and the amount of calculation and the amount of parameters are greatly reduced. At the same time, an enhanced receptive field module and a CBAM module are used in the connection process of the encoding network and the decoding network. While reducing the amount of calculation in each branch, the function of extracting different receptive fields is taken into account, and the network is made to pay more attention to channels and regions containing useful information. In the upsampling process of the decoding network, the feature maps output in the decoding network are fused to obtain more complete information. While giving full play to the advantages of deep learning, the speed of lane line detection is greatly improved, and the requirements of accuracy and real-time performance of autonomous driving are met. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flow chart of the method of the present invention;
[0027] Figure 2 It is a network model architecture diagram in the present invention;
[0028] Figure 3 It is a network structure diagram of the preprocessing module in the present invention;
[0029] Figure 4 It is a network structure diagram of each residual layer in the present invention;
[0030] Figure 5 It is a residual block network structure diagram in the present invention;
[0031] Figure 6 It is a network structure diagram of the enhanced receptive field module in the present invention;
[0032] Figure 7 It is the network structure diagram of the CBAM module in the present invention;
[0033] Figure 8 It is a network structure diagram of the feature fusion module in the present invention;
[0034] Fig. 9 It is a network structure diagram of the convolution upsampling module in the present invention;
[0035] Fig.10 It is a network structure diagram of the output module in the present invention. DETAILED DESCRIPTION
[0036] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.
[0037] Reference Figure 1 As shown, a lane line detection method based on deep learning of the present invention comprises the following steps:
[0038] Step S1: Obtain the Tusimple image dataset from the Internet;
[0039] Step S2: perform data enhancement on the lane line images in the Tusimple image dataset, and adjust the resolution of the enhanced lane line images to 512×256 (width×height), and use the adjusted images as the training dataset for the lane line detection neural network model;
[0040] The data enhancement in step S2 includes: rotation, horizontal flipping, random cropping and adding Gaussian white noise.
[0041] Step S3: Build a lane detection neural network model, determine the loss function, and use the training data set in step S2 to train the lane detection neural network model until convergence to obtain the optimal model;
[0042] like Figure 2 As shown, the neural network model consists of an encoding network, a decoding network, an enhanced receptive field module, a CBAM module and two feature fusion modules; the encoding network includes a preprocessing module and five residual layers connected in sequence, and the decoding network includes three convolutional upsampling modules and an output module connected in sequence.
[0043] like Figure 3 As shown, the preprocessing module includes: a convolution layer with a convolution kernel size of 7×7, a step size of 2, and a padding of 3, and a maximum pooling layer with a kernel of 3×3, a step size of 1, and a padding of 1. The input image resolution of the preprocessing module is 512×256 (width×height), and the width and height of the output image are halved.
[0044] like Figure 4-Figure 5As shown in the figure, each residual layer consists of two residual blocks, and each residual block consists of two branches. The first branch contains two depth-separable convolutions with a convolution kernel size of 3×3; the second branch is a convolution layer with a convolution kernel size of 1×1, and the second branch is used to ensure that the resolution and dimension of the input feature map and the output feature map are the same; the second residual layer and the third residual layer add a channel attention mechanism, and the fourth residual layer and the fifth residual layer introduce a hole convolution, and the expansion rates are 2 and 4 respectively; in the above encoding network process, the feature maps (out1, out2 and out5) obtained by outputting the first residual layer, the second residual layer and the fifth residual layer are output, and the feature map (out5) output by the fifth residual layer is enhanced by the receptive field After the module, the feature map with attention weights is obtained through the CBAM module and then enters the decoding network; the feature map (out5) output by the fifth residual layer passes through the first convolution upsampling module and then passes through the first feature fusion module, and the first feature fusion module is also connected to the feature map output by the second residual layer. The end of the first feature fusion module is connected to the second convolution upsampling module, and the feature map passing through the second convolution upsampling module passes through the second feature fusion module, and the second feature fusion module is also connected to the feature map output by the first residual layer. The end of the first feature fusion module is connected to the third convolution upsampling module, and the feature map passing through the third convolution upsampling module passes through the output module, and finally a feature map with six channels is obtained.
[0045] like Figure 6 As shown in the figure, the enhanced receptive field module consists of four parallel branches, the first branch is a 1×1 convolution, which is equivalent to the residual structure in the residual network, the second branch is a 3×3 convolution with a dilation rate of 3, the third branch is two 3×3 convolutions, and the dilation rates are 3 and 6 respectively, and the fourth branch is the global maximum pooling. The results of the second branch and the third branch are fused through a 1×1 convolution, and then fused with the first channel and the fourth channel. There is a 1×1 convolution at the input and output of the enhanced receptive field module, which is used to reduce and restore the number of channels, reduce the amount of calculation in the four branches, and speed up the network operation.
[0046] like Figure 7 As shown in the figure, the CBAM module includes a channel attention and a spatial attention. The input is multiplied by itself after the input weight is generated by the channel attention to obtain a new feature map, and then the weight of the new feature map is multiplied by itself after the spatial attention to obtain the output. The output result enters the first convolution upsampling module in the decoding network.
[0047] like Figure 8As shown, the feature fusion module includes two inputs, the first input comes from the decoding network, and the second input comes from the encoding network; the first input is calculated through spatial attention to obtain a weight with an attention mechanism, and the second input is multiplied by the weight to obtain a new feature map, which is then fused with the initial first input 1, that is, concatenated in the channel dimension, and the result is continuously output to the decoding network to enter the convolution upsampling module.
[0048] like Fig. 9 As shown, the three convolution upsampling modules each include a 1×1 ordinary convolution, an upsampling, and a depth-separable convolution with a convolution kernel size of 3×3.
[0049] like Fig.10 As shown, the output module includes a 1×1 ordinary convolution and a depthwise separable convolution with a convolution kernel size of 3×3 and 6 output channels; the depthwise separable convolution operation is followed by BatchNormalization normalization and ReLU nonlinear activation function processing.
[0050] Among them, the loss functions in the step S3 are the cross entropy loss function and the OHEM loss function. The network used in the lane line detection neural network model is a multi-classification semantic segmentation network, including background and five lane lines. The cross entropy loss function is used for training, and the maximum number of iterations is set to 100, the initial learning rate is 1e-2, and the learning rate adjustment strategy is an exponential decay adjustment strategy. The training is stopped after completing 100 training times, and the OHEM loss function is used to continue training. The samples are arranged according to the size of the cross entropy loss, and the samples with large losses are screened out. The loss is used for back propagation calculation, and the maximum number of iterations is set to 100, the initial learning rate is 1e-4, and the training is stopped when the loss value reaches a stable value.
[0051] Step S4: Load the best model parameters, input the road image into the best model, and obtain the point sets determined as different lane lines;
[0052] The road image contains lane lines and the number of lane lines does not exceed five. After the road image is input into the optimal model, the output obtained is a feature map of six channels, that is, each pixel in the feature map corresponds to six categories (background and five lane lines). After softmax is performed on the feature map, a lane line pixel probability map is obtained, and the prediction points (x, y) corresponding to each lane line are searched to form a point set [(x1, y1), (x2, y2), ... (x n ,y n )], where (x i ,y i ), i = 1, 2, ... n represents the pixel points that are divided into a lane line.
[0053] Step S5: fitting lane lines of different categories using quadratic polynomials, and superimposing the fitted lane lines on the original image to achieve visualization of lane line detection;
[0054] The quadratic polynomial y=a1x is used 2 +a2x+b, the point set [(x1,y1),(x2,y2),……(x n ,y n )] for fitting, where a1, a2, and b are the parameters to be solved, a1 is the coefficient of the quadratic term in the quadratic polynomial, a2 is the coefficient of the linear term in the quadratic polynomial, and b is the constant term in the quadratic polynomial. After solving the coefficients a1, a2, and b using the point set, the quadratic polynomial y=a1x is fitted on the input image. 2 +a2x+b draws lane lines, and selects different colors to draw different lane lines to achieve visualization of lane line detection.
[0055] The present invention has many specific application paths. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements can be made without departing from the principle of the present invention. These improvements should also be regarded as the protection scope of the present invention.
Claims
1. A lane line detection method based on deep learning, characterized in that: Here are the steps: Step S1: Obtain Tusimple image dataset; Step S2: perform data enhancement on the lane line images in the Tusimple image dataset, adjust the resolution of the enhanced lane line images to 512×256, and use the adjusted images as the training dataset for the lane line detection neural network model; Step S3: Build a lane detection neural network model, determine the loss function, and use the training data set in step S2 to train the lane detection neural network model until convergence to obtain the optimal model; Step S4: Load the best model parameters, input the road image into the best model, and obtain the point sets determined as different lane lines; Step S5: fitting lane lines of different categories using quadratic polynomials, and superimposing the fitted lane lines on the original image to achieve visualization of lane line detection; The neural network model consists of an encoding network, a decoding network, an enhanced receptive field module, a CBAM module and two feature fusion modules; the encoding network includes a preprocessing module and five residual layers connected in sequence, and the decoding network includes three convolution upsampling modules and an output module connected in sequence; Each residual layer consists of two residual blocks, and each residual block consists of two branches. The first branch contains two depth-wise separable convolutions with a convolution kernel size of 3×3; the second branch is a convolution layer with a convolution kernel size of 1×1, and the second branch is used to ensure that the resolution and dimension of the input feature map and the output feature map are the same; the second residual layer and the third residual layer add a channel attention mechanism, and the fourth residual layer and the fifth residual layer introduce a hole convolution, and the expansion rates are 2 and 4 respectively; in the above encoding network process, the feature maps obtained by the first residual layer, the second residual layer and the fifth residual layer are output, and the feature map output by the fifth residual layer passes through the enhanced receptive field module and then through the CBAM module to obtain a feature map with attention weights and then enter the decoding network; The feature map output by the fifth residual layer passes through the first convolution upsampling module and then passes through the first feature fusion module. The first feature fusion module is also connected to the feature map output by the second residual layer. The end of the first feature fusion module is connected to the second convolution upsampling module. The feature map passing through the second convolution upsampling module passes through the second feature fusion module. The second feature fusion module is also connected to the feature map output by the first residual layer. The end of the first feature fusion module is connected to the third convolution upsampling module. The feature map passing through the third convolution upsampling module passes through the output module, and finally a feature map with six channels is obtained.
2. The lane line detection method based on deep learning according to claim 1, characterized in that: The data enhancement in step S2 includes: rotation, horizontal flipping, random cropping and adding Gaussian white noise.
3. The lane line detection method based on deep learning according to claim 1, characterized in that: The preprocessing module includes: a convolution layer with a convolution kernel size of 7×7, a step size of 2, and a padding of 3, and a maximum pooling layer with a kernel of 3×3, a step size of 1, and a padding of 1. The input image resolution of the preprocessing module is 512×256, and the width and height of the output image are halved.
4. The lane line detection method based on deep learning according to claim 1, characterized in that: The enhanced receptive field module consists of four parallel branches. The first branch is a 1×1 convolution, which is equivalent to the residual structure in the residual network; the second branch is a 3×3 convolution with a dilation rate of 3; the third branch is two 3×3 convolutions with dilation rates of 3 and 6 respectively; the fourth branch is a global maximum pooling, which combines the results of the second and third branches through a 1×1 convolution and then merges them with the first and fourth channels; There is a 1×1 convolution at the input and output of the enhanced receptive field module, which is used to reduce and restore the number of channels, reduce the amount of calculation in the four branches, and speed up the network operation.
5. The lane line detection method based on deep learning according to claim 1, characterized in that: The CBAM module includes a channel attention and a spatial attention. The input is multiplied by itself after the input weight is generated by the channel attention to obtain a new feature map, and then the weight of the new feature map is multiplied by itself after the spatial attention to obtain the output. The output result enters the first convolution upsampling module in the decoding network.
6. The lane line detection method based on deep learning according to claim 5, characterized in that: The feature fusion module includes two inputs, the first input comes from the decoding network, and the second input comes from the encoding network; the first input is calculated through spatial attention to obtain a weight with an attention mechanism, and the second input is multiplied by the weight to obtain a new feature map, which is then fused with the initial first input, that is, spliced in the channel dimension, and the result is continuously output to the decoding network to enter the convolution upsampling module.
7. The lane line detection method based on deep learning according to claim 1, characterized in that: The three convolution upsampling modules each include, in sequence, a 1×1 normal convolution, an upsampling, and a depthwise separable convolution with a convolution kernel size of 3×3.
8. The lane line detection method based on deep learning according to claim 1, characterized in that: The output module includes a 1×1 ordinary convolution and a depthwise separable convolution with a convolution kernel size of 3×3 and 6 output channels; the depthwise separable convolution operation is followed by Batch Normalization and ReLU nonlinear activation function processing.
Citation Information
Patent Citations
Lane line detection method and system based on Lane SegNet
CN113343778A