Lane Line Detection Method and Computer System Based on Deep Learning Algorithm
Through the U-shaped structure neural network combined with the global information fusion module of CNN and Transformer, the problem of insufficient real-time and accuracy in lane line detection is solved, and efficient lane line detection is achieved.
Patent Information
- Application Number
- CN202210280846.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-21
AI Technical Summary
Existing deep learning algorithms have problems in lane line detection that are difficult to ensure real-time and have low prediction accuracy, especially in severe occlusion, and the existing network structure fails to effectively integrate high-level semantic information.
U-shaped neural network is adopted, combined with the global information fusion module of CNN and Transformer, and high-level semantic information is obtained through the downsampling module, and information fusion is fusion using the global information fusion module. The image size is restored through the upsampling module, reducing network parameters to improve detection speed.
It improves the accuracy and real-time nature of lane line detection, can effectively deal with severe occlusion situations, and achieves high-precision lane line detection.
Smart Images

Figure CN114677560B_ABST
Abstract
Description
Technical Field
[0001] The present invention specifically relates to a lane line detection method and a computer system based on a deep learning algorithm. Background Art
[0002] According to the classification of SAE (Society of Automotive Engineers), autonomous driving can be divided into six stages from L0 to L5. Currently, it is generally considered that L3 is the watershed between assisted driving and autonomous driving. Lane line detection plays an important role in both assisted driving and autonomous driving, mainly including: high-precision map generation, lane keeping during driving, automatic cruise, and overtaking decision-making. Lane lines are high-level visual language symbols defined in human society, which stipulate the basic norms for vehicles to drive on the road; the task of lane line detection is to segment and detect the lane lines contained in the 2D images captured by on-vehicle cameras, and certain accuracy and real-time requirements need to be met. The difficulties of lane line detection can be roughly divided into the following three points: 1. Severe occlusion: This occlusion stems from congested traffic; 2. Influence of bad weather: lighting conditions such as rain, haze, or night; 3. Characteristics of lane lines themselves: being slender objects, worn, solid and dashed lines, etc., and the problem of imbalance between positive and negative samples in the entire image caused by these characteristics.
[0003] Lane line detection is a specific problem in a specific field and generally belongs to computer vision tasks. Therefore, its development benefits from the development of computer vision technology in the general environment. In the early days, traditional image algorithms analyzed and processed images by finding the association between pixels and fusing geometric features. Although these algorithms were simple to implement, they often had low accuracy and could only be used in specific situations. In 1997, Bin Yu et al. proposed to use the Hough transform to achieve lane line detection [1], which is also a representative of traditional image algorithms in lane line detection;
[0004] In recent years, with the significant improvement in the computing power of hardware (GPU), deep learning has started to take the stage. In the field of computer vision, methods based on convolutional neural networks (CNNs) have achieved speeds and accuracies comparable to those of human eye recognition in classification, segmentation, and detection tasks, greatly breaking through the bottlenecks of previous traditional algorithms. Lane line detection based on CNNs has also achieved unprecedented real-time performance and accuracy.
[0005] Currently, there is no report on using a deep learning U-shaped network structure and fusing CNN and Transformer to solve the lane line detection task.
[0006] Chinese Patent CN112633177A discloses a lane line detection and segmentation method based on an attention spatial convolutional neural network, which mainly uses a spatial convolutional neural network based on an attention mechanism for lane line segmentation and detection.
[0007] There are some areas for improvement: 1. Real-time performance is difficult to guarantee: The way of transmitting information line by line in this attention-based spatial convolutional network results in a huge number of parameters, which leads to slow prediction of the network in the actual process and also occupies a large amount of computing resources. 2. The prediction accuracy is not high: Simply using this spatial convolution method to aggregate high-level semantic information has poor effects, easily misses a lot of important information, and performs poorly in situations such as severe occlusion where the network needs to imagine most of the occluded objects.
[0008] Compared with other computer vision tasks, the most prominent feature of lane lines is severe occlusion. Therefore, cultivating the network's imagination ability is the main direction for designing the network structure, which requires the network structure to have a powerful high-level semantic information fusion module to obtain global information. Chinese Patent CN110414386B discloses a lane line detection method based on an improved SCNN network. Although it proposes an improved version of the SCNN network, it still fails to grasp the essential problem of paying attention to fusing high-level semantic information for lane lines. Although the effect has been improved, the overall accuracy still needs to be improved. Summary of the Invention
[0009] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a lane line detection method and computer system based on deep learning algorithms.
[0010] To achieve the above purpose, the present invention provides the following technical solutions:
[0011] A lane line detection method based on deep learning algorithms, which includes the following steps:
[0012] 1) Obtain a public dataset, divide it into a training set and a test set, and obtain a text file containing path information through training;
[0013] 2) Design a U-shaped structure neural network, which includes a downsampling module for obtaining a feature map containing high-level semantic information, a global information fusion module for integrating the feature map containing high-level semantic information obtained by the downsampling module, and an upsampling module for restoring the feature map containing high-level semantic information to the size of the input image;
[0014] 3) Train the U-shaped structure neural network and save the network weight parameters with the best performance;
[0015] 4) Select the network weight parameters with the best performance saved in step 3), input the image to be predicted into the current U-shaped structure neural network, and obtain the segmented lane line image.
[0016] In step 1), first generate label image files from the label json coordinate files; then divide the training set and the test set, and generate txt text files with the paths of the training set and the test set. The generated txt text files contain the paths of the original input images to be processed, the paths of the generated label image files, and the number of lane lines included in the current images.
[0017] When generating label image files from the label json coordinate files, it is necessary to specify the width of each lane line, and when reading the coordinates of each lane line, number each lane line in the order from left to right.
[0018] The downsampling module includes four convolutional layers each composed of multiple convolutional blocks.
[0019] The global information fusion module is composed of a CNN-based cyclic accumulation module and a Transformer module. The Transformer module includes position embedding, convolutional mapping, and self-attention mechanism.
[0020] The working process of the global information fusion module is as follows:
[0021] Step 1: Combine the feature map containing high-level semantic information obtained with a position embedding module of the same size, and map it through a convolutional mapping module into three Q matrices, K matrices, and V matrices of the same size;
[0022] Step 2: Perform self-attention mechanism calculations on the Q matrix, K matrix, and V matrix and then input them into the CNN-based cyclic accumulation module;
[0023] Step 3: The CNN-based cyclic accumulation module first divides the input matrix into H rows, and then performs information accumulation and transmission from top to bottom and from bottom to top;
[0024] Step 4: Divide the input matrix into W columns, and then perform information accumulation and transmission from left to right and from right to left;
[0025] Step 5: Further fuse the information from the previous two steps through another layer of self-attention mechanism.
[0026] The upsampling module contains four upsampling layers composed of transposed convolution and bilinear interpolation methods.
[0027] Step 3) includes:
[0028] 1. Determine the training strategy and hyperparameters;
[0029] 2. Perform image augmentation on the input images;
[0030] 3. Set a strategy for saving intermediate weight parameters during training. The strategy for saving intermediate weight parameters is as follows: after each iteration cycle, the test set image is input into the current network to obtain an accuracy index, and the network weight parameter with the highest accuracy in the entire iteration cycle is selected based on the accuracy, and saved.
[0031] In step 1, the training cycle is set. In each cycle, the entire training set of images will be input into the network for training. In each iteration cycle, the image at the current network training prediction and the ground truth label are subjected to loss function and cross entropy loss function minimization calculation to achieve the current network's predicted output close to the label.
[0032] A computer system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the lane line detection method based on a deep learning algorithm when executing the computer program.
[0033] Beneficial effects of the present invention:
[0034] The present invention adopts a U-shaped network structure and adds an information fusion module to the downsampling module and the upsampling module. Due to the addition of the information fusion module, the high-level semantic information of the lane line image is enriched, thereby being able to predict lane lines that are severely occluded.
[0035] The new information fusion module (Cycle-Accumulation Transformer) proposed in the present invention deeply integrates CNN and Transformer, making the network more capable of obtaining global information, which greatly improves the accuracy of the lane line detection task; and in the cycle accumulation module (Cycle-Accumulation), the specific pace fusion method is adopted to further reduce the network parameters, which also ensures the real-time performance of the lane line detection task. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of the workflow of the present invention.
[0037] Figure 2 This is a network structure diagram of the present invention.
[0038] Figure 3 Schematic diagram of the information fusion module.
[0039] Figure 4 This is a prediction comparison chart, where the left side is the original image + label, and the right side is the original image + predicted image. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] The present invention provides a lane line detection method based on a deep learning algorithm, which includes the following steps:
[0042] 1) Obtain a public data set, divide it into a training set and a test set, and obtain a text file containing path information through training;
[0043] Before the public data set is input into the training network, generally, the label json file needs to be converted into a label image file, and then the training set and the test set are divided, and the subsequent training process reads them.
[0044] The specific steps include:
[0045] 1.1: Generate a label image file from the label json coordinate file; when generating a label image file from the label json coordinate file, the width of each lane line needs to be specified. When reading the coordinates of each lane line, each lane line needs to be numbered according to the principle from left to right.
[0046] 1.2: Generate a txt text file with information such as the paths of the training set and the test set; the txt text file includes the path of the original input image to be input, the path of the label image generated in the above step 1.1, and the number of lane lines included in the current image.
[0047] 2) Design a U-shaped structure neural network, which includes a downsampling module (Encoder) for obtaining a feature map containing high-level semantic information, a global information fusion module (Cycle-Accumulation Transformer) for integrating the feature map containing high-level semantic information obtained by the downsampling module, and an upsampling module (Decoder) for restoring the feature map containing high-level semantic information to the input image size, as Figure 2 shown;
[0048] Among them, the downsampling module (Encoder) is used to obtain high-level semantic information and is divided into four convolutional layers. Given an input image, the image passes through three of the convolutional layers, and the length and width are halved respectively, and the dimension becomes larger. At this time, it gradually changes from low-level semantic information to high-level semantic information, and finally obtains a feature map (featuremap) with high-level semantic information.
[0049] The four convolutional layers are composed of multiple convolutional blocks inside, including convolutional operations, pooling operations, normalization operations, etc.
[0050] The downsampling module can be a classic Resnet or a VGG model. And the higher the number of network layers, the stronger its ability to extract high-level semantic information. Usually, the number of network layers is selected according to the computing power.
[0051] The global information fusion module is composed of a Cycle-Accumulation module based on CNN and a Transformer module. The Transformer module includes positional embedding, convolutional projection, and self-attention mechanism. The specific fusion method of the global information fusion module is as Figure 3 shown.
[0052] The global information fusion module is used to integrate the feature map containing high-level semantic information obtained by the downsampling module. First, the feature map is added with a positional embedding module of the same size, and then mapped to three Q matrices, K matrices, and V matrices of the same size through the convolutional projection module. Then, after calculating the self-attention mechanism of the Q matrix, K matrix, and V matrix, they are input into the cycle-accumulation module; the cycle-accumulation module first divides the input matrix into H rows, and then performs information accumulation and transmission from top to bottom and from bottom to top; then the input matrix is divided into W columns, and then information accumulation and transmission from left to right and from right to left are obtained; finally, through another layer of self-attention mechanism module, the information of the previous two steps is further fused.
[0053] The cycle-accumulation module can be transmitted at a certain pace during the transmission process, which will reduce the number of network parameters and make the prediction faster.
[0054] The positional embedding module uses learnable relative position encoding.
[0055] The convolutional projection module uses trainable convolution to operate on the input tensor in a sliding window manner. After passing through the convolutional projection module, three tensors in different parameter spaces, namely the Q matrix, K matrix, and V matrix, will be generated. This operation can save a large amount of computational overhead and retain more spatial information.
[0056] The self-attention mechanism module is used to obtain the correlation inside the data or features. Its calculation depends on the Q matrix, K matrix, and V matrix obtained through the convolutional projection module.
[0057] The upsampling module contains four upsampling layers, which gradually restore the feature map with a size of 1 / 8 containing high-level semantic information to the size of the input image. At this time, the output after passing through the upsampling module already contains the lane lines to be segmented, and the loss function will be calculated with the given labels during the training phase.
[0058] The four upsampling layers are composed of transposed convolution and bilinear interpolation inside, and three of them can restore the size of the feature map by twice.
[0059] 3) Train the U-shaped neural network and save the network weight parameters with the best performance;
[0060] Training the model and saving the best weights mainly can be divided into the following three steps:
[0061] 3.1 During the training process, it is necessary to set the entire training process, set corresponding hyperparameters, such as the number of training epochs, the size of batch_size, the size of the learning rate, etc. In each iteration cycle, calculate the loss function between the image predicted by the current network training and the ground truth label.
[0062] 3.2 During the training process, it is also necessary to perform certain data augmentation operations on the dataset, using common data augmentation operations, such as random cropping, flipping, color conversion, etc.
[0063] 3.3 Set the strategy for saving intermediate weight parameters during the training process. After a certain number of iteration cycles, input the test set images into the current network to obtain performance metrics, and select and save the network weight parameters with the highest accuracy during the entire iteration cycle according to these performance metrics.
[0064] 4) Select the network weight parameters with the best performance saved in step 3), input the image to be predicted into the current U-shaped neural network, and obtain the segmented lane line image.
[0065] Select the network weight parameters of the best saved result and load them into the network at this time during the prediction phase. Input the image to be predicted, and the segmented lane line image is output after passing through the network. For more accurate prediction, further correction is performed through a post-processing algorithm after obtaining the lane line image. It mainly can be divided into the following two steps:
[0066] 4.1 Prepare the image to be predicted and put it in the same folder for input into the trained network. The image to be predicted can be obtained from the vehicle-mounted front camera of the user, and the resolution of the input image needs to be unified in advance.
[0067] 4.2 Input the lane line image output by the network into the post-processing algorithm for correction.
[0068] The present invention also provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the lane line detection method based on the deep learning algorithm are implemented.
[0069] The present invention also provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps of the lane line detection method based on the deep learning algorithm.
[0070] Embodiment:
[0071] A lane line detection method based on a deep learning U-shaped structure and integrating CNN and Transformer is as Figure 1 shown.
[0072] Step 1: Select the public dataset Tusimple and process the training set and the test set.
[0073] The pictures in the Tusimple dataset are collected from highways. The training set includes 6408 pictures, 3626 pictures in the training set with labels, and 2782 test set pictures are used to test the trained model. The resolution of each picture is 720*1280 (720 represents the height of the picture, and 1280 represents the width of the picture).
[0074] The public dataset Tusimple does not contain the label pictures required for the semantic segmentation method. Therefore, before use, it is necessary to first generate the corresponding label pictures in the training set according to the provided label json file, and divide the training set and the test set well, which will be read by the training process later. The specific steps are as follows:
[0075] 1.1: Generate label picture files from the label json coordinate files in Tusimple.
[0076] When generating label picture files from the label json coordinate files, the width of each lane line is specified as 15 pixels. When reading the coordinates of each lane line, each lane line is numbered according to the principle from left to right, and when generating the label picture files, the pixel values are set according to the numbers of each lane line. For example, if the leftmost lane line is numbered 1, then the pixel value of this lane line in the label picture is 1.
[0077] 1.2: Generate txt text files with information such as the paths of the Tusimple training set and test set.
[0078] The txt text file includes the path of the original input image to be processed, the path of the labeled image generated in the above step A1, and the number of lane lines contained in the current image. This operation needs to be performed on both the training set and the test set, because the evaluation of the subsequent network performance requires the labeled files provided by the test set images for calculation. The final format of each image in the text txt file is:
[0079] [ / clips / 0530 / 1492627171538356342_0 / 20.jpg / seg_label / 0530 / 1492627171538356342_0 / 20.png 0 1 1 1 1 0]
[0080] Step 2: Design the overall structure of the U-shaped neural network.
[0081] The overall structure of the neural network is a U-shaped structure, as Figure 1 shown. The overall network structure includes: a downsampling module (Encoder), a global information fusion module (Cycle-Accumulation Transformer), and an upsampling module (Decoder). The input image gradually obtains high-level semantic information through downsampling, and then is input into the global information fusion module for further fusion of global information and local information, and finally transmitted to the upsampling module to restore to the size of the original input image.
[0082] At the same time, the U-shaped neural network adopted in this application can add skip connections at the corresponding positions of the downsampling module and the upsampling module.
[0083] The downsampling module is used to obtain high-level semantic information and is divided into four convolutional layers. Given an image with a size of [3, 720, 1280] matrix (3: three RGB channels, 720: the length of the image, 1280: the width of the image), the image passes through the first three convolutional layers, and the length and width are halved respectively, accompanied by an increase in the number of dimensions. At this time, it gradually changes from low-level semantic information to high-level semantic information. Finally, a feature map with a size of [512, 90, 160] is obtained.
[0084] The downsampling module can adopt the classic Resnet34 pre-trained parameter model. The output channels of the four convolutional layers are 64, 128, 256, and 512 respectively. The size of the intermediate feature map obtained by the output of each layer is halved based on the previous layer, that is, the output size of the first convolutional layer is [64, 360, 640], the output size of the second convolutional layer is [128, 180, 320], the output size of the third convolutional layer is [256, 90, 160], and the output size of the third convolutional layer is [512, 90, 160].
[0085] The four convolutional layers are composed of multiple convolutional blocks inside, including convolutional operations, non-linear operations, normalization operations, etc. The basic module structure is: convolutional operation + normalization operation + non-linear operation, and when downsampling is required, a 1*1 convolution with a stride of 2 is used to halve the size.
[0086] The downsampling module can also use various networks that can obtain high-level semantic information of images, such as ResNet networks, VGG networks, mobile-net networks with different numbers of layers, and downsampling networks based on Transformer.
[0087] The global information fusion module is composed of a CNN-based cyclic accumulation module and a Transformer. The Transformer module includes position embedding, convolutional mapping, and self-attention mechanism. The specific fusion method of the global information fusion module is as Figure 3 shown.
[0088] The global information fusion module is used to integrate the feature maps containing high-level semantic information obtained by the downsampling module. First, the feature map is added with a position embedding module of the same size, and then mapped to three Q matrices, K matrices, and V matrices of the same size through the convolutional mapping module. Then, after calculating the self-attention mechanism of the Q matrix, K matrix, and V matrix, they are input into the cyclic accumulation module; the cyclic accumulation module first divides the input matrix into H rows, and then performs information accumulation and transmission from top to bottom and from bottom to top; then the input matrix is divided into W columns, and then information accumulation and transmission from left to right and from right to left are obtained; finally, through another layer of self-attention mechanism module, the information of the previous two steps is further fused.
[0089] In the process of transmission, the cyclic accumulation module first performs the fusion process from top to bottom and from bottom to top row by row, and finally performs the fusion process from left to right and from right to left column by column. Each fusion process is carried out according to a certain step, and here the step is 2 to the power of n, where n is the number of the fusion, that is, the first fusion step is 1, the second fusion step is 2, the third fusion step is 4, the fourth fusion step is 8, and so on.
[0090] The position embedding module uses learnable relative position encoding and depends on the input embedding. The formula is expressed as:
[0091] X H,W,C = X H,W,C + P H,W,C
[0092] In the above formula, X H,W,C is the current feature map representation, and P H,W,C is the added learnable relative position encoding.
[0093] The convolutional mapping module uses a trainable 1*1 convolution to operate on the input tensor in a sliding window manner. After the convolutional mapping module, three tensors with different parameter spaces, Q matrix, K matrix, and V matrix, are generated. This operation can save a lot of computational overhead and retain more spatial information. The specific formula is:
[0094]
[0095]
[0096]
[0097] In the above formula, Conv2d 1*1 Indicates the convolution operation with a convolution kernel size of 1*1. The Q matrix, K matrix, and V matrix are generated after convolution mapping.
[0098] The self-attention mechanism module is used to obtain the internal correlation of data or features. Its calculation depends on the Q matrix, K matrix and V matrix obtained by the convolution mapping module. The transpose of the Q matrix and the K matrix is used to calculate the attention matrix (attention), and finally the attention matrix and the V matrix are calculated to obtain the final output. The specific formula is expressed as:
[0099]
[0100]
[0101] In the above formula, Softmax represents the Softmax operation of the Pytorch library, Allenllon is the calculated attention matrix, and the hyperparameter d is k Set to 1.
[0102] The global information fusion module can adopt a separate CNN-based cyclic accumulation information fusion module, a separate Transformer-based information fusion module, or a module composed of the two in parallel.
[0103] The upsampling module as a whole consists of four upsampling layers, each of which is composed of deconvolution and bilinear interpolation, and restores the feature map containing high-level semantic information (size 90*160) to the input image size (720*1280). At this time, the output result of the upsampling module already contains the lane lines to be segmented, and each lane line is represented by a different pixel value. During the training phase, the loss function will be calculated with the given label.
[0104] The deconvolution operations inside the four upsampling layers use 1*3 and 3*1 row-column convolutions, and the bilinear interpolation operation uses the function provided by Pytorch. These two operations are in parallel to form an upsampling layer.
[0105] Three of the four upsampling layers will restore the feature map to twice its size, that is, gradually restore from [512, 90, 160] to [3, 720, 1680].
[0106] A learnable offset tensor is added to the layer before the output layer to optimize the obtained lane line accuracy.
[0107] The upsampling module can be a module composed of a single deconvolution operation, a module composed of a single bilinear interpolation operation, or a module formed by the parallel connection of the two.
[0108] Step 3: Train the model and save the weights with the best performance
[0109] Step 3 is to construct the elements required for training the model, including the training strategy and the strategy for saving intermediate training parameters. It mainly includes the following three points:
[0110] 3.1 Determination of the training strategy and hyperparameters;
[0111] During the training process, for the Tusimple dataset, the overall training cycle is set to 400 epochs. In each epoch, all the training set images are input into the network for training. Here, the batchsize is set to 12; for the optimizer, the SGD optimizer is used, where the maximum learning rate parameter is set to 0.02, "weight_decay" is set to 1e-4, and "momentum" is set to 0.9. Further, we adopt the "LambdaLR" strategy to dynamically adjust the learning rate.
[0112] In each iteration cycle, the loss function is calculated for the image predicted by the current network training and the ground truth label, and the cross-entropy loss function is minimized to make the predicted output of the current network and the label close enough.
[0113] 3.2 Image augmentation for the input images;
[0114] During the training process, we adopt common data augmentation strategies such as ColorJitter, RandomResize, RandomCrop, RandomRotation and GroupNormaliz. These methods are implemented through the libraries provided by the official Pytorch.
[0115] During the training process, the data augmentation operation sets a random seed, and the value of the random seed is set to 0.6. That is, only when the random number generated during the current training process is greater than 0.6 will the corresponding data augmentation operation be executed, which can increase the generalization and robustness of the entire network.
[0116] The data augmentation operation is only adopted in the training stage and will not be adopted in the testing stage.
[0117] 3.3 Set the strategy for saving intermediate weight parameters during the training process
[0118] The set saving strategy is as follows: After each iteration cycle, the test set images will be input into the current network to obtain the accuracy metric. According to the accuracy, the network weight parameters with the highest accuracy during the entire iteration cycle will be selected and saved. If the accuracy obtained in the current iteration cycle is higher than the previous one, the updated network weight parameters will be the current parameters, and so on.
[0119] For the Tusimple dataset, the accuracy metric is used to measure the performance of the current network.
[0120] For the Tusimple dataset, adopting the network structure and training strategy of the present invention can achieve an accuracy of 97% and simultaneously have a prediction speed greater than 50 FPS.
[0121] Step Four: Input the image to be predicted to predict the segmented lane lines;
[0122] Select the network weight parameters of the best result saved in Step Three and load them into the neural network during the prediction stage. The image to be predicted is obtained from the vehicle-mounted front camera of the user, and then the segmented lane line image is output after passing through the network. For more accurate prediction, after obtaining the lane line image, it is further corrected through a post-processing algorithm. The specific steps are as follows:
[0123] 4.1 The image to be predicted selected here is a segment of video captured by the front camera, and then the resolution of the image to be predicted is uniformly set to 720*1280. After processing the image to be predicted, the image can be input into the network to obtain the output lane line image.
[0124] The image to be predicted is obtained by video frame extraction, and 10 images are taken at one-second intervals to generate the dataset to be predicted.
[0125] 4.2 The lane line image output by the network is input into the post-processing algorithm for correction. The correction is mainly to complete the breakpoints of the same lane line. The corrected lane line can output coordinates or be mapped back to the original image for use by other systems of autonomous driving.
[0126] In the lane line breakpoint completion algorithm, first, sample the lane line image at a line interval of 10 pixels to determine the coordinate positions of each lane line. For the positions with gaps in the front and back intervals, perform quadratic function fitting to supplement points, and finally regenerate a new corrected lane line image.
[0127] The embodiments should not be regarded as limitations of the present invention, but any improvements based on the spirit of the present invention should be within the protection scope of the present invention.
Claims
1. A lane line detection method based on deep learning algorithm, characterized in that: It includes the following steps: 1) Obtain a public dataset, divide it into a training set and a test set, and obtain a text file containing path information through training; 2) Design a U-shaped structure neural network, which includes a downsampling module for obtaining a feature map containing high-level semantic information, a global information fusion module for integrating the feature map containing high-level semantic information obtained by the downsampling module, and an upsampling module for restoring the feature map containing high-level semantic information to the input image size; 3) Train the U-shaped structure neural network and save the network weight parameters with the best performance; 4) Select the network weight parameters with the best performance saved in step 3), input the image to be predicted into the current U-shaped structure neural network, and obtain the segmented lane line image. The global information fusion module is composed of a CNN-based cyclic accumulation module and a Transformer module. The Transformer module includes position embedding, convolutional mapping, and self-attention mechanism. The working process of the global information fusion module is as follows: Step 1: Combine the obtained feature map containing high-level semantic information with a position embedding module of the same size, and map it through a convolutional mapping module into three Q matrices, K matrices, and V matrices of the same size; Step 2: Perform self-attention mechanism calculation on the Q matrix, K matrix, and V matrix and input them into the CNN-based cyclic accumulation module; Step 3: The CNN-based cyclic accumulation module first divides the input matrix into H rows, and then performs information accumulation and transmission from top to bottom and from bottom to top; Step 4: Divide the input matrix into W columns, and then perform information accumulation and transmission from left to right and from right to left; Step 5: Further fuse the information of the previous two steps through another layer of self-attention mechanism.
2. The lane line detection method based on a deep learning algorithm according to claim 1, characterized in that: In step 1), first generate a label picture file from the label json coordinate file; then divide the training set and the test set, and generate a txt text file with the paths of the training set and the test set. The generated txt text file contains the path of the original input picture to be input, the path of the generated label picture file, and the number of lane lines contained in the current picture.
3. The lane line detection method based on a deep learning algorithm according to claim 2, characterized in that: When generating a label picture file from the label json coordinate file, it is necessary to specify the width of each lane line, and when reading the coordinates of each lane line, number each lane line according to the principle from left to right.
4. The lane line detection method based on the deep learning algorithm according to claim 1, characterized in that: The downsampling module includes four convolutional layers composed of multiple convolutional blocks.
5. The lane line detection method based on a deep learning algorithm according to claim 1, characterized in that: The upsampling module contains four upsampling layers composed of transposed convolution and bilinear interpolation methods.
6. The lane line detection method based on deep learning algorithm according to claim 1, wherein: Step 3) includes:
1. Determine the training strategy and hyperparameters; 2. Perform image augmentation on the input image; 3. Set the strategy for saving intermediate weight parameters during training. The strategy for saving intermediate weight parameters is: every time an iteration cycle passes, input the test set image into the current network to obtain an accuracy index, select the network weight parameters with the highest accuracy during the entire iteration cycle according to the accuracy, and save them.
7. The lane line detection method based on a deep learning algorithm according to claim 6, characterized in that: In step one, set the training cycle. In each cycle, the entire training set of images is input into the network for training. In each iteration cycle, the image predicted by the current network during training is used to calculate the loss function and the minimized cross-entropy loss function with the ground truth label, so as to make the predicted output of the current network close to the label.
8. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the lane line detection method based on the deep learning algorithm described in any one of claims 1 to 7 above.
Citation Information
Patent Citations
Lane detection method based on improved SCNN network
CN110414386B
Lane line detection and segmentation method based on attention space convolutional neural network
CN112633177A
Lane line detection method based on structural information
CN111242037A
Traffic sign recognition method based on dense connection and attention mechanism
CN111582029A