A Lane Detection Method Based on Transformer Encoder and Dilated Convolution
By using Transformer encoder and hollow convolution in lane line detection, global and local features are extracted, and information fusion is carried out through bidirectional weighted feature pyramids, the problem of difficulty in extracting global information in the existing technology is solved, detection accuracy and efficiency are improved, and suitable for complex traffic environments.
Patent Information
- Application Number
- CN202211193390.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-09-28
AI Technical Summary
The existing lane line detection technology is difficult to effectively extract global information of the image, resulting in low detection accuracy and low computing efficiency in complex scenarios such as light changes, occlusion and narrow lanes.
The lane line detection method based on Transformer encoder and hollow convolution is adopted. The global features are obtained through the Transformer encoder, and the multi-scale local features are extracted, and information fusion is performed through the bidirectional weighted feature pyramid. At the same time, unsupervised style migration generation adversarial networks are used to generate nighttime data, expand the data set, and improve the model's detection ability in long-tail scenarios.
It improves the accuracy and computing efficiency of lane line detection, enhances the performance of the model in complex traffic environments and night scenes, and is suitable for a variety of complex road traffic scenarios.
Smart Images

Figure CN115546750B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual lane line detection, and in particular to a lane line detection method based on a Transformer encoder and dilated convolution. Background Art
[0002] Lane detection is a challenging task because it is affected by many factors, such as lighting conditions, occlusion by other vehicles, the presence of irrelevant markings on the road, and the inherently long and narrow characteristics of the lanes themselves. In addition, considering that lane detection runs on a vehicle-based system with very limited computing resources, the computational cost of the lane detection method should also be regarded as a key indicator of overall performance. At the same time, as a basic function of the Advanced Driver Assistance System (ADAS), lane detection must meet the conditions of high accuracy, high real-time performance, and robustness. Therefore, lane detection is not only an important and complex task but also a key factor in the development of any autonomous vehicle system.
[0003] The lane line detection network framework usually adopts the form of an encoder-decoder. Currently, research on lane line recognition mainly focuses on the decoder. However, extracting clear and reliable lane line features is equally important. Extracting clear lane line features will surely reduce a lot of burden on the subsequent detection part. In most algorithms, the encoder part uses a stacked convolutional neural network to extract features from local regions of the image while downsampling the image. However, when the convolutional block extracts image features, it only operates on local pixels and ignores the global information on the image. Existing methods slice the feature map and then use sequential convolution between adjacent rows and columns to stack and transmit information on the feature map. However, since the operation of transmitting sequence information takes a long time, the inference speed is slow. At the same time, sequential information transmission between adjacent rows or columns requires multiple iterations, and some information will be lost during long-distance propagation.
[0004] The stacked convolutional neural network performs multiple downsamplings, reducing the resolution of the feature maps for post-processing and resulting in the omission of tiny lane line target information. To address the multi-scale object detection problem, the feature pyramid fuses feature maps of different scales in different ways. Currently, feature pyramids are mainly divided into unidirectional and bidirectional. FPN fuses feature maps from top to bottom by doubling the size of the upper-level feature maps and adding them to the lower level. Lizhe Liu et al. [1] used FPN to fuse multi-scale features in the lane detection network, but it lacks interpretability and has low computational efficiency. PANet [2] adds a bottom-up feature fusion on the basis of FPN, using a bidirectional fusion backbone network to ensure the diversity and integrity of features, but it cannot weigh the importance of each feature. NAS-FPN [3] uses neural architecture search to find a better cross-scale feature network topology, but it takes a lot of time during the search process, and the discovered network is irregular and difficult to interpret or modify. BiFPN [4] performs bidirectional weighted feature fusion on feature maps of different scales and uses the network to learn the magnitude of the weights to optimize feature fusion.
[0005] In addition, the diversity and quantity of traffic image data are very important for deep learning. However, in some specific driving scenarios such as occlusion, shadow, and night, the data only account for a small part of the entire driving dataset, forming long-tail data and reducing the learning effect of the deep learning neural network on this part of the data. By collecting traffic images in specific scenarios as a new dataset to solve the lane line detection problem in this scenario, this method is time-consuming and laborious and reduces the algorithm iteration efficiency. In dealing with long-tail data, Seokju Lee et al. [5] established a new dataset containing 17 lane and road marking classes, applicable to four different long-tail scenarios: rainless, rainy, heavy rain, and night. However, collecting long-tail data is a time-consuming and laborious task and does not meet the needs of rapidly developing technologies. Style transfer converts the style of one picture into that of another while keeping the content of the original picture unchanged. Gayts [6] repeatedly uses the VGG network to extract the texture information and content information of the image, so that the generated picture retains the content while having a new texture effect. Pix2Pix [7] realizes image style transfer through a generative adversarial network, which requires paired data for training. However, there are very few paired data in actual road traffic pictures, such as black and white road scene pictures with exactly the same environment and traffic flow. Therefore, the above two methods are not applicable. Cyclegan [8] ensures content invariance by introducing a cyclic consistency loss, making it unnecessary to use one-to-one corresponding pictures as input. UNIT [9] is an improvement on the basis of Cyclegan. It believes that two-domain images can be derived from their joint distribution and uses the VAE-GAN structure to retain content details. However, it is difficult to obtain paired different-style pictures in actual road traffic pictures.
[0006] References:
[0007] [1]Lizhe Liu,Xiaohao Chen,Siyu Zhu.CondLaneNet:a Top-to-down LaneDetection Framework Based on Conditional Convolution[J].arXiv preprint arXiv:2105.05003,2021.
[0008] [2]Liu S,Qi L,Qin H et al.Path Aggregation Network for InstanceSegmentation[C].IEEE Conference on Computer Vision and Pattern Recognition(CVPR),2018.
[0009] [3]Ghiasi G,Lin TY,Le QV.NAS-FPN:Learning Scalable Feature PyramidArchitecture for Object Detection[C] / / 2019 IEEE / CVF Conference on ComputerVision and Pattern Recognition(CVPR).IEEE,2019.
[0010] [4]Tan M,Pang R,Le QV.EfficientDet:Scalable and Efficient ObjectDetection[C] / / 2020 IEEE / CVF Conference on Computer Vision and PatternRecognition(CVPR).IEEE,2020.
[0011] [5]Seokju Lee,Junsik Kim,Jae Shin Yoon,et al.VPGNet:Vanishing PointGuided Network for Lane and Road Marking Detection and Recognition[C] / / 2017IEEE International Conference on Computer Vision(ICCV).IEEE,2017.
[0012] [6]Gatys LA,Ecker AS,Bethge M.Image Style Transfer UsingConvolutional Neural Networks[C] / / 2016 IEEE Conference on Computer Vision andPattern Recognition(CVPR).IEEE,2016.
[0013] [7]Phillip Isola,Jun-Yan Zhu,Tinghui Zhou et al.Image-to-ImageTranslation with Conditional Adversarial Networks[J] / / 2017 IEEE Conference onComputer Vision and Pattern Recognition(CVPR).IEEE,2017.
[0014] [8]Jun-Yan Zhu,Taesung Park,Phillip Isola et al.Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks[C] / / IEEEConference on Computer Vision and Pattern Recognition(CVPR),2017:2223-2232.
[0015] [9]Ming-Yu Liu,Thomas Breuel,Jan Kautz.Unsupervised Image-to-ImageTranslation Networks[C] / / 31st Conference on Neural Information ProcessingSystems(NIPS2017),Long Beach,CA,USA. Summary of the Invention
[0016] Aiming at the problems existing in the prior art, the present invention provides a lane line detection method based on a Transformer encoder and dilated convolution. This algorithm overcomes the limitations that stacked convolutional neural networks cannot obtain global information of images and are difficult to identify tiny lane line targets, and generates night data through style transfer, overcomes the problem of insufficient amount of long-tail data, improves the detection efficiency and accuracy of the model, and makes the model applicable to a variety of complex road traffic scenarios.
[0017] To achieve the above object, the present invention extracts local features of different scales by using dilated convolution, and uses a Transformer encoder to globally associate the slender linear structure of lane lines. Finally, a bidirectional weighted feature pyramid is used to weight and fuse local and global information to be applicable to lane line detection in complex traffic environments. In addition, the present invention also uses an unsupervised style transfer generative adversarial network to generate night driving images, improving the detection ability of the lane line detection network in night and dark traffic environments.
[0018] Specifically, a lane line detection method based on a Transformer encoder and dilated convolution provided by the present invention includes the following steps:
[0019] Use the UNIT unsupervised style transfer method to generate night traffic scene data using daytime traffic images;
[0020] Construct a backbone feature extraction network, and replace the original convolution with dilated convolution in the backbone feature extraction network to extract multi-scale local features of lane lines;
[0021] Construct a Transformer encoder, and use positional encoding and self-attention mechanism to obtain global features;
[0022] Use a bidirectional feature pyramid to perform top-down and bottom-up weighted fusion on the extracted local and global features;
[0023] Adopt a method based on instance segmentation to construct a lane line detection head;
[0024] Train the model using the dataset to make the model converge and obtain the lane line detection network parameters;
[0025] Install the model on the vehicle-mounted camera for real-time detection of lane lines to obtain a lane line instance segmentation map.
[0026] Furthermore, before performing UNIT unsupervised style transfer, it also includes the steps of: obtaining a publicly available road traffic dataset, where the dataset contains lane lines and their labels.
[0027] Furthermore, in order to cope with different traffic scenarios, the dataset should be the CULane dataset, which includes normal scenarios, congestion scenarios, turning scenarios, glare scenarios, night scenarios, lane-less scenarios, shadow scenarios, and scenarios with arrow markings on the road.
[0028] Furthermore, the UNIT unsupervised style transfer method is used to generate night traffic scenarios using daytime traffic images, including:
[0029] Let B=(X,Y), where X is the original image, Y is the label of the original image, and B is the combination of the original data and its label;
[0030] Assume B g =(X g ,Y g ), where X g is the generated image, Y g is the label of the generated image, then:
[0031] X g =G(E(X))
[0032] Y g =Y
[0033] where G is the generator; E is the encoder, and B g is the combination of the generated data and its label.
[0034] Since the style transfer only generates night images through daytime images and does not change the distribution of details such as lane lines and the environment in the images, the label of the generated image can directly use the label of the original image.
[0035] Furthermore, keep the resolution of the feature map unchanged by reducing the convolution stride of the backbone feature extraction network to 1.
[0036] Furthermore, replace the original convolution with dilated convolution in the backbone feature extraction network, including:
[0037] Modify the convolution of the last two modules of the backbone feature extraction network to dilated convolution. Assume the input Let \(W\) be the width of the input image and \(H\) be the height of the input image. After feature extraction by dilated convolution, the output feature map The size relationship between the convolution input and output is:
[0038]
[0039] where \(W\) in is the input size; \(W\) out is the output size; \(P\) is the padding number; \(K\) is the convolution kernel size; \(D\) is the convolution dilation rate; \(S\) is the convolution stride.
[0040] Furthermore, in the Transformer encoder, the feature map first passes through a convolutional layer with a kernel size of 3 and a stride of 1 to obtain the feature map embedding \(F'\), and the fixed-position encoding \(PE\) is added to it. In the self-attention module, the attention values are calculated through dot product. Finally, more features are added through residual connection without adding too much computational cost, and a single-layer convolutional network is used for further feature integration;
[0041] where the position encoding is calculated using sine and cosine with different frequencies:
[0042] \(PE(pos, 2i)=\sin(pos / 10000\) 2i / d )
[0043] \(PE(pos, 2i + 1)=\cos(pos / 10000\) 2i / d )
[0044] \(F'' = F' + PE\)
[0045] where \(pos\) is the position of the pixel; \(i\) is the current dimension; \(d\) is the total dimension size; \(F''\) is the feature map embedding after adding the position encoding, and \(PE(pos, 2i)\) is the position encoding of the pixel at position \(pos\) in the \(2i\)-th dimension.
[0046] Furthermore, in the bidirectional feature pyramid, the range of weights is constrained by fast normalization weight fusion. The fast normalization weight fusion formula is
[0047]
[0048] The output after bidirectional weighted fusion is:
[0049] \(O = conv(\omega\) io \(\cdot F\) i )
[0050] where \(\omega\) i is the initial weight of the \(i\)-th input, \(\epsilon\) is a preset extremely small number to prevent the denominator from being 0, \(\omega\) jis the weight of the j-th input, ω io is the weight of the i-th input after fast normalized weight fusion, F i is the i-th input, conv is a 3x3 convolution, and O is the fused output.
[0051] Furthermore, the total loss function includes an instance segmentation loss and a lane line presence loss.
[0052] Furthermore, in lane line detection, the instance segmentation loss is calculated by the cross entropy loss function, and the lane line presence loss is calculated by the binary cross entropy loss function;
[0053] Furthermore, when training the model, the SGD optimizer is used to optimize the network, the learning rate is set to 0.03, the momentum is set to 0.9, and the weight decay rate is 5e-4. The batch size for each training is 16, and the number of training epochs is 12.
[0054] Furthermore, at least one bidirectional feature pyramid is provided.
[0055] Compared with the prior art, the lane line detection algorithm based on Transformer and dilated convolution of the present invention has at least the following beneficial effects:
[0056] This method uses dilated convolution to extract local features of lane lines, uses a Transformer encoder to obtain global features, and strengthens feature fusion through a bidirectional weighted feature pyramid, improving the ability to extract and fuse multi-scale slender lane line features in different scenarios. In addition, an unsupervised style transfer generative adversarial network is used to augment the dataset, converting daytime-style images into nighttime, which enhances the model's ability to detect lanes in long-tail scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a schematic diagram of the overall structure of the model of the lane line detection method based on Transformer and dilated convolution in an embodiment of the present invention;
[0058] Figure 2 is a schematic diagram of the structure of the unsupervised style transfer generative adversarial network in an embodiment of the present invention;
[0059] Figure 3 is a comparison diagram of dilated convolution and ordinary convolution in an embodiment of the present invention;
[0060] Figure 4 is a schematic diagram of the structure of the Transformer encoder in an embodiment of the present invention;
[0061] Figure 5Schematic diagram of the feature fusion structure in the embodiments of the present invention;
[0062] Figure 6 Schematic flow diagram of a lane line detection method based on a Transformer encoder and dilated convolution provided by the embodiments of the present invention. Detailed implementation manners
[0063] The following will describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. The described preferred embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0064] Please refer to Figure 1 , a lane line detection method based on a Transformer encoder and dilated convolution provided by the present invention, the specific steps include:
[0065] S1. Download the publicly available road traffic dataset CULane, which is a large dataset dedicated to lane line detection and includes normal scenarios and high-challenge scenarios such as crowded, glare, shadow, ground arrow, curve, intersection, and night road. Among them, the training set contains 88,880 road traffic pictures, and the test set contains 34,680 road traffic pictures.
[0066] S2. Use the UNIT unsupervised style transfer method to generate night traffic scenes using daytime traffic images for data augmentation;
[0067] As Figure 2 shown, UNIT first co-encodes pictures in two different domains (defined as the X1 domain and the X2 domain) into the latent hidden space Z domain through the encoders E1 and E2, and then converts the data in the Z domain into the X1 and X2 domains through the generators G1 and G2. In the figure, is the X1 domain picture obtained by encoding and decoding from the X1 domain, is the X1 domain picture obtained by encoding and decoding from the X2 domain, is the X2 domain picture obtained by encoding and decoding from the X1 domain, is the X2 domain picture obtained by encoding and decoding from the X2 domain. Then calculate and the cycle consistency loss of X1, and X2, retain the detailed information of the pictures, and finally use the discriminators D1 and D2 to distinguish the authenticity of the generated pictures and real pictures, and improve the authenticity of the pictures after style transfer in an adversarial training manner.
[0068] In some embodiments of the present invention, let B=(X, Y), where X is the original image and Y is the label of the original image, and B is the combination of the original data and its label.
[0069] Style transfer only generates night images from daytime images without changing the distribution of details such as lane lines and the environment in the images. Therefore, the labels of the generated images can directly use the labels of the original images;
[0070] Assume B g =(X g , y g ), where X g is the generated image, and Y g is the label of the generated image. Then:
[0071] X g =G(E(X))
[0072] Y g =Y
[0073] where B g is the combination of the generated data and its label, G is the generator; E is the encoder.
[0074] S3. Construct a backbone feature extraction network, replace the original convolution with dilated convolution to extract multi-scale local features of lane lines;
[0075] In some embodiments of the present invention, ResNet18 is used to construct the backbone feature extraction network. Of course, in other embodiments, common networks such as VGG16 can also be used to construct the backbone feature extraction network.
[0076] In some embodiments of the present invention, in step 1, the convolutions of the last two modules of the backbone feature extraction network are modified to dilated convolutions. As Figure 3 shown, compared with ordinary convolution, dilated convolution can make the output of each convolution block contain information in a larger range, while increasing the receptive field of the convolution and preventing the feature map from being too small to lose information about small targets. Among them, assume the input i.e., X is a three-dimensional input with a size of (3, W, H). Among them, W is the width of the input image, and H is the height of the input image. After feature extraction by dilated convolution, the output feature map i.e., F has a size of (512, W / 8, H / 8). The size relationship between the convolution input and output is:
[0077]
[0078] where, W in is the input size; W oit is the output size; P is the padding number; K is the convolution kernel size; D is the convolution dilation number; S is the convolution stride.
[0079] In some embodiments of the present invention, in step 1, the convolution stride of the backbone feature extraction network is reduced to 1 to keep the resolution of the feature map unchanged.
[0080] S4. Construct a Transformer encoder to obtain global features using positional encoding and self-attention mechanism;
[0081] In some embodiments of the present invention, as Figure 3 shown, the feature map F first passes through an input embedding convolutional layer with a convolutional kernel size of 3 and a stride of 1 to obtain a feature map embedding F';
[0082] Subsequently, positional encoding PE is added. Since the feature map embedding F' and the positional encoding PE have the same dimension, the addition of positional information can be completed by adding the feature map embedding and the positional encoding. The positional encoding is calculated using sin and cos of different frequencies:
[0083] PE(pos, 2i) = sin(pos / 10000 2i / d )
[0084] PE(pos, 2i + 1) = cos(pos / 10000 2i / d )
[0085] F'' = F' + PE
[0086] where pos is the position of the pixel; i is the current dimension; d is the total dimension size. When d is odd, When d is even, F'' is the feature map embedding after adding positional encoding, and PE(pos, 2i) is the positional encoding of the pixel at position pos in the 2i-th dimension.
[0087] After the positional encoding PE, a self-attention module is added. In the self-attention module, the feature map embedding F'' after adding positional encoding is linearly transformed and resized to obtain a query vector Q, a key and a value V, where d k = 128 is the dimension of Q and K; the attention value Attention is calculated through dot product, that is, the strength of the association between pixels:
[0088]
[0089] Then, the attention value Attention is multiplied by the value V to obtain the output F of the self-attention module o :
[0090] F o = V · Attention
[0091] Moreover, there is a residual connection between the output of the input embedding convolutional layer and the self-attention module. Through the residual connection, more features are added without adding too much computational cost, and a single-layer convolutional network is further used for further feature integration.
[0092] S5. Use a bidirectional feature pyramid to perform top-down and bottom-up weighted fusion on the extracted local and global features.
[0093] At least one bidirectional feature pyramid is provided. When two or more bidirectional feature pyramids are provided, the output of the previous feature pyramid is the input of the next feature pyramid. In some embodiments of the present invention, considering real-time performance, only one bidirectional feature pyramid is provided.
[0094] Figure 1 The solid line part in the figure is the actual application, and the dotted line part is applicable but not applied considering real-time performance. Therefore, the input of the bidirectional feature pyramid is the global feature output by the top-layer Transformer encoder and the multi-scale local features directly output by the second and third-layer dilated convolutions.
[0095] Since the backbone feature extraction network is modified to make the sizes of the output feature maps of the last three layers the same, the bidirectional feature pyramid does not need to perform linear interpolation expansion or pooling reduction on the feature maps, avoiding information loss.
[0096] In some embodiments of the present invention, the range of weights is constrained by fast normalized weight fusion, so that the fused weight value ω io falls between 0 and 1, and the network automatically adjusts the size of the weights through learning. This weight fusion method can prevent training instability caused by too large weight values and run faster on the GPU.
[0097] Among them, the fast normalized weight fusion formula is
[0098]
[0099] As Figure 4 shown, the output after bidirectional weighted fusion is:
[0100]
[0101] Among them, ω i is the initial weight of the i-th input, ∈ is a preset extremely small number to prevent the denominator from being 0, ω j is the weight of the j-th input, ω io is the weight of the i-th input after fast normalized weight fusion, and F iLet the \(i\)-th input be \(x_i\), conv be a \(3\times3\) convolution, and \(O\) be the fused output.
[0102] As Figure 5 shown, three feature maps \(F1\), \(F2\), and \(F3\) are input into the bidirectional feature pyramid and fused in the direction of the arrows. For example, the fusion process of \(F5\) is as follows:
[0103]
[0104] \(\omega_1\) and \(\omega_4\) are the weights of the first input and the fourth input respectively;
[0105] In some embodiments of the present invention, \(\epsilon = 0.0001\) to prevent numerical instability.
[0106] S6. A method based on instance segmentation is used to construct a lane detection head, and a lane instance segmentation map is output through convolution;
[0107] The total loss function includes instance segmentation loss and lane presence loss. In some embodiments of the present invention, the instance segmentation loss is calculated by the cross entropy loss function, and the lane presence loss is calculated by the binary cross entropy loss function. Of course, in other embodiments, other loss functions can also be used.
[0108] The loss function formula is:
[0109]
[0110] \(L=\alpha L_{seg}\) seg +\(\beta L_{exist}\) exit
[0111] where \(L_{seg}\) seg is the instance segmentation loss; \(y\) i is the instance segmentation ground truth; \(p_i\) i is the probability of predicting the \(i\)-th lane instance; \(L_{exist}\) exit is the lane presence loss; \(q\) i is the lane presence ground truth; \(e\) i is the lane presence prediction value; \(\alpha\) and \(\beta\) are the weight coefficients of the instance segmentation loss and the lane presence loss respectively, and \(L\) is the total loss function.
[0112] S7. Use the original road traffic dataset and the dataset generated by style transfer to train the model (the lane detection network model composed of the backbone feature extraction network, Transformer encoder, bidirectional feature pyramid, and instance segmentation-based detection head), so that the model converges to obtain the lane detection network parameters.
[0113] In some embodiments of the present invention, in step 7, the network is optimized using the SGD optimizer.
[0114] The learning rate is set to 0.03;
[0115] The momentum is set to 0.9;
[0116] The weight decay rate is 0.0005;
[0117] The batch size for each training is 16;
[0118] The number of training epochs is 12.
[0119] Training is performed on a server equipped with an NVIDIA GeForce RTX2080ti graphics card.
[0120] S8. Install the network model on the vehicle-mounted camera, and real-time detection of lane lines can be achieved. This step only requires the vehicle-mounted camera to obtain road images, and then input them into the trained network model file, and the output lane line instance segmentation map will be obtained.
[0121] The lane line detection method provided by the foregoing embodiments of the present invention specifically utilizes the characteristics that the Transformer encoder can efficiently extract global features of pictures and the dilated convolution can expand the convolution receptive field and extract multi-scale local features. Based on the deep learning algorithm, taking the road traffic image as the input of the model, after local and global feature extraction, the extracted features are fused using the bidirectional weighted feature pyramid, and finally the instance segmentation detection head is used to output the lane line instance segmentation picture to achieve lane line detection. In order to improve the lane line detection ability of the model in night and dark scenes, unsupervised style transfer is used to convert the images of the daytime scene into the night and add them to the dataset. The proposed algorithm improves the accuracy and calculation efficiency of lane line feature extraction in different scenes, and can be easily integrated into other existing lane line detection algorithms for end-to-end training.
[0122] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A lane line detection method based on a Transformer encoder and dilated convolution, characterized in that, The method includes the following steps: Using the UNIT unsupervised style transfer method to generate night traffic scene data by using daytime traffic images; Constructing a backbone feature extraction network, and replacing the original convolution with dilated convolution in the backbone feature extraction network to extract multi-scale local features of lane lines; Constructing a Transformer encoder to obtain global features by using positional encoding and self-attention mechanism; Using a bidirectional feature pyramid to perform top-down and bottom-up weighted fusion on the extracted local and global features; Adopting a method based on instance segmentation to construct a lane line detection head; Training the model to make the model converge to obtain lane line detection network parameters; Installing the model on an in-vehicle camera to be used for real-time detection of lane lines to obtain a lane line instance segmentation map; Among them, replacing the original convolution with dilated convolution in the backbone feature extraction network includes: Modify the convolutions of the last two modules of the backbone feature extraction network to dilated convolutions. Assume the input Let W be the width of the input image and H be the height of the input image. After feature extraction by dilated convolution, the output feature map is F. The size relationship between the convolution input and output is as follows: Among them, W in is the input size; W out is the output size; P is the padding number; K is the convolution kernel size; D is the convolution dilation; S is the convolution stride; In the bidirectional weighted feature pyramid structure for bidirectional weighted fusion of the feature maps extracted by the feature extractor, the range of weights is constrained by fast normalization weight fusion, and the fast normalization weight fusion formula is The output after bidirectional weighted fusion is: O = conv(ω io ·F i ) where ω i is the initial weight of the i-th input, ∈ is a preset extremely small number to prevent the denominator from being zero, ω j is the weight of the j-th input, ω io is the weight of the i-th input after fast normalization weight fusion, F i is the i-th input, conv is a 3x3 convolution, and O is the fused output.
2. The lane line detection method based on a Transformer encoder and dilated convolution according to claim 1, characterized in that, Before performing the UNIT unsupervised style transfer, it also includes the step of obtaining a publicly available road traffic data set, and the data set contains lane lines and their labels.
3. The lane line detection method based on a Transformer encoder and dilated convolution according to claim 2, characterized in that, The data set includes normal scenes, congested scenes, turning scenes, glare scenes, night scenes, lane-less scenes, shadow scenes and roads, arrow-marked scenes.
4. The lane line detection method based on a Transformer encoder and dilated convolution according to claim 1, characterized in that, The using the UNIT unsupervised style transfer method to generate night traffic scenes by using daytime traffic images includes: Let B=(X, Y), where X is the original image, Y is the label of the original image, and B is the combination of the original data and its label; Hypothesis B g =(X g , Y g ), where X g is the generated image and Y g is the label of the generated image, then: X g = G(E(X)) Y g = Y Among them, G is the generator; E is the encoder, and B g is the combination of the generated data and its labels.
5. The lane line detection method based on a Transformer encoder and dilated convolution according to claim 1, characterized in that, The constructing the Transformer encoder to obtain global features by using positional encoding and self-attention mechanism includes: The feature map F first passes through a convolutional layer to obtain a feature map embedding F'; Subsequently, positional encoding PE is added, and the positional encoding is calculated using sin and cos of different frequencies: PE(pos, 2i) = sin(pos / 10000 2i / d ) PE(pos, 2i + 1) = cos(pos / 10000 2i / d ) F″ = F′ + PE Among them, pos is the position of the pixel; i is the current dimension; d is the total dimension size; F″ is the feature map embedding after adding positional encoding, and PE(pos, 2i) is the positional encoding of the pixel at position pos on the 2i-th dimension; After the positional encoding PE, a self-attention module is added. In the self-attention module, F″ is linearly transformed and resized to obtain a query vector Q, a key K, and a value V; the attention value Attention is calculated through dot product, that is, the strength of the association between pixels and pixels: Multiply the attention value Attention by the feature value V to obtain the output F o ; Finally, residual connections are used to add more features without adding too much computational cost, and a single-layer convolutional network is used for further feature integration.
6. A lane line detection method based on a Transformer encoder and dilated convolution according to claim 1, wherein, The total loss function includes instance segmentation loss and lane line presence loss.
7. A lane line detection method based on a Transformer encoder and dilated convolution according to claim 6, wherein, The instance segmentation loss is calculated by the cross entropy loss function, and the lane line presence loss is calculated by the binary cross entropy loss function. The loss function formula is: L = αL seg + βL exit Among them, y i is the ground truth of instance segmentation; p i is the probability of predicting the i-th lane line instance; q i is the ground truth of the lane line existence; e i is the predicted value of the lane line existence; L seg is the instance segmentation loss; L exit is the loss of the lane line existence; α and β are the weight coefficients of the instance segmentation loss and the loss of the lane line existence respectively, and L is the total loss function.
8. A lane line detection method based on a Transformer encoder and dilated convolution according to any one of claims 1-7, wherein, At least one bidirectional feature pyramid is set.
Citation Information
Patent Citations
Lane line detection method based on image sequence
CN113255459A
Curve detection method based on novel mixed attention module
CN114926796A