Lane line detection method and related device
By introducing the bar-perceptual feature fusion module SFF and the spatial cross attention mechanism module SCA in the lane line detection model SPLane, the problems of lane line detection accuracy and efficiency in the prior art in complex scenarios are solved, and higher detection accuracy and speed are achieved.
Patent Information
- Application Number
- CN202411362461.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-06-03
AI Technical Summary
The existing lane line detection model performs poorly in the detection accuracy and detection timeliness when facing complex scenarios, especially at night, missing lane marks, curved lane lines and busy traffic.
A lane line detection method is adopted, and the bar-perceptual feature fusion module SFF and the spatial cross attention mechanism module SCA are introduced in the image lane line detection model SPLane. This method enhances feature interaction through the bar geometric prior of lane lines and captures the spatial positional relationship of lane lines through attention mechanisms.
It effectively improves the accuracy of lane line detection results and greatly improves detection efficiency. Compared with traditional methods, the accuracy index F1 score has increased by 1.03%, and the speed index has increased by 22FPS.
Smart Images

Figure CN120088753A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a lane line detection method and related device. Background Art
[0002] Lane line detection, as a key task in the field of computer vision, plays a crucial role in intelligent transportation systems; especially in the technical system of intelligent vehicles, lane line detection occupies an important position; specifically, the front road image collected by an on-vehicle camera undergoes specialized image analysis and processing techniques to detect and accurately outline the lane line boundaries on the road surface.
[0003] Currently, existing image-based lane line detection technologies not only exhibit high recognition accuracy but also achieve high-efficiency real-time response speed when dealing with conventional and clearly structured road scenes; as shown in the appendix Figure 1 In the real-world driving environment, traditional lane line detection models often perform poorly in terms of detection accuracy and detection timeliness when facing complex scenes such as night, missing lane markings, and curved lane lines; specifically as follows:
[0004] Under nighttime conditions, due to poor lighting conditions, lane lines are difficult to be accurately identified by the detection model in the image, as shown in the appendix Figure 1 (a); in the case of the lack of clear lane line markings, even if the lane lines actually exist, their blurriness or complete absence greatly increases the detection difficulty, as shown in the appendix Figure 1 (b); in curved road areas, the irregular shape of lane lines makes feature extraction complex and further exacerbates the difficulty of positioning, as shown in the appendix Figure 1 (c); in busy traffic situations, lane lines may be blocked by surrounding vehicles or pedestrians, causing additional interference to detection, as shown in the appendix Figure 1 (d); in the above complex scenes, the visual cues provided by the image are usually rather vague, making it difficult for the model to capture the key features of lane lines, thus resulting in a large error in the lane line detection results and a long required detection time. Summary of the Invention
[0005] In view of the technical problems existing in the prior art, the present invention provides a lane line detection method and related device to solve the technical problems that existing lane line detection models often perform poorly in terms of detection accuracy and detection timeliness when facing complex scenes.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] The present invention provides a lane line detection method, including:
[0008] Obtain the lane line image to be detected;
[0009] Input the lane line image to be detected into the pre-constructed image lane line detection model SPLane for lane line detection, and output the lane line detection result; wherein, the image lane line detection model SPLane includes a backbone network, a strip perception feature fusion module SFF, an attention mechanism module, an instance heat map module, a conditional convolution module, and a prediction output module;
[0010] The backbone network is used to extract features from the lane line image to be detected and obtain a multi-level feature map;
[0011] The strip perception feature fusion module SFF is used to exchange and fuse the features in the horizontal and vertical directions of the high-level feature map in the multi-level feature map based on the strip geometric prior of the lane line, and obtain a strip perception fusion feature map;
[0012] The attention mechanism module is used to obtain local and global information at the spatial level based on the multi-level feature map and the strip perception fusion feature map and perform attention fusion to obtain an attention fusion feature map;
[0013] The instance heat map module is used to mark the starting points of lane line instances on the lane line image to be detected based on the attention fusion feature map, and obtain an instance heat map;
[0014] The conditional convolution module is used to perform conditional convolution on the attention fusion feature map to generate a Gaussian mask feature map and an offset feature map;
[0015] The prediction output module is used to obtain the lane line coordinate information in the lane line image to be detected according to the instance heat map, the Gaussian mask feature map, and the offset feature map, and output the lane line detection result.
[0016] Further, the strip perception feature fusion module SFF includes a first dilated convolution layer, a second dilated convolution layer, a third dilated convolution layer, an element addition module, a strip feature aggregation strategy module, a fusion module Concat, and a 1×1 convolution module;
[0017] The first dilated convolution layer is used to perform a 1×1 dilated convolution operation on the high-level feature map, and sequentially apply batch normalization and the ReLU activation function to obtain a feature map F 1 ;
[0018] The second dilated convolution layer is used to perform a 3×3 dilated convolution operation on the high-level feature map, and sequentially apply batch normalization and the ReLU activation function to obtain a feature map F 2 ;
[0019] The third dilated convolutional layer is used to perform a 5×5 dilated convolution operation on the high-level feature map, and then apply batch normalization and the ReLU activation function in sequence to obtain the feature map F. 3 ;
[0020] The element-wise addition module is used to perform element-wise addition on the feature map F 1 , the feature map F 2 and the feature map F 3 to obtain the feature map after dilated convolution processing.
[0021] The bar-shaped feature aggregation strategy module is used to perform feature aggregation processing on the spatial domain of the high-level feature map in the horizontal and vertical directions to obtain the feature map F 4 ;
[0022] The fusion module Concat is used to fuse the feature map after dilated convolution processing with the feature map F 4 to obtain the fused feature;
[0023] The 1×1 convolution module is used to adjust the channel dimension of the fused feature to obtain the bar-shaped perception fused feature map.
[0024] Further, the process of performing feature aggregation processing on the spatial domain of the high-level feature map in the horizontal and vertical directions to obtain the feature map F 4 is as follows:
[0025] Perform average pooling processing on the high-level feature map in the horizontal and vertical directions to obtain a feature map with a size of H×1 and a feature map with a size of 1×W;
[0026] Perform a one-dimensional dilated convolution operation on the feature map with a size of H×1 to expand the size of the feature map with a size of H×1 in the horizontal direction to obtain the horizontally expanded feature map;
[0027] Perform a one-dimensional dilated convolution operation on the feature map with a size of 1×W to expand the size of the feature map with a size of 1×W in the vertical direction to obtain the vertically expanded feature map;
[0028] Add the horizontally expanded feature map and the vertically expanded feature map pixel by pixel to obtain a new feature map;
[0029] After performing 1×1 convolution processing on the new feature map and multiplying it with the high-level feature map, obtain the feature map F 4 .
[0030] Further, the multi-level feature map further includes a low-level feature map, a mid-low-level feature map, and a mid-high-level feature map; the low-level feature map, the mid-low-level feature map, and the mid-high-level feature map are all input into the attention mechanism module to obtain an attention fusion feature map.
[0031] Further, the attention mechanism module includes a first attention module, a second attention module, and a third attention module;
[0032] The first attention module is used to generate a first-layer attention fusion feature map according to the mid-high-level feature map and the bar perception fusion feature map; wherein, the first-layer attention fusion feature map is used to generate an instance heat map through the instance heat map module;
[0033] The second attention module is used to generate a second-layer attention fusion feature map according to the mid-low-level feature map and the first-layer attention fusion feature map; wherein, the second-layer attention fusion feature map is used to generate a Gaussian mask feature map through the conditional convolution module;
[0034] The third attention module is used to generate a third-layer attention fusion feature map according to the low-level feature map and the second-layer attention fusion feature map; wherein, the third-layer attention fusion feature map is used to generate an offset feature map through the conditional convolution module.
[0035] Further, the first attention module, the second attention module, and the third attention module all adopt the spatial cross-attention module SCA;
[0036] Among them, the working process of the spatial cross-attention module SCA is specifically as follows:
[0037] Perform channel attention calculation and spatial attention calculation on the two input feature maps in sequence, and then fuse the attention calculation results to obtain the corresponding attention fusion feature map; wherein, when performing channel attention calculation, average pooling and max pooling operations are simultaneously used to aggregate the channel information of the feature map; when performing spatial attention calculation, it is used to obtain the spatial position information of the lane lines in the feature map.
[0038] The present invention also provides a lane line detection system, including:
[0039] An image acquisition module, used to acquire an image of the lane line to be detected;
[0040] A detection output module, configured to input the lane line image to be detected into a pre-constructed image lane line detection model SPLane for lane line detection, and output a lane line detection result; wherein, the image lane line detection model SPLane includes a backbone network, a strip perception feature fusion module SFF, an attention mechanism module, an instance heat map module, a conditional convolution module, and a prediction output module;
[0041] The backbone network is configured to extract features from the lane line image to be detected to obtain a multi-level feature map;
[0042] The strip perception feature fusion module SFF is configured to exchange and fuse features in the horizontal and vertical directions of the high-level feature map in the multi-level feature map based on the strip geometric prior of the lane line to obtain a strip perception fusion feature map;
[0043] The attention mechanism module is configured to obtain local and global information at the spatial level based on the multi-level feature map and the strip perception fusion feature map and perform attention blending to obtain an attention blending feature map;
[0044] The instance heat map module is configured to mark the starting points of lane line instances on the lane line image to be detected based on the attention blending feature map to obtain an instance heat map;
[0045] The conditional convolution module is configured to perform conditional convolution on the attention blending feature map to generate a Gaussian mask feature map and an offset feature map;
[0046] The prediction output module is configured to obtain the lane line coordinate information in the lane line image to be detected according to the instance heat map, the Gaussian mask feature map, and the offset feature map, and output a lane line detection result.
[0047] The present invention also provides a lane line detection device, including:
[0048] A processor, adapted to execute a computer program;
[0049] A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, the lane line detection method is executed.
[0050] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and is characterized in that when the computer program is executed by a processor, the lane line detection method is implemented.
[0051] The present invention also provides a computer program product, characterized in that the computer program product includes a computer program, and when the computer program is executed by a processor, the lane line detection method described above is implemented.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0053] In the lane line detection method provided by the present invention, by introducing a bar perception feature fusion module into the image lane line detection model SPLane, the feature interaction in the horizontal and vertical directions is strengthened through the lane line bar geometry prior, and the global context perception ability of the lane line is effectively improved, so as to effectively cope with the challenges of lane line detection in complex scenarios; in addition, in the attention mechanism module, by modeling the spatial and channel information in the lane image, the local and global information at the spatial level is obtained and attention fusion is performed to effectively capture the relative position relationship in the lane line. By using the interdependence relationship between different regions and channels in the lane image, the spatial position connection of the lane line in the shallow and deep feature maps is captured, so as to endow the model with the ability to infer the relative position of the lane line; the method of the present invention can effectively improve the accuracy of the lane line detection result, and at the same time can greatly improve the detection efficiency. Compared with the traditional CondLaneNet method, the precision index F1 score increases by 1.03%, and the speed index increases by 22FPS.
[0054] The lane line detection system, lane line detection device and computer-readable storage medium provided by the present invention have all the advantages of the above-mentioned lane line detection method. Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0056] Figure 1 It is a schematic diagram of some complex scenarios in lane line detection;
[0057] Figure 2 It is a schematic diagram of the structure of the image lane line detection model SPLane in Embodiment 1;
[0058] Figure 3 It is a schematic diagram of the structure of the bar perception feature fusion module SFF in Embodiment 1;
[0059] Figure 4 It is a schematic diagram of the principle of the bar feature aggregation strategy module in Embodiment 1;
[0060] Figure 5 Schematic diagram of the principle of the spatial cross-attention module SCA in Embodiment 1;
[0061] Figure 6 Schematic diagram of the principle of the channel attention module in Embodiment 1;
[0062] Figure 7 Schematic diagram of the principle of the spatial attention module in Embodiment 1;
[0063] Figure 8 Visualization result in a complex scenario in the CULane dataset in Embodiment 1;
[0064] Figure 9 Block diagram of the structure of the image lane detection system provided in Embodiment 2;
[0065] Figure 10 Block diagram of the structure of the image lane detection device provided in Embodiment 3. Detailed implementation manners
[0066] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention clearer and more understandable, the following specific embodiments are used to further elaborate on the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0067] Embodiment 1
[0068] This Embodiment 1 provides a lane line detection method, including the following steps:
[0069] Step 1, obtain the lane line image to be detected; wherein, the lane line image to be detected is the front road image collected by an in-vehicle camera.
[0070] Step 2, input the lane line image to be detected into a pre-constructed image lane line detection model (StripPrior Lane Detection, SPLane) for lane line detection, and output the lane line detection result. As shown in the appendix Figure 2 The image lane line detection model SPLane includes a backbone network, a strip perception feature fusion module (StripFeature Fusion, SFF), an attention mechanism module, an instance heat map module, a conditional convolution module and a prediction output module.
[0071] The backbone network is used to extract features from the lane line image to be detected, obtaining multi-level feature maps. Among them, the multi-level feature maps include low-level feature maps, mid-low-level feature maps, mid-high-level feature maps, and high-level feature maps. The strip perception feature fusion module SFF is used to exchange and fuse the features in the horizontal and vertical directions of the high-level feature map based on the strip geometric prior of the lane line, obtaining a strip perception fusion feature map. The attention mechanism module is used to obtain local and global information at the spatial level and perform attention fusion based on the multi-level feature maps and the strip perception fusion feature map, obtaining an attention fusion feature map. The instance heatmap module is used to mark the starting points of lane line instances on the lane line image to be detected based on the attention fusion feature map, obtaining an instance heatmap. The conditional convolution module is used to perform conditional convolution on the attention fusion feature map, generating a Gaussian mask feature map and an offset feature map. The prediction output module is used to obtain the lane line coordinate information in the lane line image to be detected according to the instance heatmap, the Gaussian mask feature map, and the offset feature map, and output the lane line detection result.
[0072] It should be noted that the process of inputting the lane line image to be detected into the pre-constructed image lane line detection model SPLane for lane line detection and outputting the lane line detection result is as follows:
[0073] Input the lane line image to be detected into the backbone network to achieve feature extraction, generating multi-level feature maps. Among them, the generated multi-level feature maps can capture the hierarchical features of the image at different feature extraction stages. Subsequently, input the high-level feature map into the strip perception feature fusion module SFF, and through promoting the exchange of feature information in the horizontal and vertical directions, effectively converge and integrate the overall context. Then, use the attention mechanism module to obtain local and global information at the spatial level and perform attention fusion based on the multi-level feature maps and the strip perception fusion feature map, obtaining an attention fusion feature map. Among them, a spatial cross-attention module (SCA) is introduced in the attention mechanism module to gradually refine the spatial and channel information of the high and low-level feature maps, model the lane line relationship between different spaces in the lane line image to be detected, obtain local and global information at the spatial level and perform attention fusion. After that, based on the attention fusion feature map, mark the starting points of lane line instances on the lane line image to be detected, obtaining an instance heatmap. Then, generate a Gaussian mask feature map and an offset feature map by performing conditional convolution on the attention fusion feature map. Finally, predict the lane line coordinate information in the lane line image to be detected according to the instance heatmap, the Gaussian mask feature map, and the offset feature map, and further output the lane line detection result.
[0074] In Embodiment 1, the backbone network includes a first downsampling layer, a second downsampling layer, a third downsampling layer, and a fourth downsampling layer cascaded in sequence; the first downsampling layer is used to downsample the lane line image to be detected to generate a low-level feature map; the second downsampling layer is used to downsample the low-level feature map to generate a mid-low-level feature map; the third downsampling layer is used to downsample the mid-low-level feature map to generate a mid-high-level feature map; the fourth downsampling layer is used to downsample the mid-high-level feature map to generate a high-level feature map; preferably, the high-level feature map is a high-level feature map obtained by downsampling the lane line image to be detected by 16 times.
[0075] In Embodiment 1, the bar perception feature fusion module SFF includes a first dilated convolution layer, a second dilated convolution layer, a third dilated convolution layer, an element addition module, a bar feature aggregation strategy module, a fusion module Concat, and a 1×1 convolution module, as shown in the appendix Figure 3 ; among them, the first dilated convolution layer, the second dilated convolution layer, and the third dilated convolution layer are three parallel dilated convolution layers to retain feature information while reducing the number of computational parameters; the first dilated convolution layer uses a 1×1 convolution kernel with a dilation rate of 1; the second dilated convolution layer uses a 3×3 convolution kernel with a dilation rate of 3; the third dilated convolution layer uses a 5×5 convolution kernel with a dilation rate of 5.
[0076] The first dilated convolution layer is used to perform a 1×1 dilated convolution operation on the high-level feature map and sequentially apply batch normalization and the ReLU activation function to obtain a feature map The second dilated convolution layer is used to perform a 3×3 dilated convolution operation on the high-level feature map and sequentially apply batch normalization and the ReLU activation function to obtain a feature map The third dilated convolution layer is used to perform a 5×5 dilated convolution operation on the high-level feature map and sequentially apply batch normalization and the ReLU activation function to obtain a feature map The element addition module is used to add the feature maps F 1 , the feature maps F 2 , and the feature maps F 3 element by element to obtain a feature map processed by dilated convolution; the bar feature aggregation strategy module is used to perform feature aggregation processing on the spatial domain of the high-level feature map in the horizontal and vertical directions to obtain the feature map F 4 ; the fusion module Concat is used to fuse the feature map processed by dilated convolution and the feature map F 4 to obtain a fused feature; the 1×1 convolution module is used to adjust the channel dimension of the fused feature to obtain a bar perception fusion feature map.
[0077] It should be noted that in the bar perception feature fusion module SFF, first, dilated convolution operations are performed using three convolutional kernels of sizes 1×1, 3×3, and 5×5, and batch normalization kernels and ReLU activation functions are applied in sequence to generate three feature maps of the same dimension; then, element-wise addition operations are performed on the three feature maps of the same dimension to merge multi-scale information and obtain the feature map processed by dilated convolution; in addition, using the bar feature aggregation strategy module, the spatial domain of the high-level feature map is processed in both horizontal and vertical directions to generate the feature map F 4 , to enhance the feature perception in the horizontal and vertical directions; subsequently, the feature map processed by dilated convolution and the feature map F 4 are fused to generate the fused feature; finally, to adapt to the scale requirements of the subsequent processing stage, the fused feature is adjusted in the channel dimension through a 1×1 convolutional kernel, that is, the bar perception fusion feature map is obtained.
[0078] It should also be noted that the lane line pixel points show an obvious bar-shaped distribution along the extension direction of the road, while they are relatively sparse in the direction perpendicular to the extension direction of the road; traditional global context perception modules often cannot fully exert their design advantages when dealing with tasks with obvious shape priors; for this reason, in view of the bar feature prior of the lane line, the bar perception feature fusion module based on the bar feature aggregation strategy module is introduced in Embodiment 1 of the present invention.
[0079] As shown in the appendix Figure 4 , the process of using the bar feature aggregation strategy module to perform feature aggregation processing on the spatial domain of the high-level feature map in the horizontal and vertical directions to obtain the feature map F 4 is as follows:
[0080] Average pooling processing is performed on the high-level feature map in the horizontal and vertical directions to reduce the size of the high-level feature map to H×1 and 1×W, that is, the average value of the pixel values in the pooling window is taken, and this average value is used as the feature representation after pooling, obtaining a feature map with a size of H×1 and a feature map with a size of 1×W.
[0081] Specifically, the process of performing average pooling processing on the high-level feature map in the horizontal and vertical directions is as follows:
[0082] First, use a 1×N pooling kernel to perform average pooling processing on the high-level feature map to obtain a feature map with a size of H×1; use an N×1 pooling kernel to perform average pooling processing on the high-level feature map to obtain a feature map with a size of 1×W; using 1×N and N×1 pooling kernels to perform average pooling processing on the high-level feature map can identify and associate regions scattered in the row and column directions, effectively capture long-distance bar dependencies in the image, and thus improve the prediction accuracy.
[0083] Next, perform a one-dimensional dilated convolution operation on the feature map with a size of H×1 to expand the size of the feature map with a size of H×1 in the horizontal direction to obtain an expanded feature map in the horizontal direction; perform a one-dimensional dilated convolution operation on the feature map with a size of 1×W to expand the size of the feature map with a size of 1×W in the vertical direction to obtain an expanded feature map in the vertical direction; wherein, through the one-dimensional dilated convolution operation, the output feature map after average pooling is expanded in size in the horizontal and vertical directions respectively, so that the expanded feature map in the horizontal direction and the expanded feature map in the vertical direction are restored to the same size as the high-level feature map.
[0084] Finally, add the expanded feature map in the horizontal direction and the expanded feature map in the vertical direction pixel by pixel to obtain a new feature map; after performing 1×1 convolution processing on the new feature map and multiplying it with the high-level feature map, obtain the feature map F 4 。
[0085] It should be noted that in the lane line detection task, a unique challenge is that lane lines usually have a slender bar-shaped geometric prior shape, making the space occupied by lane lines in the image relatively small, so it is difficult to be accurately recognized; the bar feature aggregation strategy module can enhance the interaction of row and column information, that is, improve the information interaction efficiency of lane line pixels in the horizontal and vertical directions, enabling the model to have global context awareness ability while further improving the detection accuracy by strengthening the feature information of rows and columns.
[0086] In this Embodiment 1, the attention mechanism module includes a first attention module, a second attention module, and a third attention module; the first attention module is used to generate a first-layer attention fusion feature map according to the middle-high level feature map and the bar perception fusion feature map; wherein, the first-layer attention fusion feature map is used to generate an instance heat map through the instance heat map module; the second attention module is used to generate a second-layer attention fusion feature map according to the middle-low level feature map and the first-layer attention fusion feature map; wherein, the second-layer attention fusion feature map is used to generate a Gaussian mask feature map through the conditional convolution module; the third attention module is used to generate a third-layer attention fusion feature map according to the low-level feature map and the second-layer attention fusion feature map; wherein, the third-layer attention fusion feature map is used to generate an offset feature map through the conditional convolution module.
[0087] The first attention module, the second attention module, and the third attention module all adopt the Spatial Cross Attention module (SCA); the Spatial Cross Attention module (SCA) is used to enhance the channel attention and spatial attention mechanisms for the shallow feature map and the deep feature map respectively, highlighting the model's attention to the spatial positions of detailed features in the shallow feature map and suppressing the interference of non-critical features, and better focusing on the overall position and feature information of the lane lines in the deep feature map, enabling the model to have the ability to infer the relative spatial positions of the lane lines, so as to improve the lane line detection performance of the model in complex scenarios.
[0088] As shown in the Figure 5 appendix, the working process of the Spatial Cross Attention module (SCA) is as follows:
[0089] Perform channel attention calculation and spatial attention calculation on the two input feature maps in sequence, and fuse the attention calculation results to obtain the corresponding attention-fused feature map; among them, when performing channel attention calculation, average pooling and max pooling operations are used simultaneously to aggregate the channel information of the feature map; when performing spatial attention calculation, it is used to obtain the spatial position information of the lane lines in the feature map.
[0090] Specifically, the implementation steps of the Spatial Cross Attention module (SCA) are as follows:
[0091] Use the first channel attention module to perform channel attention calculation on the input deep feature map Input high to generate the channel-level feature response result C high ; multiply the input shallow feature map Input low by the channel-level feature response result C high to generate the feature map F high ; use the first spatial attention module to perform spatial attention calculation on the feature map F high to generate the feature map S high ; multiply the input deep feature map Input high by the feature map S high to generate the feature map O high .
[0092] Use the second channel attention module to perform channel attention calculation on the input shallow feature map Input low to generate the channel-level feature response result C low ; multiply the input deep feature map Input high by the channel-level feature response result C low to generate the feature map F low ; use the second spatial attention module to perform spatial attention calculation on the feature map F low to generate the feature map Slow ; Multiply the input deep feature map Input low with the feature map S low to generate the feature map O low .
[0093] Add and blend the feature map O high and the feature map O low to generate the attention-blended feature map Output.
[0094] It should be noted that the input deep feature map Input high is the bar perception fusion feature map, the first-layer attention-blended feature map or the second-layer attention-blended feature map; the input shallow feature map Input low is the middle-high layer feature map, the middle-low layer feature map or the low layer feature map.
[0095] Specifically, in the first channel attention module or the second channel attention module, average pooling and max pooling operations are simultaneously used to aggregate the channel information of the feature map; as Figure 6 shown, given the size of the input feature map as After performing max pooling and average pooling operations on the given input feature map, two feature outputs with dimensions both being will be generated; then, convolutional operations with hidden layers are respectively performed on the generated features; where the activation size of the hidden layer is set to r is the downsampling ratio; preferably, r takes 1 or 2; finally, the features after the convolutional operation are output by element-wise summation.
[0096] In the first spatial attention module and the second spatial attention module, more attention is paid to the spatial position information of the lane lines in the image; as shown in the appendix Figure 7 shown, given the input feature First, perform max pooling operation and average pooling operation along the channel dimension to generate two features with dimensions of ; then, splice the two features with dimensions of to fuse the corresponding feature information to obtain a feature with a dimension of ; finally, the feature with a dimension of is output through a 1×1 convolutional kernel and Sigmoid operation.
[0097] It should be noted that the introduced spatial attention module enables the model to focus on specific spatial regions crucial for the lane detection task when processing images. Since lane lines usually exhibit linear features along the vertical direction in images and these features may appear at different spatial positions, the spatial attention module becomes an efficient tool to help the model distinguish and prioritize the processing of regions more likely to contain lane lines. As a result, the model can perform intelligent cropping when processing images, ignoring parts less likely to contain lane lines, such as the sky, buildings, and distant landscapes, and concentrating computational resources on the lane line markings on the road surface.
[0098] In addition, the spatial attention module can also assist in the detection of lane lines in complex scenarios. Specifically, when analyzing an image, if it is observed that specific regions on both sides of the image each present clearly visible lane markings and there is a large spatial span between these two regions, it can be reasonably inferred that there may also be a lane line in the middle area between these two regions. Based on this hypothesis, even if the lane line in the middle part may appear blurred due to certain reasons, such as shadows, wear, or other visual interferences, based on the reference of the clear lane lines on both sides, there is still a high confidence level to confirm its existence as a lane line.
[0099] Meanwhile, the introduced channel attention module also plays a key role in improving the performance of lane line detection. Among them, through in-depth analysis of the feature map, the channel attention can identify the key channels carrying information about the position, shape, and direction of the lane lines and assign higher weights, enabling the model to more accurately identify and locate lane lines in the face of complex and variable driving environments, such as different lighting conditions, color changes of lane lines, and possible occlusion situations.
[0100] In this Embodiment 1, the process of marking the starting point of the lane line instance on the lane line image to be detected based on the attention fusion feature map to obtain the instance heat map is specifically as follows: Based on the first-layer attention fusion feature map, mark the starting point of the lane line instance on the lane line image to be detected to obtain the instance heat map.
[0101] In this Embodiment 1, the conditional convolution module includes a first conditional convolution module and a second conditional convolution module. Use the first conditional convolution module to perform conditional convolution processing on the second-layer attention fusion feature map to generate a Gaussian mask feature map. Use the second conditional convolution module to perform conditional convolution processing on the third-layer attention fusion feature map to generate an offset feature map.
[0102] Experimental Results and Analysis:
[0103] (1) Experimental Dataset
[0104] In this Embodiment 1, in order to verify the performance of the image lane detection model SPLane, the widely used CULane image lane detection dataset is selected to conduct experimental verification.
[0105] (2) Comparison methods
[0106] In this Embodiment 1, in order to make an objective judgment on the effectiveness of the image lane detection model SPLane, advanced methods in the direction of image lane detection are selected for comparison, specifically including SCNN, ENet-SAD, RESA, LaneAF, UFLD, LaneATT, CondLaneNet, UFLDv2, LaneFormer, ADNet, PolyLaneNet, LSTR, and BezierLaneNet.
[0107] (3) Quantitative comparison experiments
[0108] On the CULane dataset, the quantitative results of the image lane detection model SPLane described in this Embodiment 1 are shown in Table 1 below; among them, the Total column in Table 1 below is a general indicator for measuring the F1 score of the model in various scenarios, and the higher the value, the better the performance of the model; in addition, FPS, as an indicator for evaluating the detection speed, reflects the ability of the model to process the number of frames per second, and the higher the value, the faster the processing speed.
[0109] It can be seen from Table 1 below that SPLane performs excellently in both accuracy and speed; specifically, in the Total column of the table, SPLane achieves the highest accuracy; in addition, in terms of the speed indicator FPS, SPLane ranks second, and this response speed can fully meet the real-time requirements of the image lane detection task. The above key indicator data proves the superiority of SPLane.
[0110] Table 1 Comparison table of the results of different models on the CULane dataset
[0111]
[0112] To further illustrate the performance of SPLane in different scenarios, Table 2 below presents the performance comparison of different models in each scenario of the CULane dataset. Among them, the F1 score is used as the performance evaluation index for eight scenarios, and a higher score means better performance. For the Cross scenario, the FP (False Positive) index is used for evaluation, and a lower value in the FP index indicates better performance. In most scenarios, SPLane shows further performance improvement compared to CondLaneNet. Especially in difficult scenarios such as Crowded, Dazzle, No line, and Night, SPLane shows the highest accuracy, with the F1 index increasing by 1.42%, 1.47%, 3.05%, and 1.59% respectively. Among them, the improvement in the No line scenario is the most significant.
[0113] Table 2 Comparison of results of different models in different scenarios in CULane
[0114]
[0115]
[0116] (4) Qualitative comparative analysis
[0117] As shown in the appendix Figure 8 shown, the appendix Figure 8 presents the visualization of the detection results of SPLane in four complex scenarios of the CULane dataset. Among them, the four complex scenarios include Dazzle, Crowded, Curve, and Night. In the visualization images of the appendix Figure 8 , the first column of images represents the original lane line image labels, the second column shows the detection results of CondLaneNet, and the last column represents the detection results of SPLane in this Example 1. The specific analysis is as follows:
[0118] Appendix Figure 8 (a) reveals the results of the qualitative analysis in the Dazzle scenario. It can be observed from the input image that in this high-light scenario, the lane line markings become discontinuous and unclear due to strong reflections. Relying on the bar perception feature fusion module in SPLane, even discontinuous bar lane line features can be perceived and enhanced. Therefore, the lane line detection results in this scenario are highly consistent with the actual annotations.
[0119] Appendix Figure 8 (b) and appendix Figure 8(d) presents the visualization results of lane line detection in Crowded and Night scenarios. In the crowded scenario, it can be found that lane lines are often covered by surrounding vehicles, and the exposed parts of the lane lines are small in size. Thanks to the perception of the spatial position information of the shallow feature map by the spatial cross-attention mechanism, and the capture of strip features and multi-scales by the strip perception feature fusion module, SPLane achieves better detection results than the baseline method in such complex scenarios; in the Night scenario, SPLane also demonstrates excellent performance, which stems from its efficient capture of lane line features in low-light environments.
[0120] Appendix Figure 8 (c) shows the visualization results in the Curve scenario. CondLaneNet fails to accurately detect the lane lines in the edge area, revealing its deficiency in accurately identifying curved lane lines. Compared with straight lane lines, curved lane lines are more difficult to identify due to their feature complexity, making the lane line detection task face greater challenges; however, the spatial cross-attention mechanism in SPLane can effectively perceive the spatial relationships in such complex scenarios, and the position and curvature of the edge lane lines can be inferred based on the shape of the identified lane lines and the relative spatial positions of the lane lines.
[0121] It can be clearly seen from the above results that the lane line detection method described in Embodiment 1 of the present invention can achieve excellent detection results in various complex scenarios such as Crowded, Night, Dazzle, and Curve, further confirming the robustness and high accuracy of the method in the lane line detection task.
[0122] When the lane line detection method described in Embodiment 1 of the present invention inputs the lane line image to be detected into the pre-constructed image lane line detection model SPLane for lane line detection, the strip perception feature fusion module SFF is used to enhance the key details in the horizontal and vertical directions of the feature map, and the attention mechanism module of the spatial cross-attention module SCA is introduced to effectively model the lane line relationships in different spaces and channels; the lane line detection method fully considers the geometric structure and spatial position relationship of the lane line strip, and through quantitative and qualitative experiments, it is verified that the method has advanced performance on the TuSimple and CULane datasets.
[0123] Embodiment 2
[0124] As shown in the appendix Figure 9As shown in the figure, Embodiment 2 of the present invention provides a lane line detection system, including an image acquisition module and a detection output module; the image acquisition module is used to acquire an image of a lane line to be detected; the detection output module is used to input the image of the lane line to be detected into a pre-constructed image lane line detection model SPLane for lane line detection, and output a lane line detection result.
[0125] In Embodiment 2 of the present invention, the image lane line detection model SPLane includes a backbone network, a strip perception feature fusion module SFF, an attention mechanism module, an instance heat map module, a conditional convolution module, and a prediction output module.
[0126] The backbone network is used to extract features from the image of the lane line to be detected to obtain a multi-level feature map; the strip perception feature fusion module SFF is used to exchange and fuse the features in the horizontal and vertical directions of the high-level feature map in the multi-level feature map based on the strip geometric prior of the lane line to obtain a strip perception fusion feature map; the attention mechanism module is used to obtain local and global information at the spatial level based on the multi-level feature map and the strip perception fusion feature map and perform attention fusion to obtain an attention fusion feature map; the instance heat map module is used to mark the starting points of lane line instances on the image of the lane line to be detected based on the attention fusion feature map to obtain an instance heat map; the conditional convolution module is used to perform conditional convolution on the attention fusion feature map to generate a Gaussian mask feature map and an offset feature map; the prediction output module is used to obtain the lane line coordinate information in the image of the lane line to be detected according to the instance heat map, the Gaussian mask feature map, and the offset feature map, and output a lane line detection result.
[0127] Embodiment 3
[0128] As shown in the appendix Figure 10 As shown in the figure, Embodiment 3 of the present invention provides a lane line detection device, including: a memory for storing a computer program; a processor for implementing the steps of the lane line detection method when executing the computer program. Alternatively, when the processor executes the computer program, it implements the functions of each module in the above lane line detection system.
[0129] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing preset functions, and the instruction segments are used to describe the execution process of the computer program in the lane line detection device.
[0130] The lane line detection device may be a computing device such as a desktop computer, notebook, palm computer, and cloud server. The lane line detection device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above are examples of the lane line detection device, which do not constitute a limitation on the lane line detection device. It may include more components than the above, or combine certain components, or different components. For example, the lane line detection device may also include input / output devices, network access devices, buses, etc.
[0131] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the lane line detection device, and connects various parts of the entire lane line detection device through various interfaces and lines.
[0132] The memory can be used to store the computer program and / or module. The processor realizes various functions of the lane line detection device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory.
[0133] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0134] Embodiment 4
[0135] Embodiment 4 further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the lane line detection method described above.
[0136] If the modules / units integrated in the lane line detection system are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0137] Based on such understanding, all or part of the processes in the above-described lane line detection method of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described lane line detection method can be implemented. Among them, the computer program includes computer program code, which can be in the form of source code, object code, executable file, or a preset intermediate form, etc.
[0138] The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0139] It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0140] Embodiment 5
[0141] Embodiment 5 provides a computer product. The computer program product includes a computer program stored in a computer-readable storage medium. The processor of the lane line detection device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the lane line detection device can execute the lane line detection method described in Embodiment 1, which will not be elaborated here.
[0142] It should be noted that those of ordinary skill in the art can understand that all or part of the processes of implementing the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments.
[0143] The lane line detection method described in the present invention realizes advanced detection accuracy while meeting the real-time detection requirements by introducing a strip perception feature fusion module SFF and a spatial cross-attention module SCA into the image lane line detection model SPLane; this method has achieved excellent performance when dealing with challenging scenarios such as nighttime, missing lane markings, and congestion; specifically, the strip perception feature fusion module SFF is used to strengthen the information exchange between rows and columns, and at the same time realize the extraction and effective combination of global context features, thereby effectively coping with the challenges of lane line detection in complex scenarios; in the spatial cross-attention module SCA, by utilizing the interdependent relationship between different regions and channels in the lane image, the spatial position connection of the lane line in the shallow and deep feature maps is captured, effectively enhancing the model's spatial position perception ability, so as to endow the model with the ability to infer the relative position of the lane line.
[0144] The above embodiments are only one of the implementation manners capable of implementing the technical solution of the present invention. The scope of protection required by the present invention is not limited to this embodiment only, but also includes any changes, substitutions, and other implementation manners that are easily conceivable by those skilled in the art within the technical scope disclosed by the present invention.
Claims
1. A lane line detection method, characterized in that: include: Obtain the lane line image to be detected; Input the lane line image to be detected into a pre-built image lane line detection model SPLane for lane line detection, and output the lane line detection result; wherein the image lane line detection model SPLane includes a backbone network, a stripe perception feature fusion module SFF, an attention mechanism module, an instance heat map module, a conditional convolution module and a prediction output module; The backbone network is used to extract features from the lane line image to be detected to obtain a multi-level feature map; The strip-perception feature fusion module SFF is used to exchange and fuse the features of the high-level feature map in the multi-level feature map in the horizontal and vertical directions based on the strip geometry prior of the lane line to obtain a strip-perception fusion feature map; The attention mechanism module is used to obtain local and global information at the spatial level and perform attention fusion based on the multi-level feature map and the strip perception fusion feature map to obtain an attention fusion feature map; The instance heat map module is used to identify the starting point of the lane line instance on the lane line image to be detected based on the attention fusion feature map to obtain an instance heat map; The conditional convolution module is used to perform conditional convolution on the attention fusion feature map to generate a Gaussian mask feature map and an offset feature map; The prediction output module is used to obtain the lane line coordinate information in the lane line image to be detected according to the instance heat map, the Gaussian mask feature map and the offset feature map, and output the lane line detection result.
2. A lane line detection method according to claim 1, characterized in that: The strip perception feature fusion module SFF includes a first hole convolution layer, a second hole convolution layer, a third hole convolution layer, an element superposition module, a strip feature aggregation strategy module, a fusion module Concat and a 1×1 convolution module; The first atrous convolution layer is used to perform a 1×1 atrous convolution operation on the high-level feature map, and sequentially apply batch normalization and ReLU activation function to obtain a feature map F1; The second atrous convolution layer is used to perform a 3×3 atrous convolution operation on the high-level feature map, and apply batch normalization and ReLU activation function in sequence to obtain a feature map F2; The third atrous convolution layer is used to perform a 5×5 atrous convolution operation on the high-level feature map, and sequentially apply batch normalization and ReLU activation function to obtain a feature map F3; The element superposition module is used to perform element-wise addition on the feature map F1, the feature map F2 and the feature map F3 to obtain a feature map processed by the dilated convolution; The strip feature aggregation strategy module is used to perform feature aggregation processing in the horizontal and vertical directions on the spatial domain of the high-level feature map to obtain a feature map F4; The fusion module Concat is used to fuse the feature map processed by the dilated convolution with the feature map F4 to obtain the fused features; The 1×1 convolution module is used to adjust the channel dimension of the fused features to obtain a strip-perceived fusion feature map.
3. A lane line detection method according to claim 2, characterized in that: The process of performing feature aggregation processing in the horizontal and vertical directions on the spatial domain of the high-level feature map to obtain the feature map F4 is as follows: Performing average pooling processing on the high-level feature map in horizontal and vertical directions to obtain a feature map with a size of H×1 and a feature map with a size of 1×W; A one-dimensional dilated convolution operation is performed on the feature map of size H×1 to expand the feature map of size H×1 in the horizontal direction to obtain a horizontally expanded feature map; A one-dimensional dilated convolution operation is performed on the feature map with a size of 1×W to expand the size of the feature map with a size of 1×W in the vertical direction to obtain a vertically expanded feature map; Adding the horizontally expanded feature map and the vertically expanded feature map pixel by pixel to obtain a new feature map; After performing 1×1 convolution processing on the new feature map, it is multiplied with the high-level feature map to obtain a feature map F4.
4. A lane line detection method according to claim 1, characterized in that: The multi-level feature map also includes a low-level feature map, a mid-low-level feature map and a mid-high-level feature map; the low-level feature map, the mid-low-level feature map and the mid-high-level feature map are all input into the attention mechanism module to obtain an attention fusion feature map.
5. A lane line detection method according to claim 4, characterized in that: The attention mechanism module includes a first attention module, a second attention module and a third attention module; The first attention module is used to generate a first-layer attention fusion feature map according to the mid- and high-level feature maps and the strip-perception fusion feature map; wherein the first-layer attention fusion feature map is used to generate an instance heat map through an instance heat map module; The second attention module is used to generate a second-layer attention fusion feature map according to the middle and low-layer feature maps and the first-layer attention fusion feature map; wherein the second-layer attention fusion feature map is used to generate a Gaussian mask feature map through a conditional convolution module; The third attention module is used to generate a third-layer attention fusion feature map based on the low-layer feature map and the second-layer attention fusion feature map; wherein the third-layer attention fusion feature map is used to generate an offset feature map through a conditional convolution module.
6. A lane line detection method according to claim 5, characterized in that: The first attention module, the second attention module and the third attention module all adopt a spatial cross attention module SCA; The working process of the spatial cross attention module SCA is as follows: Channel attention calculation and spatial attention calculation are performed on the two input feature maps in turn, and the attention calculation results are fused to obtain the corresponding attention fusion feature map; when performing channel attention calculation, average pooling and maximum pooling operations are used simultaneously to aggregate the channel information of the feature map; when performing spatial attention calculation, it is used to obtain the spatial position information of the lane lines in the feature map.
7. A lane detection system, characterized in that: include: An image acquisition module, used to acquire the lane line image to be detected; A detection output module, used to input the lane line image to be detected into a pre-built image lane line detection model SPLane for lane line detection, and output a lane line detection result; wherein the image lane line detection model SPLane includes a backbone network, a stripe perception feature fusion module SFF, an attention mechanism module, an instance heat map module, a conditional convolution module and a prediction output module; The backbone network is used to extract features from the lane line image to be detected to obtain a multi-level feature map; The strip-perception feature fusion module SFF is used to exchange and fuse the features of the high-level feature map in the multi-level feature map in the horizontal and vertical directions based on the strip geometry prior of the lane line to obtain a strip-perception fusion feature map; The attention mechanism module is used to obtain local and global information at the spatial level and perform attention fusion based on the multi-level feature map and the strip perception fusion feature map to obtain an attention fusion feature map; The instance heat map module is used to identify the starting point of the lane line instance on the lane line image to be detected based on the attention fusion feature map to obtain an instance heat map; The conditional convolution module is used to perform conditional convolution on the attention fusion feature map to generate a Gaussian mask feature map and an offset feature map; The prediction output module is used to obtain the lane line coordinate information in the lane line image to be detected according to the instance heat map, the Gaussian mask feature map and the offset feature map, and output the lane line detection result.
8. A lane line detection device, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the lane line detection method according to any one of claims 1 to 6 is executed.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the lane line detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the lane line detection method according to any one of claims 1 to 6 is implemented.