Laser radar lane line detection method using cyclic shift grouping convolution
By adopting cyclic shift packet convolution technology in the lidar lane line detection network, the problem of high computational complexity is solved, and higher detection accuracy and faster inference speed are achieved. It is suitable for low computing power platforms and meets real-time requirements.
Patent Information
- Application Number
- CN202510149187.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
Due to the large bird's-eye view size, the existing lidar lane line detection network has high computing complexity, making it difficult to meet the requirements of real-time on low-computing power platforms.
The cyclic shift grouping convolution technology is adopted, and the cross-stage cyclic shift grouping convolution module is designed through the grouping cyclic shift operator and cross-stage partial network strategy. Combined with the decoupled detection head and hyperparameter adjustment, the model is optimized to improve detection accuracy and inference speed.
It effectively reduces the computational complexity, improves detection accuracy and inference speed, and is suitable for low computing power platforms and meets real-time requirements.
Smart Images

Figure CN120071284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned driving environment perception, and particularly to a lidar lane line detection method using cyclic shift grouped convolution. Background Art
[0002] In modern transportation systems, the accurate detection of lane lines is crucial for driving assistance systems and autonomous driving technologies. With the continuous increase in traffic flow and the increasing complexity of urban roads, an effective lane line detection system has become a key component to ensure road safety and enhance the driving experience. Traditional lane line detection methods are often affected by changes in lighting, weather, and road signs, resulting in unstable detection performance. Therefore, seeking a highly accurate and adaptable lane line detection technology is an urgent need in the current transportation field.
[0003] In recent years, with the rapid development of computer vision and deep learning technologies, significant progress has also been made in the field of lane line detection. Deep learning-based methods can learn complex feature representations from large amounts of data, enabling lane line detection systems to better adapt to various complex driving scenarios. At the same time, by using high-resolution sensors and advanced image processing algorithms, lane line detection systems can perform well in low-light, adverse weather, and blurred environments.
[0004] However, the performance of camera-based lane line detection methods will significantly decline at night or under glare conditions. In the path planning of autonomous vehicles, it is necessary to convert the lane line detection results of images into a two-dimensional bird's-eye view, which will reduce the detection accuracy. Since lidar obtains environmental information through infrared signals, it has stronger robustness to changes in light and weather, so it is more suitable for dealing with problems that cannot be solved in camera-based lane line detection. In addition, compared with cameras, lidar-based detection methods do not require projection and can directly obtain the detection results under the bird's-eye view. Therefore, lidar lane line detection methods have received extensive attention in recent years and have gradually begun to be applied in practice. However, in current lidar lane line detection networks, the size of the bird's-eye view is relatively large, resulting in a high computational burden and making it difficult to meet the real-time requirements on low-computing-power platforms. Summary of the Invention
[0005] The purpose of the present invention is to provide a lidar lane line detection method using cyclic shift grouped convolution, which solves the problem of high computational complexity caused by the large size of the bird's-eye view and avoids the problem of insufficient information exchange of feature channels in grouped convolution, and has important theoretical significance and application value.
[0006] To achieve the above purpose, the present invention provides a lidar lane line detection method using cyclic shift grouped convolution, including the following steps:
[0007] Step S1: Project the point cloud to form a bird's-eye view based on the bird's-eye view encoder;
[0008] Step S2: Propose a grouped cyclic shift operator and build a cyclic shift grouped convolution therefrom;
[0009] Step S3: Design a cross-stage cyclic shift grouped convolution module based on the cross-stage partial network strategy;
[0010] Step S4: Use a decoupled detection head to predict the confidence and category of the lane line respectively;
[0011] Step S5: Introduce different hyperparameters to adjust the model so that the model achieves a good balance between detection accuracy and inference speed.
[0012] Preferably, in Step S1, projecting the point cloud to form a bird's-eye view based on the bird's-eye view encoder, the specific process is as follows:
[0013] Step S11: Use a point-based projection method to divide the point cloud in the lidar detection range into a number of cylinders. The points falling into the same cylinder belong to the same grid of the bird's-eye view, and project the point cloud onto the x-y plane to form a bird's-eye view; normalize the features including the z-coordinate value, intensity, and reflectivity in the following manner:
[0014]
[0015] where f ∈ {z-coordinate value, intensity, reflectivity}; U f and L f represent the upper and lower limits of their features respectively;
[0016] Step S12: For each feature channel of the bird's-eye view, use the average value of the corresponding features of all the points projected onto each grid to represent the value of the grid, thereby obtaining a feature map M f ;
[0017] Step S13: Generate a new feature map by normalizing the number of projected points in each grid; combine this new feature map with the existing feature map to form a bird's-eye view of size H BEV ×W BEV ×C BEV ;
[0018] Step S14: Use multiple reparameterization modules to adjust the size of the bird's-eye view and the number of feature channels, thereby obtaining the input feature map.
[0019] Preferably, in step S2, a grouped cyclic shift operator is proposed, and a cyclic shift grouped convolution is built therefrom; wherein the grouped cyclic shift operator is the core of the cyclic shift grouped convolution and is used to solve the problem of insufficient information exchange between the feature channels of different groups.
[0020] Preferably, the grouped cyclic shift operator includes three parts: cross-grouping, cyclic shift, and pooling;
[0021] When the feature map is processed by multiple grouped convolutions, the feature map generated by the previous grouped convolution is divided into different groups; within each group, the feature channels are cyclically shifted; finally, the feature channels from different groups are pooled together to form a new feature map.
[0022] Preferably, the cyclic shift grouped convolution is composed of a series connection of a grouped cyclic shift operator and a reparameterization module;
[0023] The reparameterization module consists of three branches; the first branch is composed of a series connection of a 1×1 dilated grouped convolution and a batch normalization layer; the second branch is composed of a series connection of a 3×3 dilated grouped convolution with the same dilation rate and number of groups as the first branch; the third branch consists of a batch normalization layer.
[0024] Preferably, in step S3, based on the cross-stage partial network strategy, a cross-stage cyclic shift grouped convolution module is designed;
[0025] The cross-stage cyclic shift grouped convolution module consists of two branches; one branch is composed of a series connection of multiple cyclic shift grouped convolutions, and the other branch consists of a batch normalization layer.
[0026] Preferably, the cross-stage cyclic shift grouped convolution module divides the input feature map into two parts along the channel dimension; the first half of the feature map is processed by multiple cyclic shift grouped convolutions, and the second half is transformed by batch normalization; finally, the outputs of these two branches are concatenated together, and a 1×1 convolution is used to generate the extracted features.
[0027] Preferably, in step S4, lane line detection is regarded as a multi-class segmentation problem, that is, a confidence score and a class label are assigned to each pixel of the bird's-eye view, and a decoupled detection head is used to predict the confidence and class of the lane line respectively;
[0028] The confidence score indicates whether a pixel belongs to a lane line; the class label reflects the lane line to which the pixel belongs;
[0029] The detection head consists of two decoupled convolutional layers, which are used to predict the confidence score and class label of the lane line respectively.
[0030] Preferably, the input feature map is scaled in terms of the number of feature channels through a reparameterization module, and then the scaled feature map is fed into a detection head containing two decoupled convolutional layers for processing.
[0031] Preferably, in step S5, a hyperparameter r is introduced c for scaling the number of output feature channels, where r c affects the computational complexity of the model;
[0032] A hyperparameter r is introduced g for scaling the number of groups in grouped convolution, where r g affects the information fusion between feature channels.
[0033] Therefore, the present invention adopts the above-mentioned method for lidar lane line detection using cyclic shift grouped convolution, and proposes a new grouped cyclic shift operator. Compared with the traditional operator that promotes information exchange between sub-group feature channels in grouped convolution, the grouped cyclic shift operator can make information fusion more sufficient and has a faster running speed. On this basis, a lidar lane line detection network using cyclic shift grouped convolution is built, which has higher detection accuracy and faster inference speed, greatly improving the practical application ability.
[0034] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0035] Figure 1 is a flowchart of a method for lidar lane line detection using cyclic shift grouped convolution according to the present invention;
[0036] Figure 2 is a schematic diagram of the process of generating an input feature map by a bird's-eye view encoder of point cloud according to the present invention;
[0037] Figure 3 is a schematic diagram of the grouped cyclic shift operator according to the present invention;
[0038] Figure 4 is a schematic diagram of cyclic shift grouped convolution according to the present invention;
[0039] Figure 5 is a schematic diagram of a cross-stage cyclic shift grouped convolution module according to the present invention;
[0040] Figure 6 is a schematic diagram of the decoupled detection head according to the present invention;
[0041] Figure 7 is the comprehensive comparison result of the method proposed by the present invention and the existing lidar lane line detection methods. Detailed Embodiments
[0042] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0043] As Figure 1 shown, a lidar lane line detection method using cyclic shift grouped convolution includes the following steps:
[0044] Step S1: Project the point cloud to form a bird's-eye view based on the bird's-eye view encoder;
[0045] Step S2: Propose a grouped cyclic shift operator and build a cyclic shift grouped convolution therefrom;
[0046] Step S3: Design a cross-stage cyclic shift grouped convolution module based on the cross-stage partial network strategy;
[0047] Step S4: Use a decoupled detection head to predict the confidence and category of the lane line respectively;
[0048] Step S5: Introduce different hyperparameters to adjust the model so that the model achieves a good balance between detection accuracy and inference speed.
[0049] Embodiment
[0050] Based on a lidar lane line detection method using cyclic shift grouped convolution provided by the present invention, this embodiment is tested and verified on the K-Lane dataset. K-Lane contains 15,382 frames, of which 7,687 frames are used for training and 7,695 frames are used for testing. Each dataset covers diverse road conditions and complex scenarios, including changes in lighting conditions, traffic density, and road merges and bifurcations.
[0051] Step S1: Project the point cloud to form a bird's-eye view based on the bird's-eye view encoder; wherein, the process of the bird's-eye view encoder projecting the point cloud to form a bird's-eye view is as Figure 2 shown.
[0052] Step S11: Use a point-based projection method to divide the point cloud in the lidar detection range into a number of cylinders, and regard the points falling into the same cylinder as belonging to the same grid of the bird's-eye view, and then project the point cloud onto the x-y plane to form a bird's-eye view. Specifically, the features including the z-coordinate value, intensity, and reflectivity are normalized in the following manner:
[0053]
[0054] wherein, f ∈ {z-coordinate value, intensity, reflectivity}; U f and L f respectively represent the upper and lower limits of its features.
[0055] Step S12: For each feature channel of the bird's-eye view, use the average value of the features corresponding to all points projected onto each grid to represent the value of the grid, thereby obtaining a feature map M f .
[0056] Step S13: In addition, by normalizing the number of projected points in each grid, a new feature map is generated. This new feature map is combined with the existing feature map to form a bird's-eye view with a size of H BEV ×W BEV ×C BEV .
[0057] Since the point cloud is usually sparse in the lane lines, it is crucial to use a high-resolution bird's-eye view. Therefore, the resolution of the bird's-eye view is set to 8 times the size of the label map
[0058] Step S14: Use multiple reparameterization modules to adjust the size of the bird's-eye view and the number of feature channels, thereby obtaining the input feature map
[0059] Step S2: Propose a grouped cyclic shift operator and construct cyclic shift grouped convolution therefrom
[0060] The grouped cyclic shift operator is the core of the cyclic shift grouped convolution, which is used to solve the problem of insufficient information exchange between the feature channels of different groups. As Figure 3 shown, the grouped cyclic shift operator consists of three parts: cross-grouping, cyclic shift, and pooling. When the feature map is processed by multiple grouped convolutions, the feature map generated by the previous grouped convolution is divided into different groups. Within each group, the feature channels are cyclically shifted. Finally, the feature channels from different groups are pooled together to form a new feature map
[0061] As Figure 4 shown, the cyclic shift grouped convolution consists of a series connection of a grouped cyclic shift operator and a reparameterization module. The reparameterization module consists of three branches. The first branch consists of a series connection of a 1×1 dilated grouped convolution and a batch normalization layer. The second branch consists of a series connection of a 3×3 dilated grouped convolution with the same dilation rate and number of groups as the first branch. The third branch consists of a batch normalization layer
[0062] Step S3: Based on the cross-stage partial network strategy, design a cross-stage cyclic shift grouped convolution module
[0063] The cross-stage cyclic shift grouped convolution module is as Figure 5 shown. The cross-stage cyclic shift grouped convolution module consists of two branches; one branch is composed of a series connection of multiple cyclic shift grouped convolutions, and the other branch is composed of a batch normalization layer
[0064] The cross-stage cyclic shift grouped convolution module divides the input feature map into two parts along the channel dimension. The first half of the feature map is processed through multiple cyclic shift grouped convolutions, and the second half is transformed using batch normalization. The outputs of these two branches are concatenated together, and then 1×1 convolution is used to generate the extracted features.
[0065] Step S4: Use decoupled detection heads to predict the confidence and class of the lane lines respectively.
[0066] Lane line detection is regarded as a multi-class segmentation problem, that is, a confidence score and a class label are assigned to each pixel in the bird's-eye view. The confidence score indicates whether a pixel belongs to a lane line, and the class label reflects which lane line it belongs to. The detection head consists of two decoupled convolutional layers, which are used to predict the confidence score and class label of the lane line respectively. The input feature map is scaled in terms of the number of feature channels through a reparameterization module, and then the scaled feature map is fed into the subsequent two decoupled convolutional layers.
[0067] As Figure 6 shown, the input feature map of size H head ×W head ×C head passes through the reparameterization module to scale the number of feature channels, obtaining a new feature map of size H head ×W head ×C 1 Then, this feature map is processed by a detection head containing two decoupled convolutional layers. It can be seen that its upper layer is designed to predict the confidence, and the lower layer is responsible for determining the class label.
[0068] Step S5: Introduce different hyperparameters to adjust the model, so that the model achieves a good balance between detection accuracy and inference speed.
[0069] Introduce the hyperparameter r c to scale the number of output feature channels, and r c affects the computational complexity of the model.
[0070] Introduce the hyperparameter r g to scale the number of groups in the grouped convolution, and r g affects the information fusion between feature channels.
[0071] As shown in Table 1 and Figure 7 shown, a comprehensive comparison is made quantitatively and qualitatively between the method proposed in the present invention and the current leading lidar lane line detection methods. The accuracy of the detection algorithm and the model inference speed are mainly compared under different lighting conditions and different road conditions.
[0072] As shown in Table 1, the average F1-score of the method proposed in the present invention is increased by 1.78% and 1.16% respectively compared with the LLDN-GFC lidar lane line detection method and the LLDN-RW lidar lane line detection method. Compared with LLDN-RW, the method proposed in the present invention reduces the floating-point operation amount by 39% and increases the number of frames processed per second by about 2.5 times.
[0073] It can be Figure 7 observed that the method proposed in the present invention not only has fewer false detections and smoother prediction results in most scenarios, but also performs better when dealing with long-distance or occluded lane lines. Therefore, the method proposed in the present invention has better detection effect and faster model inference speed.
[0074] Table 1 Comprehensive comparison results of the present invention and existing lidar lane line detection methods
[0075]
[0076] Therefore, the present invention adopts the above-mentioned lidar lane line detection method using cyclic shift grouped convolution, and proposes a new grouped cyclic shift operator. Compared with the traditional operator that promotes information exchange between sub-group features of grouped convolution, the grouped cyclic shift operator can make information fusion more sufficient and has a faster running speed. On this basis, a lidar lane line detection network using cyclic shift grouped convolution is built, which has higher detection accuracy and faster inference speed. The present invention not only solves the problem of high computational complexity of the lidar lane line detection network, obtains a higher frame rate, but also significantly improves the model detection effect, and has clear theoretical significance and important application value.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A laser radar lane detection method using cyclic shift group convolution, characterized in that: The following steps are involved: Step S1, projecting the point cloud into a bird's-eye view based on a bird's-eye view encoder; Step S2: Propose a grouped cyclic shift operator, and construct a cyclic shift grouped convolution based on it; Step S3: Based on the cross-stage partial network strategy, a cross-stage cyclic shift grouped convolution module is designed; Step S4, using the decoupled detection head to respectively predict the confidence and category of the lane line; Step S5: Introduce different hyperparameters to adjust the model so that the model achieves a good balance between detection accuracy and inference speed.
2. The laser radar lane line detection method using cyclic shift group convolution according to claim 1, characterized in that: In step S1, the point cloud is projected to form a bird's-eye view based on the bird's-eye view encoder. The specific process is as follows: Step S11: Use a point-based projection method to divide the point cloud of the laser radar detection range into several cylinders. Points falling into the same cylinder belong to the same grid of the bird's-eye view. Project the point cloud to the xy plane to form a bird's-eye view. Normalize the features including the z-coordinate value, intensity and reflectivity by the following method: Where f∈{z-coordinate value, intensity, reflectivity}; U f and L f Represent the upper and lower limits of its characteristics respectively; Step S12: For each feature channel of the bird's-eye view, the average value of the corresponding features of all points projected to each grid is used to represent the value of the grid, thereby obtaining a feature map M f ; Step S13: Generate a new feature map by normalizing the number of projection points in each grid; combine this new feature map with the existing feature map to form a feature map of size H BEV ×W BEV ×C BEV Aerial view of Step S14: Use multiple re-parameterization modules to adjust the size of the bird's-eye view and the number of feature channels to obtain an input feature map.
3. The laser radar lane detection method using cyclic shift group convolution according to claim 1, characterized in that: In step S2, a grouped circular shift operator is proposed, and a circular shift grouped convolution is constructed based on it; wherein, the grouped circular shift operator is the core of the circular shift grouped convolution, and is used to solve the problem of insufficient information exchange between feature channels of different groups.
4. The laser radar lane detection method using cyclic shift grouped convolution according to claim 3, characterized in that: The group cyclic shift operator includes three parts: cross grouping, cyclic shift and aggregation; When a feature map is processed by multiple grouped convolutions, the feature map generated by the previous grouped convolution is divided into different groups; within each group, the feature channels are cyclically shifted; finally, the feature channels from different groups are pooled together to form a new feature map.
5. The laser radar lane detection method using cyclic shift grouped convolution according to claim 3, characterized in that: The cyclic shift group convolution is composed of a group cyclic shift operator and a reparameterization module in series; The reparameterized module consists of three branches; the first branch consists of a 1×1 dilated grouped convolution and a batch normalization layer in series; the second branch consists of a 3×3 dilated grouped convolution in series with the same dilation rate and number of groups as the first branch; the third branch consists of a batch normalization layer.
6. The laser radar lane detection method using cyclic shift group convolution according to claim 1, characterized in that: In step S3, based on the cross-stage partial network strategy, a cross-stage cyclic shift grouped convolution module is designed; The cross-stage cyclic shift grouped convolution module consists of two branches; One branch is composed of multiple cyclic shift grouped convolutions in series, and the other branch is composed of batch normalization layers.
7. The laser radar lane detection method using cyclic shift group convolution according to claim 6, characterized in that: The cross-stage circular shift grouped convolution module divides the input feature map into two parts along the channel dimension; the first half of the feature map is processed by multiple circular shift grouped convolutions, and the second half is transformed by batch normalization; finally, the outputs of the two branches are spliced together and the extracted features are generated by 1×1 convolution.
8. The laser radar lane detection method using cyclic shift group convolution according to claim 1, characterized in that: In step S4, lane detection is regarded as a multi-class segmentation problem, that is, a confidence score and a class label are assigned to each pixel in the bird's-eye view image, and the confidence and class of the lane are predicted separately using the decoupled detection head; The confidence score indicates whether a pixel belongs to a lane line; the class label reflects the lane line to which the pixel belongs; The detection head consists of two decoupled convolutional layers, which are used to predict the confidence score and class label of the lane line respectively.
9. The laser radar lane detection method using cyclic shift group convolution according to claim 8, characterized in that: The input feature map is scaled by a reparameterization module to increase the number of feature channels, and then the scaled feature map is sent to a detection head consisting of two decoupled convolutional layers for processing.
10. The laser radar lane detection method using cyclic shift grouped convolution according to claim 1, characterized in that: In step S5, the hyperparameter r is introduced c Used to scale the number of output feature channels, r c Affects the computational complexity of the model; Introducing hyperparameter r g Used to scale the number of groups in grouped convolution, r g Affects the information fusion between feature channels.