Lane line detection method based on lane line aggregation perception and starting point prediction

By adopting a multi-dimensional feature refinement method based on starting point guidance in lane line detection, the geometric feature capture and accuracy problems of lane line detection in complex scenarios are solved, and more efficient lane line detection performance is achieved.

CN120147983APending Publication Date: 2025-06-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411447939.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In complex scenarios, existing deep learning methods are difficult to capture the complete geometric features of lane lines, and local feature quality affects detection accuracy, making it impossible to effectively deal with obscured lane lines and irregularly shaped lane lines.

Method used

The multi-dimensional feature refinement method based on starting point guidance is adopted, and features are extracted and optimized through the global feature optimization module and the starting point coordinate prediction module, and local feature refinement is used to punish the lane line perception aggregation module, and the lane prior is optimized by punishing the lane line cross-transfer and comparison loss function.

Benefits of technology

The lane line detection performance in complex scenarios is improved, the geometric feature capture capability of lane lines is enhanced, and the accuracy and robustness of detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147983A_ABST
    Figure CN120147983A_ABST
Patent Text Reader

Abstract

The invention discloses a lane line detection method based on lane line aggregation perception and starting point prediction, and the method comprises the following steps: firstly, designing a global feature optimization module and a lane line perception aggregation module, and carrying out the optimization of features from the two dimensions of global enhancement and local refinement; then, a starting point coordinate prediction module is designed to predict starting point coordinates of a lane line so as to guide generation of lane priori. And finally, designing a more universal penalty lane line intersection-to-union ratio as a loss function to evaluate a prediction result. A large number of experimental results show that the method provided by the invention has good performance in complex scene lane line detection tasks, and has strong competitiveness in the existing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and machine learning, and particularly relates to a lane line detection method in complex scenarios. Background Art

[0002] To ensure the safety of a vehicle during driving, an autonomous driving system must accurately locate the vehicle's position and navigate the vehicle along lane lines, which means that accurately perceiving lane lines is crucial. Lane line detection plays an important role in autonomous driving systems, especially in advanced driver assistance systems. Lane line detection aims to analyze two-dimensional images captured by on-vehicle cameras and accurately extract information on each lane line. Due to the slender shape of lane lines and complex road conditions, lane line detection faces significant difficulties. In addition, lane line detection algorithms require real-time data processing, which poses high requirements for the real-time performance of the algorithms. Traditional lane line detection methods first use manually designed operators to extract features, and then use techniques such as Hough transform and random sample consensus to model lane lines. The emergence of convolutional neural networks has stimulated the rapid development of lane detection, improving the efficiency and robustness of the algorithms.

[0003] However, although existing deep learning methods have made significant progress in terms of performance and speed, there is still room for optimization in complex scenarios.

[0004] In complex scenarios, some lane lines are completely occluded, and the detection network needs to have a deep understanding of context information. In addition, due to the slender and continuous nature of lane lines, it is difficult for existing deep learning methods to capture the complete geometric features of lane lines. Secondly, the quality of local features of lane lines also affects the accuracy of lane line detection. The lane lines obtained using existing deep learning methods may not be optimal because these methods perform detection on a fixed grid and may not be well aligned with lane lines of irregular shapes. Summary of the Invention

[0005] The present invention provides a complex scene lane line detection method based on starting point-guided multi-dimensional feature refinement, which is used to improve the lane line detection performance in complex scenes. This method includes: 1) Using the proposed global feature optimization module to expand the receptive field of the backbone network and extract image features and send them into the feature pyramid; 2) Then using the starting point coordinate prediction module to process the topmost feature layer of the feature pyramid, predicting the starting point coordinates and generating lane priors to form a feature representation containing lane priors; 3) Then using the lane line perception aggregation module to refine the local features of the output of the previous step, aggregate them with the global features, and then use the penalty lane line intersection over union loss function for supervision, optimize the lane priors, and pass the optimization results to the next layer of features in the feature pyramid, and re-use the lane line perception aggregation module for optimization. This process is repeated until all outputs of the feature pyramid are fully optimized. The lane priors of the last layer are the lane lines predicted by the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is the framework diagram of the method proposed in this method;

[0007] Figure 2 is the structural diagram of the global feature optimization module;

[0008] Figure 3 is the structural diagram of the lane line perception aggregation module;

[0009] Figure 4 is the schematic diagram of the penalty lane line intersection over union loss function DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings:

[0011] The purpose of the present invention is to provide a complex scene lane line detection method based on starting point-guided multi-dimensional feature refinement. The method proposed in this article aims to provide the lane line detection accuracy in complex scenes. Specifically, it includes the following steps:

[0012] S1: Select the dataset CULane as the benchmark dataset, and the number of training rounds is 36 rounds. Adopt data augmentation methods such as motion, horizontal flipping, random brightness, random contrast, random HSV adjustment, median blur, and random affine adjustment, and turn off data augmentation in the last 4 rounds. Make the convergence result closer to the real situation.

[0013] S2: The input image passes through the backbone network to generate three layers of features (X 0 , X 1 , X 2 ), and X 0 is subjected to feature extraction through the global feature optimization module to generate a feature representation X 0 ' with a large receptive field.

[0014] S3: Process X using a feature pyramid 0 ’, X 1 , X 2 to construct a high-resolution lane line feature map X 0 ', X 1 ', X 2 '.

[0015] S4: Process X 0 ' using the starting point coordinate prediction module to obtain possible starting points in the downsampled feature map by remapping and using the focal loss to supervise the heat map, and at the same time predict the offset of the starting point position and fine-tune the starting point position.

[0016] S5: Generate a lane prior based on the obtained starting point coordinates and represent it as: Represented as:

[0017]

[0018] where (x s , y s ) represents the starting point coordinates, H represents the image height, T = 40 represents the maximum number of lane prior points, and obtain a feature representation X 0 ”.

[0019] S6: Use the lane line perception aggregation module to refine the local features of X 0 ” to obtain a refined feature map and then aggregate it with the global feature map X 0 ”.

[0020] S7: Use a loss function to supervise the aggregation result of step S6, optimize the lane prior, and obtain P sl ;

[0021] S8: Combine P sl with X 1 ' and represent it as X 1 ”, use the lane line perception aggregation module to refine the local features of X 1 ” to obtain a refined feature map and then aggregate it with the global feature map X 1 ”.

[0022] S9: Use a loss function to supervise the aggregation result of step S8, optimize the lane prior, and obtain P s2 ;

[0023] S10: Combine P sl with X 2 ' and represent it as X 2 ”, use the lane line perception aggregation module to refine the local features of X 2"Perform local feature refinement to obtain a refined feature map Then aggregate it with the global feature map X 2 ".

[0024] S11: Use a loss function to supervise the aggregation result of step S10, optimize the lane prior, and obtain the final prediction result P s3 ;

[0025] Furthermore, in the said step S2, the global feature optimization process is as follows:

[0026]

[0027] Where L represents the features extracted from the last layer of the backbone network, BN represents the batch normalization layer, and GG represents the combination of the GELU activation function and global response normalization. In addition, Bra i represents multi-scale feature fusion, where Bra 0 performs no operation, and Bra 1 , Bra 2 and Bra 3 represent convolutional kernels of sizes 7×7, 11×11, and 21×21 respectively.

[0028] Furthermore, in the said step S4, the focal loss supervised proposal heatmap can be expressed as:

[0029]

[0030] Where num s represents the number of proposal starting points in the input image, p start represents the predicted value of each starting point coordinate in the proposal heatmap, and G start represents the ground truth at the predicted starting point coordinate.

[0031] The starting point fine-tuning method is:

[0032]

[0033] Where and represent the true starting point coordinates, O xy represents the true offset, and O pre represents the predicted offset. In addition, the predicted offset is trained using the smooth L1 loss and added to the coordinates of the heatmap to obtain the starting point coordinates.

[0034] Furthermore, in the said step S5, equidistant 2D points are used as the lane representation. Specifically, the lane representation is a series of points, that is, P = {(x 1 , y 1 ),..., (xN , y N )}. The y - coordinate of the points is sampled at equal intervals in the vertical direction of the image, that is where H is the height of the image. Correspondingly, the x - coordinate is associated with its respective y i ∈Y. This representation is the lane prior, and the network will predict each lane prior, which consists of four parts: (1) foreground and background probabilities. (2) The length of the lane prior. (3) The starting point of the lane line and the angle θ between the lane prior and the x - axis. (4) The horizontal offset between the prediction result and its ground truth.

[0035] Furthermore, in the steps S6, S8, and S10, the local feature refinement uses deformable convolution for local feature extraction, and uses convolution to predict the offset between a certain lane prior point on the same lane line and its N = 9 adjacent points, as follows:

[0036]

[0037] where F(p i ) represents the feature representation of the i - th lane line prior point, and Φ represents a non - linear function. Then, ΔO k is used to aggregate the features of adjacent points to the i - th lane line prior point p i :

[0038]

[0039] where w n represents the convolution weight. In summary, the refined feature map can be obtained.

[0040] The aggregation with the global feature map first uniformly selects M p = 40 lane line prior points; then, region of interest (ROI) alignment is performed to obtain the region of interest feature R i ∈R (C×Mp) ; next, the attention between the fully - connected region of interest feature R i ∈R (C ×1) and the resized and flattened feature map is used to obtain the aggregated local feature τ∈R (1 ×C) , and the specific formula is as follows:

[0041]

[0042] Finally, the aggregated local feature τ is further aggregated with the input global feature map X j ∈R (C×HW) to obtain the final output ε:

[0043]

[0044] Furthermore, the loss functions in steps S7, S9, and S11 can be written as:

[0045] L total = λ PLIoU L PLIoU + λ pro L pro + λ otfset L offset + λ cls L cls

[0046] where L cls represents the focal loss between the predicted value and the actual value, and λ PLIoU , λ pro , λ offset , and λ cls are set to 1.0, 0.3, 0.3, and 1.0 respectively. 。 where Lp LIoU represents the penalty lane intersection over union, which uses a variable lane span strategy to represent lane lines at different tilt angles:

[0047]

[0048] where is derived from the actual lane line span s and the tilt angle at the corresponding coordinates:

[0049]

[0050] In addition, the spatial correlation between non - overlapping lane lines is constructed using the penalty term shown in the following formula:

[0051]

[0052] The final penalty lane intersection over union can be expressed by the following formula:

[0053] L PLIoU = 1 - LIoU+L pen

Claims

1. A lane line detection method based on lane line aggregation perception and starting point prediction, comprising the following steps: The first step is to use global feature optimization to expand the receptive field of the backbone network; In the second step, the remapping heat map is used to predict the starting point coordinates and guide the generation of lane priors. In the third step, lane line perception aggregation is performed to refine local features while integrating global features. In the fourth step, the lane line features are optimized using the penalty lane line intersection-to-merger ratio loss function.

2. According to the method of claim 1, the global feature optimization first uses a 3×3 convolution kernel for feature extraction, then uses convolution kernels of sizes 7, 11, and 21 to extract features of different dimensions, and finally multiplies the obtained multidimensional features by the original features element by element. The process can be expressed as follows: Where L represents the feature extracted from the last layer of the backbone network, B N represents a batch normalization layer, and GG represents the combination of GELU activation function and global response normalization. i Represents multi-scale feature fusion, where Bra0 does not perform any operation, Bra1, Bra2, and Bra3 represent convolution kernels of sizes 7×7, 11×11, and 21×21, respectively.

3. The method according to claim 1, characterized in that By remapping the heatmap to obtain possible starting points in the downsampled feature map, the focus loss is used to constrain the proposed heatmap: The num s represents the number of proposed starting points in the input image, p start represents the predicted value of each starting point coordinate in the proposal heat map, G start Represents the true value of the predicted starting point coordinates. And the starting point position is fine-tuned by predicting the offset using the following formula: Among them and Represents the actual starting point coordinates, O xy Indicates the actual offset, O pre Represents the predicted offset. Finally, based on the obtained starting point, the lane prior is generated by sampling equidistantly along the vertical direction of the image: Among them (x s ,y s ) represents the starting point coordinates, H represents the image height, and T=80 represents the maximum number of lane prior points.

4. The method as claimed in claim 1, using deformable convolution to extract local features, using convolution to predict the offset between a lane prior point and its N = 9 adjacent points on the same lane line, as shown below: Among them, F(p i ) represents the feature representation of the i-th lane line prior point, and Φ represents a nonlinear function. Then, ΔO k Used to aggregate the features of adjacent points to the i-th lane line prior point p i : Among them, w n Represents the convolution weight. In summary, the refined feature map can be obtained. Then the refined feature map is aggregated with the global feature map. First, M is uniformly selected. p = 40 lane line prior points; then, ROI alignment is performed to obtain the ROI feature R i ∈R (C×Mp) ; Next, use the fully connected region of interest feature R i ∈R (C×1) and the resized and flattened feature map The attention between them is used to obtain the aggregated local features r∈R (1×C) , the specific formula is as follows: Finally, the aggregated local features τ are further combined with the input global feature map X j ∈R (C×HW) Aggregate to get the final output ε:

5. The method as claimed in claim 1, wherein a variable lane span strategy is used to represent lane lines at different tilt angles: Among them is the actual lane line span s and the inclination angle at the corresponding coordinate Derived from: In addition, the proposed penalty term as shown in the following formula is used to construct the spatial correlation between non-overlapping lane lines: The final penalty lane intersection ratio can be expressed as follows: L PLIoU =1-LIoU+L pen (12)