Rapid structure sensing lane detection model

By adjusting the loss function of the lane line detection model, introducing auxiliary segmentation and residual block preprocessing, and increasing similarity and shape loss, the problems of strong detection computing power dependence and poor recognition of large curvature paths are solved, and higher detection accuracy and robustness are achieved.

CN120164186APending Publication Date: 2025-06-17NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209509.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing lane line detection model has strong dependence on the detection computing power and has poor effect on the recognition of large curvature paths.

Method used

A fast structure-aware lane detection model is proposed. By adjusting the overall loss function, auxiliary segmentation and residual block preprocessing are introduced, similarity loss and shape loss are increased, and the generalization ability of the loss function is improved.

Benefits of technology

It significantly improves the lane line detection accuracy, improves the training accuracy and robustness of the model, reduces dependence on computing resources, and accelerates the training speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164186A_ABST
    Figure CN120164186A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of lane line identification, and discloses a fast structure sensing lane detection model, the overall loss function of the detection model is as follows: Ltotal = Lcls + alpha Lstr2 + beta Lseg, in the formula, Ltotal is segmentation loss, Lcls is classification loss, alpha and beta are weight coefficients; wherein in the # imgabs0 # formula, r is the total number of row anchor points, j is the number of row anchor points, rho is the influence coefficient of the bending weight, Lsim is the similarity loss, Lshp is the shape loss, # imgabs1 # Loci, j is the position of the ith lane line at the jth row anchor point, mu is the shape correction coefficient, h represents the number of row anchor points, and C represents the number of channels; according to the method, the problems that the existing model detection is high in calculation power dependence and poor in large-curvature path identification effect are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lane line recognition, and in particular to a fast structure-aware lane detection model. Background Technique

[0002] Real-time road detection is of great significance for the path tracking and assisted driving safety of intelligent driving vehicles. Currently, the methods for lane line detection are mainly divided into pattern recognition and deep learning. The idea of pattern recognition is to identify the edge distribution of objects in an image according to the change of color threshold, and then use the edge features of the lane line to find the position of the lane line, and then use the Canny edge detector and Hough transform to complete the data sampling and convergence. In order to improve the accuracy of the pattern recognition method, Raja Muthalagu et al. used the perspective transformation method to abstract the blurred part of the road image, which can more accurately detect distant curves and has stronger robustness to lighting conditions and road changes; however, for complex situations, such as the edge noise caused by lane line blockage, the pattern recognition method often has insufficient processing ability.

[0003] Compared with the pattern recognition method, the deep learning method has greatly improved the lane line detection effect by learning the edge information and important information of the road image. However, the large amount of computational data requirements and the lane line occlusion problem still become important problems faced by the deep learning method. The SAD (Sum of Absolute Differences) method processes the information input by multiple cameras by introducing an attention mechanism, which solves the model convergence speed problem to a certain extent, but this scheme needs to densely process data and has high requirements for the computing power of the system equipment. For the lane line occlusion problem, the SCNN (Spatial Convolutional Neural Network) method slices the output feature matrix, which is beneficial to the transmission of slender features, and this change is mainly applicable to the detection of long-distance continuous shape targets.

[0004] At the current stage, there has been great progress in the research on ultra-fast structured perception and road depth positioning of lane lines, but existing algorithms often require strong computing power support, have strong dependence on detection computing power, and have poor recognition effects on large curvature paths. Therefore, how to solve the problems of ultra-fast perception and prediction accuracy improvement of lane lines in the context of no visual cues or large data volume training is of great significance. Summary of the Invention

[0005] The present invention aims to provide a fast structure-aware lane detection model to solve the problems of strong dependence on detection computing power and poor recognition effects on large curvature paths of existing models.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A fast structure-aware lane detection model, and the overall loss function of the detection model is:

[0008] L total = L cls + αL str2 + βL seg

[0009] In the formula, L total is the segmentation loss, L cls is the classification loss, and α and β are weight coefficients;

[0010] Among them,

[0011] In the formula, r is the total number of row anchors, j is the number of row anchors, ρ is the influence coefficient of the bending weight, L sim is the similarity loss, L shp is the shape loss, Loc i,j is the position of the i-th lane line at the j-th row anchor, μ is the shape correction coefficient, h represents the number of row anchors, and C represents the number of channels.

[0012] Further, the expression of the classification loss is:

[0013]

[0014] In the formula, L ce represents the cross-entropy loss, T i,j is the coordinate label of the i-th lane line anchored in the j-th row represented by a one-hot encoding, h represents the number of row anchors, and C represents the number of channels.

[0015] Further, the expression of the similarity loss is:

[0016]

[0017] In the formula, ||·||1 represents the first-order norm of the function, h represents the number of row anchors, and C represents the number of channels.

[0018] Further, the expression of the shape loss is:

[0019]

[0020] prob i,j,: = softmax(p i,j,1:w )

[0021] In the formula, p i,j is a w-dimensional vector; prob i,jIt represents the possibility of each position, where h represents the number of row anchor points and C represents the number of channels.

[0022] The beneficial effects of the technical solution are:

[0023] A fast structure-aware lane detection model of the present invention can quickly improve the training accuracy of the detection model by means of auxiliary segmentation and residual block preprocessing. At the same time, the regularization loss function (such as the curve curvature loss regularization phase) is adjusted to supplement the generalization ability of the detection model for the curve structure loss function, effectively improving the detection accuracy of lane lines during the steering process. And the model of the present invention does not rely on established prior knowledge, and the required training data will not increase, and it can be effectively applied to real driving scenarios. The test results in actual application scenarios show that compared with advanced models, the model of the present invention can improve the lane line detection accuracy by about 10%. The hyperparameter settings of the new structure-aware formula can accelerate the training speed of the model, increasing the convergence speed of the model in the training stage by more than 3 times. Moreover, the model of the present invention has strong robustness for lane line detection in different application scenarios. Description of the Drawings

[0024] Figure 1 It is a schematic diagram of anchor points and rasterization processing in Embodiment 1 of the present invention;

[0025] Figure 2 It is a formula and a conventional segmentation diagram in Embodiment 1 of the present invention;

[0026] Figure 3 It is the segmentation semantics of the overall network structure in Embodiment 1 of the present invention;

[0027] Figure 4 It is the coordinate relationship between rows and columns of a curved lane in Embodiment 1 of the present invention;

[0028] Figure 5 It is a graph of accuracy error values obtained by testing the model of the present invention and the control group in Scenarios 1-4 in Embodiment 2 of the present invention;

[0029] Figure 6 It is a recognition result diagram of the model of the present invention and the control group in Scenarios 1-4 in Embodiment 2 of the present invention; (a1)-(a4) are the recognition results of the control group, and (b1)-(b4) are the recognition result diagrams using the model of the present invention;

[0030] Figure 7In Embodiment 3 of the present invention, the test effect diagrams of the model of the present invention and the control group are shown; A1 - A8 are the test effect diagrams of the model of the present invention under normal, crowded, strong light, shadow, no lane lines, arrow, curve, and night scenes respectively; B1 - B8 are the test effect diagrams of the control group under normal, crowded, strong light, shadow, no lane lines, arrow, curve, and night scenes respectively;

[0031] Figure 8 In Embodiment 3 of the present invention, it is a relationship curve graph between the accuracy of the control group and the model of the present invention and the number of training rounds. Detailed implementation manners

[0032] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:

[0033] Embodiment 1

[0034] First, a structure - aware segmentation model based on the row - anchor framework is established, and the weighted weights of the route similarity rule, straight - line rule, and curve rule are adaptively adjusted according to the vehicle recognition distance, so as to accurately describe the change characteristics of the vehicle - path row - anchor; then, in the learning model of the perception algorithm, an effective correction is made to the turning path and the uncertain road - blurred path by adding a road - structure loss function with a curve tendency (when the row - anchor is large, the straight - line rule term has a greater weight advantage, and vice versa, the curve - rule term has a greater advantage) to obtain an improved model.

[0035] Specifically, first, the lane lines sampled by the camera are converted into structured data available for learning and training, and are represented by anchor points. By dividing the image into a grid, the overall lane lines are divided into multiple grid - point sets containing obvious identification colors (such as white lane lines). The row vector where the grid points belonging to the same lane line are aggregated is defined as the row - anchor set, and a certain row - vector cell related to the lane line is called a row - anchor point. The column - anchor point can be obtained by the same principle. The anchor points divide the image containing the lane lines into a certain number of row - column matrices, and then a raster - grid method can divide a certain row into many cells, as shown in Figure 1 and Figure 2 shown. The rasterized input value of the lane - line condition image (anchor - point matrix) extracted by the camera is input into the neural network to extract deep features to accelerate the speed of the model. After training by the learning model, a classifier f ij is obtained, and then the lane - line detection probability at each position is output:

[0036] p i,j,: = f ij (x), S.t. i ∈ [1, C], j ∈ [1, h] (1)

[0037] Among them, x is defined as the global feature; the output p i,j, indicating that the probability that each cell of the actual i-th lane line on the j-th row of anchors is selected as a lane line is a w+1-dimensional vector.

[0038] The expression for the classification loss used for lane line detection is:

[0039]

[0040] In the formula, L ce represents the cross-entropy loss; T i,j is the coordinate label of the i-th lane line anchored in the j-th row, represented by a one-hot encoding. To be able to use prior knowledge to solve the problem without visual cues, the present invention uses a structural loss as a regularization of the classification loss. Among them, the symbol naming rules are shown in Table 1.

[0041] Table 1 Symbol naming rules

[0042] Variable Type Description Definition H Scalar Image Height W Scalar Image Width h Scalar Number of Row Anchor Points ω Scalar Number of Meshed Cells C Scalar Number of Channels X Tensor Global Feature of Image f Feature Classifier for Selecting Lane Position <![CDATA[p∈R c×h×(w+1) > Tensor Group Prediction T Tensor Group Target Prob Tensor Probability at Each Position Loc Matrix Lane Position

[0043] To enable the neural network model to learn prior road structure information, two loss functions, the similarity loss and the shape loss, are used as a structure perception method to make the lane line position points have more reasonable correlations. The loss function of the neural network learning model consists of the similarity loss L sim and the shape loss L shp in two parts. First is the similarity loss. Specifically, the predicted point positions of adjacent lane lines should be close to make the predicted lane lines more continuous. The expression for the similarity loss in the present invention is as follows:

[0044]

[0045] Among them, ||·||1 represents the first-order norm of the function.

[0046] Then, the shape loss is used to describe the shape continuity and integrity of the lane lines. Assuming that most lane lines are straight lines, a second-order difference equation is used to restrict the predicted shape of the lane lines. The expression for the shape loss is as follows:

[0047]

[0048] prob i,j,: = softmax(p i,j,1:w ) (6)

[0049] Among them, p i,j is a w-dimensional vector; prob i,j represents the possibility at each position. Note that the background grid is not counted, so it ranges from 1 to w, rather than w+1. In summary, the structural loss of the lane lines can be expressed as follows:

[0050] L str = L sim + λL shp (7)

[0051] Among them, λ is a weight coefficient used to balance the two losses.

[0052] In order to enable the model to better integrate global and local feature parameters, the neural network model adds auxiliary segmentation processing during the training stage. The segmentation semantics of the overall network model are as Figure 3 shown. The present invention uses cross-entropy as the auxiliary segmentation loss. The overall loss function can be defined as follows:

[0053] L total = L cls + αL str + βL seg (8)

[0054] Among them, L total is the segmentation loss; α and β are weight coefficients.

[0055] However, the similarity loss in the traditional structure loss function is usually defined as a function with a trend perpendicular to a straight line, while the shape loss has a trend tending to a straight line. Such a scheme has a high accuracy for straight lane line recognition, but it cannot accurately identify the turning path of real road lines. For example, the curved structure of a turn is more curved at a certain distance from the camera, and the horizontal trend is more obvious, which will cause the recognition accuracy of the traditional structure loss function for curved lane lines to decay or even fail. However, the accurate recognition of curved lines by the lane line recognition algorithm can ensure the driving safety of the vehicle when turning at high speed. Therefore, it is very necessary to improve the model's ability to recognize curves.

[0056] The prior knowledge of the traditional scheme is that the lane lines close to the camera lens are almost considered as a straight line. This assumption is reasonable within the range close to the camera near point, but the lane lines farther away should not follow this prior knowledge (that is, the farther the camera is, the higher the probability of a curve, especially when it comes to the turning process). Therefore, the grid at the bottom of the image (i.e., those parts close to the camera) is more inclined to predict a straight line, while the grid at the top of the image (i.e., those parts far from the camera) should be more inclined to predict corners. The gradual trend in the above scheme needs to be designed as a smooth transition type in the mathematical model. Therefore, the road structure loss formula with a curve tendency constructed by the present invention is defined as follows:

[0057]

[0058] Among them, Loc i,jrepresents the position of the i-th lane line at the j-th row anchor point, and μ is the shape correction coefficient. The parameter μ is a value between 0 and 1. The smaller μ is, the greater the curvature. When μ equals 1, the above formula degenerates into a loss function for a straight line shape. The intuitive understanding of the above formula is as Figure 4 shown. When the vertical coordinate of the curve is at the center point, the greater the absolute value of the horizontal distance between the two points, the more obvious the degree of road curvature. Combining the structural loss in the original text with the above structural loss, the resulting formula is calculated as follows:

[0059]

[0060] where r represents the total number of row anchor points, and j is the number of row anchor points. In this way, in the upper part of the image, j is small, and the structural loss of the straight line often occupies a small weight, while the structural loss with a curve tendency has a large weight, and the entire new structural loss has a large curve tendency; in the lower part of the image, j is large, the structural loss of the straight line often occupies a large weight, while the structural loss of the curve often occupies a small weight, and the entire new structural loss has a greater linear trend. ρ is the influence coefficient of the bending weight. Feature aggregation is performed on all loss functions, and the final loss function is as follows:

[0061] L total = L cls + αL str2 + βL seg (11)

[0062] Example 2

[0063] Recognition accuracy comparison:

[0064] The model proposed by the present invention is trained and tested using the publicly available CULane dataset. The evaluation metrics are introduced as follows: Using the evaluation metrics of the CULane dataset, that is, regarding the lane line as a line with a width of 30 pixels, calculate the intersection over union (IOU) between the predicted position and the actual position of the lane line. A predicted value with an IOU greater than 0.5 is regarded as a predicted value of TP, and the F1 measurement value is used as the evaluation metric, and its formula is as follows:

[0065]

[0066] where, TP is true positive, FP is false positive, and FN is false negative.

[0067] The range of the row anchor points for model parameter setting is 260 - 530, the step size is 10, the number of grid cells is set to 150, the size of the input image is adjusted to 288 * 800, the model is trained with the Adam optimizer, the learning rate is initialized to 0.0004, the cosine annealing learning strategy is used to reduce the loss parameters α, β, λ in equations (7), (8), (11) to 1, μ in formula (9) is set to 0.5, and ρ in formula (10) is set to 0, 0.1, and 0.5 respectively. Due to device limitations, the batch size is changed from 32 to 16. Since the batch number is reduced, it is necessary to increase the number of training epochs to 80. Reducing the minimum batch number also affects the performance of batch normalization, thus affecting the performance of the model, but it does not affect the relative performance of the improved model compared to the original model. The NVIDIA GTX 3080 Ti GPU is used for training and testing. To prevent overfitting, enhance the generalization ability of the model, and obtain the actual effect of the model, the same data augmentation method as the traditional scheme is used here, that is, the data is augmented by rotating and translating the pictures, and the lane lines are extended to maintain the lane structure.

[0068] To verify the improvement effect of the proposed scheme on detection accuracy, a traditional scheme (using the model scheme in the following literature: Qin Z, Wang H, Li X. Ultra-Fast Structure-aware Deep Lane Detection[C] / / European Conference on Computer Vision. Springer, Cham, 2020.) is introduced as a control group, and the detection results are as Figure 5 shown. Among them, to quantify the detection accuracy, the true position of the lane line in the picture is defined as the reference value, and the recognized lane line (i.e., the anchor point combination) is used as the actual value. Based on this, the detection error values at different sampling points (this value is identified by pixels) are obtained, as Figure 6 shown. Among them, a1 - a4 are the recognition results of the traditional scheme, and b1 - b4 are the recognition results of the proposed scheme.

[0069] According to Figure 5 (a1) and (a2) in, it can be found that due to the existence of the middle dotted line, the lane lines recognized by the traditional method show obvious distortion phenomena, and the detection error in this scenario is as high as 25 pixel points. Figure 5 (a3) and (a4) in show more obvious phenomena of lane line non-recognition and recognition deviation, and even misrecognition occurs at the highway exit position. This is because the loss setting for the turning scenario in the traditional scheme lacks the description of shape loss, resulting in the learning model being unable to predict the distant lane lines, and Figure 5The confusing features of multiple white lane line markings at the ramp in (a4) increase the recognition difficulty of the learning model. The model solution proposed in the present invention can effectively predict lane line anchor points by effectively adjusting the lane line loss function (i.e., the characteristic of fitting a straight line trend near and a curve at the far end), and can achieve good performance tracking both globally and locally. The specific accuracy error values are as Figure 6 shown. In comparison, in different application scenarios, the recognition errors of the traditional solutions are mostly in the range of 6 - 25 pixel points, while the model solution proposed in the present invention can constrain the maximum prediction error within 15, and even most of the sample points are within 10 pixel errors.

[0070] Example 3

[0071] In this example, different usage scenarios were introduced on the basis of the tests in Example 1 to analyze the robustness of the detection model, including vehicle congestion, night scene, no lane lines, shadows, arrows, strong light, and turning curve scenes, etc. The influence of different weight coefficients on the detection accuracy of the learning model is shown in Table 2,

[0072] Table 2 Model accuracy comparison (Note: The percentage of detection accuracy improvement is in parentheses)

[0073]

[0074] Compared with the traditional solution (control group), the model proposed in the present invention significantly improves the detection accuracy of lane lines, with a performance improvement of about 6% - 13% in different application scenarios. This shows that after setting the weight coefficient to a reasonable value, the curve regularization term of the lane lines can effectively correct the learning model and improve the detection effect of lane lines. In addition, we also tried to design the weight to be variable according to the curvature degree of the lane line and a fixed weight of 0.1, and found that using the fixed weight coefficient of 0.1 can already adapt to the curve rule of the lane line, and its detection performance has been improved to a certain extent. Figure 7 Shows the comparison of the actual recognition effects (working conditions: 1: normal, 2: crowded, 3: strong light, 4: shadow, 5: no lane lines, 6: arrow, 7: curve, 8: night). It can be seen that the model proposed in the present invention can accurately predict curve lane lines at a farther distance. Many lines with insufficient prediction accuracy of the traditional model are often more vertical than the actual lane lines, because at a long distance, the weight of similarity regularization is too large, resulting in the loss function being unable to accurately describe the curve-shaped lane lines.

[0075] In terms of the convergence speed of the training model, the improved model (the model of the present invention) requires fewer training rounds, that is, it can achieve better results than the traditional model after fewer rounds of training. Figure 8The relationship between model accuracy and the number of training epochs (the vertical lines in the figure represent error bars, calculated by collecting accuracy data around the whole number of epochs, which can reduce the contingency of the results), Figure 8 The line chart of Figure 8 shows that the model can achieve good results with only 10 epochs of training. This is because the better prior formula enables the network model to converge faster. However, the weight of the curve term regularization should not be too large, otherwise the accuracy of the model will be seriously affected (as shown in the line chart of Figure 8 when the weight of the curve regularization term is 0.5, the model performance will be significantly reduced). When approaching the vehicle, the weight of the linear rule increases, while when far away from the vehicle, only the classification accuracy is used as the training criterion, which helps to overcome the problem that the lane lines are not straight at a long distance and enables the model to dare to predict the lane lines at a long distance. It can be concluded that the model of the present invention can significantly improve the accuracy of lane prediction, effectively improve the deficiencies of the original model in curve prediction, and thus can predict lane lines at a farther distance, which proves the effectiveness of the model of the present invention and improves the reliability and effectiveness of the ultra-fast lane detection algorithm in practical applications. Among them, the test code used in the solution of the present invention is on GitHub.

[0076] In summary, in view of the fact that the structure-aware formula with a straight-line tendency in the traditional lane line model detection scheme is not conducive to identifying curved lane lines, a new ultra-fast lane structure-aware algorithm and structure-aware scheme are constructed. By using more reasonable prior information, that is, tending to straight-line prediction in the short distance, reducing straight-line prediction in the long distance, and appropriately increasing the curve tendency, the model can predict curves and straight lines more accurately. The test results show that compared with the traditional scheme, the proposed scheme determines the hyperparameters of the new structure-aware formula, and can accelerate the convergence speed of model training with a more reasonable lane line regularization scheme.

[0077] The above are only the embodiments of the present invention, and specific technical solutions or common knowledge such as well-known features in the solution are not described in detail here. It should be pointed out that for those skilled in the art, without departing from the technical solution of the present invention, several deformations and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners in the specification can be used to explain the content of the claims.

Claims

1. A fast structure-aware lane detection model, characterized in that: The overall loss function of the detection model is: L total =L cls +αL str2 +βL seg Where, L total is the segmentation loss, L cls is the classification loss, α and β are weight coefficients; in, Where r is the total number of row anchor points, j is the number of row anchor points, ρ is the influence coefficient of bending weight, and L sim is the similarity loss, L shp is the shape loss, Loc i,j is the position of the i-th lane line at the j-th row anchor point, μ is the shape correction coefficient, h is the number of row anchor points, and C is the number of channels.

2. A fast structure-aware lane detection model according to claim 1, characterized in that: The expression of classification loss is: Where, L ce represents the cross entropy loss, T i,j is the coordinate label of the i-th lane line anchored in the j-th row represented by a one-hot encoding, h represents the number of row anchors, and C represents the number of channels.

3. A fast structure-aware lane detection model according to claim 1, characterized in that: The expression of similarity loss is: Where ||·||1 represents the first-order norm of the function, h represents the number of row anchor points, and C represents the number of channels.

4. A fast structure-aware lane detection model according to claim 1, characterized in that: The expression of shape loss is: prob i,j,: =softmax(p i,j,1:w ) In the formula, p i,j is a w-dimensional vector; prob i,j represents the possibility of each position, h represents the number of row anchor points, and C represents the number of channels.