Anchor frame optimization-based 2D lane line detection method
By introducing anchor frame optimization network and PGIoU loss calculation in lane line detection, the problem of inability to effectively detect any position of the starting point of lane line in the prior art is solved, and the accuracy and effectiveness of detection are significantly improved.
Patent Information
- Application Number
- CN202311788873.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing lane line detection technology based on anchor frames cannot effectively detect lane lines such as intersection lane lines, stop lines, and separation lines, and cannot perform cross-layer optimization of feature maps, resulting in insufficient recognition accuracy and effectiveness.
A 2D lane line detection method based on anchor box optimization is proposed. The anchor box is adaptively optimized through the anchor box with scene information and the actual position of the lane line with any starting point is detected. The specific solution includes constructing an anchor box optimization network in the offline stage and using the trained network for real-time 2D lane line detection in the online stage.
The recognition accuracy and effectiveness of lane line detection are significantly improved, and the lane line at any position at the starting point can be detected, and the inference ability of the model is further improved by introducing PGIoU loss calculation.
Smart Images

Figure CN120220094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of autonomous driving, specifically a 2D lane line detection method based on anchor box optimization. Background Art
[0002] Since the starting points of prior anchor boxes are defined at the edges of images, existing anchor box-based lane line detection schemes cannot detect the actual positions of some lane lines, such as intersection lane lines, stop lines, separation lines, etc. or the prior anchor boxes cannot be optimized according to the scene, and cross-layer optimization of feature maps cannot be achieved. Existing ResNet-based lane line detection technologies have problems in effectively evaluating the overlap between predictions and ground truths and the prediction accuracy of large curvature curves. Summary of the Invention
[0003] In view of the above deficiencies of the prior art, the present invention proposes a 2D lane line detection method based on anchor box optimization, which adaptively optimizes the anchor boxes by combining scene information through an Anchors optimised Network (AONet), and detects the actual positions of lane lines with arbitrary starting points, significantly improving the recognition accuracy and effectiveness.
[0004] The present invention is realized through the following technical solutions:
[0005] The present invention relates to a 2D lane line detection method based on anchor box optimization. An anchor box optimization network AONet is constructed in the offline stage, and a Perspective Grounded IoU (PGIoU) is used to calculate the loss function to train the anchor box optimization network; in the online stage, the trained anchor box optimization network is used for real-time 2D lane line detection.
[0006] The anchor box optimization network includes: a backbone network, a confidence-starting point initialization unit, a feature pooling unit, and a lane line prediction unit, where: the backbone network performs multi-layer residual convolution operations on the input image to obtain a high-level feature map; the confidence-starting point initialization unit optimizes the prior anchor boxes according to the high-level feature map, predicts the starting points of each predefined anchor point and their corresponding confidence scores to preliminarily determine whether the anchor points belong to lane line instances or the background; the feature pooling unit performs rasterized sampling of pixel points on the high-level feature map according to the optimized anchor box positions to obtain lane line queries; the lane line prediction unit performs convolution operations on the high-level feature map to generate key-value pairs, and performs cross-layer optimization of self-attention operations in combination with the lane line queries to obtain the final lane line prediction.
[0007] The backbone network adopts a Residual Network (ResNet) or a Deep Layer Aggregation Network (DLA34).
[0008] The confidence - starting point initialization unit described above includes: two convolutional pooling layers and two fully - connected layers, where: the first and second convolutional pooling layers perform upsampling on the input high - level feature map in sequence, convert it into a one - dimensional vector through the second convolutional pooling layer and input it into the fully - connected layer to generate n 4 - dimensional vectors (pos, neg, x, y), where: n is the number of lane - line anchor boxes, pos and neg respectively represent the confidence of the ground truth or background, and (x, y) represents the starting - point coordinates.
[0009] The perspective - based basic accuracy criterion described above means that for the i - th sampling point on the lane line, its corresponding line width w in the image p =[(n - i*f) / n]*width, where: width represents the constant width of the lane - line starting point, set to 30 pixels, n represents the total number of sampling points of each lane line, and the disappearance rate f represents the visual width difference between the farthest and nearest points of the lane line. The evaluation of the prediction matching degree is obtained by calculating the intersection - over - union of the ground truth and the prediction.
[0010] The present invention relates to a lane - line detection system for implementing the above - mentioned method, including: a backbone network module, an anchor - box initialization module, an anchor - box optimization module, a lane - line query module, and a loss - calculation module, where: the backbone network module performs preliminary feature extraction on the image; the anchor - box initialization module and the anchor - box optimization module respectively perform the initialization of the lane - line anchor boxes and optimize the positions of the anchor boxes according to the feature - map information; the lane - line query module extracts the pixel points near the anchor - box positions on the feature map as queries through the feature - pooling method, performs self - attention operations with the key - values generated by the entire image, and repeats the above operations in the last three layers of the bottleneck network output feature map to obtain the cyclically optimized lane - line positions and class predictions; the loss - calculation module adopts the perspective - based basic accuracy criterion, calculates the FocalLoss and the Euclidean distance combined as the classification / regression loss of the model prediction, and backpropagates to optimize the backbone network module; the backbone network module outputs the prediction results of the lane - line confidence, lane - line category and color, and the position coordinates of the point set on the lane line. Technical effects
[0011] The present invention realizes an anchor box optimization module, uses the information of the high-level feature map to optimize the predefined anchor box, comprehensively utilizes the prior shape information of the lane line and the driving scene information, and predicts the lane line at any position of the starting point; by considering the perspective principle, designs and realizes PGIoU, defines the overlap between the lane line ground truth and the prediction, and assists the model in lane line matching and loss calculation. Compared with the prior art, the present invention improves the inference ability of the model by recalling the lane lines whose starting points are not at the edges of the picture during inference, which is reflected in the detection precision and recall rate of the current anchor box-based lane line detection model in the CityLane dataset; by introducing PGIoU for loss calculation, the inference ability of the model itself is further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 , Figure 2 is a schematic diagram of the predefined anchor box of the present invention;
[0013] Figure 3 Schematic diagram of the anchor box optimization network structure of the present invention;
[0014] Figure 4 is a schematic diagram of the working principle of the anchor box optimization module;
[0015] Figure 5 is a schematic diagram of the working principle of PGIoU;
[0016] Figure 6 is a comparison diagram of the annotation standards of CityLane (upper) and CULane (lower). DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] As Figure 3 shown, this embodiment relates to a 2D lane line detection method based on anchor box optimization, including:
[0018] Step 1: Collect and annotate the CityLane dataset covering urban and elevated scenes. This dataset is annotated based on the actual positions of the lane lines, that is, the starting points of the lane lines are truthfully annotated without extending the lane lines backward to the edges of the picture, and the annotation includes the annotation of the lane line category and color attributes. Specifically, it includes images of day, night, and rainy days in urban and elevated scenes, with a total of more than 80,000 images, among which: 15,000 images are used as the test set, and the rest are used as the training and validation sets. All images are detailedly annotated in the way of line key points.
[0019] Step 2: Initialize and use the feature pooling technology to construct the initial lane line anchor box as the basis for detection.
[0020] The initialization described above specifically includes: defining each lane line as a set of points with annotation attributes, and the predicted results of the lane line include: 1) C_lane: the predicted confidence score; 2) P_cls: the probabilities of different lane line types, corresponding to 8 common lane line types included in the dataset; 3) P_color: the probabilities of different colors, corresponding to 3 color annotations; 4) xyt: the starting coordinates and tilt angle of the lane line, used to generate the reference line; 5) l: the lane line span, indicating the vertical distance from the end point to the starting point of the lane line; 6) D_offset: the lateral distance from N sampling points on the lane line to the reference line.
[0021] The feature pooling technique mentioned above means: setting the starting points of the initial lane line anchor boxes on the left and right sides of the image such that the lane lines are farther from the main lane, as Figure 2 shown. Uniformly sample the starting points along the bottom edge of the image, and set four oriented anchor points for each point to consider different directions. Θ = 0.2πx i, i ∈ {1, 2, 3, 4}.
[0022] Step 3, construct an anchor box optimization network as Figure 3 shown, which includes a backbone network, a confidence-starting point initialization unit, a feature pooling unit, and a lane line prediction unit, and perform training and inference on the input image.
[0023] The confidence-starting point initialization unit mentioned above, based on the high-level feature map output from the bottleneck network in the backbone network, obtains the starting point estimates of N_priors anchor points and provides confidence predictions for each anchor point through two convolutional pooling layers and two fully connected layers.
[0024] The starting point estimates of the N_priors anchor points mentioned above refer to the set of sampling points on the lane line where: (x, y) represents the starting point provided by the starting point initialization module for each lane line, y_i represents the vertical sampling coordinate, and θ represents the tilt angle of the predefined anchor point, as Figure 4 shown. The goal is to use the feature information around the cropped anchor points as a query to more accurately detect the position of the lane line.
[0025] Step 4, train the anchor box optimization network, that is, perform prediction matching and calculate the loss for lane line detection based on PGIoU, as Figure 5 shown. Due to the perspective principle, the lane lines in the image appear thinner in the distance. Reducing the lane line width along the vertical coordinate can provide more accurate information for the model. This training process specifically includes:
[0026] 4.1) Enhance the image to increase the generalization of the model, and input the enhanced image into the backbone network to obtain the feature map: For the i-th sampling point on the lane line, the specific position of the ground truth lane line in the pixel coordinate system is obtained through the lane line width representation method described in PGIoU.
[0027] The enhancement process mentioned above refers to: cropping, rotation, horizontal mirroring, and Gaussian blur processing.
[0028] 4.2) During the training process, the anchor box optimization network dynamically assigns one or more predicted lanes as positive samples for each ground truth lane line according to the assignment cost C assign Specifically: the assignment cost C assign = C sim W sim + C cls W cls where: C cls is the classification loss between the prediction and the label, and C sim is the similarity loss between the predicted lane and the ground truth, and they are multiplied by predefined weights to obtain the matching loss.
[0029] 4.3) Conduct queries based on the self-attention mechanism. Specifically: use the pixel point sampling of the optimized anchor box on the feature map as the query, use the output obtained by passing the feature map through the convolutional neural network as the key value, and generate the output results including the probabilities of containing lane lines or backgrounds, lane line type classification / color classification, and the regression values of the line positions by calculating the loss function.
[0030] The loss function L AONet = W conf L conf + W xytl L xytl + W IoU LP GIoU where: W is the weight of each parameter, L conf is the focal loss between the predicted value and the label, L xytl is the smooth L1 loss of the starting point coordinates, theta angle, and lane length regression, and L PGIoU is the line IoU loss between the predicted lane and the actual ground truth.
[0031] The sum of the probabilities of the lane line and the background is 1.
[0032] Step 5, In the online stage, perform lane line detection on an image from the vehicle's front camera as shown in Figure 3 Specifically, it includes:
[0033] 5.1) Use a Residual Network (ResNet) or Deep Layer Aggregation Network (DLA34) as the backbone network for feature extraction, converting the original image input into a high-level feature representation: Unify the input image size to accelerate calculations, with the default resized size being 800x320.
[0034] The basic structure of the residual network includes residual blocks, and its input and output are: Input representation: X (l) , where: (l) represents the layer or level; Output representation: Y (l) = F(X (l) ) + X (l) , where: F(X (l) ) is the transformation function of the residual block.
[0035] The input and output of the deep layer aggregation network are: Input representation: X (l) , representing the input of the l-th layer; Output representation: Y (l) = DLA(X (l) ), where: DLA is the deep layer aggregation operation.
[0036] 5.2) Anchor box initialization and optimization: Use a confidence - starting point initialization unit to predict the starting position of each predefined anchor point and provide a confidence score for each anchor point.
[0037] 5.3) Feature pooling: Sample pixel points around the optimized anchor positions. The feature map itself is processed through different convolutional layers, and queries (query), keys (key), and values (value) are formed through convolutional operations, specifically including:
[0038] a) Generate queries: According to the adjusted anchor positions and the feature map, generate queries through convolutional operations to capture local information near the lane lines. Specifically: Q = Conv(Optimised Anchors, X feature ), where: Conv is the convolutional operation, Optimised Anchors is the optimized anchor box, and X feature is the input feature map.
[0039] b) Generate keys and values: According to the feature map, generate keys (key) and values (value) through different convolutional operations to capture global and local features. Specifically: K = Conv(X feature ), V = Conv(X feature ). The definitions of convolution and feature map are the same as in a).
[0040] c) Self-attention operation: Based on the generated query, key, and value, through the self-attention mechanism, weights are calculated according to the query, and then the weights are applied to the value to obtain the final lane line features, successfully capturing local information related to the lane line around the optimized anchor points. At the same time, self-attention operations are performed using global and local features to support the final prediction of the lane line. Specifically: Where: d k is the dimension of the key value K. Cross-level optimization is achieved through the Feature Pyramid Network (FPN) of the bottleneck network in the backbone network: The last three layers of the output of the FPN are used for self-attention operations. These outputs sequentially pass through self-attention operations from top to bottom, forming a recursive structure, and finally obtaining the final result of cross-level optimization.
[0041] After specific actual experiments, a single RTX3090Ti is used as the training environment. Based on PyTorch, ResNet / DLA34 with an FPN bottleneck network is implemented as the pre-trained backbone network. All input images are adjusted to 320×800. Random affine transformation, random horizontal flipping, and Gaussian blur are used as data augmentation methods during training. During training, the AdamW optimizer is used with an initial learning rate of 1e-3, and 30 epochs of training are performed. The coefficients for assigning costs are set as W cls = 1 and W sim = 3, and in the loss function, W conf = 6, W xytl = 0.5, W IoU = 2 are used as parameter configurations during the training process.
[0042] By collecting a large amount of on-road real-time information and annotating the lane line positions, types, and colors according to the Figure 6 standard, the CityLane dataset is obtained, which contains a total of 81,000 frames of valid annotations. Among them, 65,000 are used as the training set, and the remaining 16,000 frames are used as the test set. The results of lane line prediction are matched based on the criteria given by the Tusimple dataset, that is, a line with the number of points within 30 pixels in the horizontal direction from the true value annotation points accounting for more than 85% of the predicted line is considered a correct prediction. The accuracy, precision, recall, and F1-score are calculated based on the correct prediction results After testing the influence of different vector dimensions d on the results, the experimental data obtained are compared with the current advanced 2D lane line model CLRNet, and the following results are obtained: Method Backbone network Accuracy Precision Recall F1-score CLRNet ResNet18 0.9019 0.9455 0.8596 0.9005 The present invention ResNet18 0.9196 0.9624 0.8826 0.9208 The present invention + PGIoU ResNet18 0.9277 0.9636 0.8874 0.9239 CLRNet DLA34 0.9215 0.9483 0.8607 0.9024 The present invention DLA34 0.9451 0.9663 0.8837 0.9231 The present invention + PGIoU DLA34 0.9459 0.9674 0.8961 0.9304
[0043] The ablation experiment results of the disappearance rate in PGIoU are as follows: 0 1 / 3 1 / 2 2 / 3 Precision 0.9663 0.9674 0.9678 0.9649 Recall 0.8837 0.8951 0.8905 0.8839 F1-score 0.9231 0.9304 0.9275 0.9226
[0044] As shown in the above table, when f is set to 0, PGIoU is equivalent to LineIoU, that is, it is assumed that the width of the lane line in the image is fixed. It can be observed that when the disappearance rate is set to 1 / 3, compared with using LineIoU, the F1-Score increases by 0.73%, indicating that allowing the lane line width to change with the disappearance rate can improve the performance of the algorithm. It also shows that the advantage of the present invention over other models lies in the more accurate detection of the position and shape of the line.
[0045] Compared with the prior art, the present method optimizes the network through anchor boxes to solve the problem of detecting lane lines whose starting points are not at the edges of the picture, and significantly improves the overall detection accuracy and precision-recall rate of the model on the dataset. Specifically, the present invention is superior to CLRNet in terms of F1-Score, with an increase of 2.07%. In addition, by reconstructing the anchor points and PGIoU, the detection metrics can be further improved. Finally, the F1-Score is 2.80% higher than that of CLRNet. This can be attributed to the fact that the present invention can correctly infer lane lines with unique shapes and positions, especially those that are missed or misdetected due to their starting points not being at the edges of the image.
[0046] Those skilled in the art can make local adjustments to the above specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the constraints of the present invention.
Claims
1. A 2D lane line detection method based on anchor box optimization, characterized in that, Optimize the network by constructing anchor boxes in the offline stage, and calculate the loss function using a perspective-based basic accuracy criterion to train the anchor box optimization network; In the online stage, use the trained anchor box optimization network for real-time 2D lane line detection; The anchor box optimization network includes: a backbone network, a confidence-start point initialization unit, a feature pooling unit, and a lane line prediction unit, where: the backbone network performs multi-layer residual convolution operations on the input image to obtain a high-level feature map; the confidence-start point initialization unit optimizes the prior anchor boxes according to the high-level feature map, predicts the start point of each predefined anchor point and its corresponding confidence score to preliminarily determine whether the anchor point belongs to a lane line instance or the background; the feature pooling unit performs rasterized sampling of pixel points on the high-level feature map according to the optimized anchor box positions to obtain lane line queries; the lane line prediction unit performs convolution operations on the high-level feature map to generate key-value pairs, and performs cross-layer optimization of self-attention operations in combination with the lane line queries to obtain the final lane line prediction.
2. The 2D lane line detection method based on anchor box optimization according to claim 1, characterized in that, The confidence-start point initialization unit includes: two convolutional pooling layers and two fully connected layers, where: the first and second convolutional pooling layers sequentially perform upsampling on the input high-level feature map, convert it into a one-dimensional vector through the second convolutional pooling layer and input it into the fully connected layer to generate n 4D vectors (pos, neg, x, y), where: n is the number of lane line anchor boxes, pos and neg respectively represent the confidence of the ground truth or the background, and (x, y) represents the start point coordinates.
3. The 2D lane line detection method based on anchor box optimization according to claim 1, characterized in that, The described perspective-based basic accuracy standard means that for the i-th sampling point on the lane line, its corresponding line width w in the image p = [(n - i * f) / n] * width, where: width represents the constant width at the starting point of the lane line, set to 30 pixels, n represents the total number of sampling points for each lane line, and the disappearance rate f represents the visual width difference between the farthest and nearest points of the lane line. The evaluation of the prediction matching degree is obtained by calculating the intersection over union of the ground truth and the prediction.
4. The 2D lane line detection method based on anchor box optimization according to any one of claims 1-3, characterized in that specifically Includes: Step 1, collect and annotate the CityLane dataset covering urban and elevated scenarios, and all images are detailedly annotated in the way of line key points; Step 2, initialize and use the feature pooling technology to construct the initial lane line anchor boxes as the basis for detection; Step 3, construct an anchor box optimization network including a backbone network, a confidence-start point initialization unit, a feature pooling unit, and a lane line prediction unit, and perform training and inference on the input image; The confidence-start point initialization unit obtains the start point estimation of N_priors anchor points and provides confidence prediction for each anchor point through two convolutional pooling layers and two fully connected layers according to the high-level feature map output from the bottleneck network in the backbone network; The starting point estimation of the N_priors anchor points refers to: the set of sampling points on the lane line where: (x, y) represents the starting point provided by the starting point initialization module for each lane line, y_i represents the vertical sampling coordinate, θ represents the tilt angle of the predefined anchor point, and the goal is to use the feature information around the cropped anchor point as a query to more accurately detect the position of the lane line; Step 4, train the anchor box optimization network, that is, prediction matching and loss calculation for lane line detection based on PGIoU. Reducing the lane line width along the vertical coordinate can provide more accurate information for the model. This training process specifically includes: 4.1) Enhance the image to increase the generalization of the model, and input the enhanced image into the backbone network to obtain the feature map: for the i-th sampling point on the lane line, obtain the specific position of the ground truth lane line in the pixel coordinate system through the lane line width representation method in PGIoU; 4.2) During the training process, the anchor box optimization network assigns one or more predicted lanes as positive samples to each ground truth lane dynamically according to the assignment cost C assign Specifically: the assignment cost C assign = C sim W sim + C cls W cls where: C cls is the classification loss between the prediction and the label, and C sim is the similarity loss between the predicted lane and the ground truth, and the matching loss is obtained by multiplying them with predefined weights respectively; 4.3) Perform queries based on the self-attention mechanism, specifically: use the pixel sampling of the optimized anchor boxes on the feature map as queries, use the output obtained by passing the feature map through a convolutional neural network as keys and values, and generate an output result including the probability of containing lane lines or background, lane line type classification / color classification, and regression values of line positions by calculating the loss function. Step 5: In the online stage, perform lane line detection on an image from the vehicle's front camera as shown in Figure 3.
5. The 2D lane line detection method based on anchor box optimization according to claim 4, wherein The initialization specifically includes: defining each lane line as a set of points with annotation attributes, and the predicted results of the lane lines include: 1) C_lane: the predicted confidence score; 2) P_cls: the probabilities of different lane line types, corresponding to 8 common lane line types included in the dataset; 3) P_color: the probabilities of different colors, corresponding to 3 color annotations; 4) xyt: the starting coordinates and tilt angle of the lane line, used to generate the reference line; 5) l: the lane line span, representing the vertical distance from the end point to the starting point of the lane line; 6) D_offset: the lateral distances of N sampling points on the lane line to the reference line.
6. The 2D lane line detection method based on anchor box optimization according to claim 4, characterized in that The feature pooling technique refers to: setting the starting points of the initial lane line anchor boxes to be farther from the main lane for the lane lines on the left and right sides of the image, uniformly sampling the starting points along the bottom edge of the image, and setting four oriented anchor points for each point to consider different directions, Θ = 0.2πxi, i ∈ {1, 2, 3, 4}.
7. The 2D lane line detection method based on anchor box optimization according to claim 4, wherein The loss function L AONet = W conf L conf + W xytl L xytl + W IoU L PGIoU , where: W is the weight of each parameter, and L conf is the focal loss between the predicted value and the label, L xytl is the smooth L1 loss for the starting point coordinates, theta angle, and lane length regression, and L PGIoU is the line IoU loss between the predicted lane and the actual ground truth.
8. The 2D lane line detection method based on anchor box optimization according to claim 4, characterized in that, The specific content of step 5 includes: 5.1) Use a residual network (ResNet) or a deep layer aggregation network (DLA34) as the backbone network for feature extraction, and convert the input of the original image into a high-level feature representation. 5.2) Anchor box initialization and optimization: Use the confidence-starting point initialization unit to predict the starting positions of each predefined anchor point and provide a confidence score for each anchor point. 5.3) Feature pooling: Sample pixel points around the optimized anchor point positions. The feature map itself is processed through different convolutional layers, and queries (queries), keys (keys), and values (values) are formed through convolutional operations, specifically including: a) Generate query: According to the adjusted anchor positions and feature maps, generate a query through convolutional operations to capture local information near the lane lines. Specifically: Q = Conv(OptimisedAnchors, X feature ), where: Conv is the convolutional operation, OptimisedAnchors is the optimised anchor box, and X feature is the input feature map; b) Generate keys and values: Based on the feature map, keys (K) and values (V) are generated through different convolutional operations to capture global and local features. Specifically: K = Conv(X feature ), V = Conv(X feature ), where the convolution and the feature map are defined the same as in a); c) Self-attention operation: Based on the generated query, key, and value, through the self-attention mechanism, weights are calculated according to the query, and then the weights are applied to the value to obtain the final lane line features, successfully capturing local information related to the lane line around the optimized anchor points. At the same time, self-attention operations are carried out using global and local features to support the final prediction of the lane line. Specifically: Where: d k is the dimension of the key-value K. Cross-level optimization is achieved through the Feature Pyramid Network (FPN) of the bottleneck network in the backbone network: The last three layers of the output of the FPN are used for self-attention operations. These outputs are sequentially passed through self-attention operations from top to bottom, forming a recursive structure, and finally obtaining the final result of cross-level optimization.
9. A 2D lane line detection system based on anchor box optimization for implementing the method according to any one of claims 1-8, characterized in that, including: A backbone network module, an anchor box initialization module, an anchor box optimization module, a lane line query module, and a loss calculation module. Among them: the backbone network module performs preliminary feature extraction on the image; the anchor box initialization module and the anchor box optimization module respectively perform the initialization of the lane line anchor boxes and the optimization of the anchor box positions according to the feature map information; the lane line query module extracts the pixel points near the anchor box positions on the feature map as queries through the feature pooling method, performs self-attention operations with the keys and values generated by the entire image, and repeats the above operations in the last three layers of the bottleneck network to obtain the cyclically optimized lane line positions and class predictions. The loss calculation module uses a perspective-based basic accuracy criterion, calculates the FocalLoss and the Euclidean distance combined as the classification / regression loss of the model prediction, and backpropagates it to the backbone network module. The backbone network module outputs the predicted results of the lane line confidence, lane line category and color, and the position coordinates of the point set on the lane line.
Citation Information
Cited By
Lane line detection method for iterative optimization from local to global and application
CN121708561A