A lane detection method based on rectangular attention mechanism
By adopting a lane line detection method based on a rectangular attention mechanism in the autonomous driving system, the problem of difficulty in detecting lane lines when the road is blocked is solved, efficient and accurate lane line detection is achieved, and the real-time performance of the system is improved.
Patent Information
- Application Number
- CN202310514016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-05-09
AI Technical Summary
The existing autonomous driving system is difficult to detect lane lines when the road is blocked and has low detection efficiency.
The lane line detection method based on the rectangular attention mechanism is adopted, and the downsampled feature map is obtained through convolution operations, and the attention mechanism network is used for feature weighting, context feature information is aggregated, and lane line instances are finally output through key point detection and clustering networks.
In extreme cases, it can effectively detect hidden lane line information, improve the accuracy and efficiency of detection, reduce the use of computing resources, and speed up the model's inference speed.
Smart Images

Figure CN116630918B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and vehicle intelligent driving, and in particular to a lane line detection method, device, equipment and computer-readable storage medium based on a rectangular attention mechanism. Background Art
[0002] Lane detection is a challenging task that requires predicting complex lane topology shapes in high real-time and distinguishing different types of lanes at the same time. In the past decade, autonomous driving technology has gradually become a research hotspot in the field of computer vision and has received widespread attention from academia and industry. In order to ensure the safe driving of autonomous vehicles, the autonomous driving system needs to accurately understand the spatial information of the lane lines. Therefore, quickly calculating the shape and position information of the lane lines from the images obtained by the front camera is a crucial step in the autonomous driving system, which requires lane detection to have both high accuracy and high real-time performance.
[0003] In recent years, most of the research has regarded lane line detection as an instance segmentation or target detection problem to be solved. Most methods based on instance segmentation use multi-category classification to segment pixels into lane lines or backgrounds. Detection-based methods use the idea of anchors to predict lane lines, but there are also some methods that use the characteristics of lane lines themselves and use anchor lines to expand the feature range of anchors to predict lane line instances. However, in the process of implementing the technical solutions of the invention in the embodiments of the present application, the inventors of the present application found that the above-mentioned technologies have at least the following technical problems: when faced with some extreme situations, such as road occlusion, the detection performance of these methods is often poor. In this case, how to extract hidden lane line information from the image is crucial.
[0004] The method based on image instance segmentation will predict the categories of all pixels in the feature map. However, for the lane line detection task, the curve characteristics of the lane line itself determine that the pixels contained in the lane line account for a very small proportion of the entire feature map. Most of the predicted pixels are irrelevant to the lane line, which leads to low computational efficiency of the model during the segmentation process.
[0005] The shortcomings of the anchor-based detection method: Most lane line images only have 2-5 lane lines, but the model will predict hundreds of anchors, which leads to an obvious long-tail effect of the model. The NMS (Non-Maximum Suppression) post-processing method is required to remove redundant lane line anchors. Summary of the invention
[0006] To this end, the technical problem to be solved by the present invention is to provide a lane line detection method based on a rectangular attention mechanism to overcome the problems of difficulty in detecting lane lines and low detection efficiency when the road is blocked in existing autonomous driving.
[0007] In order to solve the above technical problems, the present invention provides a lane line detection method based on a rectangular attention mechanism, comprising: obtaining a lane line image f∈R C×H′×W′ , where C is the number of channels, H′ is the image height, and W′ is the image width; the lane line image is input into the trained and converged lane line detection model based on the rectangular attention mechanism, and a convolution operation is performed on the lane line image to obtain a downsampled feature map, and the size of the downsampled feature map is f ds ∈R C×P×W′ , where C is the number of channels, P is the image height, and W′ is the image width; the downsampled feature map is passed through the attention mechanism network to obtain Q, K, V, where Q∈R P×W′×C′ , K∈R P×C′×W′ , V∈R P×C×W′ , the Q and the K are transformed by affine transformation to generate an attention feature map A, the correlation between each point in the attention map A and each point in the same horizontal direction is calculated by the softmax layer, and the attention feature map A is multiplied with the V by matrix multiplication to obtain the attention weighted feature map f o ∈R C×P×W′ , aggregate the contextual feature information of each point on the attention-weighted feature map and all points in the same row as the point, and obtain an upsampled lane line feature map with the same size as the original input feature map through a convolution operation; input the upsampled lane line feature map into the key point detection network of the lane line detection model based on the rectangular attention mechanism, and output the coordinates of all key points in the upsampled lane line feature map including the lane line starting point; adopt a distance-based clustering network to cluster the key points into lane line instances through the offset between the key point coordinates and the starting point to which they belong.
[0008] Preferably, the lane line images in the training set of the lane line detection model based on the rectangular attention mechanism are TuSimple.
[0009] Preferably, the same convolution operation is performed twice on the lane line image by setting the size parameter, step parameter and padding value parameter of the convolution kernel to obtain a downsampled feature map.
[0010] Preferably, the loss function calculation formula of the key point detection network is as follows:
[0011]
[0012]
[0013] Among them, yx is the starting point weight parameter, L fis the balanced cross entropy loss function for binary classification, H′ is the height of the feature map, W′ is the width of the feature map, x∈[W′ l ,W′-W′ r ]and y∈[H′-H′ b ,H′] is the non-starting point area, x∈[0,W′l]or x∈[W′-W′ r ,W′]or y∈[0,H′] is the starting point area, W′ l , W′ r , H′ b They represent the left width, right width, and bottom width of the starting point area respectively, a represents the weight coefficient of the non-starting point area, and b represents the weight coefficient of the starting point area.
[0014] Preferably, the distance-based clustering network is used to cluster key points into lane line instances through the offset between the key point coordinates and the starting point to which they belong:
[0015]
[0016] Among them, H' is the height of the feature map, W' is the width of the feature map, O yx They represent the offset between the predicted point and the predicted starting point, and the offset between the actual point and the actual starting point respectively.
[0017] Preferably, in the lane line detection model based on the rectangular attention mechanism, the overall loss function formula of the network is:
[0018] L total =λ point L point +λ offset L offset
[0019] Among them, λ point , offset are the weight values of the key point and offset loss functions, L point is the loss function of the key point detection network, L offset is the loss function of the distance-based clustering network.
[0020] Preferably, a distance-based clustering network is used to cluster key points into lane line instances by the offset between the key point coordinates and their corresponding starting points, including:
[0021] Set the starting point distance threshold;
[0022] Select a key point whose coordinate offset value between it and its corresponding starting point is less than 1 and regard it as a candidate starting point of the lane line instance;
[0023] Calculate the theoretical starting points of the remaining key points according to the coordinate offset between the key points and their corresponding starting points, retain the theoretical starting points whose distances from the candidate starting points are less than the starting point distance threshold, and the points whose distances from the candidate starting points are greater than the starting point distance threshold are regarded as error points;
[0024] All the remaining starting points are concentrated in one area, the starting points include candidate starting points and theoretical starting points, and the center of the area is regarded as the actual starting point of the lane line instance;
[0025] Finally, the key points belonging to the same starting point are clustered into the same lane line instance.
[0026] The present invention also provides a lane line detection device based on a rectangular attention mechanism, comprising:
[0027] Detection sample acquisition module: obtain lane line image f∈R C×H′×W′ , where C is the number of channels, H′ is the image height, and W′ is the image width;
[0028] Image initialization module: Input the lane line image into the trained and converged lane line detection model based on the rectangular attention mechanism, perform convolution operation on the lane line image to obtain a downsampled feature map, and the downsampled feature map is f ds ∈R C×P×W′ , P = H' / 4, where C is the number of channels, P is the image height, and W' is the image width;
[0029] Feature extraction module: The downsampled feature map is passed through the attention mechanism network to obtain Q, K, V, where Q∈R P×W′×C′ , K∈R P×C′×W′ , V∈R P×C×W′ , the Q and the K are transformed by affine transformation to generate an attention feature map A, the correlation between each point in the attention map A and each point in the same horizontal direction is calculated by the softmax layer, and the attention feature map A is multiplied with the V by matrix multiplication to obtain the attention weighted feature map f o ∈R C×P×W′ , aggregate the contextual feature information of each point on the attention-weighted feature map and all points in the same row, and obtain an upsampled lane feature map with the same size as the original input feature map through a convolution operation;
[0030] Key point detection module: input the upsampled lane feature map into the key point detection network of the lane detection model based on the rectangular attention mechanism, and output the coordinates of all key points including the lane starting point in the upsampled lane feature map;
[0031] Lane tracking module: A distance-based clustering network is used to cluster key points into lane instances based on the offset between the key point coordinates and their starting points.
[0032] The present invention also provides a lane line detection device based on a rectangular attention mechanism, comprising:
[0033] A lane line image acquisition device, used for acquiring lane line images;
[0034] A host computer is connected to the lane line image acquisition device for receiving the lane line image, and when executing the computer program, implements the steps of the lane line detection method based on the rectangular attention mechanism as described above to obtain a lane line instance image corresponding to the lane line image;
[0035] Display device: connected to the host computer for displaying the lane line instance image.
[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a lane line detection method based on a rectangular attention mechanism as described in any one of the above items are implemented.
[0037] The above technical solution of the present invention has the following advantages compared with the prior art:
[0038] The lane line detection method based on the rectangular attention mechanism provided by the present invention is aimed at the difficult-to-detect lane line scenarios in autonomous driving. The lane line image is downsampled, the attention range of each key point in the lane line image is concentrated in a rectangular area, and the attention-weighted feature map is obtained through the attention mechanism network. The contextual feature information of each point on the attention-weighted feature map and all points in the same row of the point are aggregated. Without considering the global information of the entire image, the association between the unobstructed lane line and the obscured lane line in the same horizontal area can be constructed, which reduces the computing resources used by the algorithm and speeds up the model in the reasoning stage.
[0039] In addition, a loss function for lane line key point detection is proposed, which introduces the starting point weight parameter ζ yx , which increases the importance of the starting point among the key points, thereby indirectly improving the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein
[0041] Figure 1It is a flow chart of the lane line detection method based on the rectangular attention mechanism provided by the present invention;
[0042] Figure 2 It is a flowchart of the rectangular attention mechanism network to achieve feature association;
[0043] Figure 3 It is a flow chart of distance-based clustering method;
[0044] Figure 4 is the distribution of lane line starting points in the TuSimple original image;
[0045] Figure 5 The lane focal loss provided by the present invention is about W 1 ′、W r ′、H′ b Visualization of meaning. DETAILED DESCRIPTION
[0046] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0047] The present invention provides a lane line detection method based on a rectangular attention mechanism. Aiming at the difficult-to-detect lane line scene in autonomous driving, the attention range of key points in the lane line feature map is concentrated in a rectangular area, and the association between the unobstructed lane line and the obscured lane line in the same horizontal area can be constructed without considering the global information of the entire image; in addition, a loss function for lane line key point detection is proposed, and the loss function introduces a starting point weight parameter ζ yx , which increases the importance of the starting point among the key points and indirectly improves the prediction accuracy of the model.
[0048] Reference Figure 1 As shown, a flow chart of a lane line detection method based on a rectangular attention mechanism provided by an embodiment of the present invention, the specific operation steps are as follows:
[0049] Step S101: Obtain lane line image, image size f∈R C×H′×W′ , where C is the number of channels, H′ is the image height, and W′ is the image width;
[0050] In the present invention, an original image in a lane line detection is selected for testing, the image size is 1640×590 pixels, and each image has 0-5 lane lines;
[0051] The experimental environment of this test: the CPU model is Intel(R) Xeon(R) Gold6248R CPU3.00GHz with 48 cores; the GPU model is Tesla-A100 with 32G video memory.
[0052] Step S102: inputting the lane line image into a lane line detection model based on a rectangular attention mechanism that has been trained and converged, performing a convolution operation on the lane line image to obtain a downsampled feature map, the downsampled feature map;
[0053] The lane line image is input into the trained and converged lane line detection model based on the rectangular attention mechanism, and the input feature map is downsampled by 4 times through two downsampling operations. The size of the downsampled feature map is f ds ∈R C×P×W′ ,P=H′ / 4, where C is the number of channels, P is the height of the image after downsampling, and W′ is the width of the image. The two downsampling operations use the same convolution operation. By setting the convolution kernel size to (3,1), the step size to (2,1), and the padding value to (1,0), while keeping the size in the width direction unchanged, 1 / 4 downsampling in the height direction is achieved. The size of the downsampled image is 1640×147 pixels.
[0054] Step S103: extracting feature information of lane lines from the downsampled feature map through an attention mechanism network, performing weighted summation of the feature information, and constructing an association between unobstructed lane lines and obstructed lane lines in the same horizontal area;
[0055] Reference Figure 2 As shown, Figure 2 This is a flow chart of the rectangular attention mechanism network to achieve feature association:
[0056] The downsampled feature map f ds Two convolutional layers with 1×1 filters are used respectively, and the weight parameter W is learned through the network q and W k To obtain the query map (represented by Q) and the key map (represented by K), the specific implementation process is: Q = f ds ×W q , K = f ds ×W k , where Q∈R P×W′×C′ , K∈R P×C′×W′ , “×” represents matrix multiplication, and C′ is the number of channels;
[0057] Q,K generates the attention map A∈R through affine transformation operation P×W′×W′ , which means that in the feature map f dsThe correlation between any point p among the P×W′ points and the W′ points in the same horizontal direction as the point p is recalculated through the softmax layer, and the correlation between all points in the attention feature map is controlled between 0 and 1;
[0058] In addition, the feature map f ds Then learn the weight parameter W through a convolution layer with a 1×1 filter v To generate a valuemap (represented by V), the specific implementation process of V is: V = f ds ×W v , where V∈R P×C×W′ ;
[0059] The attention map A and V are operated through matrix multiplication. The weighted feature map integrates the feature information of the original input feature map into the attention map to obtain the attention-weighted feature map f o ∈R C×P×W′ , aggregate the context feature information of each point u on the feature map and all the points in the same row, and obtain the association between the unobstructed lane line and the obstructed lane line in the same horizontal area.
[0060] Step S104: performing two convolution operations on the attention feature map to obtain an upsampled lane feature map having the same size as the original input feature map;
[0061] Step S105: inputting the upsampled lane feature map into the key point detection network of the lane detection model based on the rectangular attention mechanism, and outputting the coordinates of all key points including the lane starting point in the upsampled lane feature map;
[0062] Step S106: using a distance-based clustering network to cluster the key points into lane line instances according to the offset between the key point coordinates and the starting point to which they belong;
[0063] Reference Figure 3 , Figure 3 This is a flow chart of the distance-based clustering method:
[0064] Step S301: First, set the starting point distance threshold T dis ;
[0065] Step S302: Select a key point whose coordinate offset value between it and its corresponding starting point is less than 1 and regard it as a candidate starting point P of the lane line instance c ;
[0066] Step S303: Calculate the theoretical starting points of the remaining key points according to the coordinate offset between the key points and their corresponding starting points, and retain the key points whose distance to the candidate starting points is less than the starting point distance threshold T. dis The theoretical starting point P t , points greater than the starting point distance threshold are regarded as error points;
[0067] Step S304: All the remaining starting points (including P c With P t ) is concentrated in an area, and the center of the area is regarded as the actual starting point of the lane line instance;
[0068] Step S305: In this way, each key point obtains its corresponding starting point, and finally the key points belonging to the same starting point are clustered into the same lane line instance.
[0069] Due to the design of the offset, all key points need to be clustered through their starting points to form lane line instances, which makes the prediction of the starting point very important. Therefore, the model needs to improve the accuracy of starting point prediction. The present invention proposes a key point detection loss function lane focal loss based on the distribution characteristics of the lane line starting point. The calculation formula of the key point detection loss function is as follows:
[0070]
[0071] Among them, yx is the starting point weight parameter, which improves the importance of the starting point in the prediction process according to the distribution characteristics of the lane line’s starting point. f It is a binary balanced cross entropy loss function to solve the imbalance problem between key points and non-key points. H′ is the height of the feature map and W′ is the width of the feature map.
[0072] Figure 4 The distribution of lane line starting points in the TuSimple original image. The darker the color, the more starting points there are at that location. Figure 4 Inspired by the distribution of starting points in the image, this paper assumes that the lane line starting points are concentrated in a small area within the image edge range, which is called the starting point area. The starting point weight parameter ζ yx The calculation formula is:
[0073]
[0074] Among them, x∈[W′ l ,W′-W′ r ]and y∈[H′-H′ b ,H′] is the non-starting point area, x∈[0,W′ l]or x∈[W′-W′ r ,W′]or y∈[0,H′] is the starting point area, refer to Figure 5 As shown, Figure 5 is the lane focal loss about W′ l , W′ r , H′ b Visualization of meaning, W′ l , W′ r , H′ b They represent the left width, right width, and bottom width of the starting point area respectively, a represents the weight coefficient of the non-starting point area, and b represents the weight coefficient of the starting point area.
[0075] For the distance-based clustering network, the present invention proposes a loss function L of the offset between the key point and its corresponding starting point: offset , use the starting point of each lane line to represent the lane line instance, and regress the offset between each key point and its starting point, which is specifically expressed as:
[0076]
[0077] Among them, H' is the height of the feature map, W' is the width of the feature map, O yx They represent the offset between the predicted point and the predicted starting point, and the offset between the actual point and the actual starting point respectively.
[0078] The key point detection loss function L based on the lane line starting point distribution characteristics point and the loss function L of the offset between the key point and its corresponding starting point offset By adjusting the weight coefficient, iteratively training the optimization model, and saving the network model when the combination of the two loss functions reaches the minimum value, the network model realizes the comprehensive consideration of lane line detection in key points and offset prediction. The overall loss function of the lane line detection network based on the rectangular attention mechanism is expressed as:
[0079] L total =λ point L point +λ offset L offset
[0080] Among them, λ point , offset are the weight values of the key point and offset loss functions respectively.
[0081] In this embodiment, a lane line detection model based on a rectangular attention mechanism is trained, including a rectangular attention mechanism network, a key point detection network, and a distance-based clustering network. The rectangular attention mechanism network is used to construct the association between unobstructed lane lines and obscured lane lines in the same horizontal area. The key point detection network and the distance-based clustering network cluster the key points into lane line instances through the offset between the key point coordinates and the starting points to which they belong, thereby realizing lane line detection in extreme special scenarios and improving detection accuracy.
[0082] The performance of the model of the present invention is compared with other better models on the public dataset images, including:
[0083] SCNN uses a multi-class classification method to predict the category of each pixel in the input feature map. There are n+1 categories in total, where n represents the number of lane lines and the other category is the background.
[0084] Fast-HBNet uses the original image and the flipped image to locate the lane line based on the horizontal symmetry of the lane line.
[0085] PointLaneNet uses the idea of anchors and uses each pixel in the feature map as an anchor point to predict lane lines, but the anchor points contain too few lane line features;
[0086] LaneATT uses the linear prior structure of lane lines, replaces anchor points in PointLaneNet with anchor lines, and extracts the corresponding lane line features based on equally spaced pixels in the anchor lines.
[0087] Inspired by human posture estimation, PINet regards lane line detection as a key point detection and clustering problem. It uses an hourglass network to predict key points on the lane line and predicts an embedded feature for each key point. It clusters key points whose embedded feature similarity is greater than a threshold into the same lane line instance.
[0088] Unlike PINet, which requires additional calculation of embedded features, FOLOlane predicts the offset between each key point and its adjacent key points, and then gradually extends the adjacent key points outward through the key points to achieve clustering. However, due to the dense dependencies between key points, FOLOLane may deviate from expectations due to incorrect prediction of some key points during the lane line instance construction process;
[0089] In order to avoid the situation where FOLOLane may deviate from the expected prediction due to some key point prediction errors during the lane line instance construction process, GANet indirectly clusters the key points into lane line instances by predicting the offset between the key points and their corresponding starting points;
[0090] The present invention regards lane line detection as a key point detection and clustering problem. According to the correlation between key points between lane lines in the same horizontal local area and the high dependence of all key points on the starting point, a lane line detection method based on the rectangular attention mechanism is proposed, which enhances the performance of the lane line detection algorithm for difficult-to-detect scenes and improves the inference speed of the model.
[0091] The overall prediction accuracy of the model, the prediction accuracy of some special scenarios, and the time performance comparison are shown in Table 1:
[0092] Table 1: Algorithm time performance comparison
[0093] Method Total Crowded Dazzle Shadow FPS SCNN 71.60 69.70 58.50 66.90 7.5 UFLDv2 75.90 74.90 65.70 75.30 312 LaneATT 75.11 73.32 65.69 69.58 250 ESAnet 74.20 73.10 63.10 75.10 123 Fast-HBNet 73.10 71.60 64.70 66.70 39 Bézier curve 75.57 73.20 69.20 76.74 150 PINet 74.40 72.30 66.30 68.40 25 (Ours) 77.11 76.40 68.45 78.24 89
[0094] A specific embodiment of the present invention further provides a lane line detection device based on a rectangular attention mechanism, comprising:
[0095] A lane line image acquisition device, used for acquiring lane line images;
[0096] A host computer is connected to the lane line image acquisition device for receiving the lane line image, and when executing the computer program, implements the steps of the lane line detection method based on the rectangular attention mechanism as described above to obtain a lane line instance image corresponding to the lane line image;
[0097] Display device: connected to the host computer for displaying the lane line instance image.
[0098] A specific embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the lane line detection method based on the rectangular attention mechanism are implemented.
[0099] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0100] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0101] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0103] Obviously, the above embodiments are merely examples for clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived from these are still within the protection scope of the invention.
Claims
1. A lane detection method based on rectangular attention mechanism, It is characterized in that include: Get the lane line image f∈R C×H′×W′ , where C is the number of channels, H′ is the image height, and W′ is the image width; The lane line image is input into the trained and converged lane line detection model based on the rectangular attention mechanism, and a convolution operation is performed on the lane line image to obtain a downsampled feature map, where the downsampled feature map is f ds ∈R C ×P×W′ , where C is the number of channels, P is the height of the image after downsampling, and W′ is the image width; The downsampled feature map is passed through the attention mechanism network to obtain Q, K, V, where Q∈R P×W′×C′ , K∈R P ×C′×W′ , V∈R P×C×W′ , the Q and the K are transformed by affine transformation to generate an attention feature map A, the correlation between each point in the attention map A and each point in the same horizontal direction is calculated by the softmax layer, and the attention feature map A is multiplied with the V by matrix multiplication to obtain the attention weighted feature map f o ∈R C×P×W′ , aggregate the contextual feature information of each point on the attention-weighted feature map and all points in the same row, and obtain an upsampled lane feature map with the same size as the original input feature map through a convolution operation; The upsampled lane feature map is input into the key point detection network of the lane detection model based on the rectangular attention mechanism, and the coordinates of all key points including the lane starting point in the upsampled lane feature map are output; the loss function calculation formula of the key point detection network is as follows: Among them, yx is the starting point weight parameter, L f is the balanced cross entropy loss function for binary classification, H′ is the height of the feature map, W′ is the width of the feature map, x∈[W l ′,W′-W r ′]and y∈[H′-H b ′,H′] is the non-starting point area, x∈[0,W l ′]or x∈[W′-W r ′,W′]or y∈[0,H′] is the starting point area, W l ′、W r ′、H′ b They represent the left width, right width, and bottom width of the starting point area, respectively; a represents the weight coefficient of the non-starting point area; and b represents the weight coefficient of the starting point area. A distance-based clustering network is used to cluster key points into lane line instances through the offset between the key point coordinates and their starting points.
2. According to the lane line detection method based on rectangular attention mechanism according to claim 1, It is characterized in that The lane line images in the training set of the lane line detection model based on the rectangular attention mechanism are TuSimple.
3. According to the lane line detection method based on rectangular attention mechanism according to claim 1, It is characterized in that Performing a convolution operation on the lane line image to obtain a downsampled feature map includes: performing the same convolution operation on the lane line image twice by setting a size parameter, a step parameter, and a padding value parameter of the convolution kernel to obtain a downsampled feature map.
4. The lane line detection method based on rectangular attention mechanism according to claim 1, It is characterized in that The loss function of clustering key points into lane line instances by using the distance-based clustering network through the offset between the key point coordinates and their starting points is: Among them, H' is the height of the feature map, W' is the width of the feature map, O yx They represent the offset between the predicted point and the predicted starting point, and the offset between the actual point and the actual starting point respectively.
5. The lane line detection method based on rectangular attention mechanism according to claim 1, It is characterized in that In the lane detection model based on the rectangular attention mechanism, the overall loss function formula of the network is: Among them, λ point , offset are the weight values of the key point and offset loss functions, L point is the loss function of the key point detection network, L offset is the loss function of the distance-based clustering network.
6. A lane line detection method based on a rectangular attention mechanism as described in claim 1, It is characterized in that A distance-based clustering network is used to cluster key points into lane line instances by the offset between the key point coordinates and their starting points, including: Set the starting point distance threshold; Select a key point whose coordinate offset value between it and its corresponding starting point is less than 1 and regard it as a candidate starting point of the lane line instance; Calculate the theoretical starting points of the remaining key points according to the coordinate offset between the key points and their corresponding starting points, retain the theoretical starting points whose distances from the candidate starting points are less than the starting point distance threshold, and the points whose distances from the candidate starting points are greater than the starting point distance threshold are regarded as error points; All the remaining starting points are concentrated in one area, the starting points include candidate starting points and theoretical starting points, and the center of the area is regarded as the actual starting point of the lane line instance; Finally, the key points belonging to the same starting point are clustered into the same lane line instance.
7. A lane detection device based on rectangular attention mechanism, include: Detection sample acquisition module: obtain lane line image f∈R C×H′×W′ , where C is the number of channels, H′ is the image height, and W′ is the image width; Image initialization module: Input the lane line image into the trained and converged lane line detection model based on the rectangular attention mechanism, perform convolution operation on the lane line image to obtain a downsampled feature map, and the downsampled feature map is f ds ∈R C×P×W′ , where C is the number of channels, P is the height of the image after downsampling, and W′ is the image width; Feature extraction module: The downsampled feature map is passed through the attention mechanism network to obtain Q, K, V, where Q∈R P ×W′×C′ , K∈R P×C′×W′ , V∈R P×C×W′ , the Q and the K are transformed by affine transformation to generate an attention feature map A, the correlation between each point in the attention map A and each point in the same horizontal direction is calculated by the softmax layer, and the attention feature map A is multiplied with the V by matrix multiplication to obtain the attention weighted feature map f o ∈R C×P×W′ , aggregate the contextual feature information of each point on the attention-weighted feature map and all points in the same row, and obtain an upsampled lane feature map with the same size as the original input feature map through a convolution operation; Key point detection module: input the upsampled lane feature map into the key point detection network of the lane detection model based on the rectangular attention mechanism, and output the coordinates of all key points in the upsampled lane feature map including the lane starting point; the loss function calculation formula of the key point detection network is as follows: Among them, yx is the starting point weight parameter, L f is the balanced cross entropy loss function for binary classification, H′ is the height of the feature map, W′ is the width of the feature map, x∈[W l ′,W′-W r ′]and y∈[H′-H b ′,H′] is the non-starting point area, x∈[0,W l ′]or x∈[W′-W r ′,W′]or y∈[0,H′] is the starting point area, W l ′、W r ′、H′ b They represent the left width, right width, and bottom width of the starting point area, respectively; a represents the weight coefficient of the non-starting point area; and b represents the weight coefficient of the starting point area. Lane tracking module: It uses distance-based clustering technology to cluster key points into lane instances through the offset between the key point coordinates and their starting points.
8. A lane detection device based on rectangular attention mechanism, It is characterized in that include: A lane line image acquisition device, used for acquiring lane line images; A host computer is communicatively connected to the lane line image acquisition device, receives the lane line image, and when executing the computer program, implements the steps of a lane line detection method based on a rectangular attention mechanism as described in any one of claims 1 to 6 to obtain a lane line instance image corresponding to the lane line image; Display device: connected to the host computer for displaying the lane line instance image.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a lane line detection method based on a rectangular attention mechanism as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Lane line detection method and system
CN112883807A
Detection method using fusion network based on attention mechanism, and terminal device
US11222217B1