Lane line detection method and device, electronic equipment and storage medium

By extracting and fusing features from road images, and combining 3D reference points and location information features, the accuracy problem of lane line detection in monocular images under non-flat ground scenes is solved, achieving more accurate lane line detection.

CN116863428BActive Publication Date: 2026-05-22LINKTECH NAVI TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LINKTECH NAVI TECH
Filing Date
2023-07-18
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

The accuracy of lane line detection based on monocular images in existing technologies is low, especially in non-flat ground scenes where the 3D proxy representation is severely deformed, affecting the accurate estimation of road structure.

Method used

By extracting features from road images, global and local features of lane lines are obtained, and after fusion processing, lane line query features are obtained. Combined with three-dimensional reference points and position information features, position offset is calculated to predict lane line contours.

Benefits of technology

It improves the accuracy of lane detection in monocular images, especially in non-flat terrain scenes, where it can more accurately correct the position of lane lines in three-dimensional space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863428B_ABST
    Figure CN116863428B_ABST
Patent Text Reader

Abstract

Embodiments of the application disclose a lane line detection method and device, electronic equipment and a storage medium; which can be used in the field of map or intelligent transportation or automatic driving, including: performing feature extraction processing on a road image to obtain a feature map and lane line global features, the lane line global features carrying a detection upper limit value N of the number of lane lines; obtaining lane line local features, and performing feature fusion processing on the lane line global features and the lane line local features to obtain lane line query features, the lane line local features carrying a sampling point value M used for detecting the position of the lane line; obtaining N*M three-dimensional reference points based on the lane line query features; obtaining three-dimensional position information features; obtaining N*M position offsets based on the feature map, the three-dimensional position information features, the N*M three-dimensional reference points and the lane line query features; and offsetting the corresponding three-dimensional reference points based on the position offsets to obtain position prediction points. The application can improve the accuracy of lane line detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically to a lane line detection method, apparatus, electronic device, and storage medium. Background Technology

[0002] In existing technologies, lane detection based on monocular images is typically achieved in the following two ways.

[0003] One approach is to reconstruct 3D lane lines by performing pixel-by-pixel depth estimation on the monocular image based on the 2D segmentation results. However, this method requires high-quality depth data for training and heavily relies on the accuracy of the estimated depth. Another approach is to construct a 3D surrogate representation of the monocular image by projecting it onto a bird's-eye view (BEV) using inverse perspective mapping (IPM), and then detect 3D lane lines based on this representation. However, the IPM relied upon by this approach is based on the assumption of a flat surface, which does not hold true in many real-world driving scenarios, such as uphill, downhill, or uneven terrain. This leads to distortion between the 3D surrogate representation and the original image, and this distorted 3D surrogate representation inevitably impairs the ability to accurately estimate road structure.

[0004] Therefore, existing technical solutions for lane line detection based on monocular images suffer from low accuracy. Summary of the Invention

[0005] This application provides a lane line detection method, apparatus, electronic device, and storage medium, which can improve the problem of low accuracy in lane line detection based on monocular images in the prior art.

[0006] This application provides a lane line detection method, which includes: performing feature extraction processing on a road image to obtain a feature map and global lane line features, wherein the feature map contains lane line information, and the global lane line features carry an upper limit value N for detecting the number of lane lines; acquiring local lane line features, and performing feature fusion processing on the global lane line features and the local lane line features to obtain lane line query features, wherein the local lane line features carry sampling point values ​​M for detecting lane line positions; the lane line query features carry both the detection upper limit value N and the sampling point values ​​M; and obtaining N×M three-dimensional reference points based on the lane line query features; wherein the three-dimensional reference points... The test point is used to indicate the reference position of the lane line coordinate point in the three-dimensional coordinate system; three-dimensional position information features are obtained, wherein the three-dimensional position information features are used to correct the height position of the lane line in three-dimensional space; based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points and the lane line query features, N×M position offsets are obtained, wherein the position offsets correspond one-to-one with the three-dimensional reference points; for each position offset, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point, wherein the N×M position prediction points represent: for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0007] This application embodiment also provides a lane line detection device, the device comprising:

[0008] The feature extraction unit is used to perform feature extraction processing on the road image to obtain a feature map and global lane line features. The feature map contains lane line information, and the global lane line features carry an upper limit value N for the detection of the number of lane lines.

[0009] The feature fusion unit is used to acquire local lane line features and perform feature fusion processing on the global lane line features and the local lane line features to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M.

[0010] The reference point acquisition unit is used to obtain N×M three-dimensional reference points based on the lane line query features; wherein the three-dimensional reference points are used to indicate the reference position of the lane line coordinate points in the three-dimensional coordinate system.

[0011] A location information unit is used to acquire three-dimensional location information features, wherein the three-dimensional location information features are used to correct the height position of the lane line in three-dimensional space;

[0012] The offset acquisition unit is used to obtain N×M position offsets based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points and the lane line query features, wherein the position offsets correspond one-to-one with the three-dimensional reference points;

[0013] The prediction point acquisition unit is used to offset the corresponding three-dimensional reference point based on the position offset for each position offset to obtain a position prediction point. The N×M position prediction points represent that for each of the N lane lines, the contour of the corresponding lane line is indicated by M position prediction points.

[0014] In some embodiments, the feature extraction unit includes:

[0015] The first extraction subunit is used to perform a first feature extraction process on the road image to obtain the feature map;

[0016] The second extraction subunit is used to perform a second feature extraction process on the feature map to obtain the global features of the lane line.

[0017] In some embodiments, the second extraction subunit includes:

[0018] The third extraction sub-unit is used to perform third feature extraction processing on the feature map to obtain a first feature sub-map, wherein the first feature sub-map carries the detection upper limit value N;

[0019] The fourth extraction sub-unit is used to perform a fourth feature extraction process on the feature map to obtain a second feature sub-map, wherein the second feature sub-map carries feature dimension values;

[0020] The feature fusion sub-unit is used to perform feature fusion processing on the first feature sub-map and the second feature sub-map to obtain the global features of the lane line.

[0021] In some embodiments, the feature fusion unit is specifically used to perform broadcast addition processing on the global features of the lane line and the local features of the lane line to obtain the lane line query features.

[0022] In some embodiments, the location information unit includes:

[0023] The coordinate point creation sub-unit is used to create multiple three-dimensional coordinate points located on the same horizontal plane, wherein each of the three-dimensional coordinate points has its own three-dimensional coordinate values;

[0024] The coordinate projection subunit is used to project the plurality of three-dimensional coordinate points to obtain a plurality of two-dimensional coordinate points, wherein each of the three-dimensional coordinate points has its own two-dimensional coordinate values.

[0025] A matrix creation sub-unit is used to create an initial two-dimensional matrix with a feature dimension of 3, wherein the initial two-dimensional matrix corresponds to a two-dimensional coordinate range;

[0026] The coordinate traversal subunit is used to, for the plurality of two-dimensional coordinate points, if there is a two-dimensional coordinate point whose two-dimensional coordinate value falls within the two-dimensional coordinate range, obtain the three-dimensional coordinate value corresponding to the two-dimensional coordinate point, and use the three-dimensional coordinate value to fill the feature dimension of the corresponding position point of the initial two-dimensional matrix, until all the two-dimensional coordinate points are traversed.

[0027] A two-dimensional matrix sub-unit is used to fill the position points with a preset value when the initial two-dimensional matrix still has position points that do not correspond to any two-dimensional coordinate point, so as to obtain a target two-dimensional matrix.

[0028] The three-dimensional position sub-unit is used to perform multi-layer nonlinear transformation on the target two-dimensional matrix to obtain the three-dimensional position information features.

[0029] In some embodiments, the offset acquisition unit is specifically used to update the lane line query features based on the feature map, the three-dimensional position information features, and the N×M three-dimensional reference points to obtain feature update results; and to obtain N×M position offsets based on the feature update results.

[0030] In some embodiments, the apparatus further includes:

[0031] The step jump unit is used to take the feature update result as the new lane line query feature and jump to the step: based on the lane line query feature, obtain N×M three-dimensional reference points;

[0032] The threshold is reached unit, which is used to obtain the final feature update result until the number of jumps reaches a preset value.

[0033] In some embodiments, when the lane line query feature is the result of the previous feature update, the location information unit includes:

[0034] The offset acquisition subunit is used to acquire the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result;

[0035] A transformation matrix generation sub-unit is used to generate a transformation matrix based on the rotation offset around the x-axis and the z-axis offset;

[0036] The updated matrix sub-unit is used to update the target two-dimensional matrix that is in the same round as the feature update result based on the transformation matrix, so as to obtain a new target two-dimensional matrix.

[0037] The three-dimensional position sub-unit is used to perform multi-level nonlinear transformations on the new target two-dimensional matrix to obtain the three-dimensional position information features.

[0038] In some embodiments, the offset acquisition subunit includes:

[0039] The historical acquisition subunit is used to acquire the feature map and the historical target two-dimensional matrix, wherein the historical target two-dimensional matrix is ​​a target two-dimensional matrix that is in the same round as the feature update result;

[0040] The concatenation sub-unit is used to concatenate the feature map and the historical target two-dimensional matrix to obtain the concatenation result;

[0041] The extraction result sub-unit is used to perform a fifth feature extraction process on the splicing result to obtain the feature extraction result;

[0042] The pooling result sub-unit is used to perform pooling processing on the feature extraction result to obtain the pooling result;

[0043] The offset sub-unit is used to perform multi-layer nonlinear transformation on the pooling result to obtain the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result.

[0044] In some embodiments, the offset acquisition unit includes:

[0045] The reference point projection subunit is used to project the N×M three-dimensional reference points to obtain N×M two-dimensional reference points.

[0046] The cross-attention subunit is used to perform deformable cross-attention processing on the lane line query features, the feature map, the three-dimensional position information features, and the N×M two-dimensional reference points to obtain feature update results.

[0047] In some embodiments, the cross-attention subunit includes:

[0048] The self-attention sub-unit is used to perform self-attention processing on the lane line query features to obtain the self-attention processing result.

[0049] The cross-attention subunit is used to perform deformable cross-attention processing on the self-attention processing result, the feature map, the three-dimensional position information features, and the multiple two-dimensional reference points to obtain a deformable attention processing result.

[0050] The nonlinear transformation subunit is used to perform multi-level nonlinear transformations on the deformable attention processing result to obtain the feature update result.

[0051] In some embodiments, the apparatus further includes:

[0052] The max pooling unit is used to perform max pooling on the feature update result to obtain the max pooling result.

[0053] The category acquisition unit is used to perform multi-level nonlinear transformation on the max pooling result to obtain the category to which the lane line belongs.

[0054] In some embodiments, the offset acquisition unit includes:

[0055] The visibility subunit is used to perform multi-layer nonlinear transformation on the feature update result to obtain multiple position offsets and the visibility corresponding to each position offset;

[0056] The reference point deletion subunit is used to delete the three-dimensional reference point and the position offset corresponding to the three-dimensional reference point when the visibility characterization is in an invisible state.

[0057] In some embodiments, the apparatus further includes:

[0058] The model acquisition unit is used to acquire the lane line detection model;

[0059] The model training unit is used to train the lane line detection model to obtain the trained lane line detection model.

[0060] The model application unit is used to execute the lane detection method using the trained lane detection model to determine the N×M location prediction points.

[0061] In some embodiments, the model training unit includes:

[0062] The training feature subunit is used to perform the first feature extraction process on the training road image to obtain the training feature map;

[0063] A global subunit is trained to perform a second feature extraction process on the trained feature map to obtain the global features of the trained lane lines.

[0064] The feature fusion subunit is used to randomly initialize the local features of the training lane lines and perform feature fusion processing on the global features of the training lane lines and the local features of the training lane lines to obtain the training feature fusion result, which is denoted as the training lane line query feature.

[0065] The training reference point sub-unit is used to perform a first multi-level nonlinear transformation on the training lane line query features to obtain multiple training three-dimensional reference points, and to perform projection processing on each of the training three-dimensional reference points to obtain multiple training two-dimensional reference points.

[0066] The training location information subunit is used to acquire training 3D location information features;

[0067] The training feature update subunit is used to perform deformable cross attention processing on the training lane line query features, the training feature map, the training three-dimensional position information features, and the multiple training two-dimensional reference points to obtain the training feature update result.

[0068] A training visibility subunit is used to perform a second multi-layer nonlinear transformation on the training feature update result to obtain multiple training position offsets and the training visibility corresponding to each training position offset.

[0069] The max pooling subunit is used to perform max pooling on the training feature update result to obtain the training max pooling result.

[0070] The training category sub-unit is used to perform a sixth-level nonlinear transformation on the training max pooling result to obtain the training category to which the lane line belongs.

[0071] The training segmentation result subunit is used to obtain the training lane line segmentation result based on the training feature map and the training lane line global features;

[0072] The loss function subunit is used to construct a loss function based on the training lane segmentation results, training 3D position information features, training position offset, training visibility, and training category.

[0073] The training completion sub-unit is used to determine that the lane detection model has been trained when the loss function meets the preset requirements, and to obtain the trained lane detection model.

[0074] The lane detection method provided in this application embodiment can perform feature extraction processing on a road image to obtain a feature map and global lane features. The global lane features carry an upper limit value N for detecting the number of lanes. Local lane features carrying sampling point values ​​M are obtained, where sampling points M are used to detect lane position. The global lane features and local lane features are then fused to obtain lane query features that carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lanes are obtained. Based on the feature map, three-dimensional position information features, N×M three-dimensional reference points, and lane query features, N×M position offsets are obtained, with each position offset corresponding to a three-dimensional reference point. Based on each of the N×M position offsets, the corresponding three-dimensional reference point is offset to obtain a predicted position point, resulting in a total of N×M predicted position points. The N×M location prediction points represent the contours of each of the N lane lines.

[0075] In this embodiment, the upper limit value N for the detection of the number of lanes in the road image and the sampling point value M corresponding to each lane line can be carried by global lane line features and local lane line features, respectively. The global lane line features and local lane line features are fused to obtain lane line query features that carry both the detection upper limit value N and the sampling point value M. Based on the lane line query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features for correcting the height features of the lane lines are acquired. Then, based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, N×M position offsets can be obtained; subsequently, based on the position offsets and the corresponding three-dimensional reference points, N×M position prediction points are obtained. In this embodiment, by introducing the upper limit value N for the detection of the number of lane lines, the sampling point value M corresponding to each lane line, and the three-dimensional position information features for correcting the height features of the lane lines, the accuracy of lane line detection based on monocular images can be improved. Attached Figure Description

[0076] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0077] Figure 1a This is a schematic diagram illustrating the application of the lane detection model provided in the embodiments of this application;

[0078] Figure 1b This is a schematic flowchart of a lane line detection method provided in an embodiment of this application;

[0079] Figure 1c A schematic diagram of the modules of the lane line detection model provided in an embodiment of this application is shown;

[0080] Figure 1d It shows Figure 1c A schematic diagram of the lane line perception query generator module in the image;

[0081] Figure 1e (a) shows a schematic diagram of a road image included in one specific embodiment;

[0082] Figure 1e (b) shows a schematic diagram of the lane line mask;

[0083] Figure 1f (a) shows the initial three-dimensional positions of multiple training 3D coordinate points (2) created on the same horizontal plane and the real ground (1);

[0084] Figure 1f (b) shows the adjusted 3D positions of multiple training 3D coordinate points (2) created on the same horizontal plane and the real ground (1);

[0085] Figure 1g (a) in the figure shows the quantitative results of different models on a subset of parameters on the real-world dataset OpenLane;

[0086] Figure 1g (b) in the figure shows the quantitative results of different models on the real-world dataset OpenLane regarding another part of the parameters;

[0087] Figure 1g (c) in the figure shows the quantitative results of different models on the Apollo dataset synthesized by the game engine;

[0088] Figure 1h In this context, (a) represents the true value of the three-dimensional lane line;

[0089] Figure 1h (b) in this application indicates the detection of lane lines in the embodiments of this application;

[0090] Figure 1h (c) in the text represents the optimal model for lane line detection in the current technology;

[0091] Figure 1h (d) in the figure illustrates the relationship between the detection results and the true value of lane lines in the embodiments of this application in three-dimensional space;

[0092] Figure 2 This is a schematic flowchart of a lane line detection method provided in another embodiment of this application;

[0093] Figure 3 This is a schematic diagram of a lane line detection device provided in one embodiment of this application;

[0094] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0095] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0096] This application provides a lane line detection method, apparatus, electronic device, and storage medium.

[0097] Specifically, the lane line detection device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet, smart Bluetooth device, laptop, or personal computer (PC). The server can be a single server or a server cluster consisting of multiple servers.

[0098] In some embodiments, the lane line detection device may also be integrated into multiple electronic devices, such as multiple servers, with multiple servers implementing the lane line detection method of this application.

[0099] In some embodiments, the server may also be implemented as a terminal.

[0100] Please see details Figure 1aThe method provided in this application embodiment can perform feature extraction processing on road images to obtain feature maps and global lane line features. The feature map contains lane line information, and the global lane line features carry a detection upper limit value N for the number of lane lines. Local lane line features are obtained, and the global lane line features and the local lane line features are fused to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane line query features, N×M three-dimensional reference points are obtained. These three-dimensional reference points are used to indicate... The system displays the reference position of the lane line coordinate points in a three-dimensional coordinate system; it acquires three-dimensional position information features, wherein the three-dimensional position information features are used to correct the height position of the lane line in three-dimensional space; based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, it obtains N×M position offsets, wherein the position offsets correspond one-to-one with the three-dimensional reference points; for each position offset, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point, wherein the N×M position prediction points represent: for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0101] The above method can improve the accuracy of lane line detection based on monocular images by introducing an upper limit value N for the number of lane lines, the sampling point value M corresponding to each lane line, and three-dimensional position information features used to correct the three-dimensional position of the lane lines.

[0102] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0103] The following sections provide detailed descriptions of each example. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0104] In this embodiment, a lane line detection method is provided, such as... Figure 1b As shown, this lane detection method is applied to electronic devices, and the specific process of this method may include the following steps 110 to 160:

[0105] 110. Perform feature extraction processing on the road image to obtain a feature map and global lane line features, wherein the feature map contains lane line information and the global lane line features carry an upper limit value N for detecting the number of lane lines.

[0106] A road image is an image of a road that includes lane lines. Lane lines are markings on the road surface, such as lines or road edges, that convey traffic information to road users, including guidance and restrictions.

[0107] The feature map is a high-dimensional feature containing lane line information, abstracted from a road image after feature extraction processing. The detection upper limit N refers to the upper limit of the number of lane lines that the lane line detection method provided in this embodiment can detect for a road image. If a road image actually contains more than N lane lines, the lane line detection method provided in this embodiment can detect a maximum of N lane lines. If a road image actually contains less than or equal to N lane lines, the lane line detection method provided in this embodiment can detect all lane lines; for the portion below N, the output detection result is background. For example, let's assume N is 20. If a road image actually contains 6 lane lines, the lane line detection method provided in this embodiment can detect all 6 lane lines, and the remaining 14 detection results are background.

[0108] Optionally, in one specific embodiment, step 110 may specifically include the following steps 111 to 112:

[0109] 111. Perform a first feature extraction process on the road image to obtain the feature map.

[0110] Optionally, please see details. Figure 1c It can be done Figure 1c The illustrated combined network performs a first feature extraction process on the road image to obtain a feature map. Optionally, the size of the road image can be 720×960.

[0111] In one specific implementation, the combined network includes a backbone network, a feature pyramid network (FPN), and a convolutional network. Correspondingly, the first feature extraction process may specifically include the following steps:

[0112] By processing road images using a backbone network, three feature maps at different scales can be obtained, with spatial reduction ratios of 1 / 8, 1 / 16, and 1 / 32, respectively. Correspondingly, the sizes of the three feature maps are 90×120, 45×60, and 23×30, respectively.

[0113] Subsequently, the three extracted feature maps are input into the FPN to obtain four layers of feature maps with scales of 1 / 8, 1 / 16, 1 / 32, and 1 / 64 of the road image, respectively. Correspondingly, the sizes of the four feature maps are 90×120, 45×60, 23×30, and 12×15. For the feature maps with sizes of 45×60, 23×30, and 12×15, their sizes are upsampled to 90×120 using dilated convolution.

[0114] Then, for four feature maps of the same size, the four feature maps can be fused into a single fused feature map by adding them at corresponding positions.

[0115] Next, the single fused feature map is fed into a two-layer convolutional network for convolution processing to obtain the feature map output in step 111. In the two-layer convolutional network, the kernel size of the first layer is 3, the stride is 1, and the padding is 1; the kernel size of the second layer is 1, the stride is 1, and the padding is 0.

[0116] In computer vision tasks, the network that extracts features from images is called the backbone network.

[0117] A feature map can be represented by three parameters: feature dimension C, height H, and width W. For details, please refer to [link to relevant documentation]. Figure 1d Specifically, it can be represented as: C×H×W.

[0118] 112. Perform a second feature extraction process on the feature map to obtain the global features of the lane lines.

[0119] Optionally, in one specific embodiment, step 112 may specifically include the following steps 1121 to 1123:

[0120] 1121. Perform a third feature extraction process on the feature map to obtain a first feature sub-map.

[0121] The first feature sub-map carries the detection upper limit value N. The third feature extraction process can be implemented by two connected convolutional networks. See details below. Figure 1d The feature map can be sequentially passed through a first convolutional network and a second convolutional network to obtain the corresponding processing result: the first feature sub-map. The first feature sub-map can be represented by three parameters: the detection upper bound N, the height H, and the width W. For details, please refer to [link to documentation / reference]. Figure 1d Specifically, it can be represented as: N×H×W. The first feature sub-map is composed of the activation feature maps corresponding to each lane line.

[0122] 1122. Perform a fourth feature extraction process on the feature map to obtain a second feature sub-map.

[0123] The second feature sub-map carries the feature dimension value C. The fourth feature extraction process can be implemented using two connected convolutional networks. See details below. Figure 1d The feature map can be sequentially passed through a first and third convolutional network to obtain the corresponding processing result: the second feature submap. The second feature submap can be represented by three parameters: feature dimension value C, height H, and width W. For details, please refer to [link to documentation / reference]. Figure 1d Specifically, it can be represented as: C×H×W. The second feature sub-map is the foreground feature map corresponding to the road image after abstraction.

[0124] The number of output channels of the second convolutional network is different from that of the third convolutional network, resulting in different dimensions of the output feature channels: N×H×W for N channels and C×H×W for C channels.

[0125] 1123. Perform feature fusion processing on the first feature sub-image and the second feature sub-image to obtain the global features of the lane line.

[0126] When performing feature fusion processing on the first feature sub-image and the second feature sub-image, H×W can be merged into one parameter to obtain the first feature sub-image N×(H×W) and the second feature sub-image C×(H×W).

[0127] Then we obtain the transpose of the second feature subgraph: (H×W)×C.

[0128] Then, a matrix multiplication of the transposes of the first and second feature subgraphs is performed to eliminate (H×W), resulting in the global lane line feature N×C.

[0129] 120. Obtain the local features of the lane lines, and perform feature fusion processing on the global features of the lane lines and the local features of the lane lines to obtain the lane line query features.

[0130] The lane line local features carry sampling point values ​​M for detecting the lane line position. Specifically, the lane line local features can be represented as M×C. M means that for each of the N lane lines, its position can be described by M sampling points. That is, during the detection of the position of a certain lane line, M points can predict M local coordinates, and connecting these M local coordinates together forms the outline of the lane line.

[0131] Lane line local features can be obtained through training. Optionally, trained lane line local features can be acquired. During the training phase, the trained lane line local features can be stored for later acquisition during the application phase.

[0132] The lane line query feature carries both the detection upper limit value N and the sampling point value M.

[0133] Optionally, in one specific embodiment, "performing feature fusion processing of global lane line features and local lane line features" may specifically include: performing broadcast addition processing on the global lane line features and the local lane line features to obtain the lane line query features.

[0134] The specific process of broadcast addition includes: adding a parameter to the global lane line feature N×C to obtain N×1×C; then copying the added parameter M times, and using the M value of the local lane line feature to fill the parameter obtained by copying M times, to obtain the lane line query feature (N×M)×C. For details, please refer to [link to relevant documentation]. Figure 1d (N×M) represents merging two parameters into one, changing the original three parameters N×M×C to two parameters (N×M)×C, which facilitates subsequent attention calculation.

[0135] 130. Based on the lane line query features, N×M three-dimensional reference points are obtained.

[0136] The three-dimensional reference point is used to indicate the reference position of the lane line coordinate point in the three-dimensional coordinate system.

[0137] Optionally, the lane line query features can be subjected to multi-level nonlinear transformations to obtain N×M three-dimensional reference points. See details below. Figure 1c Specifically, the first multilayer perceptron in the decoder can be used to perform multilayer nonlinear transformation on the lane line query features to obtain N×M three-dimensional reference points.

[0138] 140. Obtain three-dimensional location information features.

[0139] The three-dimensional position information features are used to correct the position of the lane line in three-dimensional space, specifically including: the height position in three-dimensional space and the offset along the road width direction.

[0140] Please see details Figure 1c In the method provided in this application embodiment, the data interaction process between the decoder and the dynamic 3D ground location embedding module can be executed multiple times. Therefore, the 3D location information features can be divided into two cases: initial acquisition and non-initial acquisition, which will be described in detail below.

[0141] Optionally, in one specific embodiment, if the three-dimensional location information features are being acquired for the first time, step 140 may specifically include the following steps A1 to A6:

[0142] A1. Create multiple 3D coordinate points located on the same horizontal plane.

[0143] Each of the three-dimensional coordinate points has its own three-dimensional coordinate values. The number of three-dimensional coordinate points should not be construed as a limitation of this application. For ease of description, let's assume the number of three-dimensional coordinate points is P. Then, P three-dimensional coordinate points can be represented as: P = {(x...} i ,y i ,z i )|i∈1,2,…,P}.

[0144] Since the P three-dimensional coordinate points are located on the same horizontal plane, the height value z of the P three-dimensional coordinate points is... i Since they are all the same, let's assume they are z. i = -1.5. It should be understood that the height value z of a three-dimensional coordinate point... i Other values ​​may also be used, but their specific values ​​should not be construed as limitations on this application.

[0145] A2. Project the plurality of three-dimensional coordinate points to obtain a plurality of two-dimensional coordinate points, wherein each of the three-dimensional coordinate points has its own two-dimensional coordinate value.

[0146] Alternatively, in one specific implementation, the projection process can be performed using the following formula:

[0147] z'×[u,v,1] T =T'×[x,y,z,1] T (1)

[0148] Where z' is the depth value of the 3D coordinate point; u,v is the 2D coordinate value (u,v) of the 2D coordinate point; T' is the camera parameter matrix, which is pre-stored in the dataset; and x,y,z is the 3D coordinate value (x,y,z) of the 3D coordinate point.

[0149] z', u, and v can be calculated using equation (1) above.

[0150] A3. Create an initial two-dimensional matrix with a feature dimension of 3, wherein the initial two-dimensional matrix corresponds to a two-dimensional coordinate range.

[0151] Let's represent the initial two-dimensional matrix as 3×H'×W'. A two-dimensional coordinate system is constructed based on this initial matrix. The origin and the directions of the coordinate axes u and v can be set by the staff based on their experience. It should be understood that the choice of the origin and the directions of the coordinate axes u and v should not be construed as a limitation of this application.

[0152] After creating the two-dimensional coordinate system, the range of two-dimensional coordinates corresponding to the initial two-dimensional matrix can be determined. For example, let's assume that the origin of the two-dimensional coordinate system is located at the upper left corner of the initial two-dimensional matrix, the direction of the u-axis is from left to right, and the direction of the v-axis is from top to bottom; then the range of the two-dimensional horizontal coordinates of the initial two-dimensional matrix is ​​[0, W'], and the range of the two-dimensional vertical coordinates is [0, H'].

[0153] A4. For the plurality of two-dimensional coordinate points, if the two-dimensional coordinate value of a two-dimensional coordinate point falls within the range of the two-dimensional coordinates, then the three-dimensional coordinate value corresponding to the two-dimensional coordinate point is obtained, and the feature dimension of the corresponding position point of the initial two-dimensional matrix is ​​filled with the three-dimensional coordinate value, until all the two-dimensional coordinate points are traversed.

[0154] Continuing with the example above, after projecting P three-dimensional coordinate points, we can obtain P two-dimensional coordinate points. Each of the P two-dimensional coordinate points has its own two-dimensional coordinate value (u,v) and the three-dimensional coordinate values ​​(x,y,z) of the corresponding three-dimensional coordinate point before projection.

[0155] Let's take any two-dimensional coordinate point s from P two-dimensional coordinate points as an example. The two-dimensional coordinate value of this two-dimensional coordinate point s is (u s ,v s The three-dimensional coordinates of the three-dimensional coordinate point corresponding to the two-dimensional coordinate point s are (x... s ,y s ,z s ).

[0156] If the two-dimensional coordinates of point s are (u s ,v s If a point s falls within the range of [0, W'] for its two-dimensional x-coordinate and [0, H'] for its two-dimensional y-coordinate in the initial two-dimensional matrix, then the three-dimensional coordinate value corresponding to the two-dimensional coordinate point s can be obtained as (x...). s ,y s ,z s ), and use this three-dimensional coordinate value (x s ,y s ,z s ) fill the initial two-dimensional matrix (u s ,v s ) Feature dimension 3 at the location.

[0157] If the two-dimensional coordinates of point s are (u s ,v s If the coordinates of point s do not fall within the range of the horizontal coordinates [0, W'] or the vertical coordinates [0, H'] of the initial two-dimensional matrix, then no further operation will be performed on that two-dimensional coordinate point s.

[0158] For each of the P two-dimensional coordinate points, the above calculation process can be performed until all P two-dimensional coordinate points have been traversed.

[0159] A5. If the initial two-dimensional matrix still has position points that do not correspond to any two-dimensional coordinate point, then fill the position points with preset values ​​to obtain the target two-dimensional matrix.

[0160] The preset value is a value set in advance. In one embodiment, the preset value can be (0,0,0). It should be understood that the preset value can also be other values, such as (1,0,1). It should be understood that the specific value of the preset value should not be construed as a limitation of this application.

[0161] After traversing all the two-dimensional coordinate points, if the initial two-dimensional matrix still has position points that do not correspond to any coordinate point, then the position points can be filled with preset values ​​to complete the feature dimensions of the initial two-dimensional matrix and obtain the target two-dimensional matrix.

[0162] A6. Perform multi-layer nonlinear transformation on the target two-dimensional matrix to obtain the three-dimensional position information features.

[0163] After obtaining the target two-dimensional matrix, the feature dimension can be increased by performing multi-level nonlinear transformations on the target two-dimensional matrix. The feature dimension is increased from 3 dimensions to 256 dimensions, resulting in three-dimensional positional information features.

[0164] In the above implementation, three-dimensional position information features can be gradually generated by creating multiple three-dimensional coordinate points located on the same horizontal plane and creating an initial two-dimensional matrix, thereby providing three-dimensional spatial position correction for subsequent lane line detection and further improving the accuracy of lane line detection.

[0165] Optionally, in another specific embodiment, if the three-dimensional location information features are not obtained for the first time, that is, if the lane line query features are the result of the previous feature update, step 140 may specifically include the following steps B1 to B4:

[0166] B1. Obtain the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result.

[0167] For 3D position information features that are not acquired for the first time, the 3D position information features of the current round can be calculated based on the rotation offset around the x-axis and the z-axis offset that are in the same round as the feature update result (i.e., the previous round).

[0168] Optionally, in one specific embodiment, step B1 may specifically include the following steps B11 to B15:

[0169] B11. Obtain the feature map and the historical target two-dimensional matrix, wherein the historical target two-dimensional matrix is ​​the target two-dimensional matrix in the same round as the feature update result.

[0170] The historical target 2D matrix is ​​in the same round as the feature update result.

[0171] For ease of description, let's assume the 3D location information feature corresponds to the k-th round. Then, the lane line query feature is the result of the previous feature update, i.e., the lane line query feature is the feature update result of the (k-1)-th round. Correspondingly, the historical target 2D matrix is ​​the target 2D matrix of the (k-1)-th round; where k is a positive integer greater than or equal to 2.

[0172] Let's assume k=5, and the 3D location information feature corresponds to the 5th round. Then, the lane line query feature is the result of the previous feature update, meaning the lane line query feature is the result of the 4th round's feature update. Correspondingly, the historical target 2D matrix is ​​the target 2D matrix of the 4th round.

[0173] B12. The feature map and the historical target two-dimensional matrix are concatenated to obtain the concatenation result.

[0174] B13. Perform a fifth feature extraction process on the splicing result to obtain the feature extraction result.

[0175] B14. Perform pooling processing on the feature extraction results to obtain the pooling results.

[0176] Pooling can specifically be average pooling. It should be understood that the specific pooling process should not be construed as a limitation of this application.

[0177] B15. Perform multi-layer nonlinear transformation on the pooling result to obtain the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result.

[0178] Alternatively, in one specific embodiment, steps B12 to B15 can be implemented using the following formula:

[0179] [Δθ x ,Δz]=MLP{AvgPool[g([X,M p ])]} (2)

[0180] Where X is the feature map; M p The historical target is a two-dimensional matrix; [X,M] p The symbol ] represents the concatenation result obtained by concatenating the feature map and the historical target two-dimensional matrix; g([X,M p]) indicates the fifth feature extraction process performed on the splicing result, resulting in the feature extraction result; AvgPool[g([X,M p ])] represents the pooling result obtained by performing average pooling on the feature extraction results; MLP{AvgPool[g([X,M p ])]} represents performing a multi-level nonlinear transformation on the pooling result; Δθ x Δz is the rotation offset around the x-axis; Δz is the z-axis offset.

[0181] In the above implementation, the feature map and the target two-dimensional matrix in the same cycle as the feature update result can be processed multiple times to obtain the rotation offset Δθ around the x-axis in the same cycle as the feature update result. x The z-axis offset Δz prepares for the calculation of subsequent non-first-time acquired 3D position information features, thereby improving the accuracy of the calculation of non-first-time acquired 3D position information features.

[0182] B2. Based on the rotation offset around the x-axis and the offset around the z-axis, generate a transformation matrix.

[0183] Continuing with the example above, we can base it on the rotation offset Δθ around the x-axis. x And the z-axis offset Δz, construct the following transformation matrix:

[0184]

[0185] B3. Based on the transformation matrix, update the target two-dimensional matrix that is in the same round as the feature update result to obtain a new target two-dimensional matrix.

[0186] According to We obtain the new target two-dimensional matrix, that is, the target two-dimensional matrix corresponding to the k-th round.

[0187] B4. Perform multi-layer nonlinear transformation on the new target two-dimensional matrix to obtain the three-dimensional position information features.

[0188] According to the formula The three-dimensional location information feature PE is calculated.

[0189] In the above implementation, the x-axis rotation offset and z-axis offset, which are in the same cycle as the feature update result, can be obtained first. Then, a transformation matrix is ​​generated based on the x-axis rotation offset and z-axis offset. The target two-dimensional matrix is ​​then updated based on the transformation matrix to obtain a new target two-dimensional matrix. Subsequently, the three-dimensional position information features are obtained based on the new target two-dimensional matrix. By continuously iterating the three-dimensional position information features using the parameters obtained in the previous cycle, the three-dimensional position information features gradually become more accurate.

[0190] 150. Based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, N×M position offsets are obtained.

[0191] The position offset corresponds one-to-one with each of the three-dimensional reference points. The position offset reflects the offset of the corresponding three-dimensional reference point.

[0192] Optionally, in one specific embodiment, step 150 may specifically include the following steps 151 to 152:

[0193] 151. Based on the feature map, the three-dimensional location information features, and the N×M three-dimensional reference points, update the lane line query features to obtain feature update results.

[0194] The feature dimensions of the feature update result are the same as those of the lane line query feature. Continuing with the example above, the lane line query feature can be represented as (N×M)×C, and the feature update result can also be represented as (N×M)×C.

[0195] Optionally, in one specific embodiment, step 151 may specifically include the following steps 1511 to 1512:

[0196] 1511. Project the N×M three-dimensional reference points to obtain N×M two-dimensional reference points.

[0197] The specific process of projecting N×M three-dimensional reference points is the same as the process of projecting multiple three-dimensional coordinate points, and both will use equation (1), so it will not be elaborated here.

[0198] 1512. Perform deformable cross-attention processing on the lane line query features, the feature map, the three-dimensional position information features, and the N×M two-dimensional reference points to obtain feature update results.

[0199] Deformable cross-attention is an attention mechanism used for image processing and computer vision tasks.

[0200] Optionally, in one specific embodiment, step 1512 may specifically include the following steps C1 to C3:

[0201] C1. Perform self-attention processing on the lane line query features to obtain the self-attention processing result.

[0202] Please see details Figure 1c , can be used Figure 1c The self-attention unit in the decoder shown performs self-attention processing on the lane line query features to obtain the self-attention processing result.

[0203] Self-attention is an attention mechanism that can be used for image data. It establishes and weights relationships between different locations in the input data, thereby better capturing long-range dependencies and important features in the image. Therefore, the results of self-attention processing better reflect the features of lane lines compared to those before self-attention processing.

[0204] C2. Perform deformable cross-attention processing on the self-attention processing result, the feature map, the three-dimensional position information features, and the multiple two-dimensional reference points to obtain the deformable attention processing result.

[0205] Please see details Figure 1c , can be used Figure 1c The deformable cross-attention unit in the decoder shown performs deformable cross-attention processing on the self-attention processing result, the feature map, the three-dimensional positional information features, and the multiple two-dimensional reference points to obtain the deformable attention processing result. The deformable cross-attention unit includes four attention heads, eight sampling points, and 256-dimensional embeddings.

[0206] The deformable cross-attention process can enhance the information exchange and fusion between the self-attention processing results, feature maps, 3D position information features, and multiple 2D reference points. This allows the deformable attention processing results to retain most of the information from the self-attention processing results while carrying the information from feature maps, 3D position information features, and multiple 2D reference points.

[0207] C3. Perform multi-layer nonlinear transformation on the deformable attention processing result to obtain the feature update result.

[0208] Please see details Figure 1c , can be used Figure 1cThe second multilayer perceptron in the decoder shown performs multilayer nonlinear transformations on the deformable attention processing results to obtain feature update results.

[0209] Through the multi-layer nonlinear transformation processing of the second multilayer perceptron, higher-level and more abstract feature representations can be gradually extracted, thereby making the feature update results more representative.

[0210] In one specific embodiment, between step 151 and step 152, the present application embodiment may further include the following steps S1 to S2:

[0211] S1. Use the updated feature result as the new lane line query feature and proceed to step 130.

[0212] S2. Until the number of jumps reaches the preset value, obtain the final feature update result.

[0213] The preset value is a pre-set value. Optionally, the preset value can be 6; it should be understood that the preset value can also be other values, such as 7, and the specific value of the preset value should not be construed as a limitation on this application.

[0214] After jumping to step 130, the updated N×M three-dimensional reference points can be obtained based on the new lane line query features determined in step S1.

[0215] Accordingly, in step 1512, deformable cross-attention processing is performed on the lane line query features, the feature map, the three-dimensional position information features, and the N×M two-dimensional reference points to obtain the feature update result, which can be achieved through the following formula:

[0216] Q k =DeformAttn(Q k-1 ,X T +PE T ,X T +PE T ,ref 2D (6)

[0217] If self-attention processing is not performed on the lane line query features, then Q k-1 This represents the feature update result obtained in the (k-1)th round; if self-attention processing is applied to the lane line query features, then Q... k-1 This represents the self-attention processing result corresponding to the feature update result obtained in the (k-1)th round. X is the feature map; PE is the 3D position information feature; ref 2D For a two-dimensional reference point; DeformAttn() is the process of deformable cross-attention processing. If step C3 is not included in one implementation, then Q kThis is the feature update result obtained in the kth round; if in another implementation, step C3 is included, then Q k This represents the deformable attention processing result obtained in the k-th round. After a preset number of iterations, the final feature update result Q can be obtained. In equation (6), the feature map X is used as... Figure 1c The K and V shown are input into the deformable cross attention unit to participate in the operation. The first X in equation (6) represents K; the second X in equation (6) represents V.

[0218] In the above implementation, after the feature update result is calculated in step 151, step 152 can be temporarily suspended. Instead, the feature update result is used as a new lane line query feature, and the process jumps to step 130 for iteration until the number of iterations reaches a preset value, thereby obtaining the final feature update result. Through the above iterative process, the feature update result can more accurately reflect the characteristics of the lane lines, thereby further improving the accuracy of lane line detection.

[0219] 152. Based on the feature update results, N×M position offsets are obtained.

[0220] Optionally, in one specific embodiment, step 152 may specifically include the following steps 1521 to 1522:

[0221] 1521. Perform multi-layer nonlinear transformation on the feature update result to obtain multiple position offsets and the visibility corresponding to each position offset.

[0222] Continuing with the example above, after obtaining the final feature update result Q, the offset at each location and the visibility corresponding to each offset can be calculated using the following formula:

[0223] [Δx,Δz,v]=MLP reg (Q) (7)

[0224] Where Δx is the offset in the road width direction; Δz is the offset in the road height direction; and v is the visibility corresponding to Δx and Δz, which indicates whether the location point can be seen in the image.

[0225] The coordinates (y-values) of the road extension direction (i.e., the direction of vehicle movement) are predefined. For example, let's assume that 20 points are uniformly sampled within a range of 3 meters to 103 meters along the road extension direction, resulting in 20 y-values. These 20 y-values ​​can be [3, 8.26, 13.53, 18.79, 24.05, 29.32, ...]. Using equation (7), the position offset Δx, Δz, and visibility v corresponding to each y-value can be predicted.

[0226] 1522. If the visibility characterization of the corresponding three-dimensional reference point is invisible, then delete the three-dimensional reference point and the position offset corresponding to the three-dimensional reference point.

[0227] Visibility has two states: visible and invisible. If the visibility is invisible, then the corresponding 3D reference point and its corresponding position offsets Δx and Δz are deleted.

[0228] In the above implementation, the position offset of each of the N×M three-dimensional reference points and the visibility of each position offset can be calculated. Based on the visibility status, the N×M three-dimensional reference points are filtered out, thereby saving subsequent computing resources.

[0229] In another implementation, the visibility v can be skipped, and the N×M position offsets can be obtained directly.

[0230] Optionally, in one specific embodiment, after step 151, the embodiments of this application may further include the following steps D1 to D2:

[0231] D1. Perform max pooling on the feature update result to obtain the max pooling result.

[0232] D2. Perform multi-level nonlinear transformation on the max pooling result to obtain the category to which the lane line belongs.

[0233] Optionally, steps D1 to D2 above can be calculated using the following formula to obtain the category to which the lane line belongs:

[0234] C=MLP cls [MaxPool(Q)] (8)

[0235] MaxPool indicates that maximum pooling is performed; MLP cls This represents a multi-level nonlinear transformation involving the D2 step; C represents the lane category, C∈R N×L L represents the number of possible categories. For example, if the lane line categories include solid white line, solid yellow line, dashed yellow line, and background, then L is 4. If the lane line category is detected as background, then the corresponding position offset and the corresponding 3D reference point can be discarded.

[0236] 160. For each of the aforementioned position offsets, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point, wherein the N×M position prediction points represent: for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0237] After obtaining the position offset corresponding to each three-dimensional reference point through the above steps, the three-dimensional reference point corresponding to the position offset can be offset to obtain the position prediction point. In the process of obtaining the position prediction point, if the visibility of the position prediction point is invisible, or the category of the lane line corresponding to the position prediction point is background, the position prediction point can be deleted.

[0238] Optionally, in one specific embodiment, before step 110, the embodiments of this application may further include the following steps X1 to X3:

[0239] X1. Obtain the lane line detection model.

[0240] X2. Train the lane line detection model to obtain the trained lane line detection model.

[0241] Steps 110 to 160 of the above processing procedure can be performed as follows: Figure 1c The lane detection model implementation is shown. Before performing step 110, it is necessary to obtain an untrained lane detection model and train it.

[0242] Optionally, in one specific embodiment, step X2 may specifically include the following steps Y1 to Y12:

[0243] Y1. Perform the first feature extraction process on the training road image to obtain the training feature map.

[0244] The first feature extraction process can be performed by... Figure 1c The combined network implementation is shown.

[0245] Y2. Perform a second feature extraction process on the training feature map to obtain the global features of the training lane lines.

[0246] The processing procedure for step Y2 is the same as that for step 112, so it will not be described in detail here.

[0247] Y3. Randomly initialize the local features of the training lane lines, and perform feature fusion processing on the global features of the training lane lines and the local features of the training lane lines to obtain the query features of the training lane lines.

[0248] Local features of the training lane lines can be obtained through random initialization.

[0249] The step of “fusing the global features of the training lane lines with the local features of the training lane lines to obtain the query features of the training lane lines” is the same as step 120, and will not be described again here.

[0250] Y4. Perform a first multi-level nonlinear transformation on the training lane line query features to obtain multiple training three-dimensional reference points, and perform projection processing on each of the training three-dimensional reference points to obtain multiple training two-dimensional reference points.

[0251] In step Y4, the process of “performing a first multi-layer nonlinear transformation on the training lane line query features to obtain multiple training three-dimensional reference points” is the same as that in step 130; the process of “projecting each of the training three-dimensional reference points to obtain multiple training two-dimensional reference points” is the same as that in step 1511, and will not be described again here.

[0252] Y5. Obtain training 3D position information features.

[0253] Step Y5 corresponds to step 140, so it will not be described again here.

[0254] Y6. Perform deformable cross-attention processing on the training lane line query features, the training feature map, the training three-dimensional position information features, and the multiple training two-dimensional reference points to obtain the training feature update result.

[0255] Step Y6 corresponds to step 1512, so it will not be described again here.

[0256] Y7. Perform a second multi-layer nonlinear transformation on the training feature update results to obtain multiple training position offsets and the training visibility corresponding to each training position offset.

[0257] Step Y7 corresponds to step 1521, so it will not be described again here.

[0258] Y8. Perform max pooling on the training feature update results to obtain the training max pooling results.

[0259] Y9. Perform a sixth-level nonlinear transformation on the training max-pooling result to obtain the training category to which the lane line belongs.

[0260] Steps Y8 to Y9 correspond to the same steps D1 to D2 mentioned above, and will not be repeated here.

[0261] Y10. Based on the training feature map and the global features of the training lane lines, the training lane line segmentation result is obtained.

[0262] Based on the global features of the training lane lines, instance segmentation can be performed on the training feature map to obtain multiple training lane line segmentation results, each of which can contain one lane line. For example, suppose the training feature map contains E lane lines, then... Figure 1c The lane segmentation unit shown divides the training feature map into E training lane segmentation results. Each training lane segmentation result contains one lane line.

[0263] Optionally, Figure 1c The lane segmentation units shown can be used only during the training phase, which helps improve training accuracy and reduce computational complexity during model application.

[0264] Y11. Based on the training lane segmentation results, training 3D position information features, training position offset, training visibility, and training category, construct a loss function.

[0265] Optionally, in one embodiment, Y11 may specifically include the following steps Z1 to Z4:

[0266] Z1. Based on the training lane line segmentation results, construct an instance segmentation auxiliary loss function.

[0267] Instance segmentation auxiliary loss function L seg It consists of three parts: the lane line existence loss function L obj Segmentation loss function L mask and classification loss function L cls Instance segmentation auxiliary loss function L seg This can be expressed by the formula:

[0268] L seg =λ o ×L obj +L mask +λ c ×L cls (9)

[0269] L mask =λ dice ×L dice +λ bce ×L bce (10)

[0270] Among them, the existence loss function L obj Supervised learning for binary classification is performed using a binary cross-entropy loss function. The classification result includes lane lines or background. The classification loss function is L. clsFocal loss is used to supervise lane line category learning. In the focal loss, the factor γ, which adjusts the weights of easy and difficult samples, is set to 2, and the category weight α is set to 0.25; λ o The value is 1; λ c The value is 2; λ dice The value is 2; λ bce The value is 5.

[0271] but

[0272] Where p is the probability value that the model determines the lane line to be of the positive class.

[0273] Segmentation loss function L mask By the Dice loss function L dice and binary cross-entropy loss function L bce It is composed of components to perform supervised learning of the predictive mask.

[0274] Dice loss function L dice for:

[0275]

[0276] in, Denotes the segmentation mask for the i-th prediction; m j This represents the truth mask for the j-th lane line.

[0277] Binary cross-entropy loss function L bce for:

[0278]

[0279] Where R represents the number of lane lines; y r The true label for the r-th lane line; This represents the output value of the model on the r-th lane.

[0280] The above instance segmentation auxiliary loss function L seg The calculation can be based on bipartite graph matching, and the bipartite matching results will be used when calculating the loss function. The bipartite graph matching algorithm calculates the matching score between each pair of ground truth and prediction based on the given scoring function, and finds an optimal scoring matching strategy to assign ground truth to each prediction. Prediction refers to the mask prediction of each lane line in the 2D image. Specifically, the mask can refer to the shape of the lane lines in the 2D image; for details, please refer to [link to relevant documentation]. Figure 1e , Figure 1e Image (a) shows a schematic diagram of a road image of one embodiment; Figure 1e L shown in (b) represents: Figure 1eThe mask corresponding to the lane lines shown in (a) is the ground truth, which refers to the mask of each lane line generated by projecting the manual annotations of the lane lines in the road image onto the two-dimensional image.

[0281] The matching score function between the i-th prediction and the true value of the j-th lane line can be defined as follows:

[0282]

[0283] Among them, c j Indicates the category of the truth value of the j-th lane line; This indicates that the i-th prediction belongs to category c. j The probability value; m j represents the segmentation mask of the i-th prediction and the ground truth mask of the j-th lane line, respectively; dice represents the dice coefficient between the prediction mask and the ground truth mask; the weight α is used to adjust the importance of the classification task and the segmentation task, and its value is 0.8.

[0284] After calculating the matching score between each pair of predictions and true values ​​using equation (*), the optimal one-to-one matching result can be obtained using the Hungarian algorithm.

[0285] Z2. Based on the training 3D position information features, construct a 3D plane update loss function.

[0286] Optionally, equation (1) can be deformed first, and then the annotations of the three-dimensional lane lines in the training road image can be projected using the deformed equation (1).

[0287] Specifically, the projection process described above can be achieved using the following formula:

[0288] l u,v =T×l x,y,z (14)

[0289] Among them, l x,y,z Indicates the labeling of three-dimensional lane lines; l u,v This indicates the annotation after projection. The projection process described above corresponds to the projection process in steps A2 to A5, so it will not be repeated here.

[0290] Through the projection process described above, after projecting the 3D lane line annotations onto the initial 2D matrix, the projection M of the 3D lane line annotations can be obtained. l .

[0291] The training target two-dimensional matrix M can be obtained based on the training three-dimensional position information features. p The training target two-dimensional matrix M is obtained. p The process is the same as steps A1 to A5, so it will not be described again here.

[0292] The 3D plane update loss function can then be calculated based on the following formula:

[0293]

[0294] Where Z is the two-dimensional matrix M of the training target. p The corresponding pixel set; Q is the projection M of the 3D lane line annotation. l The corresponding pixel set; ":" represents the three-dimensional coordinates x, y, z corresponding to coordinates u, v; M p M is the two-dimensional matrix of the training target; l The projection of the 3D lane line annotation.

[0295] In the above implementation, the loss function L can be updated using a three-dimensional plane. plane The update of multiple training 3D coordinate points located on the same horizontal plane is supervised to gradually approach the location of the real ground, thereby obtaining accurate 3D ground location information.

[0296] Please see details Figure 1f , Figure 1f (a) shows the initial three-dimensional positions of multiple training 3D coordinate points (2) created on the same horizontal plane and the real ground (1); Figure 1f (b) shows the L plane After supervision, the positions of multiple training 3D coordinate points (2) located on the same plane and the real ground (1) are determined. Figure 1f As can be seen from (b) in the figure, (1) and (2) are almost on the same plane.

[0297] Z3. Based on the training position offset, training visibility, and training category, a 3D lane detection loss function is constructed.

[0298] Alternatively, the 3D lane detection loss function can be calculated according to the following formula:

[0299] L lane =w x L x +w z L z +w v L v +w c L c (16)

[0300] Among them, L x The loss function is the offset Δx in the road width direction; w x For L x The weight value; L zThe loss function is the offset Δz in the road height direction; w z For L z The weight value; L v The loss function for visibility v; w v For L v The weight value; L c The loss function is for the category C to which the lane lines belong; w c For L c The weight value. Specifically, w x The value can be 2, w z The value can be 10, w v The value can be 1, w c It can take the value 10.

[0301] The above L x and L z The L1 loss function can be used for calculation; L v The binary cross-entropy loss function can be used to calculate it; L c Focal loss can be used, where the factor γ for adjusting the weights of easy and difficult samples is set to 2, and the class weight α is set to 0.25.

[0302] The L1 loss function is:

[0303]

[0304] Where w is the model parameter; g is the number of training position offsets; x i It is the feature of the offset of the i-th training position; y i It is the true output value of the offset at the i-th training position; f(x) i ;w) indicates that x is based on model parameters w. i The output value obtained from the prediction.

[0305] Z4. Construct the loss function based on the instance segmentation auxiliary loss function, the 3D plane update loss function, and the 3D lane detection loss function.

[0306] After calculating the instance segmentation auxiliary loss function, the 3D plane update loss function, and the 3D lane detection loss function according to steps Z1 to Z3 respectively, the overall loss function can be constructed based on the following formula:

[0307] L = w s L seg +w p L plane +w l L lane (18)

[0308] Among them, w s It is the instance segmentation auxiliary loss function L seg Weights; w p It is the three-dimensional plane update loss function L plane Weights; w l It is the 3D lane detection loss function L lane The weights. Specifically, w s The value can be 2, w p The value can be 1, w l It can take the value 1.

[0309] Y12. If the loss function meets the preset requirements, it is determined that the training of the lane detection model is complete, and the trained lane detection model is obtained.

[0310] Optionally, in one specific implementation, the loss function satisfies a preset requirement, specifically: the overall loss function (18) converges.

[0311] If the overall loss function (18) converges, it can be determined that the training of the lane detection model is complete, and the trained lane detection model is obtained.

[0312] During training, the gradient descent algorithm used to bring the overall loss function to converge can be the AdamW optimizer with a weight decay rate of 0.001. During training, a cosine annealing scheduler can also be used to improve the model's convergence and generalization ability; the learning rate can be set to 2 × 10⁻⁶. -4 The batch size used during training can be 32, meaning that 32 samples can be input into the model in one iteration.

[0313] This training process can utilize the OpenLane and Apollo datasets as training datasets. Optionally, the lane detection model can be trained for 24 epochs using the OpenLane dataset; alternatively, it can be trained for 100 epochs using the Apollo dataset. An epoch refers to the process of performing one forward and backward propagation through the model using all training samples in the training dataset. In other words, one epoch represents a complete training cycle on the entire training dataset.

[0314] The above training process can be implemented in a graphics processing unit (GPU), specifically an A100 GPU.

[0315] X3. Using the trained lane detection model, perform steps 110 to 160 above to determine the N×M predicted location points.

[0316] In the above implementation, the trained lane detection model can be used to fuse global and local lane features to obtain lane query features that carry both the detection upper limit value N and the sampling point value M. Based on the lane query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lane lines are acquired. Then, based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane query features, N×M position offsets can be obtained; subsequently, based on the position offsets and the corresponding three-dimensional reference points, N×M predicted position points are obtained.

[0317] The lane detection method provided in this application embodiment can perform feature extraction processing on a road image to obtain a feature map and global lane features. The global lane features carry an upper limit value N for detecting the number of lanes. Local lane features carrying sampling point values ​​M are obtained, where sampling points M are used to detect lane position. The global lane features and local lane features are then fused to obtain lane query features that carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lanes are obtained. Based on the feature map, three-dimensional position information features, N×M three-dimensional reference points, and lane query features, N×M position offsets are obtained, with each position offset corresponding to a three-dimensional reference point. Based on each of the N×M position offsets, the corresponding three-dimensional reference point is offset to obtain a predicted position point, resulting in a total of N×M predicted position points. The N×M predicted position points represent the following: For each of the N lane lines, M predicted position points indicate the outline of the corresponding lane line. In this embodiment, the global lane line features and local lane line features can respectively carry the detection upper limit value N of the number of lane lines in the road image and the sampling point value M corresponding to each lane line. The global lane line features and local lane line features are fused to obtain lane line query features that carry both the detection upper limit value N and the sampling point value M. Based on the lane line query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lane lines are obtained. Then, based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, N×M position offsets can be obtained; subsequently, based on the position offsets and the corresponding three-dimensional reference points, N×M predicted position points are obtained.

[0318] The embodiments of this application can improve the accuracy of lane line detection based on monocular images.

[0319] Figure 1g Figures (a) and (b) show the quantitative results of different models on the real-world OpenLane dataset. The model provided in this embodiment (i.e., LATR in the figure) achieved the highest score on this dataset, showing a significant improvement over the previous best model. On the OpenLane test set, it improved the F1 score by 11.4 points and the category accuracy by 2.5 points compared to previous models. Furthermore, this embodiment introduces a lightweight version of the LATR model, LATR-Lite, whose decoder only iterates twice. LATR-Lite achieved results comparable to the full LATR model. To ensure fair comparison, the Persformer was reimplemented for other models using the same backbone network (ResNet50) and input size, i.e., Persformer-Res50 shown in the figure. For other models, although utilizing ResNet-50 improved their F1 score from 50.5 to 53.0, it still lagged behind the LATR model provided in this embodiment. Moreover, the error rate of this embodiment is significantly reduced. Specifically, compared to Persformer, the embodiments of this application reduce the X-direction error (in meters) by 0.100 / 0.066 in the near / far range and the Z-direction error (in meters) by 0.037 / 0.037 in the near / far range. Through Figure 1g As can also be observed in (b), the embodiments of this application show an improvement of more than 10.0 in F1 scores in scenarios such as uphill and downhill, curves, and merges / splits. These results demonstrate the effectiveness and robustness of the embodiments of this application in solving this challenging problem.

[0320] Figure 1g (c) shows the quantitative results of different models on the Apollo dataset synthesized by the game engine, which contains three different scenes: Balanced Scene, Rare Subset, and Visual Variations. It can be seen that the embodiments of this application achieved state-of-the-art performance in all three scenes of this dataset, demonstrating the effectiveness and robustness of the embodiments of this application in solving this challenging problem.

[0321] In addition, please see details. Figure 1h The embodiments of this application also demonstrate the qualitative results of the embodiments of this application on the OpenLane dataset, such as... Figure 1h As shown. Figure 1h In the figure, (a), (b), and (c) represent the true value of the three-dimensional lane line, the detection of the embodiment of this application, and the detection of the optimal model in the prior art, respectively. Figure 1h (d) in the diagram illustrates the relationship between the prediction and the true value of the embodiment of this application in three-dimensional space, and it is evident that the two have a high degree of overlap. Combined with... Figure 1h As can be seen, the embodiments of this application can detect three-dimensional lane lines more accurately in various scenarios.

[0322] The method described in the above embodiments will be further described in detail below.

[0323] In this embodiment, the method of this application embodiment will be described in detail using a preset value of 6 as an example.

[0324] like Figure 2 As shown, the specific process of a lane line detection method is as follows:

[0325] 201. Perform a first feature extraction process on the road image to obtain the feature map.

[0326] 202. Perform a third feature extraction process on the feature map to obtain a first feature sub-map, wherein the first feature sub-map carries the detection upper limit value N.

[0327] 203. Perform a fourth feature extraction process on the feature map to obtain a second feature sub-map, wherein the second feature sub-map carries feature dimension values.

[0328] 204. Perform feature fusion processing on the first feature sub-image and the second feature sub-image to obtain the global features of the lane line.

[0329] 205. Obtain local features of the lane line, wherein the local features of the lane line carry sampling point values ​​M for detecting the position of the lane line.

[0330] 206. The global features of the lane line and the local features of the lane line are broadcast and added together to obtain the lane line query features. The lane line query features carry both the detection upper limit value N and the sampling point value M.

[0331] 207. Based on the lane line query features, obtain N×M three-dimensional reference points; wherein, the three-dimensional reference points are used to indicate the reference position of the lane line coordinate points in the three-dimensional coordinate system.

[0332] 208. Obtain three-dimensional position information features, wherein the three-dimensional position information features are used to correct the height position of the lane line in three-dimensional space.

[0333] 209. Based on the feature map, the three-dimensional location information features, and the N×M three-dimensional reference points, update the lane line query features to obtain the feature update result.

[0334] Optionally, in one specific embodiment, step 209 may specifically include the following steps 2091 to 2094:

[0335] 2091. Project the N×M three-dimensional reference points to obtain N×M two-dimensional reference points.

[0336] 2092. Perform self-attention processing on the lane line query features to obtain the self-attention processing result.

[0337] 2093. Perform deformable cross-attention processing on the self-attention processing result, the feature map, the three-dimensional position information features, and the multiple two-dimensional reference points to obtain the deformable attention processing result.

[0338] 2094. Perform multi-layer nonlinear transformation on the deformable attention processing result to obtain the feature update result.

[0339] 210. Use the updated feature result as the new lane line query feature and jump to step 207 until the number of jumps reaches 6, and obtain the final feature update result.

[0340] 211. Based on the feature update results, N×M position offsets are obtained, and each position offset corresponds one-to-one with the three-dimensional reference point.

[0341] Optionally, in one specific embodiment, step 211 may specifically include the following steps 2111 to 2112:

[0342] 2111. Perform multi-layer nonlinear transformation on the feature update result to obtain multiple position offsets and the visibility corresponding to each position offset.

[0343] 2112. If the visibility characterization of the corresponding three-dimensional reference point is invisible, then delete the three-dimensional reference point and the position offset corresponding to the three-dimensional reference point.

[0344] 212. Perform max pooling on the feature update result to obtain the max pooling result.

[0345] 213. Perform multi-level nonlinear transformation on the max pooling result to obtain the category to which the lane line belongs.

[0346] 214. For each of the aforementioned position offsets, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point, wherein the N×M position prediction points represent: for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0347] 215. Based on the correspondence between the location prediction point and the category to which the lane line belongs, determine the category to which the lane line to which each location prediction point belongs.

[0348] In one specific implementation, after obtaining N×M location prediction points, the location prediction points can be filtered based on their visibility status and the category to which the lane lines belong, thereby obtaining the filtered location prediction points.

[0349] The specific execution process of steps 201 to 215 has been explained in detail above, and will not be repeated here.

[0350] The lane detection method provided in this application embodiment can perform feature extraction processing on a road image to obtain a feature map and global lane features. The global lane features carry an upper limit value N for detecting the number of lanes. Local lane features carrying sampling point values ​​M are obtained, where sampling points M are used to detect lane position. The global lane features and local lane features are then fused to obtain lane query features that carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lanes are obtained. Based on the feature map, three-dimensional position information features, N×M three-dimensional reference points, and lane query features, N×M position offsets are obtained, with each position offset corresponding to a three-dimensional reference point. Based on each of the N×M position offsets, the corresponding three-dimensional reference point is offset to obtain a predicted position point, resulting in a total of N×M predicted position points. The N×M predicted position points represent the following: For each of the N lane lines, M predicted position points indicate the outline of the corresponding lane line. In this embodiment, the global lane line features and local lane line features can respectively carry the detection upper limit value N of the number of lane lines in the road image and the sampling point value M corresponding to each lane line. The global lane line features and local lane line features are fused to obtain lane line query features that carry both the detection upper limit value N and the sampling point value M. Based on the lane line query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lane lines are obtained. Then, based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, N×M position offsets can be obtained; subsequently, based on the position offsets and the corresponding three-dimensional reference points, N×M predicted position points are obtained.

[0351] The embodiments of this application can improve the accuracy of lane line detection based on monocular images.

[0352] To better implement the above methods, this application also provides a lane line detection device, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC); the server can be a single server or a server cluster composed of multiple servers. For example, in this embodiment, the method of this application embodiment will be described in detail using the example of a lane line detection device specifically integrated into a terminal of an electronic device or a server deployed in the cloud.

[0353] For example, such as Figure 3As shown, the lane line detection device may include:

[0354] The feature extraction unit 301 is used to perform feature extraction processing on the road image to obtain a feature map and global lane line features, wherein the feature map contains lane line information and the global lane line features carry an upper limit value N for detecting the number of lane lines.

[0355] The feature fusion unit 302 is used to acquire local lane line features and perform feature fusion processing on the global lane line features and the local lane line features to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M.

[0356] The reference point acquisition unit 303 is used to obtain N×M three-dimensional reference points based on the lane line query features; wherein, the three-dimensional reference points are used to indicate the reference position of the lane line coordinate points in the three-dimensional coordinate system;

[0357] The location information unit 304 is used to acquire three-dimensional location information features, wherein the three-dimensional location information features are used to correct the height position of the lane line in three-dimensional space.

[0358] The offset acquisition unit 305 is used to obtain N×M position offsets based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points and the lane line query features, wherein the position offsets correspond one-to-one with the three-dimensional reference points;

[0359] The prediction point acquisition unit 306 is used to offset the corresponding three-dimensional reference point based on the position offset for each position offset to obtain a position prediction point. The N×M position prediction points represent that for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0360] In some embodiments, the feature extraction unit 301 includes:

[0361] The first extraction subunit is used to perform a first feature extraction process on the road image to obtain the feature map;

[0362] The second extraction subunit is used to perform a second feature extraction process on the feature map to obtain the global features of the lane line.

[0363] In some embodiments, the second extraction subunit includes:

[0364] The third extraction sub-unit is used to perform third feature extraction processing on the feature map to obtain a first feature sub-map, wherein the first feature sub-map carries the detection upper limit value N;

[0365] The fourth extraction sub-unit is used to perform a fourth feature extraction process on the feature map to obtain a second feature sub-map, wherein the second feature sub-map carries feature dimension values;

[0366] The feature fusion sub-unit is used to perform feature fusion processing on the first feature sub-map and the second feature sub-map to obtain the global features of the lane line.

[0367] In some embodiments, the feature fusion unit 302 is specifically used to perform broadcast addition processing on the global features of the lane line and the local features of the lane line to obtain the lane line query features.

[0368] In some embodiments, the location information unit 304 includes:

[0369] The coordinate point creation sub-unit is used to create multiple three-dimensional coordinate points located on the same horizontal plane, wherein each of the three-dimensional coordinate points has its own three-dimensional coordinate values;

[0370] The coordinate projection subunit is used to project the plurality of three-dimensional coordinate points to obtain a plurality of two-dimensional coordinate points, wherein each of the three-dimensional coordinate points has its own two-dimensional coordinate values.

[0371] A matrix creation sub-unit is used to create an initial two-dimensional matrix with a feature dimension of 3, wherein the initial two-dimensional matrix corresponds to a two-dimensional coordinate range;

[0372] The coordinate traversal subunit is used to, for the plurality of two-dimensional coordinate points, if there is a two-dimensional coordinate point whose two-dimensional coordinate value falls within the two-dimensional coordinate range, obtain the three-dimensional coordinate value corresponding to the two-dimensional coordinate point, and use the three-dimensional coordinate value to fill the feature dimension of the corresponding position point of the initial two-dimensional matrix, until all the two-dimensional coordinate points are traversed.

[0373] A two-dimensional matrix sub-unit is used to fill the position points with a preset value when the initial two-dimensional matrix still has position points that do not correspond to any two-dimensional coordinate point, so as to obtain a target two-dimensional matrix.

[0374] The three-dimensional position sub-unit is used to perform multi-layer nonlinear transformation on the target two-dimensional matrix to obtain the three-dimensional position information features.

[0375] In some embodiments, the offset acquisition unit is specifically used to update the lane line query features based on the feature map, the three-dimensional position information features, and the N×M three-dimensional reference points to obtain feature update results; and to obtain N×M position offsets based on the feature update results.

[0376] In some embodiments, the apparatus further includes:

[0377] The step jump unit is used to take the feature update result as the new lane line query feature and jump to the step: based on the lane line query feature, obtain N×M three-dimensional reference points;

[0378] The threshold is reached unit, which is used to obtain the final feature update result until the number of jumps reaches a preset value.

[0379] In some embodiments, when the lane line query feature is the result of the previous feature update, the location information unit includes:

[0380] The offset acquisition subunit is used to acquire the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result;

[0381] A transformation matrix generation sub-unit is used to generate a transformation matrix based on the rotation offset around the x-axis and the z-axis offset;

[0382] The updated matrix sub-unit is used to update the target two-dimensional matrix that is in the same round as the feature update result based on the transformation matrix, so as to obtain a new target two-dimensional matrix.

[0383] The three-dimensional position sub-unit is used to perform multi-level nonlinear transformations on the new target two-dimensional matrix to obtain the three-dimensional position information features.

[0384] In some embodiments, the offset acquisition subunit includes:

[0385] The historical acquisition subunit is used to acquire the feature map and the historical target two-dimensional matrix, wherein the historical target two-dimensional matrix is ​​a target two-dimensional matrix that is in the same round as the feature update result;

[0386] The concatenation sub-unit is used to concatenate the feature map and the historical target two-dimensional matrix to obtain the concatenation result;

[0387] The extraction result sub-unit is used to perform a fifth feature extraction process on the splicing result to obtain the feature extraction result;

[0388] The pooling result sub-unit is used to perform pooling processing on the feature extraction result to obtain the pooling result;

[0389] The offset sub-unit is used to perform multi-layer nonlinear transformation on the pooling result to obtain the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result.

[0390] In some embodiments, the offset acquisition unit includes:

[0391] The reference point projection subunit is used to project the N×M three-dimensional reference points to obtain N×M two-dimensional reference points.

[0392] The cross-attention subunit is used to perform deformable cross-attention processing on the lane line query features, the feature map, the three-dimensional position information features, and the N×M two-dimensional reference points to obtain feature update results.

[0393] In some embodiments, the cross-attention subunit includes:

[0394] The self-attention sub-unit is used to perform self-attention processing on the lane line query features to obtain the self-attention processing result.

[0395] The cross-attention subunit is used to perform deformable cross-attention processing on the self-attention processing result, the feature map, the three-dimensional position information features, and the multiple two-dimensional reference points to obtain a deformable attention processing result.

[0396] The nonlinear transformation subunit is used to perform multi-level nonlinear transformations on the deformable attention processing result to obtain the feature update result.

[0397] In some embodiments, the apparatus further includes:

[0398] The max pooling unit is used to perform max pooling on the feature update result to obtain the max pooling result.

[0399] The category acquisition unit is used to perform multi-level nonlinear transformation on the max pooling result to obtain the category to which the lane line belongs.

[0400] In some embodiments, the offset acquisition unit includes:

[0401] The visibility subunit is used to perform multi-layer nonlinear transformation on the feature update result to obtain multiple position offsets and the visibility corresponding to each position offset;

[0402] The reference point deletion subunit is used to delete the three-dimensional reference point and the position offset corresponding to the three-dimensional reference point when the visibility characterization is in an invisible state.

[0403] In some embodiments, the apparatus further includes:

[0404] The model acquisition unit is used to acquire the lane line detection model;

[0405] The model training unit is used to train the lane line detection model to obtain the trained lane line detection model.

[0406] The model application unit is used to execute the lane detection method using the trained lane detection model to determine the N×M location prediction points.

[0407] In some embodiments, the model training unit includes:

[0408] The training feature subunit is used to perform the first feature extraction process on the training road image to obtain the training feature map;

[0409] A global subunit is trained to perform a second feature extraction process on the trained feature map to obtain the global features of the trained lane lines.

[0410] The feature fusion subunit is used to randomly initialize the local features of the training lane lines and perform feature fusion processing on the global features of the training lane lines and the local features of the training lane lines to obtain the training feature fusion result, which is denoted as the training lane line query feature.

[0411] The training reference point sub-unit is used to perform a first multi-level nonlinear transformation on the training lane line query features to obtain multiple training three-dimensional reference points, and to perform projection processing on each of the training three-dimensional reference points to obtain multiple training two-dimensional reference points.

[0412] The training location information subunit is used to acquire training 3D location information features;

[0413] The training feature update subunit is used to perform deformable cross attention processing on the training lane line query features, the training feature map, the training three-dimensional position information features, and the multiple training two-dimensional reference points to obtain the training feature update result.

[0414] A training visibility subunit is used to perform a second multi-layer nonlinear transformation on the training feature update result to obtain multiple training position offsets and the training visibility corresponding to each training position offset.

[0415] The max pooling subunit is used to perform max pooling on the training feature update result to obtain the training max pooling result.

[0416] The training category sub-unit is used to perform a sixth-level nonlinear transformation on the training max pooling result to obtain the training category to which the lane line belongs.

[0417] The training segmentation result subunit is used to obtain the training lane line segmentation result based on the training feature map and the training lane line global features;

[0418] The loss function subunit is used to construct a loss function based on the training lane segmentation results, training 3D position information features, training position offset, training visibility, and training category.

[0419] The training completion sub-unit is used to determine that the lane detection model has been trained when the loss function meets the preset requirements, and to obtain the trained lane detection model.

[0420] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0421] The lane detection method provided in this application embodiment can perform feature extraction processing on a road image to obtain a feature map and global lane features. The global lane features carry an upper limit value N for detecting the number of lanes. Local lane features carrying sampling point values ​​M are obtained, where sampling points M are used to detect lane position. The global lane features and local lane features are then fused to obtain lane query features that carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lanes are obtained. Based on the feature map, three-dimensional position information features, N×M three-dimensional reference points, and lane query features, N×M position offsets are obtained, with each position offset corresponding to a three-dimensional reference point. Based on each of the N×M position offsets, the corresponding three-dimensional reference point is offset to obtain a predicted position point, resulting in a total of N×M predicted position points. The N×M predicted position points represent the following: For each of the N lane lines, M predicted position points indicate the outline of the corresponding lane line. In this embodiment, the global lane line features and local lane line features can respectively carry the detection upper limit value N of the number of lane lines in the road image and the sampling point value M corresponding to each lane line. The global lane line features and local lane line features are fused to obtain lane line query features that carry both the detection upper limit value N and the sampling point value M. Based on the lane line query features, N×M three-dimensional reference points are obtained. Three-dimensional position information features used to correct the height features of the lane lines are obtained. Then, based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, N×M position offsets can be obtained; subsequently, based on the position offsets and the corresponding three-dimensional reference points, N×M predicted position points are obtained.

[0422] The embodiments of this application can improve the accuracy of lane line detection based on monocular images.

[0423] This application also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0424] In some embodiments, the lane line detection device may also be integrated into multiple electronic devices, such as multiple servers, with multiple servers implementing the lane line detection method of this application.

[0425] In this embodiment, the electronic device will be used as an example for detailed description, such as... Figure 4 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:

[0426] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0427] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 401.

[0428] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0429] The electronic device also includes a power supply 403 that supplies power to the various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0430] The electronic device may also include an input module 404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0431] The electronic device may also include a communication module 405. In some embodiments, the communication module 405 may include a wireless module, through which the electronic device can perform short-range wireless transmission, thereby providing users with wireless broadband internet access. For example, the communication module 405 can be used to help users send and receive emails, browse web pages, and access streaming media.

[0432] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0433] Feature extraction is performed on the road image to obtain a feature map and global lane line features. The feature map contains lane line information, and the global lane line features carry an upper limit value N for detecting the number of lane lines. Local lane line features are then obtained, and the global lane line features and the local lane line features are fused to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane line query features, N×M three-dimensional reference points are obtained. These three-dimensional reference points are used to indicate the lane line positions. The reference position of the marker is marked in the three-dimensional coordinate system; three-dimensional position information features are obtained, wherein the three-dimensional position information features are used to correct the height position of the lane line in three-dimensional space; based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points and the lane line query features, N×M position offsets are obtained, wherein the position offsets correspond one-to-one with the three-dimensional reference points; for each position offset, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point, wherein the N×M position prediction points represent: for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0434] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0435] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0436] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the lane detection methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0437] Feature extraction is performed on the road image to obtain a feature map and global lane line features. The feature map contains lane line information, and the global lane line features carry an upper limit value N for detecting the number of lane lines. Local lane line features are then obtained, and the global lane line features and the local lane line features are fused to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane line query features, N×M three-dimensional reference points are obtained. These three-dimensional reference points are used to indicate the lane line positions. The reference position of the marker is marked in the three-dimensional coordinate system; three-dimensional position information features are obtained, wherein the three-dimensional position information features are used to correct the height position of the lane line in three-dimensional space; based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points and the lane line query features, N×M position offsets are obtained, wherein the position offsets correspond one-to-one with the three-dimensional reference points; for each position offset, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point, wherein the N×M position prediction points represent: for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

[0438] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0439] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0440] Since the instructions stored in the storage medium can execute the steps of any lane line detection method provided in the embodiments of this application, the beneficial effects that any lane line detection method provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0441] The above provides a detailed description of a lane line detection method, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A lane line detection method, characterized in that, The method includes: The road image is processed by feature extraction to obtain a feature map and global lane line features. The feature map contains lane line information, and the global lane line features carry an upper limit value N for the detection of the number of lane lines. Local lane line features are obtained, and the global lane line features are fused with the local lane line features to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M. Based on the lane line query features, N×M three-dimensional reference points are obtained; wherein, the N×M three-dimensional reference points are used to indicate the reference positions of M sampling points of each lane line in the three-dimensional coordinate system; Acquire three-dimensional position information features, wherein the three-dimensional position information features are used to correct the height position of the lane line in three-dimensional space; Based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features, N×M position offsets are obtained, wherein each position offset corresponds one-to-one with a three-dimensional reference point. For each of the aforementioned position offsets, the corresponding three-dimensional reference point is offset based on the position offset to obtain a position prediction point. The N×M position prediction points represent that for each of the N lane lines, the outline of the corresponding lane line is indicated by M position prediction points.

2. The method as described in claim 1, characterized in that, The feature extraction process of the road image to obtain feature maps and global lane line features includes: The road image is subjected to a first feature extraction process to obtain the feature map; The feature map is subjected to a second feature extraction process to obtain the global features of the lane lines.

3. The method as described in claim 2, characterized in that, The second feature extraction process on the feature map to obtain the global features of the lane lines includes: The feature map is subjected to a third feature extraction process to obtain a first feature sub-map, wherein the first feature sub-map carries the detection upper limit value N; The feature map is subjected to a fourth feature extraction process to obtain a second feature sub-map, wherein the second feature sub-map carries feature dimension values; The first feature sub-image and the second feature sub-image are subjected to feature fusion processing to obtain the global features of the lane line.

4. The method as described in claim 1, characterized in that, The step of fusing the global features and local features of the lane lines to obtain lane line query features includes: The global features and local features of the lane lines are broadcast and added together to obtain the lane line query features.

5. The method as described in claim 1, characterized in that, The acquisition of three-dimensional location information features includes: Create multiple three-dimensional coordinate points located on the same horizontal plane, wherein each of the three-dimensional coordinate points has its own three-dimensional coordinate values; Projecting the plurality of three-dimensional coordinate points yields a plurality of two-dimensional coordinate points, wherein each of the three-dimensional coordinate points has its own two-dimensional coordinate values. Create an initial two-dimensional matrix with a feature dimension of 3, wherein the initial two-dimensional matrix corresponds to a two-dimensional coordinate range; For the plurality of two-dimensional coordinate points, if there exists a two-dimensional coordinate point whose two-dimensional coordinate value falls within the two-dimensional coordinate range, then the three-dimensional coordinate value corresponding to the two-dimensional coordinate point is obtained, and the feature dimension of the corresponding position point of the initial two-dimensional matrix is ​​filled with the three-dimensional coordinate value, until all the two-dimensional coordinate points are traversed. If the initial two-dimensional matrix still has position points that do not correspond to any two-dimensional coordinate point, then fill the position points with preset values ​​to obtain the target two-dimensional matrix; The target two-dimensional matrix is ​​subjected to multi-layer nonlinear transformation to obtain the three-dimensional position information features.

6. The method as described in claim 1, characterized in that, The method of obtaining N×M position offsets based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points, and the lane line query features includes: Based on the feature map, the three-dimensional location information features, and the N×M three-dimensional reference points, the lane line query features are updated to obtain the feature update results. Based on the feature update results, N×M position offsets are obtained.

7. The method as described in claim 6, characterized in that, After updating the lane line query features based on the feature map, the three-dimensional position information features, and the N×M three-dimensional reference points to obtain the feature update result, and before obtaining the N×M position offsets based on the feature update result, the method further includes: The updated feature is used as the new lane line query feature, and the process jumps to step: based on the lane line query feature, N×M three-dimensional reference points are obtained. The final feature update result is obtained only after the number of jumps reaches the preset value.

8. The method as described in claim 7, characterized in that, When the lane line query feature is the result of the previous feature update, the step of obtaining the three-dimensional location information feature includes: Obtain the rotation offset around the x-axis and the z-axis offset that are in the same cycle as the feature update result; Based on the rotation offset around the x-axis and the offset around the z-axis, a transformation matrix is ​​generated; Based on the transformation matrix, the target two-dimensional matrix that is in the same round as the feature update result is updated to obtain a new target two-dimensional matrix; The new target two-dimensional matrix is ​​subjected to multi-level nonlinear transformation to obtain the three-dimensional position information features.

9. The method as described in claim 8, characterized in that, The step of obtaining the x-axis rotation offset and z-axis offset that are in the same cycle as the feature update result includes: Obtain the feature map and the historical target two-dimensional matrix, wherein the historical target two-dimensional matrix is ​​the target two-dimensional matrix that is in the same round as the feature update result; The feature map and the historical target two-dimensional matrix are concatenated to obtain the concatenation result; The splicing result is subjected to a fifth feature extraction process to obtain the feature extraction result; The feature extraction results are then subjected to pooling to obtain the pooling result; The pooling result is subjected to a multi-layer nonlinear transformation to obtain the x-axis rotation offset and z-axis offset, which are in the same cycle as the feature update result.

10. The method as described in claim 6, characterized in that, The lane line query features are updated based on the feature map, the three-dimensional location information features, and the N×M three-dimensional reference points to obtain feature update results, including: The N×M three-dimensional reference points are projected to obtain N×M two-dimensional reference points; The lane line query features, the feature map, the three-dimensional location information features, and the N×M two-dimensional reference points are subjected to deformable cross-attention processing to obtain the feature update results.

11. The method as described in claim 10, characterized in that, The process of performing deformable cross-attention processing on the lane line query features, the feature map, the three-dimensional location information features, and the N×M two-dimensional reference points to obtain feature update results includes: The lane line query features are subjected to self-attention processing to obtain the self-attention processing result; The self-attention processing result, the feature map, the three-dimensional position information features, and the multiple two-dimensional reference points are subjected to deformable cross-attention processing to obtain a deformable attention processing result. The deformable attention processing result is subjected to multi-layer nonlinear transformation to obtain the feature update result.

12. The method as described in claim 6, characterized in that, After updating the lane line query features based on the feature map, the three-dimensional location information features, and the N×M three-dimensional reference points to obtain the feature update result, the method further includes: The feature update result is subjected to max pooling to obtain the max pooling result; The maximum pooling result is subjected to a multi-level nonlinear transformation to obtain the category to which the lane line belongs.

13. The method as described in claim 6, characterized in that, Based on the feature update result, N×M position offsets are obtained, including: The feature update result is subjected to a multi-layer nonlinear transformation to obtain multiple position offsets and the visibility corresponding to each position offset.

14. The method as described in claim 1, characterized in that, Before performing feature extraction processing on the road image to obtain the feature map and global lane line features, the method further includes: Obtain the lane line detection model; The lane line detection model is trained to obtain a trained lane line detection model. The lane detection method as described in claim 1 is executed using the trained lane detection model to determine the N×M predicted location points.

15. The method as described in claim 14, characterized in that, The step of training the lane detection model to obtain the trained lane detection model includes: The training road images are subjected to the first feature extraction process to obtain the training feature map; The training feature map is subjected to a second feature extraction process to obtain the global features of the training lane lines; Randomly initialize the local features of the training lane lines, and perform feature fusion processing on the global features of the training lane lines and the local features of the training lane lines to obtain the training feature fusion result, which is denoted as the training lane line query feature. The training lane line query features are subjected to a first multi-level nonlinear transformation to obtain multiple training three-dimensional reference points, and each training three-dimensional reference point is projected to obtain multiple training two-dimensional reference points. Obtain training 3D position information features; Deformable cross-attention processing is performed on the training lane line query features, the training feature map, the training three-dimensional position information features, and the multiple training two-dimensional reference points to obtain the training feature update results. The training feature update results are subjected to a second multi-layer nonlinear transformation to obtain multiple training position offsets and the training visibility corresponding to each training position offset; The training feature update results are subjected to max pooling to obtain the training max pooling results; The training max pooling results are subjected to a sixth-level nonlinear transformation to obtain the training category to which the lane lines belong. Based on the training feature map and the global features of the training lane lines, the training lane line segmentation result is obtained; Based on the training lane segmentation results, training 3D position information features, training position offset, training visibility, and training category, a loss function is constructed. If the loss function meets the preset requirements, it is determined that the training of the lane detection model is complete, and the trained lane detection model is obtained.

16. A lane line detection device, characterized in that, The device includes: The feature extraction unit is used to perform feature extraction processing on the road image to obtain a feature map and global lane line features. The feature map contains lane line information, and the global lane line features carry an upper limit value N for the detection of the number of lane lines. The feature fusion unit is used to acquire local lane line features and perform feature fusion processing on the global lane line features and the local lane line features to obtain lane line query features. The local lane line features carry sampling point values ​​M for detecting lane line positions. The lane line query features carry both the detection upper limit value N and the sampling point values ​​M. The reference point acquisition unit is used to obtain N×M three-dimensional reference points based on the lane line query features; wherein, the N×M three-dimensional reference points are used to indicate the reference positions of M sampling points of each lane line in the three-dimensional coordinate system. A location information unit is used to acquire three-dimensional location information features, wherein the three-dimensional location information features are used to correct the height position of the lane line in three-dimensional space; The offset acquisition unit is used to obtain N×M position offsets based on the feature map, the three-dimensional position information features, the N×M three-dimensional reference points and the lane line query features, wherein the position offsets correspond one-to-one with the three-dimensional reference points; The prediction point acquisition unit is used to offset the corresponding three-dimensional reference point based on the position offset for each position offset to obtain a position prediction point. The N×M position prediction points represent that for each of the N lane lines, the contour of the corresponding lane line is indicated by M position prediction points.

17. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps in the lane line detection method as described in any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the lane line detection method according to any one of claims 1 to 15.