Lane line detection method, device, equipment, vehicle, medium and product
By collecting multi-focal lane line images on the vehicle and directly generating 3D coordinates using specific detection models, the problems of low accuracy and complex calculation of lane line 3D information in the prior art are solved, and more efficient and accurate lane line detection is achieved to meet the accuracy and real-time requirements of autonomous driving.
Patent Information
- Application Number
- CN202510087840.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the lane line 3D information output using network models has low accuracy and complex calculations, making it difficult to meet the accuracy and real-time requirements in autonomous driving.
By installing multiple image acquisition devices on the vehicle, we collect lane line images with different focal lengths, and using detection models including query vector generation module, 3D position encoding generation module, transformer module and 3D lane line detection module, we directly generate the 3D coordinates of the lane line to avoid the conversion between 2D and 3D coordinates.
It improves the accuracy and efficiency of lane line detection, reduces the error caused by coordinate conversion, and meets the requirements of automatic driving for accuracy and real-time.
Smart Images

Figure CN120014577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a lane line detection method, device, equipment, vehicle, medium and product. Background Art
[0002] With the rapid development of autonomous driving, lane detection is crucial in autonomous driving. The network model based on deep learning is the core tool for lane detection, which enables vehicles to maintain normal driving and make path decisions in complex environments. Therefore, it is very important to ensure the accuracy of lane information output by the network model.
[0003] In the prior art, the three-dimensional (3D) information of lane lines output by a deep learning-based network model for lane line detection is usually first output by the network model to predict the two-dimensional (2D) information of lane lines, and then a mapping relationship between the world coordinate system and the camera coordinate system is established to map the 2D lane lines to the 3D coordinate system to obtain the 3D information of lane lines.
[0004] However, the 3D lane line information output by the network model in the existing technology has low accuracy and complex calculation, which makes it difficult to meet the requirements of accuracy and real-time performance in autonomous driving. Summary of the invention
[0005] One of the purposes of the present invention is to provide a lane line detection method, device, equipment, vehicle, medium and product to solve the problem in the prior art that the lane line 3D information output by the network model has low accuracy, complex calculation, and is difficult to meet the accuracy and real-time requirements in autonomous driving.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A lane line detection method, characterized in that the method comprises:
[0008] A plurality of lane line images are collected by using a plurality of image acquisition devices installed on the vehicle, each of the image acquisition devices having a different focal length;
[0009] Inputting the multiple lane line images into a lane line detection model, and obtaining 3D coordinates of the lane lines output by the lane line detection model, wherein the lane line detection model includes a query vector generation module, a 3D position code generation module, a transformer module, and a 3D lane line detection module;
[0010] When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line according to the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; the converter module integrates the image features and 3D position code of each lane line image into the first query vector to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line according to the second query vector.
[0011] Through the above technical means, during the driving process of the vehicle, multiple lane line images with different focal lengths are collected by multiple image acquisition devices installed on the vehicle to ensure the accuracy and completeness of the lane lines represented by the lane line images. When the lane line detection model processes multiple lane line images and obtains the 3D coordinates of the lane lines, there is no need to convert between 2D coordinates and 3D coordinates, which improves processing efficiency and reduces errors caused by too many conversion operations, thereby ensuring the accuracy of the generated 3D coordinates and meeting the requirements of autonomous driving for accuracy and real-time performance.
[0012] Further, the transformer module includes a self-attention submodule, a deformable cross-attention submodule and a FFN submodule;
[0013] When the lane line detection model generates the 3D coordinates of the lane line, the self-attention submodule calculates the attention score between each element in the first query vector and other elements, and updates the elements according to the attention score to obtain a third query vector;
[0014] The deformable cross attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain a fourth query vector;
[0015] The FFN submodule updates the fourth query vector using a nonlinear activation function to obtain the second query vector.
[0016] According to the above technical means, by combining the self-attention submodule, the deformable cross-attention submodule and the FNN submodule in a network connection mode, the learning ability of the global feature dependency is enhanced, and the lane line detection model is further enabled to extract effective information from features of different scales.
[0017] Furthermore, the query vector generation module also includes a convolution submodule and an MLP submodule;
[0018] When the lane line detection model generates the 3D coordinates of the lane line, the convolution submodule splices and fuses the image features of each lane line image to generate a first splicing feature;
[0019] The MLP submodule generates a first query vector according to the first concatenated feature.
[0020] According to the above technical means, a query vector generation module is composed of a convolution sub-module and an MLP sub-module, which combines the spatial capture capability of the convolution operation and the nonlinear mapping advantage of the MLP sub-module, so that the generated first query vector carries rich semantic information, reduces the computational complexity, and improves the response speed.
[0021] Furthermore, the lane line detection model further includes a feature extraction module, which extracts image features of each lane line image, and flattens and splices the image features of each lane line image to generate a second splicing feature;
[0022] The 3D position code generation module processes the image features of each sample lane line image to generate a 3D position code for each sample lane line image, and flattens and splices the 3D position code for each sample lane line image to generate a spliced 3D position feature, wherein the format of the second spliced feature is the same as the format of the spliced 3D position feature.
[0023] According to the above technical means, the format of the second stitching feature is adjusted to be consistent with the format of the stitching 3D position feature, so as to facilitate the subsequent feature superposition operation.
[0024] Furthermore, the deformable cross attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain a fourth query vector, including:
[0025] The deformable cross-attention sub-module incorporates the second splicing feature and the splicing 3D position feature into the third query vector to obtain the fourth query vector.
[0026] According to the above technical means, since the stitched 3D position feature contains the position information of the lane line in three-dimensional space, the second stitching feature is superimposed with the stitching 3D position feature, so that the position information of the lane line in three-dimensional space can be integrated into the features of the lane line. This enables the lane line detection model to better understand the positional relationship between each pixel point in the lane line image in the three-dimensional scene, which helps to more accurately detect and identify the 3D coordinates of the lane line.
[0027] Furthermore, the method further comprises:
[0028] Obtain the lane line category and obstacle information output by the lane line detection model.
[0029] According to the above technical means, the lane line detection model not only has the function of identifying the 3D coordinates of the lane line, but also can have the function of identifying the lane line category and obstacle information, thereby improving the comprehensiveness of the acquired lane line information.
[0030] Furthermore, before collecting a plurality of lane line images by using a plurality of image acquisition devices installed on the vehicle, the method further includes:
[0031] Acquire a plurality of sample lane line images in front of a sample vehicle, and a sample label of each sample lane line image, wherein the sample label includes a sample 2D coordinate and a sample 3D coordinate of the sample lane line, and the plurality of sample lane line images are acquired by a sample image acquisition device with different focal lengths installed in the sample vehicle;
[0032] The initial model is trained according to the multiple sample lane line images and the label of each sample lane line image to obtain the lane line detection model.
[0033] According to the above technical means, model training is performed by using multiple sample lane line images collected by sample image acquisition devices with different focal lengths and sample labels of each sample lane line image to obtain a lane line detection model, laying the foundation for the subsequent acquisition of the 3D coordinates of the lane line based on the lane line detection model.
[0034] Furthermore, the training of the initial model according to the multiple sample lane line images and the label of each sample lane line image to obtain the lane line detection model includes:
[0035] Processing the plurality of sample lane line images by using the initial model to obtain predicted 3D coordinates of the sample lane lines generated by the initial model;
[0036] Training model parameters in the initial model according to the predicted 3D coordinates of the sample lane line and the corresponding sample 3D coordinates;
[0037] Project the predicted 3D coordinates of the sample lane line onto the corresponding sample lane line image to obtain the predicted 2D coordinates of each sample lane line;
[0038] According to the predicted 2D coordinates of each sample lane line and the corresponding sample 2D coordinates, the model parameters in the initial model are trained until a training cutoff condition is met, thereby obtaining the lane line detection model.
[0039] According to the above technical means, by performing error feedback between the predicted 2D coordinates and the sample 2D coordinates, and performing error feedback again between the predicted 3D coordinates and the sample 3D coordinates, the optimization of the model parameters is achieved, thereby further improving the accuracy and robustness of the lane line position prediction of the lane line detection model.
[0040] Furthermore, the sample label includes lane line category and obstacle information.
[0041] According to the above technical means, by obtaining detailed sample labels, richer lane line scenarios can be provided during model training, and the shape, position and characteristics of lane lines can be better understood, so that the lane line detection model can recognize lane lines.
[0042] A lane line detection device, comprising:
[0043] A collection module, used for collecting multiple lane line images by using multiple image collection devices installed on the vehicle, each image collection device having a different focal length;
[0044] An input module, used to input the multiple lane line images into a lane line detection model to obtain the 3D coordinates of the lane lines output by the lane line detection model, wherein the lane line detection model includes a query vector generation module, a 3D position code generation module, a transformer module and a 3D lane line detection module;
[0045] When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line according to the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; the converter module integrates the image features and 3D position code of each lane line image into the first query vector to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line according to the second query vector.
[0046] Further, the transformer module includes a self-attention submodule, a deformable cross-attention submodule and a FFN submodule;
[0047] When the lane line detection model generates the 3D coordinates of the lane line, the self-attention submodule calculates the attention score between each element in the first query vector and other elements, and updates the elements according to the attention score to obtain a third query vector;
[0048] The deformable cross attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain a fourth query vector;
[0049] The FFN submodule updates the fourth query vector using a nonlinear activation function to obtain the second query vector.
[0050] Furthermore, the query vector generation module also includes a convolution submodule and an MLP submodule;
[0051] When the lane line detection model generates the 3D coordinates of the lane line, the convolution submodule splices and fuses the image features of each lane line image to generate a first splicing feature;
[0052] The MLP submodule generates a first query vector according to the first concatenated feature.
[0053] Furthermore, the lane line detection model further includes a feature extraction module, which extracts image features of each lane line image, and flattens and splices the image features of each lane line image to generate a second splicing feature;
[0054] The 3D position code generation module processes the image features of each sample lane line image to generate a 3D position code for each sample lane line image, and flattens and splices the 3D position code for each sample lane line image to generate a spliced 3D position feature, wherein the format of the second spliced feature is the same as the format of the spliced 3D position feature.
[0055] Furthermore, the deformable cross attention submodule updates the third query vector according to the image features and 3D position encoding of each lane line image to obtain a fourth query vector, including:
[0056] The deformable cross-attention sub-module incorporates the second splicing feature and the splicing 3D position feature into the third query vector to obtain the fourth query vector.
[0057] Furthermore, the lane line detection device also includes an acquisition module, which is used to:
[0058] Obtain the lane line category and obstacle information output by the lane line detection model.
[0059] Furthermore, before collecting a plurality of lane line images by using a plurality of image acquisition devices installed on the vehicle, the lane line detection device further includes a training module for:
[0060] Acquire a plurality of sample lane line images in front of a sample vehicle, and a sample label of each sample lane line image, wherein the sample label includes a sample 2D coordinate and a sample 3D coordinate of the sample lane line, and the plurality of sample lane line images are acquired by a sample image acquisition device with different focal lengths installed in the sample vehicle;
[0061] The initial model is trained according to the multiple sample lane line images and the label of each sample lane line image to obtain the lane line detection model.
[0062] Furthermore, the lane line detection device also includes a training module, which is specifically used to:
[0063] Processing the plurality of sample lane line images by using the initial model to obtain predicted 3D coordinates of the sample lane lines generated by the initial model;
[0064] Training model parameters in the initial model according to the predicted 3D coordinates of the sample lane line and the corresponding sample 3D coordinates;
[0065] Project the predicted 3D coordinates of the sample lane line onto the corresponding sample lane line image to obtain the predicted 2D coordinates of each sample lane line;
[0066] According to the predicted 2D coordinates of each sample lane line and the corresponding sample 2D coordinates, the model parameters in the initial model are trained until a training cutoff condition is met, thereby obtaining the lane line detection model.
[0067] Furthermore, the sample label includes sample lane line category and sample obstacle information.
[0068] An electronic device includes: a processor, a memory, and computer-executable instructions stored in the memory and executable on the processor, wherein the processor is used to implement the above-mentioned model training method and lane line detection method when executing the computer-executable instructions.
[0069] A vehicle comprises: a vehicle body, a processor, a memory and computer execution instructions stored in the memory and executable on the processor, wherein the processor is used to implement the above-mentioned model training method and lane line detection method when executing the computer execution instructions.
[0070] A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to implement the above-mentioned model training method and lane line detection method when executed by a processor.
[0071] A computer program product includes a computer program, characterized in that when the computer program is executed by a processor, it is used to implement the above-mentioned model training method and lane line detection method.
[0072] Beneficial effects of the present invention:
[0073] (1) During the driving process of the vehicle, multiple lane line images with different focal lengths are collected by multiple image acquisition devices installed on the vehicle to ensure the accuracy and completeness of the lane lines represented by the lane line images.
[0074] (2) When the lane line detection model processes multiple lane line images and obtains the 3D coordinates of the lane line, there is no need to convert between 2D coordinates and 3D coordinates, which improves processing efficiency and reduces errors caused by too many conversion operations, thereby ensuring the accuracy of the generated 3D coordinates and meeting the requirements of autonomous driving for accuracy and real-time performance.
[0075] (3) During the model training process, the predicted 3D coordinates and the sample 3D coordinates are used for error feedback to achieve model parameter optimization. Then, the predicted 2D coordinates and the sample 2D coordinates are used for error feedback again to further optimize the model parameters, thereby further improving the accuracy and robustness of the lane line position prediction of the lane line detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 1 ;
[0077] Figure 2 A schematic diagram of viewing angles of image acquisition devices with different focal lengths provided by an embodiment of the present invention;
[0078] Figure 3 A schematic diagram of the structure of a query vector generation module provided in an embodiment of the present invention;
[0079] Figure 4 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 2 ;
[0080] Figure 5 A schematic diagram of 2D coordinate points of lane lines provided in an embodiment of the present invention;
[0081] Figure 6 A schematic diagram of the network structure of the initial model provided by an embodiment of the present invention;
[0082] Figure 7 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 3 ;
[0083] Figure 8 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 4 ;
[0084] Fig. 9 A schematic diagram of the structure of a lane line detection device provided by an embodiment of the present invention;
[0085] Fig.10 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention
[0086] Fig.11 A schematic structural diagram of a vehicle provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0087] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances shall be provided for users to choose to authorize or refuse.
[0089] Before introducing the present invention, the application background of the present invention is first explained.
[0090] With the rapid development of autonomous driving, lane detection is crucial in autonomous driving systems. Lane detection can help vehicles stay in lanes, reduce the risk of deviation and collision, and by detecting lanes in real time, autonomous driving systems can plan the best driving path and provide users with a good driving experience.
[0091] The deep learning-based network model is the core tool for realizing lane line detection, so that vehicles can maintain normal driving and make path decisions in complex environments. Therefore, ensuring the accuracy of lane line information output by the network model is an urgent problem to be solved.
[0092] In the prior art, the 3D lane line information output by the deep learning-based network model for lane line detection is usually obtained by first inputting the labeled lane line data into the deep learning model (such as a convolutional neural network), then extracting features in the image using a multi-layer structure, and performing feature learning, and finally outputting the 2D coordinate position of the lane line in the image; then, the 2D coordinate position is projected into the 3D world through the Inverse Perspective Mapping (IPM) method to obtain the 3D coordinate position of the lane line.
[0093] However, the prior art has the following technical problems:
[0094] 1. The existing technology requires two steps to obtain the 3D coordinates of the lane line, that is, first obtain the 2D coordinates, and then convert the 2D coordinates into 3D coordinates. The whole process is complicated and the processing efficiency is low. In addition, the conversion process between 2D coordinates and 3D coordinates will bring errors, resulting in low processing accuracy.
[0095] 2. The IPM method is strictly based on the assumption of flat ground. In actual driving, it is difficult to avoid ups and downs and bumps, which makes the accuracy of the 3D coordinates calculated by the IPM method low.
[0096] In summary, the 3D lane line information output by the existing technology using the network model has low accuracy and complex calculation, which makes it difficult to meet the accuracy and real-time requirements of autonomous driving.
[0097] Based on this, the technical concept of the present invention is as follows: a lane line detection model can be pre-trained, and the lane line detection model can directly generate the 3D coordinates of the lane line based on the input lane line image, without the need to convert between 2D coordinates and 3D coordinates, thereby avoiding the problems of low processing accuracy and low efficiency caused by coordinate conversion. At the same time, the 3D coordinate generation process does not need to be implemented through the IPM method, which further improves the accuracy of the acquired 3D coordinates.
[0098] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0099] Figure 1 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 1 ;like Figure 1 As shown, the lane line detection method may include the following steps:
[0100] S11. Capture multiple lane line images using multiple image acquisition devices installed on the vehicle.
[0101] The execution subject of the embodiment of the present invention is an electronic device or a vehicle. The electronic device can be a terminal device, such as a laptop computer, a desktop computer, a tablet computer, etc., and can also be a server. In practical applications, the electronic device can be a terminal device, a server, or a vehicle, which can be determined according to actual conditions and is not specifically limited.
[0102] Exemplarily, the image acquisition device may be a front camera for acquiring a lane line image of the road in front of the vehicle, wherein the lane line image includes the lane line in front of the vehicle. It should be understood that the image acquisition device may also be other devices with image acquisition functions, and the embodiments of the present invention do not specifically limit this.
[0103] It should be understood that the images of the multiple lane line images have the same size.
[0104] Among them, the focal length of each image acquisition device is different.
[0105] Figure 2 A schematic diagram of the viewing angle of an image acquisition device with different focal lengths provided by an embodiment of the present invention, such as Figure 2 As shown, Figure 2 There are two image acquisition devices in total. The image acquisition device with a shorter focal length can obtain lane line images with a short focal field of view (FOV) of 120°, which can cover the road information within 150m in front of the vehicle. The image acquisition device with a longer focal length can obtain lane line images with a long focal field of view (FOV) of 30°, which can cover the road information within 150m-250m in front of the vehicle.
[0106] Among them, the size of the short-focus image collected by the image acquisition device with a shorter focal length and the size of the long-focus image collected by the image acquisition device with a longer focal length are both 2160*3840.
[0107] S12: Input multiple lane line images into a lane line detection model to obtain the 3D coordinates of the lane lines output by the lane line detection model.
[0108] Among them, the lane line detection model includes a query vector generation module, a 3D position code generation module, a transformer module and a 3D lane line detection module.
[0109] When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line based on the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; the converter module integrates the image features and 3D position code of each lane line image into the first query vector to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line based on the second query vector.
[0110] The first query vector is used to represent the position of the lane line in the lane line image. For example, Figure 5 As shown, for Figure 5 In the lane line image shown, the first query vector has a size of 6*20*256, where 6 is the number of lane lines, 20 is the number of position points, and 256 is the number of channels.
[0111] Furthermore, the converter module integrates the image features and 3D position codes of each lane line image into the first query vector, so that the obtained second query vector integrates the feature information of the image (such as the color, shape and other features of the lane line) and the 3D position code information (the position relationship of the lane line in three-dimensional space, etc.), so as to facilitate the subsequent processing of the second query vector by the 3D lane line detection module to generate the 3D coordinates of the lane line to obtain the 3D position information of the lane line.
[0112] In practical applications, the query vector generation module can be implemented as a Query generator, the 3D position encoding generation module can be implemented as a 3D position encoding (PE) generator, the transformer module can be implemented as a transformer network, and the 3D lane line detection module can be implemented as a 3D lane line detection head.
[0113] It should be noted that when the execution subject is a vehicle, the vehicle needs to include a memory and a processor. The memory is used to store the lane line detection model and related network parameters to enable the vehicle to accurately detect the lane line in front of the road; the processor is used to process the lane line image input by the image acquisition device, and input its image processing results into the lane line detection model, and finally output the detection results of the model.
[0114] The lane line detection method provided by the embodiment of the present invention collects multiple lane line images through multiple image acquisition devices installed on the vehicle, and each image acquisition device has a different focal length; then the multiple lane line images are input into the lane line detection model to obtain the 3D coordinates of the lane line output by the lane line detection model, and the lane line detection model includes a query vector generation module, a 3D position code generation module, a converter module and a 3D lane line detection module. When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line according to the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; the converter module is used to update the first query vector according to the image features and 3D position code of each lane line image to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line according to the second query vector.
[0115] In this technical solution, since the lane line has a long span when it is close to the vehicle, and the width of the lane line gradually becomes thinner when it is close to the vehicle. Therefore, multiple lane line images with different focal lengths are collected by multiple image acquisition devices installed on the vehicle to ensure the accuracy and completeness of the lane lines represented by the lane line images. When the lane line detection model processes multiple lane line images and obtains the 3D coordinates of the lane lines, there is no need to convert between 2D coordinates and 3D coordinates, which improves processing efficiency and reduces errors caused by too many conversion operations, thereby ensuring the accuracy of the generated 3D coordinates and meeting the requirements of autonomous driving for accuracy and real-time performance.
[0116] Optionally, in some embodiments, the lane line detection model further includes a feature extraction module, which extracts image features of each lane line image, and flattens and splices the image features of each lane line image to generate a second spliced feature. The 3D position code generation module processes the image features of each sample lane line image, generates a 3D position code of each sample lane line image, and flattens and splices the 3D position code of each sample lane line image to generate a spliced 3D position feature. The format of the second spliced feature is the same as the format of the spliced 3D position feature.
[0117] Optionally, the feature extraction module can be implemented as a convolutional neural network.
[0118] For example, it is assumed that the telephoto image and the short-focus image acquired by the image acquisition device are first scaled and cropped to obtain two images with a resolution of 512*960 to ensure that the size of the input model is appropriate, where 512 represents the number of pixels in the height of the image, and 960 represents the number of pixels in the width, that is, the image contains 512*960 pixels. The feature extraction module extracts features from the cropped telephoto image and the cropped short-focus image to obtain two feature maps of size H*W*C (assuming 16*30*256) (the short-focus feature map is denoted as F1 and the telephoto feature map is denoted as F2), where H and W are the height and width of the feature, respectively, and C is the number of channels of the feature. Further, the feature extraction module flattens the two image features F1 and F2 into 480*256 respectively, and then splices them together to obtain a second spliced feature of 960*256.
[0119] It can be understood that by extracting and fusing multi-scale features of long-focus images and short-focus images, effective integration of local image information and global image information is achieved.
[0120] Furthermore, the 3D position coding generation module is used to build the connection between 2D space and 3D space. For example, given the image feature F im , each pixel (u, v) can be defined as a series of points {pk (u,v)=(u*d k ,v*d k ,d k ,1) T ,k=1,2,…,d}, where d is the number of points sampled along the depth direction, then the corresponding point coordinates projected in the 3D world are in is the transformation matrix between the i-th camera coordinate system and the vehicle coordinate system, is the intrinsic conversion matrix of the i-th camera. Then the position encoding of this pixel can be obtained as Where φ is MLP. By performing the above operation on each pixel in the image feature, the 3D PE of each pixel can be obtained, and then the 3D position encoding of each sample lane line image can be obtained.
[0121] Afterwards, similar to the generation method of the second splicing feature, the 3D position code of each sample lane line image is flattened and spliced together to obtain a spliced 3D position feature with a size of 960*256.
[0122] In the above embodiment, the format of the second stitching feature is adjusted to be consistent with the format of the stitching 3D position feature, so as to facilitate the subsequent feature superposition operation.
[0123] Optionally, in some embodiments, the query vector generation module further includes a convolution submodule and a multi-layer perceptron (MLP) submodule. When the lane line detection model generates the 3D coordinates of the lane line, the convolution submodule splices and fuses the image features of each lane line image to generate a first splicing feature. The MLP submodule generates a first query vector based on the first splicing feature.
[0124] It can be understood that the first query vector generated in this way can extract the most significant feature information for lane line detection from the original features, reduce unnecessary feature interference, and improve the output efficiency of the model.
[0125] After the convolutional neural network extracts features, the convolution submodule concatenates and further fuses the two feature maps, which not only enhances the model's ability to capture multi-scale features of the image, but also improves the richness of feature expression. The final output concatenated feature map is then passed through the MLP submodule to generate the first query vector, which helps the model identify lane lines more accurately in multi-scale and complex scenes.
[0126] Figure 3 Schematic diagram of the structure of the query vector generation module provided in an embodiment of the present invention. Figure 3As shown, the query vector generation module includes: a convolution submodule and an MLP submodule. Continuing with the above example, the convolution submodule splices the feature map F1 and the feature map F2 to obtain a 16*60*256 spliced feature map, and then fuses the spliced map again to obtain a 8*30*256 fused feature map. The fused feature map is straightened to generate a first spliced feature; finally, the convolution submodule outputs the first spliced feature to the MLP submodule, and the MLP submodule generates the first query vector by resizing.
[0127] In the above embodiment, the query vector generation module is composed of a convolution submodule and an MLP submodule, which combines the spatial capture capability of the convolution operation and the nonlinear mapping advantage of the MLP submodule, so that the generated first query vector carries rich semantic information, reduces computational complexity, and improves response speed.
[0128] Optionally, in some embodiments, the transformer module includes a self-attention submodule, a deformable cross-attention submodule, and a feed-forward network (FFN) submodule. When the lane line detection model generates the 3D coordinates of the lane line, the self-attention submodule calculates the attention score between each element and other elements in the first query vector, and updates the elements according to the attention score to obtain a third query vector. The deformable cross-attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain a fourth query vector. The FFN submodule updates the fourth query vector using a nonlinear activation function to obtain a second query vector.
[0129] Among them, the self-attention submodule can capture the relationship between different elements in the first query vector, and adjust its own representation according to the correlation between them, so as to better represent the structural information of the entire lane line. The deformable cross-attention submodule updates the elements of the first query vector, so that the obtained third query vector integrates the feature information of the image (such as the color and shape of the lane line) and the 3D position encoding information (the position relationship of the lane line in three-dimensional space, etc.). The FFN submodule uses a nonlinear activation function to process the nonlinear features in the fourth query vector, which can better express the type, width, shape and other attributes of the vehicle.
[0130] In a specific implementation, the deformable cross attention submodule integrates the second splicing feature and the splicing 3D position feature into the third query vector to obtain a fourth query vector.
[0131] In the above example, the second spliced feature of 960*256 is the input k and v of the deformable cross attention submodule, and the spliced 3D position feature of size 960*256 is the input k pos of the deformable cross attention submodule.
[0132] It should be noted that in the transformer module, the first query vector is first input into the self-attention submodule, which uses a multi-head self-attention mechanism to divide the self-attention into multiple "heads", each head independently performs attention calculations, so that features can be learned from multiple different subspaces, global information and long-distance dependencies can be obtained, and then the input of the deformable cross-attention submodule (the third query vector) can be obtained. The specific calculation formula of the multi-head attention mechanism is:
[0133] MultiHead(Q′,K′,V′)=Concat(head1,head2,…,head h )W o
[0134] head i =Attention(Q i ,K i ,V i )
[0135]
[0136] Q=K,V
[0137] Among them, W o It is a full connection operation. Q is a vector extracted from the first query vector, which represents the features that need to be paid attention to at present. The key (Key, K) is a vector corresponding to the first query vector, which is formed by different linear transformations. The value (Value, V) is another vector corresponding to the first query vector, which is also formed by a certain linear transformation.
[0138] Afterwards, in order to increase the calculation speed while ensuring accuracy, the scope of attention calculation is narrowed, and only the sampling points within the range around the reference point are focused on for feature extraction. Therefore, the second spliced feature, the spliced 3D position feature, and the third query vector are input into the deformable cross attention submodule together. The deformable cross attention submodule only uses a part of the key points for attention calculation. It can calculate more detailed attention weights by interacting with the second spliced feature, the spliced 3D position feature, and the third query vector, effectively improving the model's ability to express features and its understanding of contextual information. The specific deformable attention calculation formula is:
[0139]
[0140] Among them, p q is the reference point of 2D; Δp mqk is the sampling offset; W m and W′ m is a learnable weight; A mqk is the attention weight; z q is the third query vector after being processed by the deformable attention module; x is the image feature; K is the number of sampled points, which can be set to 4.
[0141] Finally, in order to increase the expressive power of the model and enhance the ability to obtain local information, the fourth query vector processed by the cross attention module can be input into the FFN submodule and processed using the nonlinear activation function in the FFN submodule, namely the Rectified Linear Unit (ReLU), to obtain the second query vector. The specific calculation formula of the sub-FNN module can be:
[0142] FFN(x)=max(0,xW1+b1)W2+b2
[0143] Among them, W1 and W2 are weight matrices, b1 and b2 are bias terms, and max(0,x) represents the ReLU activation function.
[0144] It is understandable that by inputting the first query vector into the self-attention submodule and utilizing the multi-head self-attention mechanism, the model can learn features from multiple subspaces, capture global information and long-distance dependencies, and provide input for the deformable cross-attention submodule. Subsequently, by narrowing the attention calculation range and focusing on the sampling points around the reference point, the calculation speed can be effectively improved while maintaining accuracy. The second splicing feature, 3D position feature, and the third query vector processed by the self-attention submodule further enhance the ability to obtain local information. Finally, the fourth query vector after cross-attention processing is input into the FFN submodule to generate the second query vector, further improving the expression ability of the model. This method comprehensively improves the model's understanding and processing capabilities of features and increases the stability of the model.
[0145] In the above embodiment, by combining the self-attention submodule, the deformable cross-attention submodule and the FNN submodule in a network connection mode, the learning ability of the global feature dependency is enhanced, and the lane line detection model is further enabled to extract effective information from features of different scales. In addition, since the spliced 3D position feature contains the position information of the lane line in three-dimensional space, the second spliced feature is superimposed with the spliced 3D position feature, and the position information of the lane line in three-dimensional space can be integrated into the features of the lane line, which enables the lane line detection model to better understand the position relationship of each pixel in the lane line image in the three-dimensional scene, which helps to more accurately detect and identify the 3D coordinates of the lane line.
[0146] Optionally, in some embodiments, the lane line category and obstacle information output by the lane line detection model may also be obtained.
[0147] The 3D lane detection head consists of two branch structures, namely the regression branch and the classification branch. The regression branch is used to generate 3D coordinates, namely the y-direction offset, z-direction offset, and obstacle information of the 3D lane image pixel points; the classification branch is used to obtain the category of each lane and obstacle information. Therefore, when the 3D lane detection head outputs the 3D coordinates of the lane, it can also output the lane category and obstacle information of the lane.
[0148] The specific lane line categories may be solid lines, dotted lines, dividing lines, etc., and the obstacle information is whether the lane line is blocked by obstacles, causing the lane line to be invisible.
[0149] In this implementation, the lane line detection model not only has the function of identifying the 3D coordinates of the lane line, but also can have the function of identifying the lane line category and obstacle information, thereby improving the comprehensiveness of the acquired lane line information.
[0150] Next, the model training process is explained. It should be understood that the execution subject of the model training process and the execution subject of the lane line detection method can be the same or different, and there is no specific limitation on this.
[0151] Figure 4 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 2 .like Figure 4 As shown, before collecting multiple lane line images by multiple image acquisition devices installed on the vehicle, the method also includes the following steps:
[0152] S41, obtaining a plurality of sample lane line images in front of a sample vehicle and a sample label of each sample lane line image.
[0153] The sample labels include sample 2D coordinates and sample 3D coordinates of the sample lane lines, and the multiple sample lane line images are collected by sample image acquisition devices with different focal lengths installed in the sample vehicle.
[0154] Exemplarily, the model training process can be implemented in Ubuntu 22.04 and Pytorch deep learning framework. It should be understood that the model training process can also be implemented in other model frameworks, and the embodiment of the present invention does not specifically limit this.
[0155] Furthermore, sample lane line images can be collected by sample image acquisition devices with different focal lengths in multiple vehicle use scenarios. Afterwards, 80% of the sample lane line images are used for model training, and 20% of the sample lane line images are used for model verification.
[0156] In practical applications, the above ratios (80% and 20%) can be predetermined according to actual conditions, and the embodiment of the present invention does not impose any specific limitation on this.
[0157] Optionally, in order to make the output lane line information better show the actual position and shape of the lane line and improve the practicality of the lane line, the input sample label also includes sample lane line category and sample obstacle information.
[0158] It can be understood that by obtaining sample lane line images corresponding to images with different focal lengths, as well as detailed sample labels, the model can provide richer lane line scenes during training, better understand the shape, position, and characteristics of the lane lines, and enable the lane line detection model to have the ability to identify lane line types and obstacle information, further improving the model output accuracy and robustness.
[0159] Figure 5 A schematic diagram of 2D coordinate points of lane lines provided in an embodiment of the present invention; Figure 5 As shown, the vehicle coordinate system takes the ground at the center of the vehicle's rear axle as the origin, the vehicle's forward direction is the positive X direction, the Y direction is perpendicular to the left side of the vehicle, and the Z direction follows the right-hand vertical upward rule. Each black dot represents a position point in a lane line.
[0160] For example, there are 6 lanes to be detected, the detection distance is 250m, and each lane has 20 points evenly distributed in the positive X direction. Then the coordinates of each black dot are Among them, i represents the lane line number, and j represents the number of the dot distributed along the positive direction of X.
[0161] S42: Train the initial model according to the multiple sample lane line images and the label of each sample lane line image to obtain a lane line detection model.
[0162] In the above embodiment, model training is performed based on multiple sample lane line images collected by sample image acquisition devices with different focal lengths and sample labels of each sample lane line image to obtain a lane line detection model, laying the foundation for the subsequent acquisition of the 3D coordinates of the lane line based on the lane line detection model.
[0163] Figure 6 A schematic diagram of the network structure of the initial model provided in the embodiment of the present invention; Figure 6 As shown, this embodiment, based on any of the above embodiments, describes in detail the network structure of the initial model, and the initial model includes: a convolutional neural network, a query generator, a 3D PE generator, a transformer network and a 3D lane line detection head.
[0164] Among them, the convolutional neural network adopts the residual network (Residual Network, ResNet) 50+ Feature Pyramid Network (FPN) structure, which has the advantages of fast speed and the ability to integrate semantic information of different scales; the transformer network includes self-attention submodule, deformable cross-attention submodule and FFN submodule.
[0165] It should be noted that the convolutional neural network is used to extract the sample image features of each sample lane line image, and flatten and splice the sample image features of each sample lane line image to generate a second sample splicing feature; the 3D PE generator is used to process the sample image features of each sample lane line image, generate a sample 3D position code for each sample lane line image, and then flatten and splice the sample 3D position code of each sample lane line image to generate a sample splicing 3D position feature. The query vector generation module is used to generate the sample image features of each sample lane line image by splicing, fusing, stretching, and resizing to generate a first sample query vector. The sample splicing 3D position feature is used to update the first sample query vector according to the second sample splicing feature and the sample splicing 3D position feature to generate a second sample query vector; the 3D lane line detection head is used to process the second sample query vector to generate a predicted 3D coordinate for each sample lane line.
[0166] It should be understood that the transformer network is repeated N times, the output of the last layer is used as the second query vector, and the 3D lane line detection head is connected to generate the predicted 3D coordinates of each sample lane line.
[0167] exist Figure 6 On the basis of Figure 7 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 3 .
[0168] like Figure 7 As shown, S42 can be implemented by the following steps:
[0169] S71. Process multiple sample lane line images using an initial model to obtain predicted 3D coordinates of the sample lane lines generated by the initial model.
[0170] It should be understood that the method of generating the predicted 3D coordinates of the sample lane line is consistent with the method of generating the 3D coordinates of the lane line based on multiple lane line images during the model reasoning process, and will not be repeated here.
[0171] S72. Train the model parameters in the initial model according to the predicted 3D coordinates of the sample lane line and the corresponding sample 3D coordinates.
[0172] S73, projecting the predicted 3D coordinates of the sample lane line onto the corresponding sample lane line image to obtain the predicted 2D coordinates of each sample lane line;
[0173] Optionally, the predicted 2D coordinates of each sample lane line can be obtained by projecting the camera's internal and external parameter matrix to the image coordinates.
[0174] S74. According to the predicted 2D coordinates of each sample lane line and the corresponding sample 2D coordinates, the model parameters in the initial model are trained until the training cutoff condition is met to obtain a lane line detection model.
[0175] It should be understood that the error value of the current model for lane line detection can be obtained through the predicted coordinates (predicted 2D coordinates and predicted 3D coordinates) of each sample lane line and the corresponding sample coordinates (sample 2D coordinates and sample 3D coordinates), and then the relevant network parameters in the model are tuned to reduce the error and improve the accuracy of the model. In other words, the predicted 3D coordinates and the sample 3D coordinates are fed back for error, and on the basis of optimizing the model parameters, the predicted 2D coordinates and the sample 2D coordinates are fed back for error again, and the model parameters are further optimized, further improving the accuracy and robustness of the lane line detection model for predicting the lane line position.
[0176] It should be noted that in addition to calculating the error of the predicted coordinates of each sample lane line and the corresponding sample coordinates, it is also necessary to calculate the error of the lane line type and obstacle information output by the 3D lane line detection head. The specific loss function calculation process is:
[0177]
[0178] Where y=(c,p 3d ,p 2d ) is the sample lane line category, sample 3D coordinates, sample 2D coordinates, is the lane line category, 3D coordinate, and 2D coordinate predicted by the model, γ cls is a hyperparameter that balances different loss functions, L cls is the classification loss function, L reg is the regression loss function.
[0179] Optionally, the average accuracy value of multiple categories can be used as the model evaluation indicator. That is, when the average accuracy value is greater than the set threshold, the model training can be stopped, the current model parameters can be saved, and the lane line detection model can be obtained. The number of training times threshold or model parameter update threshold can also be set. The number of model parameter updates indicates the number of model training times. When the number of model training times reaches the set training times threshold, or the number of model parameter updates does not change with the increase of model training times, it means that the model training meets the training cutoff condition, the current model parameters are saved, and the lane line detection model is obtained.
[0180] It can be understood that by comparing the predicted coordinates of each sample lane line with the corresponding sample coordinates, the predicted lane line category and the lane line category of the sample, the model can train the parameters in the initial model according to the preset feedback until the preset training preset conditions are met. The lane line detection model is effectively optimized, the model prediction accuracy and robustness are improved, and the lane line detection model is finally obtained. It can not only identify lane lines more accurately, but also improve the application performance in complex environments, thereby enhancing the overall safety and reliability of the autonomous driving system.
[0181] It can be understood that by obtaining the 3D coordinates of the lane lines contained in multiple image acquisition devices output by the lane line detection model, the autonomous driving system can more accurately identify the road structure, thereby improving the perception of the surrounding environment, ensuring that the vehicle can follow the correct lane during driving, and improving driving safety and comfort.
[0182] Figure 8 Schematic diagram of the process of the lane line detection method provided by the embodiment of the present invention Figure 4 ;like Figure 8 As shown, the following steps are included:
[0183] S81. Acquire multiple sample lane line images in front of a sample vehicle and a sample label of each sample lane line image.
[0184] S82. Train the initial model according to the multiple sample lane line images and the label of each sample lane line image to obtain a lane line detection model.
[0185] S83. Collect multiple lane line images by installing multiple image acquisition devices on the vehicle, each of which has a different focal length, and input the multiple lane line images into a lane line detection model to obtain the 3D coordinates of the lane lines contained in the multiple lane line images output by the lane line detection model.
[0186] It should be noted that the specific implementation process is illustrated by way of example in the above embodiments, and the embodiments of the present invention will not be described in detail herein.
[0187] Fig. 9 A schematic diagram of the structure of a lane line detection device provided in an embodiment of the present invention; Fig. 9 As shown, the lane line detection device 90 includes:
[0188] The acquisition module 91 is used to acquire multiple lane line images by using multiple image acquisition devices installed on the vehicle, and each image acquisition device has a different focal length;
[0189] An input module 92, used to input multiple lane line images into a lane line detection model, and obtain the 3D coordinates of the lane line output by the lane line detection model, wherein the lane line detection model includes a query vector generation module, a 3D position code generation module, a transformer module, and a 3D lane line detection module;
[0190] When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line based on the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; the converter module integrates the image features and 3D position code of each lane line image into the first query vector to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line based on the second query vector.
[0191] Furthermore, the transformer module includes a self-attention submodule, a deformable cross-attention submodule, and a FFN submodule;
[0192] When the lane detection model generates the 3D coordinates of the lane, the self-attention submodule calculates the attention score between each element in the first query vector and other elements, and updates the elements according to the attention score to obtain the third query vector;
[0193] The deformable cross attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain the fourth query vector;
[0194] The FFN submodule updates the fourth query vector using a nonlinear activation function to obtain a second query vector.
[0195] Furthermore, the query vector generation module also includes a convolution submodule and an MLP submodule;
[0196] When the lane detection model generates the 3D coordinates of the lane, the convolution submodule splices and fuses the image features of each lane image to generate a first splicing feature;
[0197] The MLP submodule generates a first query vector according to the first concatenated feature.
[0198] Furthermore, the lane line detection model also includes a feature extraction module, which extracts image features of each lane line image, and flattens and splices the image features of each lane line image to generate a second splicing feature;
[0199] The 3D position code generation module processes the image features of each sample lane line image to generate a 3D position code for each sample lane line image, and flattens and splices the 3D position code for each sample lane line image to generate a spliced 3D position feature. The format of the second spliced feature is the same as the format of the spliced 3D position feature.
[0200] Furthermore, the deformable cross attention submodule incorporates the second splicing feature and the splicing 3D position feature into the third query vector to obtain a fourth query vector.
[0201] Furthermore, the lane line detection device 90 further includes an acquisition module, which is used to:
[0202] Get the lane line category and obstacle information output by the lane line detection model.
[0203] Furthermore, before collecting multiple lane line images by multiple image acquisition devices installed on the vehicle, the lane line detection device 90 also includes a training module for:
[0204] Acquire multiple sample lane line images in front of the sample vehicle and a sample label of each sample lane line image, wherein the sample label includes a sample 2D coordinate and a sample 3D coordinate of the sample lane line, and the multiple sample lane line images are acquired by a sample image acquisition device with different focal lengths installed in the sample vehicle;
[0205] The initial model is trained according to multiple sample lane line images and a label of each sample lane line image to obtain a lane line detection model.
[0206] Furthermore, the lane line detection device 90 also includes a training module, which is specifically used for:
[0207] Processing multiple sample lane line images through the initial model to obtain predicted 3D coordinates of the sample lane lines generated by the initial model;
[0208] The model parameters in the initial model are trained according to the predicted 3D coordinates of the sample lane lines and the corresponding sample 3D coordinates;
[0209] Project the predicted 3D coordinates of the sample lane line onto the corresponding sample lane line image to obtain the predicted 2D coordinates of each sample lane line;
[0210] According to the predicted 2D coordinates of each sample lane line and the corresponding sample 2D coordinates, the model parameters in the initial model are trained until the training cutoff condition is met to obtain the lane line detection model.
[0211] Furthermore, the sample label includes sample lane line category and sample obstacle information.
[0212] The lane line detection device provided in an embodiment of the present invention can be used to execute the lane line detection method in any of the above embodiments. Its implementation principle and technical effects are similar and will not be repeated here.
[0213] Fig.10 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Fig.10 As shown, the electronic device 10 may include: a processor 101, a memory 102, and computer-executable instructions stored in the memory 102 and executable on the processor 101. When the processor 101 executes the computer-executable instructions, the model training method and lane line detection method provided in any of the above embodiments are implemented. Optionally, the above-mentioned components of the electronic device 10 may be connected via a system bus.
[0214] The memory 102 may be a separate storage unit or a storage unit integrated in a processor. The number of processors may be one or more.
[0215] Optionally, the electronic device 10 may further include a communication interface for interacting with other devices.
[0216] The electronic device provided in an embodiment of the present invention can be used to execute the model training method and lane line detection method provided in any of the above method embodiments. The implementation principles and technical effects are similar and will not be repeated here.
[0217] Fig.11 The schematic diagram of the structure of the vehicle provided by the embodiment of the present invention is shown in FIG. Fig.11 As shown, the vehicle may include: a processor 111, a memory 112, a vehicle body 113, and computer-executable instructions stored in the memory 112 and executable on the processor 111. When the processor 111 executes the computer-executable instructions, the model training method and lane line detection method provided in any of the above embodiments are implemented. Optionally, the above-mentioned components of the vehicle may be connected via a system bus.
[0218] The memory 112 may be a separate storage unit or a storage unit integrated in a processor. The number of processors may be one or more.
[0219] Optionally, the vehicle may also include a communication interface for interacting with other devices.
[0220] The vehicle provided in an embodiment of the present invention can be used to execute the model training method and lane line detection method provided in any of the above method embodiments. The implementation principles and technical effects are similar and will not be repeated here.
[0221] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in the present invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0222] The system bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0223] All or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions. The above-mentioned program can be stored in a readable memory. When the program is executed, the steps of the above-mentioned method embodiments are executed. The above-mentioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc and any combination thereof.
[0224] An embodiment of the present invention provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed on a computer, the computer executes the above-mentioned model training method and lane line detection method.
[0225] The computer-readable storage medium mentioned above can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0226] Optionally, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0227] An embodiment of the present invention also provides a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, the above-mentioned model training method and lane line detection method can be implemented.
[0228] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A lane line detection method, characterized in that: The method comprises: A plurality of lane line images are collected by using a plurality of image acquisition devices installed on the vehicle, each of the image acquisition devices having a different focal length; Inputting the multiple lane line images into a lane line detection model, and obtaining 3D coordinates of the lane lines output by the lane line detection model, wherein the lane line detection model includes a query vector generation module, a 3D position code generation module, a transformer module, and a 3D lane line detection module; When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line according to the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; the converter module integrates the image features and 3D position code of each lane line image into the first query vector to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line according to the second query vector.
2. The method according to claim 1, characterized in that The transformer module includes a self-attention submodule, a deformable cross-attention submodule and a feed-forward network FFN submodule; When the lane line detection model generates the 3D coordinates of the lane line, the self-attention submodule calculates the attention score between each element in the first query vector and other elements, and updates the elements according to the attention score to obtain a third query vector; The deformable cross attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain a fourth query vector; The FFN submodule updates the fourth query vector using a nonlinear activation function to obtain the second query vector.
3. The method according to claim 1, characterized in that The query vector generation module also includes a convolution submodule and a multi-layer perceptron MLP submodule; When the lane line detection model generates the 3D coordinates of the lane line, the convolution submodule splices and fuses the image features of each lane line image to generate a first splicing feature; The MLP submodule generates a first query vector according to the concatenated features.
4. The method according to claim 2, characterized in that: The lane line detection model also includes a feature extraction module, which extracts image features of each lane line image, and flattens and splices the image features of each lane line image to generate a second splicing feature; When the lane line detection model generates the 3D coordinates of the lane line, the 3D position code generation module processes the image features of each sample lane line image to generate a 3D position code for each sample lane line image, and flattens and splices the 3D position code for each sample lane line image to generate a spliced 3D position feature, wherein the format of the second spliced feature is the same as the format of the spliced 3D position feature.
5. The method according to claim 4, characterized in that The deformable cross attention submodule incorporates the image features and 3D position encoding of each lane line image into the third query vector to obtain a fourth query vector, including: The deformable cross-attention submodule incorporates the second splicing feature and the splicing 3D position feature into the third query vector to obtain the fourth query vector.
6. The method according to any one of claims 1 or 5, characterized in that: The method further comprises: Obtain the lane line category and obstacle information output by the lane line detection model.
7. The method according to any one of claims 1 or 5, characterized in that: Before collecting a plurality of lane line images by using a plurality of image acquisition devices installed on the vehicle, the method further includes: Acquire a plurality of sample lane line images in front of a sample vehicle, and a sample label of each sample lane line image, wherein the sample label includes a sample 2D coordinate and a sample 3D coordinate of the sample lane line, and the plurality of sample lane line images are acquired by a sample image acquisition device with different focal lengths installed in the sample vehicle; The initial model is trained according to the multiple sample lane line images and the label of each sample lane line image to obtain the lane line detection model.
8. The method according to claim 7, characterized in that The training of the initial model according to the multiple sample lane line images and the label of each sample lane line image to obtain the lane line detection model includes: Processing the plurality of sample lane line images by using the initial model to obtain predicted 3D coordinates of the sample lane lines generated by the initial model; Training model parameters in the initial model according to the predicted 3D coordinates of the sample lane line and the corresponding sample 3D coordinates; Project the predicted 3D coordinates of the sample lane line onto the corresponding sample lane line image to obtain the predicted 2D coordinates of each sample lane line; According to the predicted 2D coordinates of each sample lane line and the corresponding sample 2D coordinates, the model parameters in the initial model are trained until a training cutoff condition is met, thereby obtaining the lane line detection model.
9. The method according to claim 7, characterized in that: The sample label includes a sample lane line category and sample obstacle information.
10. A lane line detection device, characterized in that: include: A collection module, used for collecting multiple lane line images by using multiple image collection devices installed on the vehicle, each image collection device having a different focal length; An input module, used to input the multiple lane line images into a lane line detection model to obtain the 3D coordinates of the lane lines output by the lane line detection model, wherein the lane line detection model includes a query vector generation module, a 3D position code generation module, a transformer module and a 3D lane line detection module; When the lane line detection model generates the 3D coordinates of the lane line, the query vector generation module generates a first query vector for the lane line according to the image features of each lane line image; the 3D position code generation module generates a 3D position code corresponding to each lane line image; The converter module integrates the image features and 3D position encoding of each lane line image into the first query vector to obtain a second query vector, and the 3D lane line detection module generates the 3D coordinates of the lane line according to the second query vector.
11. An electronic device, comprising: A processor, a memory, and a computer-executable instruction stored in the memory and executable on the processor, wherein the processor is used to implement the method according to any one of claims 1 to 7 when executing the computer-executable instruction.
12. A vehicle comprising: A vehicle body, a processor, a memory, and computer-executable instructions stored in the memory and executable on the processor, wherein the processor is used to implement the method according to any one of claims 1 to 7 when executing the computer-executable instructions.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it is used to implement the method according to any one of claims 1 to 7.