Lane detection method, electronic device, storage medium and computer program product
By using the Transformer model and dynamic convolution kernel in lane line detection, the problem of difficulty in dealing with complex lane lines in the existing technology is solved, and accurate lane line detection in complex scenarios is achieved.
Patent Information
- Application Number
- CN202310126302.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-02-08
AI Technical Summary
Existing lane line detection methods are difficult to deal with lane lines of complex topological structures, such as curves, forks, dense lanes, etc., as well as blocked lane lines.
The encoder and decoder modules in the Transformer model are used to convolve image features through the heat map dynamic convolution kernel and the offset map dynamic convolution kernel to capture global context information, thereby achieving accurate detection of complex lane lines.
In complex application scenarios, it can accurately detect lane lines, meet the needs of complex application scenarios, and avoid dependence on hand-designed components.
Smart Images

Figure CN116343147B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and more specifically to a lane line detection method, electronic device, storage medium and computer program product. Background Art
[0002] Lane detection is an important task in automated driving systems (ADS) and advanced driving assistant systems (ADAS), and is crucial for downstream tasks such as driving route planning, lane keeping, and adaptive cruise control.
[0003] Existing lane detection methods rely on deep neural networks and hand-crafted components (such as line anchors, parametric curves, predefined key points, segmentation clustering, etc.) to detect lane lines. Although these hand-crafted components can help the modeling process, they can only handle simpler straight line scenes and cannot handle lane lines with complex topological structures (such as curves, forks, dense lanes, etc.) and occluded lane lines. Therefore, a new lane detection solution is needed to solve the above technical problems. Summary of the invention
[0004] The present application is proposed in view of the above problems. The present application provides a lane line detection method, an electronic device, a storage medium and a computer program product.
[0005] According to one aspect of the present application, a lane line detection method is provided, including: acquiring an image to be processed; performing a lane line detection operation on the image to be processed to obtain a lane line detection result, wherein the lane line detection operation includes: performing feature extraction on the image to be processed to obtain initial image features; encoding the initial image features through an encoder module in a converter model to obtain encoded image features; comprehensively decoding the encoded image features and lane line query features through a decoder module in the converter model to obtain decoded image features, wherein the lane line query features include query features corresponding one-to-one to at least one lane line template; performing feature conversion on the decoded image features to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel, wherein the heat map dynamic convolution kernel includes a heat map feature corresponding one-to-one to at least one lane line template, and the offset map dynamic convolution kernel The kernel includes an offset feature corresponding to at least one lane line template one by one; the initial image feature or the encoded image feature is convolved with the heat map dynamic convolution kernel to obtain a heat map set, the heat map set includes at least one heat map corresponding to at least one lane line template one by one, each heat map is used to indicate the lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is the prediction probability that the local area is a lane point on the corresponding lane line template; the initial image feature or the encoded image feature is convolved with the offset map dynamic convolution kernel to obtain an offset map set, the offset map set includes at least one offset map corresponding to at least one lane line template one by one, each offset map is used to indicate the offset of each local area on the image to be processed relative to the corresponding lane line template; the lane line detection result is obtained based on the heat map set and the offset map set.
[0006] Exemplarily, the local area is a single pixel, the pixels in each heat map correspond one-to-one to the pixels in the image to be processed, and any pixel in each heat map is used to indicate the lane point prediction probability of the corresponding pixel in the image to be processed, the pixels in each offset map correspond one-to-one to the pixels in the image to be processed, and any pixel in each offset map is used to indicate the offset of the corresponding pixel in the image to be processed relative to the corresponding lane line template in the target direction, the target direction is the row direction and / or the column direction, and the lane line detection result is obtained based on the heat map set and the offset map set, including: for each lane line template in at least one lane line template, based on the heat map and offset map corresponding to the lane line template, voting for each pixel in the jth group of pixels in the image to be processed, and determining the voting score of each pixel in the jth group of pixels, the jth group of pixels is a group of pixels extending along the target direction, j=1,2,…,Num 1 , Num 1Represents the total number of groups of pixels extending along the target direction of the image to be processed; determines the pixel with the highest voting score in the j-th group of pixels as the group predicted lane point corresponding to the j-th group of pixels; for the image to be processed, based on the group predicted lane points corresponding to each group of pixels, determines the initial lane point set corresponding to the lane line template; based on the initial lane point set corresponding to at least one lane line template, determines the initial lane line corresponding to at least one lane line template; based on the initial lane line corresponding to at least one lane line template, determines at least one final lane line to obtain a lane line detection result.
[0007] Exemplarily, based on the heat map and the offset map corresponding to the lane line template, voting for each pixel in the j-th group of pixels in the processed image to determine the voting score of each pixel in the j-th group of pixels includes: calculating the position of the pixel prediction lane point corresponding to the k-th pixel based on the offset corresponding to the k-th pixel in the j-th group of pixels in the offset map corresponding to the lane line template, where k=1, 2, ..., Num 2 , Num 2 represents the total number of pixels in the j-th group of pixels; for any target pixel in the j-th group of pixels, determine the anchor pixel whose corresponding pixel predicted lane point position coincides with the target pixel; based on the heat map corresponding to the lane line template, determine the lane point prediction probability corresponding to each anchor pixel; based on the lane point prediction probability corresponding to each anchor pixel, determine the voting score corresponding to the target pixel.
[0008] Exemplarily, based on the lane point prediction probabilities corresponding to each anchor pixel, determining the voting score corresponding to the target pixel includes: adding the lane point prediction probabilities corresponding to each anchor pixel to obtain the voting score corresponding to the target pixel.
[0009] Exemplarily, before obtaining the lane line detection result based on the heat map set and the offset map set, the method also includes: performing classification based on the decoded image features to obtain a lane line range set, the lane line range set including at least one set of lane line range information corresponding one-to-one to at least one lane line template, each set of lane line range information being used to indicate the coverage range of the corresponding lane line template on the image to be processed in a direction perpendicular to the target direction; based on the initial lane point set corresponding to each of the at least one lane line templates, determining the initial lane line corresponding to each of the at least one lane line templates, including: for each lane line template in the at least one lane line template, selecting a group of predicted lane points from the corresponding initial lane point set that are within the coverage range corresponding to the lane line template; and composing the initial lane line corresponding to the lane line template based on the selected group of predicted lane points.
[0010] Exemplarily, before obtaining the lane line detection result based on the heat map set and the offset map set, the method also includes: performing classification based on the decoded image features to obtain a lane line score set, the lane line score set including at least one group of lane line scores corresponding one-to-one to at least one lane line template, each group of lane line scores being used to indicate the lane line prediction probability, the lane line prediction probability being the prediction probability of the corresponding lane line template being included in the image to be processed; determining at least one final lane line based on the initial lane line corresponding to each of the at least one lane line templates to obtain the lane line detection result, including: selecting the initial lane line corresponding to the lane line template whose lane line prediction probability is higher than the target probability threshold as at least one final lane line to obtain the lane line detection result.
[0011] Exemplarily, the target direction is the row direction and the column direction. When the target direction is the row direction, at least one lane line template included in the lane line query feature is at least one first lane line template, and the at least one final lane line determined is at least one first final lane line. When the target direction is the column direction, at least one lane line template included in the lane line query feature is at least one second lane line template, and the at least one final lane line determined is at least one second final lane line. Before obtaining the lane line detection result based on the heat map set and the offset map set, the method also includes: performing classification based on the decoded image features to obtain a lane line score set, the lane line score set including at least one group of lane line scores corresponding to at least one lane line template one by one, each group of lane line scores being used to indicate the lane line prediction probability, and the lane line prediction probability is the prediction probability that the corresponding lane line template is included in the image to be processed; based on the heat map set and the offset map set, The lane line detection result is obtained by shifting the map set, and also includes: determining a first comprehensive score of at least one first final lane line based on the lane line score corresponding to at least one first final lane line, and the lane line score corresponding to each first final lane line is the lane line score corresponding to the first lane line template corresponding to the first final lane line; determining a second comprehensive score of at least one second final lane line based on the lane line score corresponding to at least one second final lane line, and the lane line score corresponding to each second final lane line is the lane line score corresponding to the second lane line template corresponding to the second final lane line; comparing the first comprehensive score with the second comprehensive score, if the first comprehensive score is greater than the second comprehensive score, selecting at least one first final lane line, and if the second comprehensive score is greater than the first comprehensive score, selecting at least one second final lane line; determining the selected final lane line as the new lane line detection result.
[0012] Exemplarily, the lane line detection operation is performed through a lane line detection model, and the method also includes: obtaining a sample image and corresponding annotation information, the annotation information is used to indicate the position of at least one real lane line in the sample image; performing a lane line detection operation on the sample image to obtain a lane line prediction result, wherein the lane line prediction result includes at least one predicted lane line obtained for the sample image detection; based on at least one real lane line and at least one predicted lane line, determining an optimal single-shot function using a bipartite matching loss algorithm; calculating the prediction loss of the lane line detection model based on the optimal single-shot function; and optimizing the parameters in the lane line detection model based on the prediction loss.
[0013] Exemplarily, optimizing parameters in the lane detection model based on the prediction loss includes optimizing the parameters in the lane detection model and the lane query features together based on the prediction loss.
[0014] According to another aspect of the present application, an electronic device is also provided, including a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used to execute the above-mentioned lane line detection method when the processor is running.
[0015] According to another aspect of the present application, a storage medium is provided, on which program instructions are stored, wherein the program instructions are used to execute the above-mentioned lane line detection method when running.
[0016] According to yet another aspect of the present application, a computer program product is provided, the computer program product comprising a computer program, wherein the computer program is used to execute the above lane line detection method when running.
[0017] According to the lane line detection method, electronic device, storage medium and computer program product of the embodiment of the present application, the encoder module and decoder module of the Transformer are used, and the lane line query feature is used to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel containing features corresponding to each lane line template and with global context information. Subsequently, the heat map dynamic convolution kernel can be used to convolve the initial image features or the encoded image features to obtain a heat map set, and the offset map dynamic convolution kernel can be used to convolve the initial image features or the encoded image features to obtain an offset map set, and then the lane line detection result is obtained based on the heat map set and the offset map set. Since the above method uses the heat map dynamic convolution kernel and the offset map dynamic convolution kernel for convolution, it can effectively capture global context information, so it can obtain accurate lane line detection results in various complex application scenarios such as curves, forks, dense lanes, and lane lines are blocked, meeting the needs of complex application scenarios. In addition, the dynamic convolution kernel can approximate various types of complex lane lines, so the use of the above dynamic convolution kernel can also avoid dependence on manually designed components. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0019] Figure 1 A schematic block diagram showing an example electronic device for implementing the lane line detection method and apparatus according to an embodiment of the present application;
[0020] Figure 2 A schematic flow chart of a lane line detection method according to an embodiment of the present application is shown;
[0021] Figure 3 A schematic diagram showing a lane detection model according to an embodiment of the present application is shown;
[0022] Figure 4 The heat map B corresponding to the i-th lane line template according to one embodiment of the present application is shown. i With offset map Z i Schematic diagram of relevant information;
[0023] Figure 5 A schematic diagram showing an attention map according to an embodiment of the present application;
[0024] Figure 6 A schematic block diagram showing a lane line detection device according to an embodiment of the present application; and
[0025] Figure 7 A schematic block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0026] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, image processing, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.
[0027] In order to make the purpose, technical solutions and advantages of the present application more obvious, the example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.
[0028] The embodiments of the present application provide a lane line detection method, electronic device, storage medium and computer program product. According to the lane line detection method of the embodiments of the present application, the position of the lane line can be accurately detected in complex application scenarios (curves, forks, dense lanes, lane lines are blocked, etc.). The lane line detection technology according to the embodiments of the present application can be applied to any field involving lane line detection.
[0029] First, refer to Figure 1 An example electronic device 100 for implementing the lane line detection method and apparatus according to an embodiment of the present application is described.
[0030] like Figure 1 As shown, the electronic device 100 includes one or more processors 102 and one or more storage devices 104. Optionally, the electronic device 100 may also include an input device 106, an output device 108, and an image acquisition device 110, and these components are interconnected via a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structures of the electronic device 100 shown are merely exemplary and non-limiting. The electronic device may also have other components and structures as required.
[0031] The processor 102 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic array (PLA), and a microprocessor. The processor 102 can be a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), or one or a combination of other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 100 to perform desired functions.
[0032] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may run the program instructions to implement the client functions (implemented by the processor) in the embodiments of the present application described below and / or other desired functions. Various applications and various data may also be stored in the computer-readable storage medium, such as various data used and / or generated by the application, etc.
[0033] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.
[0034] The output device 108 can output various information (such as images and / or sounds) to the outside (such as a user), and can include one or more of a display, a speaker, etc. Optionally, the input device 106 and the output device 108 can be integrated together and implemented using the same interactive device (such as a touch screen).
[0035] The image acquisition device 110 can acquire images and store the acquired images in the storage device 104 for use by other components. The image acquisition device 110 can be a separate camera or a camera in a mobile terminal, etc. It should be understood that the image acquisition device 110 is only an example, and the electronic device 100 may not include the image acquisition device 110. In this case, other devices with image acquisition capabilities can be used to acquire images and send the acquired images to the electronic device 100.
[0036] Exemplarily, the example electronic device for implementing the lane line detection method and apparatus according to the embodiments of the present application can be implemented on a device such as a personal computer, a terminal device, an attendance machine, a panel machine, a camera or a remote server, etc. The terminal device includes but is not limited to: a tablet computer, a mobile phone, a PDA (Personal Digital Assistant), a touch-screen all-in-one machine, a wearable device, etc.
[0037] Next, we will refer to Figure 2 A lane line detection method according to an embodiment of the present application is described. Figure 2 FIG. 2 is a schematic flow chart of a lane line detection method 200 according to an embodiment of the present application. Figure 2 As shown, the lane line detection method 200 includes the following steps S210 and S220.
[0038] Step S210, obtaining an image to be processed.
[0039] The image to be processed may be a road image, which may be a road image acquired by an image acquisition device disposed on an object such as a moving vehicle, a road or a building. In an embodiment of the present application, the image to be processed may contain at least lane lines. The image to be processed may be an original image acquired by an image acquisition device (such as the above-mentioned image acquisition device 110), or an image obtained after preprocessing the original image acquired by the image acquisition device. Preprocessing may include normalization, scaling, smoothing and other processing. Preprocessing may also include an operation of extracting a partial image area containing lane lines from the original image acquired by the image acquisition device to obtain the image to be processed.
[0040] The image to be processed may come from an external device and be transmitted by the external device to the electronic device 100 for lane line detection. In addition, the image to be processed may also be acquired by the electronic device 100 itself. For example, the electronic device 100 may use an image acquisition device 110 (e.g., an independent camera) to acquire the image to be processed. The image acquisition device 110 may transmit the acquired image to be processed to the processor 102, and the processor 102 may perform lane line detection.
[0041] Step S220, performing a lane line detection operation on the image to be processed to obtain a lane line detection result, wherein the lane line detection operation may include the following steps S221, S222, S223, S224, S225, S226 and S227. The lane line detection operation of the embodiment of the present application can be implemented based on a lane line detection model. Exemplarily, the lane line detection model may include a feature extraction module, a dynamic convolution kernel module and a lane detection module. The dynamic convolution kernel module may include an encoder module, a decoder module and a feature conversion module. The encoder module and the decoder module can be implemented using the encoder and decoder in the transformer model. Figure 3 A schematic diagram of a lane detection model according to an embodiment of the present application is shown below. Figure 3 Provide explanation.
[0042] Step S221, extracting features from the image to be processed to obtain initial image features.
[0043] For example, the image to be processed X can be expressed as Among them, H 0 , W 0 and C 0 Respectively represent the height, width and number of channels of the image to be processed X. For example, when the image to be processed X is an RGB image, the number of channels C 0 Can be 3. Figure 3 , the image to be processed X can be input into the feature extraction module, and the feature extraction module can be used to extract features of the image to be processed to obtain the initial image features F, F∈R H×W×C Wherein H, W and C represent the height, width and number of channels of the initial image feature F, respectively. By way of example and not limitation, the feature extraction module can be implemented using a Convolutional Neural Networks Backbone (CNN backbone).
[0044] Step S222, encoding the initial image features through the encoder module in the converter model to obtain encoded image features.
[0045] Exemplarily, according to the obtained initial image feature F, after encoding by the encoder module in the Transformer model, the encoded image feature M can be obtained. In the encoder module, a self-attention mechanism operation can be performed to calculate the correlation between the feature vector at each position in the initial image feature and the feature vector at each position in the entire initial image feature. Therefore, the encoded image feature obtained by the encoder module can be a feature fused with global context information. Each position in the above initial image feature can correspond to a local area on the image to be processed, and the local area can be an area containing a single pixel or multiple pixels. The above correlation can be similarity.
[0046] Step S223, comprehensively decoding the encoded image features and the lane line query features through the decoder module in the converter model to obtain decoded image features, wherein the lane line query features may include query features that correspond one-to-one to at least one lane line template.
[0047] For example, refer to Figure 3 , for the decoder module in the Transformer model, the encoded image feature M and the lane line query feature S can be used as the input of the decoder module to obtain its corresponding decoded image feature T. In one embodiment, the lane line query feature S can be a predefined feature sequence. The lane line query feature can be represented as S∈R L ×C , which consists of L learnable feature vectors of length C. The L learnable feature vectors of length C are the query features corresponding to the L lane templates. The query features can also be called query feature vectors. Where L represents the number of lane templates, and C represents the number of channels of the lane query feature S, which is the same as the C value of the initial image feature F in the previous embodiment. For the decoded image feature T, there exists T∈R L×C . That is, the decoded image feature T may also include L feature vectors of length C corresponding to the L lane line templates one by one, and each feature vector may be referred to as a decoded feature vector. In the decoder module, a cross-attention mechanism operation may be performed between the encoded image feature M and the lane line query feature S to calculate the correlation between the feature vector at each position in the encoded image feature M and the feature vector at each position in the lane line query feature S (i.e., the feature vector of length C corresponding to each lane line template). Therefore, the decoded image feature obtained by the decoder module may include global context information corresponding to each lane line template. Through the global context information corresponding to each lane line template, the lane lines corresponding to each lane line template may be predicted. Exemplarily, the lane line query feature S may be obtained by pre-training, wherein each lane line template represents a template of a lane line of different modes, and the above modes may include position and / or shape, etc.
[0048] Step S224, perform feature conversion on the decoded image features to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel, the heat map dynamic convolution kernel includes a heat map feature corresponding one-to-one to at least one lane line template, and the offset map dynamic convolution kernel includes an offset feature corresponding one-to-one to at least one lane line template.
[0049] By way of example and not limitation, reference is made to Figure 3 , the multilayer perceptron (MLP) can be used to transform the decoded image features T to obtain the corresponding heat map dynamic convolution kernel K b and offset map dynamic convolution kernel K z . Heatmap dynamic convolution kernel K b and offset map dynamic convolution kernel K z Each contains its own corresponding feature information. Heat map dynamic convolution kernel K b Includes heat map features that correspond to different lane templates. For example, there may be K b ∈R L×C , represents the heat map dynamic convolution kernel K b It includes a feature vector of length C (i.e., heat map feature) corresponding to each of the L lane templates. z Includes offset map features that correspond one-to-one to different lane templates. Similarly, there may be K z ∈R L×C , represents the dynamic convolution kernel K of the offset map z It includes a feature vector (i.e., offset map feature) of length C that corresponds one-to-one to the L lane line templates.
[0050] It can be understood that since the decoded image feature T contains the global context information corresponding to each lane line template, the converted heat map dynamic convolution kernel K b and offset map dynamic convolution kernel K z The global context information corresponding to each lane line template is also included in the heat map dynamic convolution kernel K b and offset map dynamic convolution kernel K z It can be understood as a new lane line query feature that is more suitable for the current image to be processed by integrating the image features of the current image to be processed based on the initial lane line query feature S. That is, different from the lane line query feature S shared by all images, the heat map dynamic convolution kernel K b and offset map dynamic convolution kernel K z Contains specific image information related to the current image to be processed. It is worth noting that S, K b and K zIt can approximate various types of complex lane lines, thus eliminating the need for hand-designed components. b and offset map dynamic convolution kernel K z By convolving the initial image feature F or the encoded image feature M as the convolution kernel respectively, the correlation between the global context information corresponding to each lane line template and each position in the initial image feature F or the encoded image feature M can be calculated, so as to further predict whether the image to be processed has the lane lines corresponding to each lane line template.
[0051] Step S225, convolve the initial image features or the encoded image features using a heat map dynamic convolution kernel to obtain a heat map set, where the heat map set includes at least one heat map corresponding to at least one lane line template, and each heat map is used to indicate the lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is the prediction probability that the local area is a lane point on the corresponding lane line template.
[0052] For example, the heat map dynamic convolution kernel K can be used b Convolution (i.e., dynamic convolution) is performed on the initial image feature F or the encoded image feature M to obtain a corresponding heat map set B. The heat map set B may include L heat maps corresponding to the L lane line templates. In one embodiment, upsampling may be performed after convolution to obtain a heat map of the same size as the image to be processed X. For example, when the encoded image feature M is a feature sequence of dimension HW×C, if a dynamic convolution kernel K is used, b To convolve the coded image feature M, the coded image feature M can be reshaped into M′ first, and then the reshaped coded image feature M′ is convolved to obtain the corresponding heat map set. For the reshaped coded image feature M′, M′∈R H×W×C. Each heat map is used to indicate the lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is the prediction probability that the local area is a lane point on the corresponding lane line template. Among them, each local area can contain multiple pixels, in which case the size of the heat map is smaller than the image to be processed X. For example, the size of the heat map set B can be L×H×W, indicating that it contains L heat maps of size H×W, and the size of each heat map is consistent with the size of each feature map in the initial image feature F or the encoded image feature M. Each heat map contains H×W pixels, which correspond one-to-one to the H×W local areas on the image to be processed X. The pixel value of each pixel on the heat map can indicate the prediction probability that the corresponding local area on the image to be processed X is a lane point on the corresponding lane line template. In addition, each local area can be a single pixel, in which case the heat map is the same size as the image to be processed X. For example, the initial heat map set B of size L×H×W can be upsampled by, for example, interpolation to obtain a heat map of size L×H 0 ×W 0 The new heat map set B contains L heat maps of size H 0 ×W 0 Each heat map contains H 0 ×W 0 pixels, and H on the image to be processed X 0 ×W 0 The pixels correspond one to one, and the pixel value of each pixel on the heat map can represent the predicted probability that the corresponding pixel on the image to be processed X is the lane point on the corresponding lane line template.
[0053] Step S226, convolve the initial image features or the encoded image features using the offset map dynamic convolution kernel to obtain an offset map set, where the offset map set includes at least one offset map corresponding to at least one lane line template, and each offset map is used to indicate the offset of each local area on the image to be processed relative to the corresponding lane line template.
[0054] Exemplarily, the implementation of step S226 is similar to that of step S225. With reference to the above description of "using the dynamic convolution kernel of the heat map to convolve the initial image features or the encoded image features to obtain a heat map set", the implementation of this step can be understood. For the sake of brevity, it will not be repeated here. Each offset map can be used to indicate the offset of any pixel (or pixel point) in each local area on the image to be processed relative to the lane point in the same row or column as the pixel on the lane line template. For the same local area, the offsets corresponding to the pixels inside it are equal, that is, the same offset value can be predicted for the same local area. The local area here is similar to the description above, and may contain multiple pixels or only a single pixel.
[0055] Step S227, obtaining lane line detection results based on the heat map set and the offset map set.
[0056] Exemplarily, a lane line detection result can be obtained based on the acquired heat map set B and offset map set Z. The lane line detection result may include N predicted lane lines corresponding to N lane line templates obtained for the image to be processed. The number N of predicted lane lines obtained may be arbitrary, and it may depend on the actual detection situation. For example, N may be a value of 0 or greater than 0, which is at most equal to the number of at least one lane line template included in the lane line query feature S.
[0057] According to the lane line detection method provided in the embodiment of the present application, the encoder module and the decoder module of the Transformer are used, and the lane line query features are used to obtain the heat map dynamic convolution kernel and the offset map dynamic convolution kernel containing the features corresponding to each lane line template and with global context information. Subsequently, the heat map dynamic convolution kernel can be used to convolve the initial image features or the encoded image features to obtain a heat map set, and the offset map dynamic convolution kernel can be used to convolve the initial image features or the encoded image features to obtain an offset map set, and then the lane line detection result is obtained based on the heat map set and the offset map set. Since the above method uses the heat map dynamic convolution kernel and the offset map dynamic convolution kernel for convolution, it can effectively capture the global context information, so it can obtain accurate lane line detection results in various complex application scenarios such as curves, forks, dense lanes, and lane lines are blocked, meeting the needs of complex application scenarios. In addition, the dynamic convolution kernel can approximate various types of complex lane lines, so the use of the above dynamic convolution kernel can also avoid dependence on manually designed components.
[0058] Exemplarily, the lane line detection method according to the embodiment of the present application can be implemented in a device, apparatus or system having a memory and a processor.
[0059] The lane line detection method according to the embodiment of the present application can be deployed at an image acquisition end, for example, can be deployed at a personal terminal or a server end with an image acquisition function.
[0060] Alternatively, the lane line detection method according to the embodiment of the present application can also be deployed in a distributed manner on the server side (or cloud) and the personal terminal. For example, an image can be acquired on the client side, and the client side transmits the acquired image to the server side (or cloud), and the server side (or cloud) performs lane line detection.
[0061] Exemplarily, the local area is a single pixel, the pixels in each heat map correspond one-to-one to the pixels in the image to be processed, and any pixel in each heat map is used to indicate the lane point prediction probability of the corresponding pixel in the image to be processed, the pixels in each offset map correspond one-to-one to the pixels in the image to be processed, and any pixel in each offset map is used to indicate the offset of the corresponding pixel in the image to be processed relative to the corresponding lane line template in the target direction, the target direction is the row direction and / or the column direction, and the lane line detection result is obtained based on the heat map set and the offset map set, which may include: for each lane line template in at least one lane line template, based on the heat map and offset map corresponding to the lane line template, voting for each pixel in the jth group of pixels in the image to be processed, and determining the voting score of each pixel in the jth group of pixels, the jth group of pixels is a group of pixels extending along the target direction, j=1,2,…,Num 1 , Num 1 Represents the total number of groups of pixels extending along the target direction of the image to be processed; determines the pixel with the highest voting score in the j-th group of pixels as the group predicted lane point corresponding to the j-th group of pixels; for the image to be processed, based on the group predicted lane points corresponding to each group of pixels, determines the initial lane point set corresponding to the lane line template; based on the initial lane point set corresponding to at least one lane line template, determines the initial lane line corresponding to at least one lane line template; based on the initial lane line corresponding to at least one lane line template, determines at least one final lane line to obtain a lane line detection result.
[0062] In one embodiment, the local area contains only a single pixel. Each heat map, each offset map and the size of the image to be processed are the same, that is, the pixels in each heat map or each offset map correspond one-to-one to the pixels in the image to be processed. Any pixel in each heat map can represent the predicted probability that the corresponding pixel in the image to be processed is a lane point on the corresponding lane line template (i.e., the lane point prediction probability). Any pixel in each offset map can represent the offset of the corresponding pixel in the image to be processed relative to the corresponding lane line template in the target direction. The target direction can be a row direction and / or a column direction. For ease of description and understanding, the following mainly takes the target direction as the row direction as an example for explanation. It can be understood that the method for determining the final lane line when the target direction is the column direction is similar to the method for determining the final lane line when the target direction is the row direction, and the implementation scheme of the former can be understood with reference to the latter.
[0063] For at least one lane template, an initialized counting matrix W can be used. Count H on the image to be processed 0 ×W 0 The voting score of each pixel in pixels. The vote matrix W can be regarded as L pixels of size H 0 ×W 0The element at position (h,w) in the i-th two-dimensional matrix represents the heat map B corresponding to the i-th lane line template. i and offset map Z i Voting is performed to obtain the voting score of the pixel at the position (h, w) on the image to be processed. Where h = 1, 2, 3, ..., H 0 ; w=1,2,3,…,W 0 . For example, the vote counting matrix W can be initialized to a zero matrix. For the jth group of pixels, if the group of pixels extends along the row direction (i.e., arranged in a row), then it is the jth row of pixels; if the group of pixels extends along the column direction (i.e., arranged in a column), then it is the jth column of pixels. The following takes the jth group of pixels representing the jth row of pixels as an example. At this time, Num 1 Indicates the total number of groups of pixels extending along the row direction of the image to be processed. For example, if the image to be processed has a total of 1000 rows of pixels, then Num 1 is 1000. According to the voting score statistics of each pixel in the j-th row of pixels, the pixel with the highest voting score is determined as the group predicted lane point corresponding to the j-th row of pixels. For each row of pixels, one of the pixels can be optionally selected as the group predicted lane point.
[0064] For the image to be processed, based on the group predicted lane points corresponding to each of the 1000 rows of pixels, these 1000 group predicted lane points can be determined as the initial lane point set O corresponding to the lane line template. Based on the initial lane point sets corresponding to different lane line templates, the lane points in any initial lane point set are connected to obtain the corresponding initial lane lines. Thus, L lane line templates can obtain L initial lane lines. In addition, optionally, any initial lane point set can also be filtered, for example, based on the following lane line range information, the lane points within the coverage range corresponding to the lane line template are used as lane points on the initial lane line, and other lane points are discarded, thereby obtaining the initial lane lines corresponding to each of the L lane line templates.
[0065] In addition, at least one final lane line can also be determined based on the initial lane line corresponding to each of the at least one lane line template. In one example, the initial lane line corresponding to each of the L lane line templates can be directly determined as the final lane line to obtain a lane line detection result. In another example, the initial lane lines can also be screened, and the initial lane lines that meet the requirements (for example, the following lane line prediction probability is higher than the target probability threshold) are selected as the final lane lines to obtain the lane line detection result. It can be understood that the lane line detection result may include at least one final lane line determined.
[0066] According to the above technical solution, based on the heat map and offset map corresponding to the lane line template, the group predicted lane point corresponding to the jth group of pixels in the image to be processed is determined by voting for each pixel in the jth group of pixels. Then, based on the group predicted lane points corresponding to each group of pixels, the lane line detection result is obtained. This method can improve the accuracy of the detected group predicted lane points, thereby improving the accuracy of the lane line detection result.
[0067] Exemplarily, based on the heat map and the offset map corresponding to the lane line template, voting for each pixel in the j-th group of pixels in the processed image to determine the voting score of each pixel in the j-th group of pixels may include: calculating the position of the pixel prediction lane point corresponding to the k-th pixel based on the offset corresponding to the k-th pixel in the j-th group of pixels in the offset map corresponding to the lane line template, where k=1, 2, ..., Num 2 , Num 2 represents the total number of pixels in the j-th group of pixels; for any target pixel in the j-th group of pixels, determine the anchor pixel whose corresponding pixel predicted lane point position coincides with the target pixel; based on the heat map corresponding to the lane line template, determine the lane point prediction probability corresponding to each anchor pixel; based on the lane point prediction probability corresponding to each anchor pixel, determine the voting score corresponding to the target pixel.
[0068] Num 2 It can be understood that when the jth group of pixels is the jth row of pixels, Num 2 It can be the number of columns of the image to be processed. Conversely, when the jth group of pixels is the jth column of pixels, Num 2 Can be the number of rows of the image to be processed.
[0069] Figure 4 The heat map B corresponding to the i-th lane line template according to one embodiment of the present application is shown. i With offset map Z i Schematic diagram of related information. Figure 4 The region 410 shown is based on the heat map B i According to the embodiment of this invention, the part belonging to the lane line can be regarded as the foreground, and the part not belonging to the lane line can be regarded as the background. For any pixel on the image to be processed, at least based on the heat map B iThe corresponding lane point prediction probability on determines whether the pixel is a lane point, that is, whether it belongs to the lane line. Exemplarily, if the lane point prediction probability corresponding to any pixel is greater than the target probability threshold (which can be called the first target probability threshold), it can be determined that the pixel belongs to the foreground, otherwise it is determined that the pixel belongs to the background. The first target probability threshold can be set to any suitable size as needed. It can be any preset value between 0 and 1. For example, the first target probability threshold can be equal to 0.8. In addition, it is also possible to optionally determine that pixels within the coverage range corresponding to the lane line template belong to lane points based on the following lane line range, and others are deemed not to belong to lane points. In the case where the above-mentioned target direction is the row direction, the coverage range can be from the starting row R i0 To the end of row R i1 Therefore, we can determine the range of Figure 4 A foreground area 410 is shown.
[0070] For the offset map Z corresponding to the i-th lane template i The kth pixel in the jth row of pixels X jk , the offset Z corresponding to the pixel ijk By adding it to its column coordinate k, we can obtain the horizontal coordinate n of the pixel predicted lane point corresponding to the pixel. Figure 4 P* shown ij Represents the group of predicted lane points corresponding to the j-th row of pixels for the i-th lane line template.
[0071] The pixel whose position of the corresponding pixel predicted lane point coincides with the target pixel can be used as the anchor pixel. For example, assuming that the pixel predicted lane point calculated based on the pixel in the 9th row and 7th column is the pixel in the 9th row and 10th column, then when the pixel in the 9th row and 10th column is used as the target pixel, its corresponding anchor pixel includes the pixel in the 9th row and 7th column. For any target pixel in the jth row of pixels, such as the pixel in the 9th row and 4th column, assume that the anchor pixel corresponding to the pixel is the pixel in the 2nd column, 5th column and 6th column in the 9th row. Then the pixels in the 2nd column, 5th column and 6th column can be used as the anchor pixels of the current target pixel. Then, the target pixel can be voted for by the lane point prediction probability corresponding to the above anchor pixel on the heat map to obtain the voting score corresponding to the target pixel.
[0072] If the pixel is determined to be a lane point simply based on the lane point prediction probability corresponding to any pixel on the heat map, the error will be relatively large, because there may be multiple pixels in the same group of pixels with the same probability. According to the above technical solution, the anchor pixels in the same group (same row or same column) of each target pixel can be determined based on the offset map corresponding to any lane line template, and then the target pixel is voted on whether it belongs to a lane point based on the lane point prediction probability corresponding to each anchor pixel, and the voting score corresponding to the target pixel is determined. This method can improve the accuracy of the determined lane point because the target pixel is voted based on the information of each anchor pixel.
[0073] Exemplarily, determining the voting score corresponding to the target pixel based on the lane point prediction probabilities corresponding to each anchor pixel may include: adding the lane point prediction probabilities corresponding to each anchor pixel to obtain the voting score corresponding to the target pixel.
[0074] In one embodiment, according to the heat map B corresponding to the current lane line template i , the lane point prediction probabilities corresponding to the anchor pixels (e.g., the pixels corresponding to the 2nd, 5th, and 6th columns of the 9th row, respectively) can be added together. By formula: i92 +B i95 +B i96 , the lane point prediction probabilities corresponding to the three anchor pixels are added together to obtain the voting score corresponding to the target pixel (the 4th pixel in the 9th row). Of course, direct addition (i.e., summing) is only an example, and other methods can also be used to calculate the voting score, such as averaging the lane point prediction probabilities corresponding to the anchor pixels. Optionally, the above-mentioned summing or averaging of the lane point prediction probabilities corresponding to the anchor pixels can be a weighted summation or weighted averaging method based on weights.
[0075] The predicted probabilities of lane points corresponding to the anchor pixels are added together to obtain the voting score corresponding to the target pixel. This algorithm is relatively simple and fast.
[0076] Exemplarily, before obtaining the lane line detection result based on the heat map set and the offset map set, the method may also include: performing classification based on the decoded image features to obtain a lane line range set, the lane line range set including at least one set of lane line range information corresponding to at least one lane line template, each set of lane line range information being used to indicate the coverage range of the corresponding lane line template on the image to be processed in a direction perpendicular to the target direction; determining the initial lane line corresponding to at least one lane line template based on the initial lane point set corresponding to each of the at least one lane line templates, which may include: for each lane line template in the at least one lane line template, selecting a group of predicted lane points from the corresponding initial lane point set that are within the coverage range corresponding to the lane line template; and composing the initial lane line corresponding to the lane line template based on the selected group of predicted lane points.
[0077] In one embodiment, the decoded image features T can be classified using MLP (e.g. Figure 3 As shown), to obtain the lane line range set RG ( Figure 3 denoted as “R” in the figure). For example, RG∈R L×2 . The lane line range set RG may include at least one set of lane line range information corresponding to different lane line templates. Exemplarily, the lane line range information corresponding to any lane line template may indicate the coverage of the lane line template in a direction perpendicular to the target direction. For example, when the target direction is the row direction, a set of lane line range information may include two data, which are respectively used to represent the row number of the starting row and the row number of the ending row corresponding to the lane line template on the image to be processed X, and the range covered between the starting row and the ending row is the coverage of the lane line template in the column direction. For example, the coverage of the current lane line template a in the column direction is from the 100th row to the 500th row. It can be understood that when the target direction is the column direction, a set of lane line range information may include two data, which are respectively used to represent the column number of the starting column and the column number of the ending column corresponding to the lane line template on the image to be processed X.
[0078] For each lane line template in different lane line templates, for example, lane line template a, a group of predicted lane points located between the 100th and 500th rows are selected from its corresponding initial lane point set, and the line connected by this part of the group of predicted lane points can be used as the initial lane line corresponding to the lane line template a.
[0079] According to the above technical solution, a set of predicted lane points is selected based on the lane line range information corresponding to the lane line template to determine the initial lane line corresponding to the lane line template. In this way, the initial lane point set can be preliminarily screened to ensure the accuracy of the determined initial lane line.
[0080] Exemplarily, before obtaining the lane line detection result based on the heat map set and the offset map set, the method also includes: performing classification based on the decoded image features to obtain a lane line score set, the lane line score set including at least one group of lane line scores corresponding one-to-one to at least one lane line template, each group of lane line scores being used to indicate the lane line prediction probability, the lane line prediction probability being the prediction probability of the corresponding lane line template being included in the image to be processed; determining at least one final lane line based on the initial lane line corresponding to each of the at least one lane line templates to obtain the lane line detection result, which may include: selecting the initial lane line corresponding to the lane line template whose lane line prediction probability is higher than the target probability threshold as at least one final lane line to obtain the lane line detection result.
[0081] In one embodiment, the decoded image feature T can be classified using MLP to obtain a lane score set C. For example, C∈R L×2 . The lane line score set C includes at least one set of lane line scores corresponding to different lane line templates. Optionally, the lane line score can be represented by two probability values. These two probability values can represent the foreground probability and the background probability, respectively. Among them, the foreground indicates that the image to be processed contains lane lines, and the background indicates that the image to be processed does not contain lane lines. It can be understood that the lane line score can also be represented by a probability value, for example, the lane line score is represented by the probability value of the foreground. Because the background probability = 1-foreground probability. The lane line score can be used to determine whether a lane line is hit.
[0082] The lane line prediction probability is the foreground probability. In the case where the lane line score is expressed as a probability value, the initial lane line corresponding to the lane line template whose lane line prediction probability is higher than the target probability threshold (which can be called the second target probability threshold) can be selected as the final lane line. The second target probability threshold can be set to any suitable value as needed, which can be any pre-set value between 0 and 1. For example, the second target probability threshold can be equal to 0.6. Based on the determined final lane line, it can be used as the lane line detection result.
[0083] According to the above technical solution, through the lane line score set, the initial lane line corresponding to the lane line template whose lane line prediction probability is higher than the target probability threshold is selected as at least one final lane line. In this way, the final lane line can be effectively selected to improve the accuracy of the obtained lane line detection result.
[0084] Exemplarily, the target direction is the row direction and the column direction. When the target direction is the row direction, at least one lane line template included in the lane line query feature is at least one first lane line template, and the at least one final lane line determined is at least one first final lane line. When the target direction is the column direction, at least one lane line template included in the lane line query feature is at least one second lane line template, and the at least one final lane line determined is at least one second final lane line. Before obtaining the lane line detection result based on the heat map set and the offset map set, the method also includes: performing classification based on the decoded image features to obtain a lane line score set, the lane line score set including at least one group of lane line scores corresponding to at least one lane line template one by one, each group of lane line scores being used to indicate the lane line prediction probability, and the lane line prediction probability is the prediction probability that the corresponding lane line template is included in the image to be processed; based on the heat map set and the offset map set, The lane line detection result is obtained by shifting the map set, and also includes: determining a first comprehensive score of at least one first final lane line based on the lane line score corresponding to at least one first final lane line, and the lane line score corresponding to each first final lane line is the lane line score corresponding to the first lane line template corresponding to the first final lane line; determining a second comprehensive score of at least one second final lane line based on the lane line score corresponding to at least one second final lane line, and the lane line score corresponding to each second final lane line is the lane line score corresponding to the second lane line template corresponding to the second final lane line; comparing the first comprehensive score with the second comprehensive score, if the first comprehensive score is greater than the second comprehensive score, selecting at least one first final lane line, and if the second comprehensive score is greater than the first comprehensive score, selecting at least one second final lane line; determining the selected final lane line as the new lane line detection result.
[0085] In one example, the lane line detection operation can be performed only on the basis that the target direction is the row direction to obtain at least one first final lane line, and the final lane line obtained at this time can be regarded as the final lane line detection result. In another example, the lane line detection operation can be performed only on the basis that the target direction is the column direction to obtain at least one second final lane line, and the final lane line obtained at this time can be regarded as the final lane line detection result. In yet another example, the lane line detection operation can be performed on the basis that the target direction is the row direction to obtain at least one first final lane line, and the lane line detection operation can also be performed on the basis that the target direction is the column direction to obtain at least one second final lane line. Then, the final lane line detection result can be determined based on the two lane lines. In the third example, optionally, at least one first final lane line and at least one second final lane line can both be regarded as the final lane line detection result, that is, these lane lines can be not screened. Optionally, the comprehensive scores of the two lane lines can be calculated and compared, and the final lane line with a high comprehensive score can be regarded as the lane line detection result, which is equivalent to updating the lane line detection result obtained based on the first final lane line and the second final lane line, so a new lane line detection result is obtained.
[0086] Exemplarily, based on the lane line score corresponding to at least one first final lane line, determining the first comprehensive score of at least one first final lane line includes: summing or averaging the lane line scores corresponding to at least one first final lane line to obtain the first comprehensive score. Exemplarily, based on the lane line score corresponding to at least one second final lane line, determining the second comprehensive score of at least one second final lane line includes: summing or averaging the lane line scores corresponding to at least one second final lane line to obtain the second comprehensive score. Optionally, summing or averaging the lane line scores may be a weighted summation or weighted averaging based on weights.
[0087] It can be understood that the lane line patterns corresponding to the lane line templates in the two cases where the target direction is the row direction and the target direction is the column direction may be different. The first lane line template mainly corresponds to the longitudinal lane line, and the second lane line template mainly corresponds to the transverse lane line. The longitudinal direction refers to the height direction on the image, and the transverse direction refers to the width direction on the image. In the fields of autonomous driving, the lane lines collected by the image acquisition device on the vehicle are mainly longitudinal lane lines. Even if the angle between the lane line and the transverse direction is small, the lane line can be detected by the detection method when the target direction is the row direction. However, in extreme cases, such as when the lane line is horizontal or approximately horizontal with the transverse direction, for example, when the angle between the lane line and the transverse direction is less than the predetermined angle threshold, the detection method with the target direction being the column direction can be used for lane line detection. In the case where transverse lane lines and longitudinal lane lines exist at the same time, the above two methods can be combined to determine the lane line, and the detection results of the two methods can be used together as the final lane line detection result. In some cases, for example, if the current lane line only has horizontal lane lines or vertical lane lines, but it is impossible to confirm which one is contained in the image to be processed, both methods can be used for detection, and the result with the highest lane line score is selected as the final lane line detection result.
[0088] Through the above scheme, the detection results of the two detection modes of target direction as row direction and target direction as column direction can be combined, which is helpful to accurately detect the lane line in the correct direction when the extension direction of the lane line is unknown.
[0089] Exemplarily, encoding the initial image features through the encoder module in the converter model to obtain the encoded image features can include: position encoding the image to be processed to obtain position encoding features; adding the initial image features to the position encoding features and flattening the addition result to obtain input features; inputting the input features into the encoder module to perform self-attention mechanism operations to obtain image encoding features.
[0090] Exemplarily, the position coding may include any one of conditional position coding, learnable absolute position coding, sine-cosine function coding, and relative position coding. In one embodiment, the sine-cosine function may be used to perform position coding on the two-dimensional position coordinates of each local area (which may include a single pixel or multiple pixels) in the processed image X to obtain its corresponding position coding feature E. Figure 3 As shown, the position encoding feature (position embedding) E can be obtained through position encoding, E∈R H×W×C That is, the size of the position encoding feature E is H×W×C, which is consistent with the size of the initial image feature F. By adding the feature values of the corresponding positions in the initial image feature F and the position encoding feature E, and flattening the addition result, the input feature I can be obtained, I∈RHW×C . The input feature at this time contains C channels, and each channel contains HW elements. For example, the length (H), width (W), and number of channels (C) of the initial image feature F and the position encoding feature E are 100, 100, and 50 respectively, then the number of channels of the input feature I is 50, and each channel contains 10,000 elements. The input feature I is input into the encoder module for self-attention mechanism operation, and the most relevant input feature corresponding to each feature in the input feature I can be obtained. According to the most relevant input feature corresponding to each feature in the input feature I, the corresponding image encoding feature M can be obtained, M∈R HW×C .
[0091] As mentioned above, through the operation of the self-attention mechanism, the image encoding features obtained can be associated with the global context information, so the feature information contained therein is more comprehensive and can be applied to complex scenarios such as lane lines being occluded or overlapped.
[0092] Exemplarily, comprehensively decoding the encoded image features and the lane line query features through a decoder module in the converter model to obtain the decoded image features may include: inputting the encoded image features and the lane line query features into the decoder module to perform a cross-attention mechanism operation to obtain the decoded image features.
[0093] In one embodiment, the expression for the cross-attention mechanism operation is:
[0094]
[0095] Where Q is the query feature sequence obtained by linearly transforming the lane line query feature S; K and V are the key and value sequences obtained by linearly transforming the encoded image feature M; Q i and K j are the features of the i-th and j-th positions of Q and K respectively; Q i Transpose; A represents an attention map, which can represent Q i and K j The pairwise relationship A i,j ; T i is the i-th decoded feature vector in the decoded image feature T, and its sum with the i-th query feature vector S i Correspondingly; exp represents the exponential function; g is a nonlinear transformation function. Figure 5 FIG. 4 is a schematic diagram of an attention map according to an embodiment of the present application. In one embodiment, the lane line query feature S may include 80 query feature vectors, among which S is selected. 27 , S 35 , S 49 , S73 The four corresponding attention maps are visualized. Figure 5 It can be seen that each query feature vector S i can correspond to a lane line at a specific location, and the lane points in the lane line can be highlighted. The cross attention mechanism can be used to query the feature vector S i Find the most relevant lane line pixels, so by weighted summing the features of these pixels, we can get the feature vector S that matches the query i The corresponding decoded feature vector T i .
[0096] The role of the cross-attention mechanism operation has been described above and will not be repeated here.
[0097] Exemplarily, the lane line detection operation is performed through a lane line detection model, and the method may further include: obtaining a sample image and corresponding annotation information, the annotation information being used to indicate the position of at least one real lane line in the sample image; performing a lane line detection operation on the sample image to obtain a lane line prediction result, wherein the lane line prediction result may include at least one predicted lane line obtained for the sample image detection; based on at least one real lane line and at least one predicted lane line, determining an optimal single-shot function using a bipartite matching loss algorithm; calculating the prediction loss of the lane line detection model based on the optimal single-shot function; and optimizing the parameters in the lane line detection model based on the prediction loss.
[0098] In one embodiment, the lane line detection model used in this embodiment can be obtained by training a training data set. The training data set may include multiple sample images and annotation information (Groundtruth) corresponding to the multiple sample images. The method of obtaining the sample image is similar to the method of obtaining the image to be processed in the previous embodiment. For the sake of brevity, it will not be repeated here. The annotation information can be used to indicate the position of at least one real lane line in the sample image, thereby obtaining the real lane line set Where N represents the number of real lane lines in the current sample image.
[0099] A person skilled in the art can understand how the lane line detection operation is implemented by reading the above steps S221 to S227. For the sake of brevity, it will not be described here. Through the lane line detection operation, the lane line prediction result corresponding to each sample image can be obtained. The lane line prediction result includes one or more predicted lane lines contained in each sample image. Thus, the predicted lane line set can be obtained. Where L represents the number of predicted lane lines detected for the current sample image. Based on the real lane line set y* and the predicted lane line set y, the optimal single injection function can be determined using the bipartite matching loss algorithm. Figure 3, shows a calculation module (loss calculation module) for calculating the bipartite matching loss.
[0100] Using the bipartite matching loss algorithm to determine the optimal single-shot function, we can first calculate the pairwise matching loss L match (y i ,y j *). For example, in an embodiment where a lane line range set and a lane line score set are obtained by detection, the pairwise matching loss L match (y i ,y j *) can include the target loss L obj , heat map loss L heat , offset map loss L off and lane line range loss L rng In one embodiment,
[0101] Target loss L obj The formula is:
[0102]
[0103] Among them, C i0 and C i1 are the probabilities of the i-th predicted lane line belonging to the background and foreground, respectively, and C j0 * and C j1 * are the binary labels of the background and foreground corresponding to the jth real lane line. In one embodiment, C j0 * = 0, C j1 * = 1. The target loss L in the current calculation obj Negative samples, i.e., background probability, are utilized in .
[0104] Heatmap loss L heat The formula is:
[0105]
[0106] Among them, B ikm represents the lane point prediction probability (i.e., foreground prediction probability) of the pixel in the kth row and mth column on the heat map corresponding to the i-th predicted lane line, P jk0 * is the horizontal coordinate of the kth lane point of the jth real lane line, R j0 * and R j1 * are the starting and ending row numbers of the jth real lane line. It should be noted that after dynamic convolution based on the heat map dynamic convolution kernel, the row-by-row pixels in the convolution result can be normalized by the softmax function, and then the heat map set B is obtained. Heat map loss L heat Can only be in R j0 * and Rj1 * Calculated between.
[0107] Offset map loss L off The expression is:
[0108]
[0109] Among them, Z ikm is the offset corresponding to the pixel in the kth row and mth column on the offset map corresponding to the i-th predicted lane line, Z ikm +m is the horizontal coordinate of the kth lane point predicted by the pixel (i.e. the lane point predicted by the pixel above). off It is also possible to only j0 * and R j1 * Calculated between.
[0110] Lane line range loss L rng The expression is:
[0111]
[0112] Among them, R i0 and R i1 are the row numbers of the starting and ending rows of the i-th predicted lane line.
[0113] The final pairwise matching loss L match (y i ,y j *) is expressed as:
[0114]
[0115] Among them, λ obj , heat , off , rng L obj , L heat , L off , L rng These weights can be set to any size as needed.
[0116] The goal of bipartite matching is to find an optimal injective function z: Minimizing the above matching loss, we can get the following expression:
[0117]
[0118] Where z(j) is the index of the predicted lane line assigned to the jth real lane line. This expression can be solved using the Hungarian algorithm, which is not described here. The optimal single-shot function is used to represent the predicted lane line that matches each real lane line. For each real lane line, there is a unique predicted lane line that matches it.
[0119] According to the optimal single-shot function z, the final prediction loss loss can be calculated as follows:
[0120]
[0121] in, It is the index set of the predicted lane lines that match the real lane lines. Among them C i0 is the background probability in the lane score corresponding to the i-th predicted lane. According to the above prediction loss calculation formula, it can be seen that when calculating the target loss L obj When using the negative sample, i.e., the lane score corresponding to the predicted lane line that does not match the real lane line, the other losses only use the positive sample, i.e., the information corresponding to the predicted lane line that matches the real lane line. For example, assuming that there are 3 real lane lines and 10 predicted lane lines (this number can be consistent with the number of lane line templates L mentioned above), then 3 predicted lane lines that match the 3 real lane lines can be determined from the 10 predicted lane lines, and then When calculating, use the three predicted lane lines matched above. The remaining 7 predicted lane lines are used for calculation. In addition, it should be noted that the above-mentioned bisection matching loss algorithm is only an example and not a limitation of the present application. For example, any predicted lane line and any real lane line can be combined into pairs to calculate the prediction loss. For example, 3 real lane lines and 10 predicted lane lines can form 30 pairs to calculate the prediction loss.
[0122] After training the parameters in the initial lane line detection model, what is obtained is the lane line detection model used in the previous embodiment. The lane line prediction results and the annotation information of multiple sample images can be substituted into the above-mentioned prediction loss calculation function to perform loss calculation and obtain the prediction loss. Subsequently, the parameters in the initial lane line detection model can be optimized using the back propagation and gradient descent algorithms based on the calculated prediction loss. The optimization of the parameters can be iteratively performed until the lane line detection model reaches a convergence state. When the training is completed, the obtained lane line detection model can be used for subsequent lane line detection, and this stage can be called the testing or reasoning stage of the model.
[0123] According to the above technical solution, during the training of the lane detection model, a binary match is performed between the predicted lane and the real lane, and the prediction loss function is calculated based on the determined optimal single-shot function. This ensures that there are no redundant and repeated prediction results, and there is no need to use the non-maximum suppression (NMS) algorithm for deduplication, thereby improving the training efficiency of the lane detection model.
[0124] Exemplarily, optimizing parameters in the lane detection model based on the prediction loss includes optimizing the parameters in the lane detection model and the lane query features together based on the prediction loss.
[0125] The lane line query feature used in step S223 can be learned from the image under the supervision of the real lane line. Therefore, when the parameters in the lane line detection model are optimized based on the prediction loss, the lane line query feature can also be optimized together with the parameters in the lane line detection model.
[0126] According to another aspect of the present application, a lane line detection device is provided. Figure 6 A schematic block diagram of a lane detection device 600 according to an embodiment of the present application is shown.
[0127] like Figure 6 As shown, the lane line detection device 600 according to the embodiment of the present application includes an acquisition module 610 and a detection module 620. The detection module 620 may include an extraction submodule 621, an encoding submodule 622, a decoding submodule 623, a conversion submodule 624, a first convolution submodule 625, a second convolution submodule 626 and an acquisition submodule 627. Each module may respectively execute the above Figure 2 The following only describes the main functions of the various components of the lane line detection device 600, and omits the details already described above.
[0128] The acquisition module 610 is used to acquire the image to be processed. The acquisition module 610 can be composed of Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0129] The detection module 620 is used to perform lane line detection on the image to be processed to obtain a lane line detection result. The detection module 620 can be composed of Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0130] Specifically, the extraction submodule 621 is used to extract features from the image to be processed to obtain initial image features.
[0131] The encoding submodule 622 is used to encode the initial image features through the encoder module in the converter model to obtain encoded image features.
[0132] The decoding submodule 623 is used to comprehensively decode the encoded image features and the lane line query features through the decoder module in the converter model to obtain the decoded image features, wherein the lane line query features include query features that correspond one-to-one to at least one lane line template.
[0133] The conversion submodule 624 is used to perform feature conversion on the decoded image features to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel. The heat map dynamic convolution kernel includes heat map features that correspond one-to-one to at least one lane line template, and the offset map dynamic convolution kernel includes offset features that correspond one-to-one to at least one lane line template.
[0134] The first convolution submodule 625 is used to convolve the initial image features or the encoded image features with the heat map dynamic convolution kernel to obtain a heat map set, wherein the heat map set includes at least one heat map corresponding to at least one lane line template, and each heat map is used to indicate the lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is the prediction probability that the local area is a lane point on the corresponding lane line template.
[0135] The second convolution submodule 626 is used to convolve the initial image features or the encoded image features using the offset map dynamic convolution kernel to obtain an offset map set, where the offset map set includes at least one offset map corresponding to at least one lane line template, and each offset map is used to indicate the offset of each local area on the image to be processed relative to the corresponding lane line template.
[0136] The acquisition submodule 627 is used to obtain lane line detection results based on the heat map set and the offset map set.
[0137] Figure 7 A schematic block diagram of an electronic device 700 according to an embodiment of the present application is shown. The electronic device 700 includes a memory 710 and a processor 720 .
[0138] The memory 710 stores computer program instructions for implementing corresponding steps in the lane line detection method according to the embodiment of the present application.
[0139] The processor 720 is used to run the computer program instructions stored in the memory 710 to perform the corresponding steps of the lane line detection method according to the embodiment of the present application.
[0140] In one embodiment, the computer program instructions are used by the processor 720 to execute the following steps when the processor 720 is running: obtaining an image to be processed; performing a lane line detection operation on the image to be processed to obtain a lane line detection result, wherein the lane line detection operation includes: performing feature extraction on the image to be processed to obtain initial image features; encoding the initial image features through an encoder module in the converter model to obtain encoded image features; comprehensively decoding the encoded image features and lane line query features through a decoder module in the converter model to obtain decoded image features, wherein the lane line query features include query features that correspond one-to-one to at least one lane line template; performing feature conversion on the decoded image features to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel, wherein the heat map dynamic convolution kernel includes a heat map feature that corresponds one-to-one to at least one lane line template, and the offset map The dynamic convolution kernel includes an offset feature corresponding to at least one lane line template; the initial image feature or the encoded image feature is convolved with the heat map dynamic convolution kernel to obtain a heat map set, the heat map set includes at least one heat map corresponding to at least one lane line template, each heat map is used to indicate the lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is the prediction probability that the local area is a lane point on the corresponding lane line template; the initial image feature or the encoded image feature is convolved with the offset map dynamic convolution kernel to obtain an offset map set, the offset map set includes at least one offset map corresponding to at least one lane line template, each offset map is used to indicate the offset of each local area on the image to be processed relative to the corresponding lane line template; the lane line detection result is obtained based on the heat map set and the offset map set.
[0141] Exemplarily, the electronic device 700 may further include an image acquisition device 730. The image acquisition device 730 is used to acquire the image to be processed. The image acquisition device 730 is optional, and the electronic device 700 may not include the image acquisition device 730. At this time, the processor 720 may obtain the image to be processed in other ways, such as obtaining the image to be processed from an external device or from the memory 710.
[0142] In addition, according to an embodiment of the present application, a storage medium is also provided, on which program instructions are stored, and when the program instructions are executed by a computer or a processor, they are used to execute the corresponding steps of the lane line detection method of the embodiment of the present application, and are used to implement the corresponding modules in the lane line detection device according to the embodiment of the present application. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.
[0143] In one embodiment, when the program instructions are executed by a computer or processor, the computer or processor may implement the various functional modules of the lane line detection device according to the embodiment of the present application, and / or may execute the lane line detection method according to the embodiment of the present application.
[0144] In one embodiment, the program instructions are used to execute the following steps when running: obtaining an image to be processed; performing a lane line detection operation on the image to be processed to obtain a lane line detection result, wherein the lane line detection operation includes: performing feature extraction on the image to be processed to obtain initial image features; encoding the initial image features through an encoder module in a converter model to obtain encoded image features; comprehensively decoding the encoded image features and lane line query features through a decoder module in the converter model to obtain decoded image features, wherein the lane line query features include query features that correspond one-to-one to at least one lane line template; performing feature conversion on the decoded image features to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel, wherein the heat map dynamic convolution kernel includes a heat map feature that corresponds one-to-one to at least one lane line template, and the offset map dynamic convolution kernel The kernel includes an offset feature corresponding to at least one lane line template one by one; the initial image feature or the encoded image feature is convolved with the heat map dynamic convolution kernel to obtain a heat map set, the heat map set includes at least one heat map corresponding to at least one lane line template one by one, each heat map is used to indicate the lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is the prediction probability that the local area is a lane point on the corresponding lane line template; the initial image feature or the encoded image feature is convolved with the offset map dynamic convolution kernel to obtain an offset map set, the offset map set includes at least one offset map corresponding to at least one lane line template one by one, each offset map is used to indicate the offset of each local area on the image to be processed relative to the corresponding lane line template; the lane line detection result is obtained based on the heat map set and the offset map set.
[0145] In addition, according to an embodiment of the present application, a computer program product is also provided. The computer program product includes a computer program, and the computer program is used to execute the above-mentioned lane line detection method 200 when running.
[0146] Each module in the electronic device according to the embodiment of the present application can be implemented by running computer program instructions stored in a memory by a processor of an electronic device that implements lane line detection or lane line detection according to the embodiment of the present application, or can be implemented when computer instructions stored in a computer-readable storage medium of a computer program product according to the embodiment of the present application are executed by a computer.
[0147] In addition, according to an embodiment of the present application, a computer program is also provided, which is used to execute the above-mentioned lane line detection method 200 when running.
[0148] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application to this. Those of ordinary skill in the art may make various changes and modifications therein without departing from the scope and spirit of the present application. All these changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0149] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0151] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0152] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various application aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present application should not be interpreted as reflecting the following intention: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with features less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby explicitly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.
[0153] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0154] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0155] The various component embodiments of the present application may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) may be used in practice to implement some or all of the functions of some modules in the lane detection device according to the embodiment of the present application. The present application may also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application may be stored on a computer-readable medium, or may be in the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0156] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and that those skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets should not be constructed as a limitation to the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of multiple such elements. The present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim that lists several devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names.
[0157] The above is only a specific implementation method or description of a specific implementation method of the present application, and the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. The protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A lane line detection method, comprising: Get the image to be processed; Performing a lane line detection operation on the image to be processed to obtain a lane line detection result, wherein the lane line detection operation includes: Extracting features of the image to be processed to obtain initial image features; Encoding the initial image features through an encoder module in a converter model to obtain encoded image features; Comprehensively decoding the encoded image features and the lane line query features through a decoder module in the converter model to obtain decoded image features, wherein the lane line query features include query features corresponding to at least one lane line template; Performing feature conversion on the decoded image features to obtain a heat map dynamic convolution kernel and an offset map dynamic convolution kernel, wherein the heat map dynamic convolution kernel includes a heat map feature corresponding one-to-one to the at least one lane line template, and the offset map dynamic convolution kernel includes an offset feature corresponding one-to-one to the at least one lane line template; Convolving the initial image features or the encoded image features using the heat map dynamic convolution kernel to obtain a heat map set, wherein the heat map set includes at least one heat map corresponding to the at least one lane line template, each heat map is used to indicate a lane point prediction probability corresponding to each local area on the image to be processed, and the lane point prediction probability is a prediction probability that the local area is a lane point on the corresponding lane line template; Convolving the initial image feature or the encoded image feature using the offset map dynamic convolution kernel to obtain an offset map set, wherein the offset map set includes at least one offset map corresponding to the at least one lane line template, and each offset map is used to indicate the offset of each local area on the image to be processed relative to the corresponding lane line template; The lane line detection result is obtained based on the heat map set and the offset map set.
2. The method of claim 1, wherein: The local area is a single pixel, the pixels in each heat map correspond one-to-one to the pixels in the image to be processed, and any pixel in each heat map is used to indicate the lane point prediction probability of the corresponding pixel in the image to be processed, the pixels in each offset map correspond one-to-one to the pixels in the image to be processed, and any pixel in each offset map is used to indicate the offset of the corresponding pixel in the image to be processed relative to the corresponding lane line template in the target direction, and the target direction is the row direction and / or the column direction, The obtaining of the lane line detection result based on the heat map set and the offset map set includes: For each lane line template in the at least one lane line template, Based on the heat map and the offset map corresponding to the lane line template, voting for each pixel in the j-th group of pixels in the image to be processed to determine the voting score of each pixel in the j-th group of pixels, where the j-th group of pixels is a group of pixels extending along the target direction, j=1, 2, ..., Num1, where Num1 represents the total number of groups of pixels of the image to be processed extending along the target direction; Determine the pixel with the highest voting score in the j-th group of pixels as the group predicted lane point corresponding to the j-th group of pixels; For the image to be processed, based on the group predicted lane points corresponding to each group of pixels, determine an initial lane point set corresponding to the lane line template; Determining an initial lane line corresponding to each of the at least one lane line templates based on an initial lane point set corresponding to each of the at least one lane line templates; Based on the initial lane line corresponding to each of the at least one lane line templates, at least one final lane line is determined to obtain the lane line detection result.
3. The method of claim 2, wherein: The step of voting for each pixel in the j-th group of pixels in the image to be processed based on the heat map and the offset map corresponding to the lane line template to determine the voting score of each pixel in the j-th group of pixels includes: Based on the offset corresponding to the kth pixel in the jth group of pixels in the offset map corresponding to the lane line template, calculate the position of the pixel predicted lane point corresponding to the kth pixel, where k=1, 2, ..., Num2, Num2 represents the total number of pixels in the jth group of pixels; For any target pixel in the j-th group of pixels, determine an anchor pixel whose position of the corresponding pixel predicted lane point coincides with the target pixel; Determine the lane point prediction probability corresponding to each of the anchor point pixels based on the heat map corresponding to the lane line template; Based on the lane point prediction probabilities corresponding to the anchor pixels, the voting scores corresponding to the target pixels are determined.
4. The method of claim 3, wherein: The determining the voting score corresponding to the target pixel based on the lane point prediction probability corresponding to each of the anchor pixels includes: The lane point prediction probabilities corresponding to the anchor pixels are added together to obtain the voting score corresponding to the target pixel.
5. The method of claim 2, wherein: Before obtaining the lane line detection result based on the heat map set and the offset map set, the method further includes: Classify based on the decoded image features to obtain a lane line range set, the lane line range set including at least one set of lane line range information corresponding to the at least one lane line template, each set of lane line range information is used to indicate the coverage of the corresponding lane line template on the image to be processed in a direction perpendicular to the target direction; The determining, based on the initial lane point set corresponding to each of the at least one lane line templates, the initial lane line corresponding to each of the at least one lane line templates comprises: For each lane line template in the at least one lane line template, Selecting a group of predicted lane points located within the coverage range corresponding to the lane line template from the corresponding initial lane point set; An initial lane line corresponding to the lane line template is formed based on the selected group of predicted lane points.
6. The method of claim 2, wherein: Before obtaining the lane line detection result based on the heat map set and the offset map set, the method further includes: Classify based on the decoded image features to obtain a lane line score set, the lane line score set including at least one group of lane line scores corresponding to the at least one lane line template, each group of lane line scores is used to indicate a lane line prediction probability, and the lane line prediction probability is a prediction probability that the corresponding lane line template is included in the image to be processed; The determining at least one final lane line based on the initial lane line corresponding to each of the at least one lane line templates to obtain the lane line detection result includes: An initial lane line corresponding to a lane line template whose lane line prediction probability is higher than a target probability threshold is selected as the at least one final lane line to obtain the lane line detection result.
7. The method of claim 2, wherein: The target direction is a row direction and a column direction. When the target direction is a row direction, at least one lane line template included in the lane line query feature is at least one first lane line template, and the at least one final lane line determined is at least one first final lane line. When the target direction is a column direction, at least one lane line template included in the lane line query feature is at least one second lane line template, and the at least one final lane line determined is at least one second final lane line. Before obtaining the lane line detection result based on the heat map set and the offset map set, the method further includes: Classify based on the decoded image features to obtain a lane line score set, the lane line score set including at least one group of lane line scores corresponding to the at least one lane line template, each group of lane line scores is used to indicate a lane line prediction probability, and the lane line prediction probability is a prediction probability that the corresponding lane line template is included in the image to be processed; The obtaining of the lane line detection result based on the heat map set and the offset map set also includes: Determine a first comprehensive score of the at least one first final lane line based on the lane line score corresponding to the at least one first final lane line, wherein the lane line score corresponding to each first final lane line is the lane line score corresponding to the first lane line template corresponding to the first final lane line; Determine a second comprehensive score of the at least one second final lane line based on the lane line score corresponding to the at least one second final lane line, wherein the lane line score corresponding to each second final lane line is the lane line score corresponding to the second lane line template corresponding to the second final lane line; Comparing the first comprehensive score with the second comprehensive score, if the first comprehensive score is greater than the second comprehensive score, selecting the at least one first final lane line, and if the second comprehensive score is greater than the first comprehensive score, selecting the at least one second final lane line; The selected final lane line is determined as the new lane line detection result.
8. The method according to any one of claims 1 to 7, wherein: The lane line detection operation is performed by a lane line detection model, and the method further includes: Acquire a sample image and corresponding annotation information, where the annotation information is used to indicate the position of at least one real lane line in the sample image; Performing the lane line detection operation on the sample image to obtain a lane line prediction result, wherein the lane line prediction result includes at least one predicted lane line obtained by detecting the sample image; Determine an optimal injective function based on the at least one real lane line and the at least one predicted lane line using a bipartite matching loss algorithm; Calculating the prediction loss of the lane detection model based on the optimal single-shot function; Parameters in the lane detection model are optimized based on the prediction loss.
9. The method of claim 8, wherein: The optimizing the parameters in the lane detection model based on the prediction loss includes: Parameters in the lane line detection model and the lane line query features are optimized together based on the prediction loss.
10. An electronic device comprising a processor and a memory, wherein: The memory stores computer program instructions, which are used by the processor to execute the lane line detection method according to any one of claims 1 to 9 when the processor is running the computer program instructions.
11. A storage medium having program instructions stored thereon, wherein: The program instructions are used to execute the lane line detection method as described in any one of claims 1 to 9 when running.
12. A computer program product, the computer program product comprising a computer program, wherein: The computer program is used to execute the lane line detection method as described in any one of claims 1 to 9 when running.