Lane line detection method and device, equipment, storage medium and program product

By acquiring image features at different depths and using a layered cross attention decoder for feature fusion, the problem of difficult feature selection and difficult detection in complex environments in lane line detection is solved, and high accuracy and robust lane line detection is achieved.

CN120182932APending Publication Date: 2025-06-20INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510161380.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing lane line detection technology faces the dilemma of feature selection and the problem of difficulty in detection in complex environments.

Method used

By acquiring n-layer image features at different depths, and using a layered cross attention decoder to perform cross attention operations with trained lane line instances, different layer features are fused to generate multiple lane line instance features.

Benefits of technology

This method can effectively identify and locate lane lines in complex environments, improve detection accuracy and robustness, and solve the problems of difficulty in feature selection and detection in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182932A_ABST
    Figure CN120182932A_ABST
Patent Text Reader

Abstract

The invention provides a lane line detection method, device and equipment, a storage medium and a program product, and is applied to the technical field of target detection. The method comprises the following steps: acquiring a target scene image after enhancement processing; obtaining n layers of image features with different feature depths according to the target scene image; the n layers of image features are input into layered cross attention decoders, cross attention operation is carried out on the n layers of image features and the trained lane line instances to obtain features of different layers, the features of the different layers are fused to obtain multiple lane line instance features, and one layer of image feature corresponds to one layered cross attention decoder; performing cross attention operation on each lane line instance feature in the plurality of lane line instance features and other lane line instance features to obtain updated features, and determining at least one piece of lane line information according to the updated features; and determining the lane line information of which the confidence coefficient is higher than a target threshold value in the at least one piece of lane line information as a lane line detection result of the target scene image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly to a lane line detection method, device, equipment, storage medium and program product. Background Art

[0002] At present, with the booming development of intelligent transportation systems, lane line detection technology is crucial. Its core goal is to accurately identify and locate the lane lines on which vehicles are driving in an open road environment, and at the same time track the driving lanes where vehicles are located under complex road conditions. This technology is widely used in multiple fields such as autonomous driving and traffic management, and is of great significance for improving road safety and traffic efficiency, and has received high attention from the academic and industrial communities in recent years.

[0003] Currently, the research paradigms for lane line detection are rich and diverse. Mainly based on the type of lane representation, they can be divided into the following categories: methods based on key point representation, analogous to human pose estimation, regarding lane points as key points for detection, and then grouping them to form lane instances; methods based on row representation, representing lane instances as a set of x coordinates at fixed rows, and using a row-by-row detection algorithm to detect using the prior of the lane line shape; methods based on parameters, achieving lane detection by parametric modeling and regression of lane curves; methods based on segmentation representation, which estimate the possibility of the existence of lane lines from the pixel level from bottom to top.

[0004] However, there are many problems with these existing technical solutions. For methods based on key point representation, the computational cost of post-processing grouped lane points is high; for methods based on row representation, it is difficult to distinguish at the instance level, and clustering methods are difficult to directly apply; for methods based on parameters, they are sensitive to predicted parameters, and errors in higher-order coefficients are likely to cause lane shape deviations; for methods based on segmentation representation, they face the problem of instance-level discrimination, rely on post-processing such as clustering and non-maximum suppression (NMS), and have poor detection effects under complex lighting and occlusion conditions.

[0005] Generally speaking, the existing lane line detection technology faces two major challenges: one is the dilemma of feature selection. Although low-semantic features have accurate positioning, they are easily confused with ground signs, while high-semantic features can accurately identify the lane line category, but the positioning is fuzzy; the other is that under complex lighting and occlusion environments, the difficulty of lane line detection is extremely high, seriously affecting the accuracy and reliability of detection. Summary of the Invention

[0006] The present invention provides a lane line detection method, device, equipment, storage medium and program product to solve the problems of difficult feature selection in lane line detection technology and great detection difficulty in complex environments in the prior art.

[0007] The present invention provides a lane line detection method, including: obtaining a target scene image after enhancement processing; obtaining n layers of image features with different feature depths according to the target scene image; inputting the n layers of image features into a hierarchical cross-attention decoder, performing cross-attention operation with trained lane line instances to obtain features of different layers, fusing the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross-attention decoder; performing cross-attention operation on each lane line instance feature in the multiple lane line instance features with other lane line instance features to obtain updated features, and determining at least one lane line information according to the updated features; determining the lane line detection result of the target scene image for the lane line information with a confidence level higher than a target threshold in the at least one lane line information.

[0008] According to a lane line detection method provided by the present invention, the obtaining n layers of image features with different feature depths according to the target scene image includes: inputting the target scene image into an encoder to obtain n layers of image features with different feature depths, where the encoder includes a backbone network and a pyramid network, and the pyramid network is used for upsampling and fusing n layers of feature maps with different depths.

[0009] According to a lane line detection method provided by the present invention, before obtaining the target scene image after enhancement processing, the method further includes: determining n layers of image features of a first scene image in a training set; performing cross-attention operation on the n layers of image features of the first scene image and m learnable instance queries in a lane line detection model to obtain features of different layers, fusing the features of different layers to obtain m lane line instance features; performing cross-attention operation on each lane line instance feature in the m lane line instance features with other lane line instance features to obtain m updated features, and determining multiple lane line information according to the m updated features; performing one-to-one matching on the multiple lane line information and the ground truth data in the training set through the Hungarian matching algorithm to obtain lane line matching information; calculating a total loss function based on the lane line matching information, and training the lane line detection model based on the total loss function until a preset training end condition is satisfied.

[0010] According to a lane line detection method provided by the present invention, the total loss function includes a first loss function, a second loss function, and a third loss function. The first loss function is used to calculate the class loss of the lane line, the second loss function is used to calculate the lane line point set loss, and the third loss function is used to calculate the starting point coordinate loss.

[0011] According to a lane line detection method provided by the present invention, the enhancement processing operation on the target scene image includes a cropping operation, a scaling operation, and a perspective transformation operation.

[0012] The present invention also provides a lane line detection device, including the following modules: an acquisition module and a processing module; the acquisition module is used to acquire the target scene image after enhancement processing; the processing module is used to obtain n layers of image features with different feature depths according to the target scene image; input the n layers of image features into a hierarchical cross-attention decoder, perform cross-attention operation with the trained lane line instances to obtain features of different layers, fuse the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross-attention decoder; perform cross-attention operation on each lane line instance feature in the multiple lane line instance features with other lane line instance features to obtain updated features, and determine at least one lane line information according to the updated features; determine the lane line detection result of the target scene image for the lane line information with a confidence level higher than the target threshold in the at least one lane line information.

[0013] According to a lane line detection device provided by the present invention, the processing module is used to input the target scene image into an encoder to obtain n layers of image features with different feature depths, and the encoder includes a backbone network and a pyramid network, and the pyramid network is used for upsampling and fusion of n layers of feature maps with different depths.

[0014] According to a lane line detection device provided by the present invention, the processing module is used to determine n layers of image features of the first scene image in the training set; perform cross-attention operation on the n layers of image features of the first scene image and m learnable instance queries in the lane line detection model to obtain features of different layers, fuse the features of different layers to obtain m lane line instance features; perform cross-attention operation on each lane line instance feature in the m lane line instance features with other lane line instance features to obtain m updated features, and determine multiple lane line information according to the m updated features; perform one-to-one matching on the multiple lane line information and the data ground truth in the training set through the Hungarian matching algorithm to obtain lane line matching information; calculate the total loss function based on the lane line matching information, and train the lane line detection model based on the total loss function until the preset training end condition is met.

[0015] According to a lane line detection device provided by the present invention, the total loss function includes a first loss function, a second loss function and a third loss function, the first loss function is used to calculate the class loss of the lane line, the second loss function is used to calculate the lane line point set loss, and the third loss function is used to calculate the starting point coordinate loss.

[0016] According to a lane line detection device provided by the present invention, the enhancement processing operation on the target scene image includes a cropping operation, a scaling operation and a perspective transformation operation.

[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the lane line detection method as described in any one of the above is implemented.

[0018] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the lane line detection method as described in any one of the above is implemented.

[0019] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the lane line detection method as described in any one of the above is implemented.

[0020] The lane line detection method, device, equipment, storage medium, and program product provided by the present invention obtain n-layer image features with different depths, so that the feature map contains both high-level semantic information and low-level position information; through hierarchical cross-attention operation, the query of each lane line instance can interact with the features of the corresponding layer, fully mining the relevant information of the lane line instances in different layer features; by fusing the features of different layers to obtain lane line instance features, it can ensure that features of different scales and depths are comprehensively utilized, making the lane line instance features have both the discriminability of high-semantic features and the sensitivity to position of low-semantic features; through lane line cross-attention operation, the relationship between lane lines can be captured, and in a complex environment, the occluded part of the lane line can be inferred based on the information of adjacent or relevant lane line instances, thus effectively solving the occlusion problem and improving the robustness of lane line detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0022] Figure 1 is one of the schematic flowcharts of the lane line detection method provided by the present invention; Figure 2 is another schematic flowchart of the lane line detection method provided by the present invention; Figure 3 is yet another schematic flowchart of the lane line detection method provided by the present invention; Figure 4 is the schematic structural diagram of the lane line detection device provided by the present invention; Figure 5 is the schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions and advantages of this application clearer, the following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts fall within the scope of protection of this application.

[0024] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0025] It should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of this application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0026] To facilitate a clear description of the technical solutions in the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order.

[0027] Some exemplary embodiments are described in the embodiments of this application for the purpose of illustration. It should be understood that this application can be implemented in other ways not specifically shown in the drawings.

[0028] Such as Figure 1As shown in the figure, an embodiment of the present application provides a lane line detection method, which can be applied to a lane line detection device. The lane line detection method may include S101 - S105: S101. The lane line detection device determines n - layer image features of the first - scene image in the training set.

[0029] Optionally, the lane line detection device may first perform enhancement processing operations on the first - scene image in the training set. The enhancement processing operations may include cropping operations, scaling operations, and perspective transformation operations.

[0030] Specifically, the cropping operation includes randomly cropping and saving the image within the range specified by the user in the first - scene image; such random cropping can generate images with different perspectives and local contents, simulating the situation where only part of the scene may be captured in the actual scene and increasing the diversity of data. The perspective transformation operation includes performing random perspective transformation on the first - scene image so that the performance of the lane lines in the image is more in line with the actual situation; this can make the lane lines in the image look more similar to the lane lines in the real world under different perspectives.

[0031] It should be noted that by performing various transformations on the original data of the training set and simulating more different actual scenes, the lane line detection model can learn richer features during training, thereby reducing the overfitting phenomenon.

[0032] It should be noted that the first - scene image , where H, W, and C represent height, width, and the number of channels respectively.

[0033] Optionally, after performing the enhancement processing operation on the first - scene image, the lane line detection device may input the first - scene image into the encoder of the lane line detection model to obtain n - layer image features of the first - scene image. The feature depths of the n - layer image features are different. The n - layer image features can include both high - level semantic information and low - level position information. The high - level semantic information can help the lane line detection model understand the meaning and function of the lane lines in the entire scene, while the low - level position information can accurately locate the specific position of the lane lines in the image. The organic combination of the two greatly improves the expression ability of the feature map for the lane line features, thereby providing a better feature basis for subsequent tasks. Optionally, the above - mentioned encoder includes a backbone network and a pyramid network. Among them, the backbone network adopts a ResNet structure, and the pyramid network is used for upsampling and fusing n - layer feature maps with different depths.

[0034] For example, taking n as 3. In the lane line detection task, the lane line detection device may input the first - scene image into the encoder to obtain image features image features and image features .

[0035] S102. The lane line detection device performs a cross-attention operation on the n-layer image features of the first scene image and m learnable instance queries in the lane line detection model to obtain features of different layers, and fuses the features of different layers to obtain m lane line instance features.

[0036] Optionally, the lane line detection model may further include n hierarchical cross-attention decoders, which are composed of a cross-attention layer and a fully connected layer. The lane line detection device may input the preset m learnable instance queries L and n layers of image features into the corresponding hierarchical cross-attention decoders respectively, and one layer of image features corresponds to one hierarchical cross-attention decoder. In this process, cross-attention operations are performed to obtain features of different layers. For example, multiple learnable queries are pre-set as queries for each lane line instance, and each hierarchical cross-attention decoder uses the same query. For the i-th lane line of the n-th layer decoder, its feature representation can be obtained. , Finally, the features of different layers of each lane line instance are fused to obtain m lane line instance features.

[0037] like Figure 2 As shown in the figure, in order to enhance the expressiveness and generalization ability of the lane detection model, the image features can be added to the hierarchical cross attention decoder. , image features and image features Add a set of identical learnable lane line instance queries, the shape of this query matrix is ​​(Line_num, embedding_dim). And the query will be convolved without sharing parameters before being used as the input of the layered cross attention decoder at different levels, so that the preset lane line instance query can adapt to features of different scales.

[0038] Specifically, in each layer of the hierarchical cross-attention decoder, the lane detection device can combine the preset learnable query Q with the image feature Perform crisscross attention operation: ; Among them, Q is the query matrix (i.e., the preset learnable query vector), K and V are the features from the image, respectively. The key and value matrices extracted from Is the dimension of the key vector, used to scale the dot-product result.

[0039] For each layer of image features , the query Q of each lane line instance interacts with the features of this layer, and then obtains multiple feature representations: .

[0040] Taking n as 3 as an example. The feature fusion process of the i-th lane line instance at different layers can be expressed as: ; Among them, represents the feature fusion function.

[0041] S103. The lane line detection device performs cross-attention operation on each lane line instance feature among the m lane line instance features to obtain m updated features, and determines multiple lane line information according to the m updated features.

[0042] Optionally, as Figure 2 shown, the lane line detection model may further include a lane line cross-attention module. The lane line detection device may input the m lane line instance features into the lane line cross-attention module. For each lane line instance feature, the lane line cross-attention module may perform cross-attention operation on it and other lane line instance features to obtain updated features, and further obtain multiple lane line information through convolution and fully connected layers. The lane line information may include lane line attributes and point coordinate sets.

[0043] Specifically, the features of each lane line instance in the lane line cross-attention module The interaction process between them can be expressed as: .

[0044] S104. The lane line detection device performs one-to-one matching of the multiple lane line information with the data ground truth in the training set to obtain lane line matching information.

[0045] Specifically, taking m as 4 as an example. The lane line detection device may perform one-to-one Hungarian matching on the lane line information and the ground truth in the training set. By solving a bipartite graph matching problem, the predicted lane lines and the real lane lines are optimally assigned to minimize the distance error between the lane line information.

[0046] S105. The lane line detection device calculates the total loss function based on the lane line matching information, and trains the lane line detection model based on the total loss function until the preset training end condition is satisfied.

[0047] Optionally, the total loss function includes a first loss function, a second loss function, and a third loss function. The first loss function is used to calculate the class loss of the lane line, the second loss function is used to calculate the lane line point set loss, and the third loss function is used to calculate the starting point coordinate loss.

[0048] Specifically, the first loss function can be expressed as: ; where represents the predicted probability of the correct class, and are adjustment parameters.

[0049] The second loss function can be expressed as: ; ; where P represents the point set of the predicted lane line, and G represents the point set of the true lane line.

[0050] The third loss function can be expressed as: ; where is the predicted starting point coordinate, is the true starting point coordinate.

[0051] The total loss function can be expressed as: ; where , , are weight hyperparameters used to balance the contributions of different loss terms to the total loss.

[0052] As Figure 3 shown, an embodiment of the present application provides a lane line detection method, which can be applied to a lane line detection device. The lane line detection method may include S301 - S305: S301. The lane line detection device acquires the enhanced target scene image.

[0053] Optionally, the enhancement processing operations on the target scene image include a cropping operation, a scaling operation, and a perspective transformation operation.

[0054] S302. The lane line detection device obtains n - layer image features with different feature depths according to the target scene image.

[0055] Optionally, the lane line detection device obtains n layers of image features with different feature depths from the target scene image, including: inputting the target scene image into an encoder to obtain n layers of image features with different feature depths.

[0056] S303. The lane line detection device inputs the n layers of image features into a hierarchical cross-attention decoder, performs cross-attention operations with the trained lane line instances, obtains features of different layers, and fuses the features of different layers to obtain multiple lane line instance features.

[0057] Among them, one layer of image features corresponds to one hierarchical cross-attention decoder.

[0058] S304. The lane line detection device performs cross-attention operations on each lane line instance feature in the multiple lane line instance features with other lane line instance features, obtains updated features, and determines at least one lane line information according to the updated features.

[0059] S305. The lane line detection device determines the lane line information with a confidence level higher than the target threshold in the at least one lane line information as the lane line detection result of the target scene image.

[0060] In the embodiments of the present application, by obtaining n layers of image features with different depths, the feature map contains both high-level semantic information and low-level position information; through hierarchical cross-attention operations, the query of each lane line instance can interact with the corresponding layer features, fully mining the relevant information of the lane line instances in different layer features; by fusing the features of different layers to obtain lane line instance features, it can ensure that features of different scales and depths are comprehensively utilized, so that the lane line instance features have both the discriminability of high-semantic features and the sensitivity to position of low-semantic features; through lane line cross-attention operations, the relationship between lane lines can be captured. In a complex environment, the lane lines in the occluded part can be inferred based on the information of adjacent or relevant lane line instances, thus effectively solving the occlusion problem and improving the robustness of lane line detection.

[0061] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of the method. To implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0062] In the lane line detection method provided by the embodiments of the present application, the execution subject may be a lane line detection device, or a control module for lane line detection in the lane line detection device. In the embodiments of the present application, taking the lane line detection device as an example to execute the lane line detection method, the lane line detection device provided by the embodiments of the present application is described.

[0063] It should be noted that the embodiments of the present application can divide the functional modules of the lane line detection device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. Optionally, the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0064] As Figure 4 shown, the embodiments of the present application provide a lane line detection device 400. The lane line detection device 400 includes: an acquisition module 401 and a processing module 402. The acquisition module 401 is configured to acquire an enhanced target scene image; the processing module 402 is configured to obtain n-layer image features with different feature depths according to the target scene image; input the n-layer image features into a hierarchical cross-attention decoder, perform cross-attention operations with trained lane line instances to obtain features of different layers, fuse the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross-attention decoder; perform cross-attention operations on each lane line instance feature in the multiple lane line instance features with other lane line instance features to obtain updated features, and determine at least one lane line information according to the updated features; determine the lane line information with a confidence level higher than the target threshold in the at least one lane line information as the lane line detection result of the target scene image.

[0065] Optionally, the processing module 402 is configured to input the target scene image into an encoder to obtain n-layer image features with different feature depths, and the encoder includes a backbone network and a pyramid network, and the pyramid network is used for upsampling and fusing n-layer feature maps with different depths.

[0066] Optionally, the processing module 402 is configured to determine n-layer image features of the first scene image in the training set; perform cross-attention operations on the n-layer image features of the first scene image and m learnable instance queries in the lane line detection model to obtain features of different layers, and fuse the features of different layers to obtain m lane line instance features; perform cross-attention operations on each lane line instance feature among the m lane line instance features and other lane line instance features to obtain m updated features, and determine multiple lane line information according to the m updated features; perform one-to-one matching on the multiple lane line information and the data ground truth in the training set through the Hungarian matching algorithm to obtain lane line matching information; calculate a total loss function based on the lane line matching information, and train the lane line detection model based on the total loss function until a preset training end condition is satisfied.

[0067] Optionally, the total loss function includes a first loss function, a second loss function, and a third loss function. The first loss function is used to calculate the class loss of the lane line, the second loss function is used to calculate the lane line point set loss, and the third loss function is used to calculate the starting point coordinate loss.

[0068] Optionally, the enhancement processing operation on the target scene image includes a cropping operation, a scaling operation, and a perspective transformation operation.

[0069] In the embodiment of the present application, by obtaining n-layer image features of different depths, the feature map contains both high-level semantic information and low-level position information; through hierarchical cross-attention operations, the queries of each lane line instance can interact with the corresponding layer features, fully mining the relevant information of the lane line instances in different layer features; by fusing the features of different layers to obtain lane line instance features, it can ensure that features of different scales and depths are comprehensively utilized, so that the lane line instance features have both the discriminability of high-semantic features and the sensitivity to position of low-semantic features; through lane line cross-attention operations, the relationship between lane lines can be captured. In a complex environment, the lane lines of the occluded part can be inferred based on the information of adjacent or relevant lane line instances, thereby effectively solving the occlusion problem and improving the robustness of lane line detection.

[0070] Figure 5 An entity structure diagram of an electronic device is illustrated, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a lane line detection method, which includes: obtaining an enhanced target scene image; obtaining n layers of image features with different feature depths according to the target scene image; inputting the n layers of image features into a hierarchical cross-attention decoder, performing cross-attention operations with trained lane line instances to obtain features of different layers, fusing the features of different layers to obtain multiple lane line instance features, where one layer of image features corresponds to one hierarchical cross-attention decoder; performing cross-attention operations on each lane line instance feature among the multiple lane line instance features with other lane line instance features to obtain updated features, and determining at least one lane line information according to the updated features; and determining the lane line detection result of the target scene image for the lane line information with a confidence level higher than the target threshold among the at least one lane line information.

[0071] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.

[0072] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the lane line detection method provided by the above-mentioned various methods. The method includes: obtaining an enhanced target scene image; obtaining n layers of image features with different feature depths according to the target scene image; inputting the n layers of image features into a hierarchical cross-attention decoder, performing cross-attention operations with trained lane line instances to obtain features of different layers, fusing the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross-attention decoder; performing cross-attention operations on each lane line instance feature among the multiple lane line instance features with other lane line instance features to obtain updated features, and determining at least one lane line information according to the updated features; determining the lane line information with a confidence level higher than a target threshold in the at least one lane line information as the lane line detection result of the target scene image.

[0073] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the lane line detection method provided by the above-mentioned various methods. The method includes: obtaining an enhanced target scene image; obtaining n layers of image features with different feature depths according to the target scene image; inputting the n layers of image features into a hierarchical cross-attention decoder, performing cross-attention operations with trained lane line instances to obtain features of different layers, fusing the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross-attention decoder; performing cross-attention operations on each lane line instance feature among the multiple lane line instance features with other lane line instance features to obtain updated features, and determining at least one lane line information according to the updated features; determining the lane line information with a confidence level higher than a target threshold in the at least one lane line information as the lane line detection result of the target scene image.

[0074] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lane line detection method, characterized in that: include: Acquire the enhanced processed target scene image; Obtaining n layers of image features with different feature depths according to the target scene image; Input the n-layer image features into a hierarchical cross attention decoder, perform cross attention operation with the trained lane line instance, obtain features of different layers, fuse the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross attention decoder; Performing a cross-attention operation on each lane line instance feature of the multiple lane line instance features and other lane line instance features to obtain an updated feature, and determining at least one lane line information according to the updated feature; The lane line information with a confidence level higher than a target threshold in the at least one lane line information is determined as the lane line detection result of the target scene image.

2. The lane line detection method according to claim 1, characterized in that: The step of obtaining n layers of image features with different feature depths according to the target scene image includes: The target scene image is input into an encoder to obtain n layers of image features with different feature depths. The encoder includes a backbone network and a pyramid network. The pyramid network is used for upsampling and fusing n layers of feature maps with different depths.

3. The lane line detection method according to claim 1, characterized in that: Before acquiring the enhanced target scene image, the method further includes: Determine n layers of image features of the first scene image in the training set; Performing cross-attention operations on n-layer image features of the first scene image and m learnable instance queries in the lane line detection model to obtain features of different layers, and fusing the features of different layers to obtain m lane line instance features; Performing a cross-attention operation on each lane line instance feature of the m lane line instance features and other lane line instance features to obtain m updated features, and determining a plurality of lane line information according to the m updated features; The lane line information is matched one-to-one with the true value of the data in the training set by the Hungarian matching algorithm to obtain lane line matching information; A total loss function is calculated based on the lane line matching information, and the lane line detection model is trained based on the total loss function until a preset training end condition is met.

4. The lane line detection method according to claim 2, characterized in that: The total loss function includes a first loss function, a second loss function and a third loss function. The first loss function is used to calculate the category loss of the lane line, the second loss function is used to calculate the lane line point set loss, and the third loss function is used to calculate the starting point coordinate loss.

5. The lane line detection method according to any one of claims 1 to 4, characterized in that: The enhancement processing operation on the target scene image includes a cropping operation, a scaling operation and a perspective transformation operation.

6. A lane line detection device, characterized in that: include: Acquisition module and processing module; The acquisition module is used to acquire the enhanced target scene image; The processing module is used to obtain n layers of image features with different feature depths according to the target scene image; Input the n-layer image features into a hierarchical cross attention decoder, perform cross attention operation with the trained lane line instance, obtain features of different layers, fuse the features of different layers to obtain multiple lane line instance features, and one layer of image features corresponds to one hierarchical cross attention decoder; Each lane line instance feature among the multiple lane line instance features is cross-attentionally operated with other lane line instance features to obtain an updated feature, and at least one lane line information is determined based on the updated feature; the lane line information among the at least one lane line information whose confidence level is higher than a target threshold is determined as the lane line detection result of the target scene image.

7. The lane line detection device according to claim 6, characterized in that: The processing module is used to input the target scene image into an encoder to obtain n layers of image features with different feature depths. The encoder includes a backbone network and a pyramid network. The pyramid network is used to upsample and fuse n layers of feature maps with different depths.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the lane line detection method as described in any one of claims 1 to 5 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the lane line detection method according to any one of claims 1 to 5 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the lane line detection method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Target detection method and device based on registration instance segmentation, equipment and medium

    CN121121162A