Traffic signal light detection method and device, electronic equipment and storage medium

By detecting the pose and feature information of candidate traffic lights in images acquired from vehicles, and combining navigation paths and vehicle coordinates, a comprehensive similarity algorithm is used to improve the accuracy of traffic light detection. This solves the problem of insufficient detection accuracy at low cost, and enhances the safety and user experience of autonomous driving.

CN117011825BActive Publication Date: 2026-01-09GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310765699.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2026-01-09
Estimated Expiration
2043-06-26

AI Technical Summary

Technical Problem

Existing traffic light detection methods are low in accuracy due to their low cost, especially those based on image acquisition, which are not accurate enough and affect the safety of autonomous driving.

Method used

By detecting images acquired from vehicles, the pose and feature information of candidate traffic lights are determined. Combining navigation paths and vehicle coordinates, a comprehensive similarity algorithm is used to improve detection accuracy and reduce the precision requirements of image acquisition equipment. Low-precision map information is used to determine the detection results of three-dimensional traffic lights.

Benefits of technology

It improves the accuracy of traffic light detection, reduces the probability of autonomous vehicles violating traffic regulations, and enhances driving safety and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011825B_ABST
    Figure CN117011825B_ABST
Patent Text Reader

Abstract

The application provides a traffic signal lamp detection method and device, electronic equipment and storage medium. The method comprises the following steps: detecting an image obtained by a vehicle to obtain detection information of a candidate traffic signal lamp; determining two-dimensional target traffic signal lamp information in a navigation path, and obtaining candidate ground truth information of the target traffic signal lamp in the image based on the target traffic signal lamp information and a vehicle coordinate; determining comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information; determining target detection information of the detection information of the candidate traffic signal lamp with the highest comprehensive similarity to the candidate ground truth information, and determining target pose information of the candidate ground truth information with the highest similarity to the target detection information. The application can improve the accuracy of traffic signal lamp detection results based on two-dimensional information of the traffic signal lamp at a low cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent driving, and more particularly to a traffic signal lamp detection method and device, an electronic device and a storage medium. BACKGROUND

[0002] Automatic driving is an important development direction of automobiles, and the traffic environment perception technology of vehicles is crucial to the development of automatic driving technology. In the traffic environment perception technology of vehicles, the detection of traffic signal lamps is an important part of environment perception. Existing traffic signal lamp detection schemes can be performed by a detection method based on vehicle-to-road communication, a detection method based on the state of surrounding vehicles, a detection method based on GPS navigation, and a detection method based on vision, etc. Among them, the detection method based on vision is mainly realized by the way of "multi-sensor fusion + high-precision map", but due to the high cost, long cycle and great technical difficulty of making a high-precision map, the detection efficiency of traffic signal lamps is low, and if only the images collected by the image collection sensor are recognized, the accuracy of the detection result is low. Therefore, how to improve the accuracy of the detection result of traffic signal lamps on the basis of low cost has become a problem to be solved. SUMMARY

[0003] In view of this, the embodiments of the present application provide a traffic signal lamp detection method and device, an electronic device and a storage medium to improve the above problems.

[0004] According to a first aspect of the embodiments of the present application, a traffic signal lamp detection method is provided, which comprises: detecting an image obtained by a vehicle to obtain detection information of a candidate traffic signal lamp, the detection information comprising pose information and feature information; determining two-dimensional target traffic signal lamp information in a navigation path of the vehicle in a navigation map according to the navigation path, and obtaining candidate ground truth information of a target traffic signal lamp in the image based on the two-dimensional target traffic signal lamp information and the vehicle coordinates, the candidate ground truth information comprising candidate ground truth pose information and candidate ground truth feature information; determining a comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information according to a similarity between the pose information of the candidate traffic signal lamp and the candidate ground truth pose information and a similarity between the feature information of the candidate traffic signal lamp and the candidate ground truth feature information; determining the detection information with the highest comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information as target detection information, and determining the candidate ground truth pose information with the highest similarity to the target detection information as target pose information.

[0005] According to a first aspect of the embodiment of the present application, the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information is determined according to the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate true value feature information, which includes: determining a candidate feature block according to the candidate true value pose information, and determining a target feature block according to the detection information of the candidate traffic signal lamp; determining the feature similarity and the pose similarity between the candidate feature block and the target feature block; and determining the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information based on the pose similarity and the feature similarity. In this embodiment, a plurality of candidate feature blocks corresponding to a plurality of candidate true value information and a plurality of target feature blocks corresponding to a plurality of detection information are determined first, and then the feature similarity and the pose similarity between the plurality of candidate feature blocks and the plurality of target feature blocks are determined, and further the similarity between the plurality of candidate true value information and the plurality of detection information is determined according to the feature similarity and the pose similarity between the plurality of candidate feature blocks and the plurality of target feature blocks. In this way, the detection result of the target traffic signal lamp can be determined based on the similarity between the plurality of candidate true value information and the plurality of detection information, and the detection accuracy of the traffic signal lamp is improved by improving the matching accuracy between the plurality of candidate true value information and the plurality of detection information.

[0006] According to a first aspect of the embodiment of the present application, the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information is determined according to the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate true value feature information, which includes: determining a candidate feature block according to the candidate true value pose information, and determining a target feature block according to the detection information of the candidate traffic signal lamp; determining the feature similarity and the pose similarity between the candidate feature block and the target feature block; and determining the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information based on the pose similarity and the feature similarity. In this embodiment, a plurality of candidate feature blocks corresponding to a plurality of candidate true value information and a plurality of target feature blocks corresponding to a plurality of detection information are determined first, and then the feature similarity and the pose similarity between the plurality of candidate feature blocks and the plurality of target feature blocks are determined, and further the similarity between the plurality of candidate true value information and the plurality of detection information is determined according to the feature similarity and the pose similarity between the plurality of candidate feature blocks and the plurality of target feature blocks. In this way, the detection result of the target traffic signal lamp can be determined based on the similarity between the plurality of candidate true value information and the plurality of detection information, and the detection accuracy of the traffic signal lamp is improved by improving the matching accuracy between the plurality of candidate true value information and the plurality of detection information.

[0007] According to a first aspect of the embodiment of the present application, the matching value between the detection information of the candidate traffic signal lamp and the candidate true value information is determined based on the pose similarity and the feature similarity, comprising: determining a current passable direction indicated by the target traffic signal lamp; determining whether the current passable direction is the same as a target passable direction indicated by the navigation path; if the same, determining a matching additional value, and determining the matching value between the detection information of the candidate traffic signal lamp and the candidate true value information based on the pose similarity, the feature similarity and the matching additional value; or if different, determining the matching value between the detection information of the candidate traffic signal lamp and the candidate true value information based on the pose similarity and the feature similarity. The embodiment determines the matching additional value by determining whether the current passable direction indicated by the target traffic signal lamp is the same as the target passable direction indicated by the navigation path, so as to reasonably punish the mis-detection when the detection model detects the image of the target traffic signal lamp, to improve the matching strategy when the perception mis-detects, to improve the accuracy of the matching value, and to further improve the accuracy of the detection result of the target traffic signal lamp.

[0008] According to a first aspect of the embodiment of the present application, the image obtained by the vehicle is detected to obtain the detection information of the candidate traffic signal lamp, comprising: performing grayscale processing on the image obtained by the vehicle to obtain a grayscale image corresponding to the image obtained by the vehicle; determining a grayscale value corresponding to each pixel according to the grayscale image, and performing cluster analysis on the grayscale value to obtain a brightness threshold interval corresponding to the grayscale image; determining a detection model corresponding to the brightness threshold interval from a plurality of detection models according to the brightness threshold interval corresponding to the grayscale image, wherein the plurality of detection models are obtained by individually training image data of different brightness threshold intervals obtained by the vehicle; inputting the image obtained by the vehicle into the detection model corresponding to the brightness threshold interval, and the detection model corresponding to the brightness threshold interval is used to detect the image obtained by the vehicle to obtain the detection information of the candidate traffic signal lamp. The embodiment determines a plurality of brightness threshold intervals by performing grayscale processing on the image of the target traffic signal lamp, so as to be able to train a plurality of detection models based on the image data corresponding to the determined plurality of brightness threshold intervals, and then determine the corresponding detection model based on the trained plurality of detection models and the brightness threshold interval corresponding to the image, so as to detect the image of the candidate traffic lamp according to the determined detection model, to avoid the high mis-detection probability caused by uneven illumination, and to ensure the accuracy of the detection result.

[0009] According to a first aspect of the embodiment of the present application, the candidate ground truth information of the obtained image of the target traffic signal lamp in the image is obtained based on the target traffic signal lamp information and the vehicle coordinates, including: obtaining attribute information of the traffic signal lamp, wherein the attribute information includes a plurality of specification information and a plurality of installation height information; determining candidate ground truth box information based on the plurality of specification information and the plurality of installation height information, the candidate ground truth box information including size information and installation height information; determining three-dimensional coordinate information of the target traffic lamp in the vehicle coordinate system based on the candidate ground truth box information, the target traffic lamp information, and the vehicle coordinates; and projecting the three-dimensional coordinate information of the target traffic signal lamp in the vehicle coordinate system to the image to obtain the candidate ground truth information of the target traffic signal lamp in the image. In this embodiment, a plurality of candidate specification information and a plurality of candidate installation height information are determined in the obtained attribute information of the traffic signal lamp according to the preset size information and the preset height information, and then the candidate ground truth box information is determined based on the plurality of candidate specification information and the plurality of candidate installation height information. Then, the three-dimensional coordinate information of the target traffic lamp in the vehicle coordinate system is determined based on the candidate ground truth box information, the target traffic signal lamp information, and the vehicle coordinates, so as to determine the candidate ground truth information of the target traffic signal lamp in the image. Since the obtained attribute information of the traffic signal lamp is based on the real data analysis of the general traffic signal lamp on the market, the accuracy of the candidate ground truth information is ensured, and the accuracy of the detection result of the target traffic signal lamp is ensured.

[0010] According to a second aspect of the embodiment of the present application, a traffic signal lamp detection device is provided, including: a detection information determination module, configured to detect an image obtained by a vehicle to obtain detection information of a candidate traffic signal lamp, the detection information including pose information and feature information; a candidate ground truth information determination module, configured to determine two-dimensional target traffic signal lamp information in a navigation path of the vehicle in a navigation map based on the navigation path, and obtain candidate ground truth information of a target traffic signal lamp in the image based on the two-dimensional target traffic signal lamp information and vehicle coordinates, the candidate ground truth information including candidate ground truth pose information and candidate ground truth feature information; a comprehensive similarity determination module, configured to determine a comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information according to a similarity between the pose information of the candidate traffic signal lamp and the candidate ground truth pose information and a similarity between the feature information of the candidate traffic signal lamp and the candidate ground truth feature information; and a target information determination module, configured to determine target detection information as the detection information in the candidate traffic signal lamp with the highest comprehensive similarity with the candidate ground truth information, and determine target pose information as the candidate ground truth pose information with the highest similarity with the target detection information.

[0011] According to a third aspect of the embodiments of the present application, an electronic device is provided, comprising: a processor; a memory, the memory having computer readable instructions stored thereon, the computer readable instructions, when executed by the processor, implementing the traffic signal detection method as described above.

[0012] According to a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided, having computer readable instructions stored thereon, the computer readable instructions, when executed by a processor, implementing the traffic signal detection method as described above.

[0013] In the scheme of the present application, the detection result of the three-dimensional traffic signal is determined based on the information of the two-dimensional traffic signal on the low-precision map, the precision requirement of the image acquisition device of the vehicle is reduced, the cost of obtaining the information of the traffic signal is reduced, the comprehensive similarity is determined based on the pose information and the feature information, and then the detection result of the traffic signal is determined, so as to improve the accuracy of the detection result of the traffic signal, and further reduce the probability of the vehicle violating the traffic rules due to the automatic driving function, thereby effectively improving the driving safety performance of the automatic driving vehicle and improving the user experience.

[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0015] The drawings herein are incorporated into the specification and form a part of the specification, show embodiments consistent with the present application, and together with the specification serve to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 is a flowchart of the traffic signal detection method according to an embodiment of the present application.

[0017] Figure 2 is a flowchart of the traffic signal detection method according to another embodiment of the present application.

[0018] Figure 3 is a flowchart of the specific steps of step 250 according to an embodiment of the present application.

[0019] Figure 4 is a flowchart of the traffic signal detection method according to still another embodiment of the present application.

[0020] Figure 5 is a flowchart of the training process of the detection model according to an embodiment of the present application. is a flowchart of the training process of the detection model according to an embodiment of the present application.

[0021] Figure 6 is a flow chart of a method for detecting a traffic signal according to another embodiment of the present application.

[0022] Figure 7 is a block diagram of a device for detecting a traffic signal according to an embodiment of the present application.

[0023] Figure 8 is a hardware structure diagram of an electronic device according to an embodiment of the present application.

[0024] The specific embodiments have been shown by way of illustration in the drawings and more details will be described in the following, these drawings and the written description are not intended to limit the scope of the inventive concept in any way, but to explain the inventive concept to those skilled in the art through specific embodiments. DETAILED DESCRIPTION

[0025] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that this disclosure will be thorough and complete, and will fully convey the inventive concept to those skilled in the art.

[0026] It should be noted that the terms "first", "second", and the like, used in the description and the claims of the present application as well as the above-described drawings imply similar objects without necessarily implying mentioned particular order or sequence. It should be understood that the use of the data thus used is interchangeable so that the embodiments of the application described herein can be implemented in other than the order discussed herein. Additionally, the terms "comprising" and "including" and any variations thereof are intended to cover a non-exclusive inclusion, for example, a process, method, system, product or apparatus that comprises a list of steps or units not necessarily comprising only those steps or units that are clearly listed or that are inherent to such process, method, product or apparatus. Further, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures, devices, and operations have not been shown or described in detail to avoid obscuring aspects of the application.

[0027] Further, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures, devices, and operations have not been shown or described in detail to avoid obscuring aspects of the application.

[0028] The block diagrams shown in the drawings are merely functional entities and do not necessarily have to correspond to physically independent entities. That is, the functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices. The flowcharts shown in the drawings are merely exemplary and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further divided, and some operations / steps can be combined or partially combined, so that the actual execution order can be changed according to actual conditions.

[0029] Please refer to Figure 1 , Figure 1 A traffic signal lamp detection method provided by an embodiment of the application is shown. In specific embodiments, the traffic signal lamp detection method can be applied to a traffic signal lamp detection device 600 as shown in Figure 7 and an electronic device 700 configured with the traffic signal lamp detection device 600. Figure 8 The specific flow of the embodiment will be described below. It can be understood that the method can be executed by an electronic device with computing processing capability, such as a server, a smartphone in communication connection with a vehicle, a smart wearable device, a vehicle processor, etc., which is not specifically limited here. The traffic signal lamp detection method will be described in detail below with respect to the flow shown in Figure 1 The traffic signal lamp detection method can specifically include the following steps:

[0030] Step 110: detecting an image acquired by a vehicle to obtain detection information of a candidate traffic signal lamp, the detection information including pose information and feature information.

[0031] As one way, the image of the candidate traffic signal lamp can be collected in real time by an image collection sensor of the vehicle, and the collected image can be taken as an input of a detection model, so as to detect the image of the target traffic signal lamp by the detection model to obtain the detection information of the candidate traffic signal lamp. Optionally, when the detection model detects the image of the candidate traffic signal lamp, a plurality of feature maps of the image of the candidate traffic signal lamp can be determined first, and then the plurality of feature maps are detected respectively to obtain a plurality of detection information. Optionally, the detection model can be constructed by a plurality of neural networks, such as a fully connected neural network, a convolutional neural network, a recurrent neural network, etc., and the neural network of the detection model can be set according to actual needs.

[0032] In order to ensure the accuracy of the detection information, the detection model needs to be trained in advance. Specifically, a sample set is constructed in advance, which includes a plurality of sample images and label information of the sample images, wherein the label information of the sample images is used to indicate the detection information indicated by the detection task corresponding to the sample image. Among them, the sample set includes images of a plurality of traffic lights.

[0033] During the training process, a sample image is input into the detection model for feature extraction, obtaining a plurality of feature maps of the sample image, and then the plurality of feature maps are detected to output a plurality of sample detection information corresponding to the detection task corresponding to the sample image. It can be understood that the plurality of sample detection information indicates whether the sample image includes the object indicated by the detection task corresponding to the sample image, and the detection information indicated by the detection task corresponding to the sample image.

[0034] Then, the loss value of the loss function is calculated based on the label information of the sample image and the sample detection information corresponding to the sample image. If the loss value does not converge, the parameters of the detection model are adjusted in reverse, and the sample detection information is output again by the detection model with adjusted parameters for the sample image, and the loss value of the loss function is calculated again until the loss value converges.

[0035] For each sample image, the above process is repeated, and when the training end condition is reached, the training of the detection model is ended. Then, the detection model is used to perform online detection on the image of the target traffic light, which can ensure the accuracy of the detection information.

[0036] As another way, a detection model based on a learning algorithm can be used to recognize the image obtained by the vehicle, that is, during the detection process of the detection model on the image of the candidate traffic light, the image of the candidate traffic light and the corresponding plurality of detection information are used as training data of the detection model, so as to train the detection model, so that the detection model detects the traffic light more accurately.

[0037] Step 120, determining two-dimensional target traffic light information in the navigation path of the vehicle in the navigation map based on the navigation path of the vehicle in the navigation map, and obtaining candidate ground truth information of the target traffic light in the image based on the two-dimensional target traffic light information and the vehicle coordinates, the candidate ground truth information includes candidate ground truth pose information and candidate ground truth feature information.

[0038] As a manner, the navigation path can be a path planned by the navigation system based on the destination input by the user and the current location of the vehicle. Alternatively, the navigation path determined by the navigation system can be a plurality of paths, and the path selected by the user in the path planned by the navigation system can be determined as the navigation path. Alternatively, the navigation system can be a navigation system of the vehicle or a navigation system in an electronic device in communication connection with the vehicle, wherein the electronic device in communication connection with the vehicle can be a smart phone, a tablet computer, a smart watch, etc., and the communication connection can be a Bluetooth connection, ZigBee, WiFi, etc., which are not specifically limited here.

[0039] As a manner, the map can be a local map corresponding to the navigation path, and alternatively, the map can be a two-dimensional map including low-precision two-dimensional information such as position information and direction information of the traffic signal light passed in the navigation path, and further, the target traffic signal light matched with the navigation path can be determined from a plurality of two-dimensional traffic signal lights in the map. Alternatively, the determined target traffic signal light information includes low-precision two-dimensional information.

[0040] As a manner, the navigation information of the navigation system is updated in real time according to the driving state of the vehicle, and the navigation system can obtain the driving information (for example, driving speed, driving direction, current position information of the vehicle) of the vehicle in real time, and update the navigation information according to the driving information of the vehicle. Alternatively, the navigation system can update the navigation information based on a preset frame rate, which is related to the calculation resource speed of the vehicle and the communication speed between nodes of the vehicle, for example, the preset frame rate can be set to 10 Hz. Alternatively, the control system of the vehicle can control the speed and driving direction of the vehicle according to the real-time updated navigation information. For example, if the navigation information indicates that the current driving needs to be decelerated, the control system of the vehicle controls the vehicle to decelerate according to the navigation information.

[0041] As a manner, after the two-dimensional target traffic signal light is determined according to the navigation path, the candidate ground truth information of the target traffic signal light in the image can be determined according to the two-dimensional target traffic signal light information and the vehicle coordinates. Alternatively, the two-dimensional target traffic signal light information can be projected into the vehicle coordinate system according to the coordinates of the vehicle in the world coordinate system, so as to obtain the two-dimensional target traffic signal light information in the vehicle coordinate system, and the candidate ground truth information of the installation height and size of the traffic signal light can be selected according to the market research and literature analysis results, the candidate ground truth box information is generated, the three-dimensional coordinate information of the target traffic light in the vehicle coordinate system is obtained according to the candidate ground truth box information and the two-dimensional target traffic signal light information in the vehicle coordinate system, and finally the three-dimensional target traffic signal light information is converted into the image coordinate system to obtain the candidate ground truth pose information and the candidate ground truth feature information.

[0042] As a manner, the candidate value information can be obtained by sending a request of querying current market high frequency adopted traffic signal lamp specification and installation height information to the cloud by the navigation system, sending the queried traffic signal lamp specification and installation height information to the navigation system by the cloud, and obtaining the candidate true value information of the traffic signal lamp of different specification and installation height combination by the navigation system according to the received traffic signal lamp specification and installation height.

[0043] Step 130, determining the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information according to the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate true value feature information.

[0044] As a manner, the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information can be determined by calculating the cosine distance between the pose information of the candidate traffic signal lamp and the candidate true value pose information, and the similarity between the feature information of the candidate traffic signal lamp and the candidate true value feature information can be determined by calculating the cosine distance between the feature information of the candidate traffic signal lamp and the candidate true value feature information. Alternatively, at least one of the Pearson correlation coefficient, the Jaccard similarity coefficient and the Tanaka coefficient between the pose information of the candidate traffic signal lamp and the candidate true value pose information can be calculated to determine the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information.

[0045] As a manner, the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information can be determined by calculating the average similarity of the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate true value feature information. Alternatively, the corresponding weight of the similarity between different information can be set in advance, and then the comprehensive similarity between the detection information and the candidate true value information is determined according to the weight.

[0046] Step 150, determining the detection information with the highest comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information as the target detection information, and determining the candidate true value pose information with the highest similarity to the target detection information as the target pose information.

[0047] As a manner, in order to determine the target detection information corresponding to the current state of the target traffic signal light in the plurality of detection information, the detection information with the highest comprehensive similarity between the plurality of candidate true value information in the plurality of detection information can be determined as the target detection information, so as to determine the target pose information of the target traffic signal light based on the detection information with the highest similarity. Among them, the plurality of detection information can include the indicated passable direction, color, category and the like of the target traffic signal light. Optionally, since the plurality of candidate true value pose information respectively includes the installation information and the specification information of the corresponding traffic signal light, in order to accurately determine the installation information and the specification information of the candidate traffic signal light, the candidate true value pose information with the highest similarity with the target detection information is determined as the target pose information, so as to facilitate determining the height, angle and the like information of the target traffic signal light based on the target pose information.

[0048] As a manner, when the target detection information and the target pose information are determined, the state information corresponding to the target detection information is determined as the state information of the target traffic signal light. Optionally, the state information can include the current category, color and passable direction of the traffic signal light.

[0049] Optionally, after the target pose information and the target detection information of the candidate traffic signal light are determined, the control system of the vehicle can control the driving state of the vehicle based on the target pose information and the target detection information of the target traffic signal light. For example, if the target pose information and the target detection information of the target traffic signal light indicate that the straight direction of the traffic signal light at the straight intersection 500 meters in front is red, the control system of the vehicle controls the vehicle to slow down or brake based on the target detection information and the target pose information.

[0050] In the embodiment of the present application, the detection information of the candidate traffic signal light is obtained by detecting the image obtained by the vehicle, and then the two-dimensional target traffic signal light corresponding to the navigation path in the navigation map is determined according to the vehicle, so as to determine the candidate true value information of the target traffic signal light in the image according to the two-dimensional target traffic signal light and the vehicle coordinates, wherein the candidate true value information includes candidate true value pose information and candidate true value feature information, so as to determine the target pose information and the target detection information based on the comprehensive similarity between the candidate true value pose information and the candidate feature information and the detection information of the candidate traffic signal light. Thereby, the accuracy requirement of the image acquisition device of the vehicle is reduced, the cost of obtaining the information of the traffic signal light is reduced, the detection result of the traffic signal light is determined based on the similarity, so as to improve the accuracy of the detection result of the traffic signal light, and further reduce the probability of the vehicle violating the traffic rules due to the automatic driving function, thereby effectively improving the driving safety performance of the automatic driving vehicle and improving the user experience.

[0051] Please refer toFigure 2 , Figure 2 A traffic signal detection method is shown. The following will be described in detail with respect to the flow shown in the figure. The traffic signal detection method can include the following steps. Figure 2

[0052] In step 210, the image acquired by the vehicle is detected to obtain detection information of a candidate traffic signal, which includes pose information and feature information.

[0053] In step 220, two-dimensional target traffic signal information in a navigation path of the vehicle in a navigation map is determined, and candidate ground truth information of a target traffic signal in the image is obtained based on the two-dimensional target traffic signal information and the vehicle coordinates, which includes candidate ground truth pose information and candidate ground truth feature information.

[0054] In step 230, a candidate feature block is determined according to the candidate ground truth pose information, and a target feature block is determined according to the detection information of the candidate traffic signal.

[0055] As a manner, when the image acquired by the vehicle is detected, the target feature block of the candidate traffic signal in the image can be determined based on a feature map layer of the image, and then the candidate ground truth pose information of the target traffic signal is projected into the image to determine the candidate feature block of the target traffic signal in the image. Optionally, when the image acquired by the vehicle is detected, the detection information of the candidate traffic signal and the position of the target feature block in the image can be included in the output detection result, wherein the position can be the coordinate position information of the target feature block in the image coordinate system of the image. Optionally, the candidate feature block of the target traffic signal in the image can be determined by projecting the candidate ground truth pose information of the target traffic signal into the image.

[0056] In step 240, the feature similarity and the pose similarity between the candidate feature block and the target feature block are determined.

[0057] As a manner, the feature similarity between the plurality of candidate feature blocks and the plurality of target feature blocks can be determined by calculating the cosine distance between the plurality of candidate feature blocks and the plurality of target feature blocks. Optionally, the cosine distance between the feature blocks can be calculated by the formula , wherein feature t is the candidate feature block; feature p is the target feature block; d(feature t ,feature p ) is the cosine distance between the candidate feature block and the target feature block; x i ​is a one-dimensional feature vector corresponding to the candidate feature block; y i is a one-dimensional feature vector corresponding to the target feature block, N is the dimension number of the one-dimensional feature vector, when the dimensions of the candidate feature block and the target feature block are inconsistent, the same dimension is aligned by filling 0 values, and N is the dimension of the aligned one-dimensional vector.

[0058] As one way, the pose similarity between the plurality of candidate feature blocks and the plurality of target feature blocks can be determined by calculating the Euclidean distance between the plurality of candidate feature blocks and the plurality of target feature blocks. Optionally, before calculating the Euclidean distance between the plurality of candidate feature blocks and the plurality of target feature blocks, the offset proportion of the length and width of the region where each candidate feature block is located in the length and width of the image of the candidate traffic signal lamp is determined, and the offset proportion of the length and width of the region where each target feature block is located in the length and width of the image of the candidate traffic signal lamp is determined. Optionally, the four corner point coordinate information of each candidate feature block and each target feature block in the image is determined with the upper left corner of the image of the candidate traffic signal lamp as the coordinate origin, and then the offset proportions of the length and width corresponding to the four corner points of the target traffic signal lamp image are determined, and then the Euclidean distance between the plurality of candidate feature blocks and the plurality of target feature blocks can be determined. Optionally, the Euclidean distance between the plurality of candidate feature blocks and the plurality of target feature blocks can be calculated by the following formula:

[0059]

[0060] wherein, d(location t ,location p ) is the Euclidean distance between the candidate feature block and the target feature block, is the horizontal coordinate offset ratio of the upper left corner point of the candidate feature block; is the horizontal coordinate offset ratio of the upper left corner point of the target feature block; is the vertical coordinate offset ratio of the upper left corner point of the candidate feature block; is the vertical coordinate offset ratio of the upper left corner point of the target feature block; is the horizontal coordinate offset ratio of the lower left corner point of the candidate feature block; is the horizontal coordinate offset ratio of the lower left corner point of the target feature block; is the vertical coordinate offset ratio of the lower left corner point of the candidate feature block; is the vertical coordinate offset ratio of the lower left corner point of the target feature block; is the horizontal coordinate offset ratio of the upper right corner point of the candidate feature block; is the horizontal coordinate offset ratio of the upper right corner point of the target feature block; is the vertical coordinate offset ratio of the upper right corner point of the candidate feature block; a horizontal coordinate offset ratio of a top-right corner point of the target feature block; a horizontal coordinate offset ratio of a bottom-right corner point of the candidate feature block; a horizontal coordinate offset ratio of a bottom-right corner point of the target feature block; a horizontal coordinate offset ratio of a bottom-right corner point of the candidate feature block; a horizontal coordinate offset ratio of a bottom-right corner point of the target feature block.

[0061] Step 250, determining a comprehensive similarity between the detection information of the candidate traffic signal and the candidate true value information based on the pose similarity and the feature similarity.

[0062] As a manner, a weight ratio of the pose similarity and the feature similarity can be preset, and then the comprehensive similarity between the candidate true value information and the detection information of the candidate traffic signal is determined based on the weight ratio. For example, if the weight ratio of the feature similarity is preset as 2 / 3 and the weight ratio of the pose similarity is preset as 1 / 3, the similarity between the plurality of candidate true value information and the plurality of detection information can be determined by the formula

[0063] In some embodiments, as shown in FIG. 2, the step 250 comprises: Figure 3

[0064] Step 251, determining a matching value between the detection information of the candidate traffic signal and the candidate true value information based on the pose similarity and the feature similarity.

[0065] As a manner, the matching value between the plurality of candidate true value information and the plurality of detection information of the candidate traffic signal can be determined based on the pose similarity and the feature similarity by a similarity matching formula, so as to determine the comprehensive similarity between the plurality of candidate true value information and the plurality of detection information according to the matching value. Optionally, the similarity matching formula is wherein, is a first regular coefficient, which is used for the matching value calculation between the plurality of candidate true value information and the plurality of detection information of the candidate traffic signal.

[0066] ​In some embodiments, the step 251 comprises: determining a current passable direction indicated by the target traffic signal light; determining whether the current passable direction is same as a target passable direction indicated by the navigation path; if same, determining a matching additional value, and determining a matching value between the detection information of the candidate traffic signal light and the candidate true value information based on the pose similarity, the feature similarity and the matching additional value; or if different, determining a matching value between the detection information of the candidate traffic signal light and the candidate true value information based on the pose similarity and the feature similarity.

[0067] As a manner, the current passable direction indicated by the target traffic signal light can be determined according to the state information of the traffic signal light corresponding to each detection information in the detection information.

[0068] As a manner, in order to avoid large similarity difference caused by false detection when detecting the detection model, a logical function is introduced to reasonably punish the calculation of the matching value, and the expression of the logical function is Sign(class t ==class p ), the meaning of the logical expression is that when class t ==class p , the function is true and the value is 1, otherwise the value is 0. Wherein, class t is the current passable direction indicated by the target traffic signal light; class p is the target passable direction indicated by the navigation path. That is, the function is used to determine whether the current passable direction indicated by the target traffic signal light corresponding to each detection information is same as the target passable direction indicated by the navigation path, and the calculation of the matching value is reasonably punished based on the result.

[0069] As a manner, the second regularization coefficient can be set in advance, the second regularization coefficient is multiplied by the value of the logical function as the matching additional value, when the current passable direction indicated by the target traffic signal light corresponding to a detection information is same as the target passable direction indicated by the navigation path, the matching additional value obtained by multiplying the second regularization coefficient and the value of the logical function is added to the value of the similarity matching formula to obtain the matching value between the detection information and the candidate true value information. Optionally, when it is determined that the current passable direction indicated by the target traffic signal light is same as the target passable direction indicated by the navigation path, the matching value between the detection information and the candidate true value information is determined based on the formula Wherein, d(bbox t ,bbox p ) is the matching value between the candidate true value information and the detection information; is the second regularization coefficient; is the matching additional value. Wherein, With The setting can be made based on the detection performance of the detection model, but the guarantee For example, the

[0070] Alternatively, when it is determined that the current passable direction indicated by the target traffic light is different from the target passable direction indicated by the navigation path, the logical function takes a value of 0, and no matching additional value is added, and the matching value between the plurality of detection information and the plurality of candidate true value information can be directly determined according to the formula shown in step 250.

[0071] In step 252, the average similarity between the detection information of the candidate traffic light and the candidate true value information is determined according to the matching value.

[0072] As a way, the matching value between each detection information and each candidate true value information can be determined, and then the sum of the matching values is determined, and the average similarity between each detection information and the plurality of candidate true value information is determined based on the sum of the matching values, which can be determined by the formula to determine the average similarity, wherein P is the average similarity; K is the number of candidate true value information; d(bbox t (bbox p ) is the matching value between each detection information and the plurality of candidate true value information.

[0073] In step 253, the difference between the average similarity and the preset value is determined, and the absolute value of the difference is determined as the comprehensive similarity between the detection information of the candidate traffic light and the candidate true value information.

[0074] As a way, since the determined average similarity can include the Euclidean distance and the cosine distance, the greater the average similarity, the less similar the corresponding candidate true value information and the detection information, and in order to determine the comprehensive similarity between each detection information and the plurality of candidate true value information, the preset value is subtracted from the average similarity, so as to obtain the comprehensive similarity between the detection information and the candidate true value information. The preset value can be 1, and the corresponding calculation formula can be: Wherein, confidence(bbox p ) is the comprehensive similarity.

[0075] Please continue to refer to Figure 2 In step 260, the detection information between the detection information of the candidate traffic light and the candidate true value information with the highest comprehensive similarity is determined as the target detection information, and the candidate true value pose information with the highest similarity to the target detection information is determined as the target pose information.

[0076] The specific step descriptions of steps 210-220 and step 260 can refer to steps 110-120 and step 140, which will not be described again here.

[0077] In this embodiment, the corresponding candidate feature block and the detection information of the candidate traffic signal lamp are determined through the candidate true value information, and then the feature similarity and the pose similarity between the candidate feature block and the target feature block are determined according to the candidate feature block and the target feature block. Then, the comprehensive similarity between the candidate true value information and the detection information of the candidate traffic signal lamp can be determined according to the feature similarity and the pose similarity between the feature blocks, and then the detection result of the target traffic signal lamp can be determined based on the comprehensive similarity between the candidate true value information and the detection information of the candidate traffic signal lamp. The detection accuracy of the traffic signal lamp is improved by improving the matching accuracy between multiple candidate true value information and multiple detection information.

[0078] Please refer to Figure 4 , Figure 4 A traffic signal lamp detection method provided by an embodiment of the application is shown. The following will be described in detail with respect to the flow shown in Figure 4 The traffic signal lamp detection method can specifically include the following steps:

[0079] In step 310, the image obtained by the vehicle is subjected to grayscale processing to obtain a grayscale image corresponding to the image obtained by the vehicle.

[0080] As a way, over-bright or over-dark lighting will cause the detection model to easily misdetect the color / direction of the traffic signal lamp when detecting the image of the candidate traffic signal lamp. Since the image brightness is related to the gray value, the higher the gray value, the brighter the image, and the stronger the corresponding lighting, the gray value of each pixel of the image of the target traffic signal lamp can be subjected to grayscale processing through the formula Gray = R*0.299 + G*0.587 + B*0.114, so as to obtain the gray value of each pixel of the image corresponding to the image, and then obtain the grayscale image of the target traffic signal lamp, wherein R, G and B correspond to red, green and blue in three primary colors, respectively.

[0081] In step 320, the gray value corresponding to each pixel is determined according to the grayscale image, and the gray value is subjected to cluster analysis to obtain a brightness threshold interval corresponding to the grayscale image.

[0082] As a manner, the different luminance threshold intervals can be determined by clustering analysis on the gray values, wherein each luminance threshold interval corresponds to a luminance grade, so as to determine the misjudgment in detecting the target traffic signal lamp in the same image due to the uneven illumination caused by the angle and the image acquisition sensor of the vehicle.

[0083] At step 330, the detection model corresponding to the luminance threshold interval of the gray image is determined from a plurality of detection models according to the luminance threshold interval of the gray image, wherein the plurality of detection models are obtained by separately training the image data of different luminance threshold intervals acquired by the vehicle.

[0084] As a manner, the detection model can be a learning-based detection model, that is, the detection model can be learned based on the images and detection information of the traffic signal lamp acquired by the vehicle in real time, so as to improve the detection ability of the detection model by learning the traffic signal lamp acquired in different situations.

[0085] Optionally, the corresponding detection model can be set for each luminance threshold interval, and then the detection model corresponding to each luminance threshold interval is separately trained. When applied offline, the corresponding detection model can be determined according to the luminance threshold interval corresponding to the gray image, so as to ensure the accuracy of detecting the gray image and improve the accuracy of the detection information.

[0086] As another manner, the detection model can be trained based on a plurality of groups of luminance data under different illuminations. As shown in Figure 5 The training of the detection model includes a preprocessing procedure, a training procedure and a testing procedure, wherein the preprocessing procedure refers to determining the illumination grade of a plurality of groups of traffic signal lamp images under different illuminations, performing image gray processing, luminance distribution statistics and other steps, obtaining a plurality of different luminance threshold intervals, and then taking these data as prior information for determining the luminance grade of the traffic signal lamp. Optionally, a plurality of groups of traffic signal lamp images under different illuminations can be acquired in advance, and the traffic signal lamps under different illuminations are grayed to obtain a plurality of groups of gray images. Then, the gray values of each group of gray images are counted respectively, and the luminance distribution curve corresponding to each group of gray images is determined according to the gray values of the gray images. Then, a plurality of groups of luminance threshold intervals are determined based on the luminance distribution curves corresponding to each group of gray images, wherein the group luminance difference in each group of luminance threshold intervals is the largest, and the group luminance difference is the smallest. Optionally, the corresponding traffic signal lamp images under a plurality of different illuminations can all be images under uniform illumination.

[0087] After determining the multiple sets of different brightness threshold intervals, a target detection CNN method is used to construct a detection model based on images of different brightness levels, and the detection model is trained according to the multiple sets of different brightness threshold intervals and the multiple sets of traffic signal lamp images under different illuminations corresponding to the multiple sets of different brightness threshold intervals. Optionally, the detection model can be trained and applied through a ResNet50-based SSD detection algorithm.

[0088] Optionally, after determining the multiple sets of different brightness threshold intervals, the brightness can be graded according to the multiple sets of different brightness threshold intervals, and each brightness grade and the corresponding brightness threshold interval can be used as prior information for determining the brightness level of the image of the target traffic signal lamp to test the trained detection model. The number of brightness levels is related to the imaging state of the image sensor. If the image sensor is extremely sensitive to light intensity, multiple brightness grades can be set for the same brightness threshold interval to improve its generalization ability under different illuminations.

[0089] The training process refers to training the detection model based on the data sets corresponding to different brightness levels determined in the preprocessing process to determine the model parameters corresponding to different brightness levels. The testing process refers to testing the detection model based on part of the data sets in the data sets corresponding to different brightness levels determined in the preprocessing process to verify whether the detection model can accurately determine the model parameters corresponding to different brightness levels for detection. Optionally, after determining the multiple sets of different brightness threshold intervals, the detection model can be individually trained according to the brightness levels corresponding to the multiple sets of different brightness threshold intervals and the corresponding images, and different model weights can be set for training the detection model for different brightness levels. Thus, when testing the detection model, the model weight of the corresponding detection model can be determined according to the brightness level in the test data, and the image of the corresponding brightness level can be detected. The training data set composed of the multiple sets of different brightness threshold intervals, the brightness levels corresponding to the multiple sets of different brightness threshold intervals, and the images can be split to obtain a training data set and a test data set, so as to select the parameters of the detection model with the best fitting ability as the model parameters corresponding to each brightness level by comprehensively training the set loss and the test set generalization ability, save the model parameters of different brightness levels, and then detect the image of the target traffic signal lamp based on the model parameters corresponding to the brightness of the image of the target traffic signal lamp to improve the detection ability of the detection model. Optionally, the data sets of different brightness levels can be split into a training data set and a test data set at a ratio of 4:1. If the data of the images under different illuminations is sufficient, the data set can also be split at a ratio of 1:1. The splitting ratio can be determined according to the actual data distribution and quantity, which is not specifically limited herein.

[0090] Please continue to refer to Figure 4, step 340, input the image obtained by the vehicle into the detection model corresponding to the brightness threshold interval, the detection model corresponding to the brightness threshold interval is used to detect the image obtained by the vehicle to obtain the detection information of the candidate traffic signal lamp.

[0091] As a way, after determining the detection model corresponding to the gray image, the image obtained by the vehicle is input into the detection model, so as to avoid the situation that the image cannot be accurately detected due to exposure or too dark light, and ensure the accuracy of the detection information of the candidate traffic signal lamp.

[0092] Step 350, determine the two-dimensional target traffic signal lamp information in the navigation path of the vehicle in the navigation map according to the navigation path of the vehicle in the navigation map, and obtain the candidate ground truth information of the target traffic signal lamp in the image based on the two-dimensional target traffic signal lamp information and the vehicle coordinates, the candidate ground truth information includes candidate ground truth pose information and candidate ground truth feature information.

[0093] Step 360, determine the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information according to the similarity between the pose information of the candidate traffic signal lamp and the candidate ground truth pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate ground truth feature information.

[0094] Step 370, determine the target detection information as the detection information with the highest comprehensive similarity between the candidate traffic signal lamp and the candidate ground truth information, and determine the target pose information as the candidate ground truth pose information with the highest similarity to the target detection information.

[0095] In this embodiment, a plurality of brightness threshold intervals are determined by gray processing of the target traffic signal lamp, so as to train a plurality of detection models based on the determined plurality of brightness threshold intervals, and then determine the corresponding detection model based on the trained plurality of detection models and the brightness threshold interval corresponding to the image, so as to detect the image of the candidate traffic light according to the determined detection model, avoid the situation that the false detection probability is large due to uneven illumination, and ensure the accuracy of the detection result.

[0096] Please refer to Figure 6 , Figure 6 The traffic signal lamp detection method provided by an embodiment of the application is shown. The following will be described in detail with respect to the flow shown in Figure 6 The traffic signal lamp detection method can specifically include the following steps:

[0097] Step 410, detecting the image obtained by the vehicle to obtain the detection information of the candidate traffic signal lamp, the detection information includes pose information and feature information.

[0098] In step 420, attribute information of the traffic signal light is acquired, wherein the attribute information includes multiple specification information and multiple installation height information.

[0099] As one way, the vehicle can send an instruction to acquire attribute information of the traffic signal light to a cloud server, and the cloud server can perform literature retrieval in the cloud and search in a market research website to statistically obtain specification information and installation information of the traffic signal light currently used in the market. The installation information can include installation height of the traffic signal light, and the specification information includes size information of length, width and height of different categories of traffic lights. Alternatively, the acquired attribute information of the traffic signal light is related to the application scenario. If the attribute information of the traffic signal light is common to each city, the literature retrieval range is set to be nationwide cities; if the attribute information of the traffic signal light is for a specified area, the retrieval range is set to be the specified area, so as to determine the attribute information of the traffic signal light in the specified area.

[0100] In step 430, candidate ground truth box information is determined based on the multiple specification information and the multiple installation height information, and the candidate ground truth box information includes size information and installation height information.

[0101] As one way, the size information and the installation height information can be set according to actual needs. The size information can be a radius, for example, [0.25m, 0.29m, 0.3m]. The installation height information of different categories of traffic signal lights is different. For example, the installation height information of the crossroad signal light is [3.0m, 3.1m, 3.2m, 3.3m, 3.4m, 3.5m]; the installation height information of the pedestrian crossing signal light is [2.0m, 2.1m, 2.2m, 2.3m, 2.4m, 2.5m]; the installation height information of the lane signal light is [5.5m, 5.7m, 5.9m, 6.1m, 6.3m, 6.5m, 6.7m, 6.9m, 7.1m, 7.3m, 7.5m]; and the installation height information of the non-motor vehicle lane signal light is [2.5m, 2.6m, 2.7m, 2.8m, 2.9m, 3.0m]. Alternatively, the number of candidate ground truth box information can be set to determine a certain number of candidate specification information and candidate installation information.

[0102] In step 440, three-dimensional coordinate information of the target traffic light in the vehicle coordinate system is determined based on the candidate ground truth box information, the target traffic light information and the vehicle coordinate.

[0103] As a manner, after the plurality of candidate installation information and the plurality of specification information are determined, permutation and combination are performed to obtain a plurality of sets of candidate ground truth box information, the candidate ground truth box information is matched with the target traffic light information, and then the candidate box information matched with the target traffic signal lamp is converted to a vehicle coordinate system according to the vehicle coordinate based on the conversion relationship between a Universal Transverse Mercator Grid System (UTM) coordinate system and the vehicle coordinate system, wherein the UTM coordinate system is a projection coordinate, the coordinate is represented by using a grid-based method, is a plane coordinate converted from a spherical latitude and longitude coordinate through a projection algorithm, that is, an XY coordinate, and the coordinate origin of the UTM coordinate system is located at the intersection of the prime meridian and the equator, the positive east direction is the positive direction of the x axis (UTM Easting), and the positive north direction is the positive direction of the y axis (UTM Northing).

[0104] Step 450, projecting the three-dimensional coordinate information of the target traffic signal lamp in the vehicle coordinate system to the image to obtain candidate ground truth information of the target traffic signal lamp in the image.

[0105] As a manner, the plurality of sets of candidate box information in the vehicle coordinate system can be further converted to the image coordinate system according to the vehicle coordinate and the conversion relationship between the vehicle coordinate system and the image coordinate system to obtain the candidate ground truth information of the target traffic signal lamp in the image.

[0106] Step 460, determining the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information according to the similarity between the pose information of the candidate traffic signal lamp and the candidate ground truth pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate ground truth feature information.

[0107] Step 470, determining the detection information with the highest comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information as target detection information, and determining the candidate ground truth pose information with the highest similarity to the target detection information as target pose information.

[0108] The specific step descriptions of steps 410- and steps 460- step 470 can refer to steps 110 and steps 130- step 140, and will not be described here.

[0109] In the embodiment, the three-dimensional coordinate information of the target traffic light in the vehicle coordinate system is determined based on the candidate ground truth information, the target traffic signal lamp information and the vehicle coordinates, so as to determine the candidate ground truth information of the target signal lamp in the image. Since the attribute information of the acquired traffic signal lamp is obtained based on the real data analysis of the general traffic signal lamp on the market, the accuracy of the candidate ground truth information is ensured, and the accuracy of the detection result of the target traffic signal lamp is ensured.

[0110] Figure 7 is a block diagram of a traffic signal lamp detection device according to an embodiment of the present application, as shown in the figure, the traffic signal lamp detection device 600 includes a detection information determination module 610, a candidate ground truth information determination module 620, a comprehensive similarity determination module 630 and a target information determination module 640. Figure 7

[0111] The detection information determination module 610 is used for detecting the image obtained by the vehicle to obtain the detection information of the candidate traffic signal lamp, and the detection information includes the pose information and the feature information. The candidate ground truth information determination module 620 is used for determining the two-dimensional target traffic signal lamp information in the navigation path of the vehicle in the navigation map based on the navigation path, and obtaining the candidate ground truth information of the target traffic signal lamp in the image based on the two-dimensional target traffic signal lamp information and the vehicle coordinates, and the candidate ground truth information includes the candidate ground truth pose information and the candidate ground truth feature information. The comprehensive similarity determination module 630 is used for determining the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate ground truth information according to the similarity between the pose information of the candidate traffic signal lamp and the candidate ground truth pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate ground truth feature information. The target information determination module 640 is used for determining the target detection information as the detection information with the highest comprehensive similarity between the candidate ground truth information in the detection information of the candidate traffic signal lamp, and determining the target pose information as the candidate ground truth pose information with the highest similarity between the target detection information.

[0112] ​In some embodiments, the comprehensive similarity determination module 630 comprises: a feature block determination submodule, configured to determine a candidate feature block according to the candidate ground truth pose information, and determine a target feature block according to the detection information of the candidate traffic signal light; a similarity determination submodule, configured to determine a feature similarity and a pose similarity between the candidate feature block and the target feature block; and a comprehensive similarity determination submodule, configured to determine a comprehensive similarity between the detection information of the candidate traffic signal light and the candidate ground truth information based on the pose similarity and the feature similarity.

[0113] In some embodiments, the comprehensive similarity determination submodule comprises: a matching value determination unit, configured to determine a matching value between the detection information of the candidate traffic signal light and the candidate ground truth information based on the pose similarity and the feature similarity; a comprehensive similarity determination unit, configured to average similarity determination unit, configured to determine an average similarity between the detection information of the candidate traffic signal light and the candidate ground truth information according to the matching value; determine a difference value between the average similarity and a preset value, and determine an absolute value of the difference value as the comprehensive similarity between the detection information of the candidate traffic signal light and the candidate ground truth information.

[0114] In some embodiments, the matching value determination unit comprises: a passable direction determination subunit, configured to determine a current passable direction indicated by the target traffic signal light; a judgment subunit, configured to determine whether the current passable direction is the same as a target passable direction indicated by the navigation path; a matching value first determination subunit, configured to, if the current passable direction is the same as the target passable direction, determine a matching additional value, and determine a matching value between the detection information of the candidate traffic signal light and the candidate ground truth information based on the pose similarity, the feature similarity and the matching additional value; or a matching value second determination subunit, configured to, if the current passable direction is different from the target passable direction, determine a matching value between the detection information of the candidate traffic signal light and the candidate ground truth information based on the pose similarity and the feature similarity.

[0115] In some embodiments, the detection information determination module 610 comprises: a greyscale processing submodule, configured to perform greyscale processing on the image acquired by the vehicle to obtain a greyscale image corresponding to the image acquired by the vehicle; a clustering analysis submodule, configured to determine a greyscale value corresponding to each pixel according to the greyscale image, and perform clustering analysis on the greyscale value to obtain a luminance threshold interval corresponding to the greyscale image; a detection model determination submodule, configured to determine, according to the luminance threshold interval corresponding to the greyscale image, a detection model corresponding to the luminance threshold interval from a plurality of detection models, wherein the plurality of detection models are obtained by separately training image data of different luminance threshold intervals acquired by a vehicle; and a detection submodule, configured to input the image acquired by the vehicle into the detection model corresponding to the luminance threshold interval, so that the detection model corresponding to the luminance threshold interval is used to detect the image acquired by the vehicle to obtain detection information of a candidate traffic signal lamp.

[0116] In some embodiments, the candidate ground truth information determination module 620 comprises: an attribute information acquisition submodule, configured to acquire attribute information of a traffic signal lamp, wherein the attribute information comprises a plurality of specification information and a plurality of installation height information; a candidate ground truth box information determination submodule, configured to determine candidate ground truth box information based on the plurality of specification information and the plurality of installation height information, wherein the candidate ground truth box information comprises size information and installation height information; a three-dimensional coordinate information determination submodule, configured to determine three-dimensional coordinate information of the target traffic signal lamp in a vehicle coordinate system based on the candidate ground truth box information, the target traffic signal information, and a vehicle coordinate; and a candidate ground truth information determination submodule, configured to project the three-dimensional coordinate information of the target traffic signal lamp in the vehicle coordinate system to the image to obtain candidate ground truth information of the target traffic signal lamp in the image.

[0117] According to an aspect of some embodiments of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the method in any of the above embodiments.

[0118] According to an aspect of some embodiments of the present application, an electronic device is also provided, as shown in the accompanying drawings. Figure 8 As shown in the accompanying drawings, the vehicle 700 comprises a processor 710 and one or more memories 720, the one or more memories 720 are configured to store program instructions executed by the processor 710, and the processor 710 executes the program instructions to implement the above-mentioned traffic signal lamp detection method.

[0119] Further, the processor 710 can include one or more processing cores. The processor 710 runs or executes instructions, programs, code sets or instruction sets stored in the memory 720, and calls data stored in the memory 720. Alternatively, the processor 710 can be implemented in at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 710 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processor (GPU), and a modem. Among them, the CPU is mainly used to process operating systems, user interfaces, and application programs; the GPU is used to be responsible for rendering and drawing display content; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor, but be realized by a separate communication chip.

[0120] According to an aspect of the present application, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device. The above computer readable storage medium carries computer readable instructions, which, when executed by a processor, implement the method in any of the above embodiments.

[0121] It should be noted that the computer-readable medium in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carrying computer-readable program code in a baseband or as a part of a carrier wave. Such a propagated data signal can take on various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium that can send, propagate or transmit a program for use by or in connection with an instruction execution system, device or apparatus. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination thereof.

[0122] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and they can also be executed in reverse order, depending on the involved functions. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0123] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope of the application being indicated by the following claims.

[0124] It is to be understood that the application is not limited to the precise details of construction and the exemplary embodiments described above and illustrated in the drawings. Various modifications and changes can be made thereunto without departing from the scope of the application. The scope of the application is indicated only by the appended claims.

Claims

1. A method of detecting a traffic signal, characterized by, The method comprises: detecting an image obtained by a vehicle to obtain detection information of a candidate traffic signal, the detection information comprising pose information and feature information; determining two-dimensional target traffic signal information in a navigation path of the vehicle in a navigation map, and obtaining attribute information of the traffic signal; wherein the attribute information comprises a plurality of specification information and a plurality of installation height information; wherein the specification information comprises size information of the length, width and height of traffic lights of different categories; arranging and combining the plurality of specification information and the plurality of installation height information to obtain a plurality of groups of candidate ground truth box information; each group of candidate ground truth box information comprises a size information and an installation height information; determining three-dimensional coordinate information of a target traffic signal in a vehicle coordinate system based on the two-dimensional target traffic signal information, the plurality of groups of candidate ground truth box information and the vehicle coordinates; projecting the three-dimensional coordinate information of the target traffic signal in the vehicle coordinate system to the image to obtain candidate ground truth information of the target traffic signal in the image, the candidate ground truth information comprising candidate ground truth pose information and candidate ground truth feature information; determining a comprehensive similarity between the detection information of the candidate traffic signal and the candidate ground truth information according to a similarity between the pose information of the candidate traffic signal and the candidate ground truth pose information and a similarity between the feature information of the candidate traffic signal and the candidate ground truth feature information; determining the detection information of the candidate traffic signal with the highest comprehensive similarity between the detection information of the candidate traffic signal and the candidate ground truth information as target detection information, and determining the candidate ground truth pose information with the highest similarity to the target detection information as target pose information.

2. The method of claim 1, wherein, The determination of the comprehensive similarity between the detection information of the candidate traffic signal and the candidate ground truth information according to the similarity between the pose information of the candidate traffic signal and the candidate ground truth pose information and the similarity between the feature information of the candidate traffic signal and the candidate ground truth feature information comprises: determining a candidate feature block according to the candidate ground truth pose information, and determining a target feature block according to the detection information of the candidate traffic signal; determining a feature similarity and a pose similarity between the candidate feature block and the target feature block; determining the comprehensive similarity between the detection information of the candidate traffic signal and the candidate ground truth information based on the pose similarity and the feature similarity.

3. The method of claim 2, wherein, The determination of the comprehensive similarity between the detection information of the candidate traffic signal and the candidate ground truth information based on the pose similarity and the feature similarity comprises: determining a matching value between the detection information of the candidate traffic signal and the candidate ground truth information based on the pose similarity and the feature similarity; determining an average similarity between the detection information of the candidate traffic signal and the candidate ground truth information according to the matching value; and determine a difference between the average similarity and a preset value, and determine an absolute value of the difference as a comprehensive similarity between the detection information of the candidate traffic signal light and the candidate true value information.

4. The method of claim 3, wherein, The determining the matching value between the detection information of the candidate traffic signal light and the candidate true value information based on the pose similarity and the feature similarity comprises: determining a current passable direction indicated by the target traffic signal light; determining whether the current passable direction is same as a target passable direction indicated by the navigation path; if the current passable direction is same as the target passable direction, determining a matching additional value, and determining the matching value between the detection information of the candidate traffic signal light and the candidate true value information based on the pose similarity, the feature similarity and the matching additional value; or if the current passable direction is different from the target passable direction, determining the matching value between the detection information of the candidate traffic signal light and the candidate true value information based on the pose similarity and the feature similarity.

5. The method according to any one of claims 1 to 4, characterized in that, The detecting the image obtained by the vehicle to obtain the detection information of the candidate traffic signal light comprises: performing grayscale processing on the image obtained by the vehicle to obtain a grayscale image corresponding to the image obtained by the vehicle; determining a grayscale value corresponding to each pixel according to the grayscale image, and performing clustering analysis on the grayscale value to obtain a luminance threshold interval corresponding to the grayscale image; determining a detection model corresponding to the luminance threshold interval from a plurality of detection models according to the luminance threshold interval corresponding to the grayscale image, wherein the plurality of detection models are obtained by separately training image data of different luminance threshold intervals obtained by the vehicle; inputting the image obtained by the vehicle into the detection model corresponding to the luminance threshold interval, and the detection model corresponding to the luminance threshold interval is used to detect the image obtained by the vehicle to obtain the detection information of the candidate traffic signal light.

6. A traffic signal detection apparatus characterized by comprising: The device comprises: a detection information determination module configured to detect an image obtained by a vehicle to obtain detection information of a candidate traffic signal light, wherein the detection information comprises pose information and feature information; a candidate true value information determination module configured to determine two-dimensional target traffic signal light information in a navigation path of the vehicle in a navigation map, and obtain attribute information of a traffic signal light, wherein the attribute information comprises a plurality of specification information and a plurality of installation height information, wherein the specification information comprises size information of lengths, widths and heights of different types of traffic lights, the plurality of specification information and the plurality of installation height information are arranged and combined to obtain a plurality of groups of candidate true value box information, each group of the candidate true value box information comprises one size information and one installation height information, three-dimensional coordinate information of a target traffic signal light in a vehicle coordinate system is determined based on the two-dimensional target traffic signal light information, the plurality of groups of the candidate true value box information and the vehicle coordinate, and the three-dimensional coordinate information of the target traffic signal light in the vehicle coordinate system is projected to the image to obtain candidate true value information of the target traffic signal light in the image, wherein the candidate true value information comprises candidate true value pose information and candidate true value feature information. The comprehensive similarity determination module is configured to determine the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information according to the similarity between the pose information of the candidate traffic signal lamp and the candidate true value pose information and the similarity between the feature information of the candidate traffic signal lamp and the candidate true value feature information. The target information determination module is configured to determine the detection information with the highest comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information as target detection information, and determine the candidate true value pose information with the highest similarity to the target detection information as target pose information.

7. The apparatus of claim 6, wherein, The comprehensive similarity determination module comprises: A feature block determination submodule configured to determine a candidate feature block according to the candidate true value pose information, and determine a target feature block according to the detection information of the candidate traffic signal lamp; A similarity determination submodule configured to determine a feature similarity and a pose similarity between the candidate feature block and the target feature block; A comprehensive similarity determination submodule configured to determine the comprehensive similarity between the detection information of the candidate traffic signal lamp and the candidate true value information based on the pose similarity and the feature similarity.

8. An electronic device, comprising: The electronic device comprises: A processor; A memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the method of any one of claims 1 to 6.

9. A computer readable storage medium, characterized in that, The computer readable storage medium stores program code, which can be called and executed by the processor to implement the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Traffic signal lamp state identification method, device and equipment and storage medium

    CN111310708A

  • Traffic signal lamp identification method and device, computer equipment and storage medium

    CN114639085A